Product overview
Real-Time Speech to Text
Agora's Real-Time Speech to Text (STT) transcribes live voice streams to deliver closed captions and transcription for enhanced accessibility. With advanced features like silent audio removal, it optimizes performance and reduces costs.
Transcribed text can be translated into multiple languages in real-time or used as input for AI models like GPT, seamlessly connecting real-time engagement with AI-powered applications.
Product Features
Live transcription for RTC
Integrated with Agora’s voice and video service, live transcription and captions improve accessibility for your audience. Perfect for meetings, live streaming, lectures, interviews, live shopping, and more.
Real-time translation (Beta)
Break down language barriers with live speech-to-text translation to multiple languages during real-time communication or live streaming. The high accuracy translation text, delivered with ultra low latency, can be integrated with LLMs for enhanced capabilities.
Cloud-based STT
Cloud-based service converts voice to text for active or specific hosts and then distributes the text to all participants in the channel for further processing. The service does not depend on the client's device performance and network conditions.
Speaker labeling
Label each transcribed text with the speaker's UID. Separate transcription of each host ensures accuracy even when multiple hosts are talking simultaneously.
Caption recording
Upload the transcriptions as .vtt files to cloud storage, then play back audio or video recordings with closed captions (CC). The timestamps in the .vtt file ensure that the text is perfectly synchronized with the audio or video, so that it appears exactly where it was generated.
Multi-language support
Real-time transcription supports all major languages and dialects, and each channel can support audio to text transcription for up to two languages simultaneously. Real-time translation supports translation of up to two source languages into five target languages with support for 30+ languages.