Intelligent transcription with Gemini 3.5 Transcribe

| Source: Google DeepMind Blog

Tags: Gemini, Google DeepMind, speech-to-text, transcription, ASR, voice AI, Gemini API

Google DeepMind launched Gemini 3.5 Transcribe, achieving 2.6% word error rate for pre-recorded audio and 4.0% for real-time streaming — now available via the Gemini API with speaker attribution, word-level timestamps, and 85-language support for developers building voice agents and call analytics pipelines.

Details

Google DeepMind's Gemini 3.5 Transcribe processes raw audio directly into formatted, polished text rather than raw transcripts. The model achieves a 2.6% WER for pre-recorded audio and 4.0% WER for real-time streaming as measured by Artificial Analysis, with strong performance in noisy environments and accurate capture of alphanumeric entities like postal codes and order IDs. Two separate APIs are available: the Live API (gemini-3.5-transcribe-live) for sub-second latency bidirectional streaming suited to interactive voice apps, and the Interactions API (gemini-3.5-transcribe) for pre-recorded audio processing with speaker attribution and word-level timestamps. Access is through Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform. Key capabilities include smart disfluency handling — auto-removing filler words and cleaning self-corrections such as 'let's meet Tuesday, no, Wednesday' — custom vocabulary support for specialized jargon, automatic detection across 85+ languages, and function calling that can delegate tasks like image generation to other Gemini models (currently in the macOS Gemini app). Google notes the model already powers consumer features including Rambler on Android and voice capabilities in the macOS Gemini app, providing real-world validation before the developer API release.