Google has unveiled Gemini 3.5 Transcribe, its most precise speech-to-text model yet. The new model converts raw audio directly into accurate, polished, formatted text, handling background noise, complex jargon, and disfluency cleanup that trips up conventional speech recognition systems. It automatically detects 85+ languages and regional dialects, and filters out filler words such as ums and ahs to capture natural speaking style and intent.

The model ships through two developer-facing APIs. The Live API delivers real-time bidirectional streaming with sub-second latency for voice agents and live captioning via gemini-3.5-transcribe-live, while the Interactions API handles pre-recorded audio with speaker attribution and word-level timestamps. It marks a major step up from Google's previous transcription model, Chirp 3, with improved word error rates and significantly better latency.

Gemini 3.5 Transcribe already powers consumer features including Rambler on Android, which turns spoken thoughts into well-formatted text and strips out filler words, plus new voice capabilities in the Gemini app on macOS that pair speech with screen context. Developers can access the model through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. For founders building voice agents, real-time captioning tools, or post-call analytics pipelines, the model removes a huge chunk of transcription engineering from day one.