
Intelligent transcription with Gemini 3.5 Transcribe
Google introduced Gemini 3.5 Transcribe, a speech-to-text model that converts raw audio into polished, formatted text. It is designed for high-precision real-time transcription and intelligent voice interactions.
Why it matters
This allows users to dictate text and use voice commands more naturally by automatically removing filler words and correcting speech. Developers can use it to build more accurate and responsive voice agents or captioning tools.
The details
The model is available via the Live API for real-time streaming and the Interactions API for pre-recorded audio. It can identify up to three speakers in recordings and improves time to final transcription by 70% over the Chirp 3 model. Integration includes the Gemini macOS app and the Rambler feature on Gboard for Android.
What's next
The model is coming soon to Chrome for web field dictation and to Gemini Enterprise for Customer Experience.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.