
Intelligent transcription with Gemini 3.5 Transcribe
Google introduced Gemini 3.5 Transcribe, a speech-to-text model that converts raw audio into polished, formatted text. The model is available for developers via API and integrated into various Google products.
Why it matters
Users and developers can now transcribe voice interactions more accurately in noisy environments while automatically removing filler words. This enables more natural voice-driven interfaces and efficient voice-to-text productivity.
The details
Artificial Analysis measured the Word Error Rate at 4.0% for streaming and 2.6% for non-streaming. The model can delegate complex tasks, such as image generation, to other Gemini models via function calls. It is available through two separate APIs: the Live API for real-time streaming and the Interactions API for pre-recorded audio. Integration with Gboard's Rambler feature allows users to filter filler words and edit text using voice.
What's next
The model is coming soon to Chrome, allowing users to talk to type in any web field.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.