Intelligent transcription with Gemini 3.5 Transcribe
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions.
Key points
- Intelligent transcription with Gemini 3.5 Transcribe
- Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
- Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.
- Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
Sources (1)
- [1]Intelligent transcription with Gemini 3.5 TranscribeGoogle DeepMind Blog · Aug 26, 05:01 PM
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions.
Intelligent transcription with Gemini 3.5 Transcribe
Extractive summary: sentences quoted from the sources.