Google has launched Gemini 3.5 Transcribe, a specialised speech-to-text model designed to convert spoken audio into structured, formatted text while improving the accuracy and readability of transcripts. The new model is designed to automatically filter out filler words, handle mid-sentence self-corrections and recognise specialised terminology and jargon. Google is positioning Gemini 3.5 Transcribe as a solution for both real-time voice applications and recorded-audio transcription. Two APIs for Different Use Cases Gemini 3.5 Transcribe is available through two developer interfaces: the Live API and the Interactions API. The Live API is designed for continuous, two-way audio streaming, delivering sub-second latency for applications such as voice agents and real-time conversational experiences. The Interactions API is targeted at recorded audio and includes features such as speaker attribution and word-level timestamps, giving developers greater control over how transcripts are processed and presented. Supports More Than 85 Languages The model supports more than 85 languages,…