Google Debuts “Gemini 3.5 Transcribe” to Revolutionize Speech-to-Text Accuracy
Google has officially pulled the curtain back on a significant upgrade to its speech recognition capabilities with the release of Gemini 3.5 Transcribe. According to the tech giant, this new model marks a major leap forward from its predecessor, “Chirp 3,” offering substantial improvements in multilingual performance and a notable reduction in word error rates.
Designed to streamline the often tedious process of transcribing audio, the model introduces a “natural editing” feature that allows users to adjust text using only their voice. Furthermore, the system is engineered to automatically polish transcripts by formatting text and stripping out conversational filler words like “um” and “uh.”
Tailored for Precision
One of the standout features of Gemini 3.5 Transcribe is its ability to handle specialized terminology. Users can provide a customized vocabulary to the model, ensuring that industry-specific jargon and unique proper nouns are spelled correctly the first time. This preemptive adaptation effectively eliminates the need for manual post-transcription edits.
For content creators and researchers, the tool also offers robust utility: it can differentiate and attribute speech for up to three distinct speakers in pre-recorded audio and provides highly granular word-level timestamps.
Rollout and Availability
Google is wasting no time in getting the technology into the hands of users. Starting today, Gemini 3.5 Transcribe is rolling out to all macOS Gemini app users in English. Additionally, the “Rambler” dictation feature on Android is now expanding to select countries and languages.
For the developer community, the model is available in public preview via the Gemini API in both AI Studio and Antigravity. While currently limited to specific platforms, Google has confirmed that support for the Chrome browser is on the horizon.
Clarification on Gemini 3.5 Live
While early reports surrounding today’s announcement hinted at the arrival of “Gemini 3.5 Live” and “Gemini 3.5 Live Experimental,” Google has since clarified that only the Transcribe model is launching today.
Initial information suggested that the “Live” models were designed to refine how Gemini handles mid-sentence interruptions, rapid language switching, and live visual processing. The “Experimental” iteration was teased as being capable of narrating its own reasoning steps in real-time while tackling complex tasks. However, Google has retracted the release window for these specific features and has not yet provided a revised launch date.
