🇮🇳
स्वतंत्रता दिवस की हार्दिक शुभकामनाएं! 🇮🇳 Happy Independence Day! | Har Ghar Tiranga | देश के 80वें स्वतंत्रता दिवस पर आज़ादी का अमृत महोत्सव मनाएं! - Celebrate the 80th Independence Day of India!

Google launches Gemini 3.5 Transcribe, which powers Rambler

Google launches Gemini 3.5 Transcribe, which powers Rambler

Google Unveils Gemini 3.5 Transcribe: The Next Evolution in Intelligent Speech-to-Text

Google has officially taken the wraps off Gemini 3.5 Transcribe, positioning the new technology as its most precise speech-to-text model to date. Designed to move beyond the limitations of traditional transcription, the model is already being integrated into several of the company’s core products to redefine how users interact with their devices through voice.

Beyond Simple Dictation

Unlike conventional speech recognition models that often struggle with background noise, complex jargon, or stuttering, Gemini 3.5 Transcribe processes raw audio directly into polished, formatted text. The model is built to understand natural speaking styles, effectively filtering out “ums,” “ahs,” and other filler words while intelligently managing self-corrections—such as a user changing their mind mid-sentence—to deliver a clean, professional output.

According to Google, the model’s performance represents a significant leap over the “Chirp 3” model introduced in 2025. Data from Artificial Analysis indicates that time-to-final-transcription has improved by an impressive 70%, with the model demonstrating remarkable accuracy in both streaming and non-streaming environments.

Key Technical Advancements

Google highlighted several areas where Gemini 3.5 Transcribe excels:

  • Precision and Accuracy: The model achieves a Word Error Rate (WER) of just 4.0% for streaming and 2.6% for non-streaming, excelling at capturing difficult alphanumeric entities like order IDs and postal codes.
  • Custom Vocabulary: Users can leverage specialized jargon, with the model adapting its recognition capabilities to specific, unique spellings and technical terminology.
  • Global Reach: It features automatic language detection for over 85 languages, offering robust support for diverse regional accents and dialects.
  • Multi-Speaker Identification: The system can accurately attribute speech to up to three individual speakers in pre-recorded audio, complete with precise timestamps.

Integrating Voice into Workflows

The power of Gemini 3.5 Transcribe lies in its ability to do more than just record text; it can execute tasks. By utilizing “function calling,” the model can delegate complex operations—such as image generation or deep file analysis—to other specialized Gemini models.

This functionality is already appearing in the wild, notably in the “Speak to Window” feature within the Gemini app for macOS and the “Rambler” integration on Android’s Gboard. Additionally, the model is being deployed within Google Antigravity, where it uses screen context and chat history to provide high-accuracy transcription for documents and agent-based workflows.

What’s Next for Chrome

Google confirmed that the model is slated for a rollout within the Chrome browser. This update will allow users to dictate replies, draft posts, and prompt Gemini directly in any web field, making it easier than ever to navigate the web using voice commands.

As Google continues to embed Gemini 3.5 Transcribe across its ecosystem, the barrier between spoken intent and digital execution appears to be narrowing, signaling a future where voice is the primary interface for complex productivity tasks.

Leave a Reply

Your email address will not be published. Required fields are marked *