Search Agent Sky
← Back to search Memory Lane Recent answers
Cited source trail
Gemini-3.5-Transcribe
Original question
Investigate this live topic: Gemini-3.5-Transcribe. Start with https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe and browse beyond it. Summarize what changed, why it matters, and cite the strongest sources.
Sources checkedblog.googlearstechnica.comdeepmind.google9to5google.com
Next step

Research any question with live sources, then publish the cited answer as a free shareable link.

## Gemini 3.5 Transcribe — Investigation Summary **What it is:** On Aug 26, 2026, Google launched **Gemini 3.5 Transcribe**, its "most precise speech-to-text model yet," built for intelligent voice interactions. Unlike conventional speech recognition that struggles with background noise, jargon, and disfluency cleanup, it converts raw audio directly into accurate, polished, formatted text. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe)) ### What changed vs. the prior model (Chirp 3) - **~70% faster** time-to-final-transcription (per Artificial Analysis). ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [Ars Technica](https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/)) - **Lower word error rate:** 4.0% streaming / 2.6% non-streaming (Artificial Analysis); 5.50% streaming / 5.04% non-streaming on the FLEURS multilingual benchmark — improved over Chirp 3's 7.32% live error rate. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [Ars Technica](https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/)) - **Smart transcription:** handles self-corrections ("let's meet Tuesday—no, Wednesday"), strips filler words ("ums"/"ahs"), auto-formats text, and supports natural voice editing. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [9to5Google](https://9to5google.com/2026/08/26/gemini-3-5-transcribe/)) - **New capabilities:** custom vocabulary for jargon, **85+ languages**, multi-speaker attribution with word-level timestamps (up to 3 speakers; 3+ experimental), and **function calling** to delegate tasks (image generation, file analysis) to other Gemini models. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [Google DeepMind](https://deepmind.google/models/gemini-audio/ai-transcription/)) ### Two developer APIs - **Live API** (`gemini-3.5-transcribe-live`): real-time bidirectional streaming, sub-second latency for interactive voice apps. - **Interactions API** (`gemini-3.5-transcribe`): pre-recorded audio with speaker attribution and timestamps. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [the-decoder](https://the-decoder.com/googles-gemini-3-5-transcribe-turns-speech-to-text-in-85-languages-while-auto-correcting-your-verbal-stumbles/)) ### Rollout - **Live now:** Gboard "Rambler" on Android (Pixel 11), Gemini app on macOS (English), Google Antigravity, Google AI Studio Build mode. - **Coming soon:** Chrome (talk-to-type in any web field). - **Developers/enterprises:** public preview in the Gemini API via Google AI Studio and Gemini Enterprise Agent Platform. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [9to5Google](https://9to5google.com/2026/08/26/gemini-3-5-transcribe/)) ### Why it matters 1. **A paradigm shift in voice input** — the model *edits your speech as you talk* rather than literally transcribing it, capturing intent and cleaning up disfluencies. Ars Technica flags the tradeoff: the AI technically changes your wording, which "may not be appropriate for all situations." ([Ars Technica](https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/)) 2. **Voice as an interface, not just dictation** — with function calling and sub-second latency, it powers voice agents that *execute tasks* (summarize files, generate images, search) across Google surfaces. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe)) 3. **Broad ecosystem adoption** — platforms (Agora, LiveKit, LangChain, Pipecat, Vercel) and companies (vivo, IntelliTek Health, Lingopal) are already integrating it as a real-time transcription layer. ([Google blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe), [Google DeepMind](https://deepmind.google/models/gemini-audio/ai-transcription/)) ### Strongest sources (all opened) - [Google blog — Introducing Gemini 3.5 Transcribe](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe) (official primary) - [Google DeepMind — Gemini Audio AI transcription](https://deepmind.google/models/gemini-audio/ai-transcription/) (official model page, performance data + partner testimonials) - [Ars Technica — Google announces Gemini 3.5 Transcribe](https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/) (independent analysis incl. caveats) - [9to5Google — Google launches Gemini 3.5 Transcribe](https://9to5google.com/2026/08/26/gemini-3-5-transcribe/) (product rollout detail)

Sources visited

1navigatehttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe
2get_text
3searchGemini 3.5 Transcribe announcement Google speech-to-text
4navigatehttps://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/
5navigatehttps://deepmind.google/models/gemini-audio/ai-transcription/
6navigatehttps://9to5google.com/2026/08/26/gemini-3-5-transcribe/
Shared by Proto Státis · Aug 28, 2026

Ask your own question

Search live sources for a question you choose. You can publish the cited answer as a free link.