Google announced Gemini 3.5 Transcribe, a speech-to-text model, on August 26, 2026 local time. Rather than simply converting audio into words, it strips out self-corrections and filler sounds such as um and uh, and returns text that is already formatted. Developers get access through two separate APIs, and the model is already running behind consumer-facing features in Gboard on Android and the Gemini app on macOS.

From transcription to cleanup

Conventional speech recognition has long struggled with background noise, industry jargon, and the messiness of natural speech. Gemini 3.5 Transcribe targets exactly those weak points, converting raw audio directly into polished, formatted text.

Self-correction handling is the clearest example. When a speaker says something like let us meet Tuesday, no, Wednesday, the model resolves the intent instead of transcribing both options verbatim. Filler removal and automatic formatting work the same way, cutting down the manual cleanup pass that usually follows a transcript.

Custom vocabulary is the other piece. Register your own terms in advance and the model adapts to internal jargon and unusual spellings. Google says it also captures alphanumeric strings such as postal codes and order IDs accurately in noisy, real-world environments.

Two APIs, split by use case

Availability is divided into two tracks that map to different developer workflows.

For real-time work there is the Live API, using the model name gemini-3.5-transcribe-live. It offers continuous bidirectional streaming with sub-second latency, aimed at voice agents and live captioning.

Pre-recorded audio goes through the Interactions API with gemini-3.5-transcribe. It separates meeting recordings and call logs by speaker and returns word-level timestamps. Speaker attribution is officially supported for up to three speakers, and support beyond that is still experimental. The model automatically detects and transcribes over 85 languages, including regional accents and dialects.

In the Gemini app on macOS, the model can also use function calling to hand heavier work such as image generation and file analysis to other Gemini models. Voice input is clearly shifting from a way to enter text toward a way to drive an application.

A clear step up from Chirp 3

On the numbers, Artificial Analysis measures an average Word Error Rate of 4.0% for streaming and 2.6% for non-streaming use. On the multilingual FLEURS benchmark across a set of top languages and locales, the model records 5.50% WER in streaming mode and 5.04% in non-streaming mode.

Compared with Chirp 3, Google's previous transcription model, time to final transcription improves by 70%. For real-time applications, that reduction in waiting is arguably as significant as the accuracy gain.

Keep in mind that WER is a lower-is-better metric that shifts with test conditions and audio sources. The published figures are best read as results under a specific measurement setup.

Where you can use it now, and what is coming

For consumers, the entry point on Android is Rambler, a new Gboard feature. It turns spoken input into formatted text and lets you edit, fix misspellings, and change writing style by voice. It is rolling out in select countries and languages.

The Gemini app on macOS supports English and pairs voice commands with screen context. Summarizing local files, reusing text across apps, and generating images right at the cursor are all meant to work by voice alone. Chrome support is coming soon, which will allow dictation into any web field.

Developer and enterprise access is in public preview. Developers can try the model through the Gemini API in Google AI Studio and Google Antigravity, while enterprises can use the Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience planned. Developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents already support it through the Live API, so teams can offload the real-time streaming infrastructure.

Summary

Gemini 3.5 Transcribe is less about raw recognition scores and more about removing the work that normally comes after transcription. A 2.6% non-streaming WER, support for more than 85 languages, and a 70% reduction in time to final transcription against Chirp 3 give it a solid practical case. Since it is already live in Gboard and the Gemini app on macOS, the fastest way to judge it is to talk at it and see how much of your own speech it tidies up.