Google has announced the release of Gemini 3.5 Live Translate, a new audio model that converts spoken words into another language as speech almost in real time[1]. The model automatically detects more than 70 languages and generates natural-sounding translated speech while preserving the speaker's intonation, pacing and pitch[1]. It is rolling out gradually through the Google Translate app, Google Meet and a developer API[1].

How continuous "translate while you speak" works

Gemini 3.5 Live Translate is Google's latest audio model for live speech-to-speech translation[1]. According to Google, it automatically detects 70+ languages from the incoming audio and produces smooth, natural translated speech[1].

Many earlier systems were turn-by-turn, waiting for the speaker to finish before translating[1]. By contrast, 3.5 Live Translate generates speech continuously[1]. It balances the trade-off between waiting for more context to improve quality and translating immediately to stay in sync with the speaker, so there are no awkward pauses and the output stays just a few seconds behind the speaker throughout the session[1].

Google notes that translation has been part of its machine learning work since its early days, and that it now translates more than a trillion words every month for billions of users across its products[1].

A phased rollout to developers, enterprises and everyone

Gemini 3.5 Live Translate is rolling out across Google products starting the same day, with availability differing by user type[1].

For developers, it is available in public preview through the Gemini Live API and Google AI Studio[1]. For enterprises, a private preview begins this month in Google Meet[1]. And for everyone, it is available through the Google Translate app on Android and iOS[1].

Robust to noise, with support for multilingual input

A key strength for developers is that the model processes speech as it is streamed[1]. It handles multilingual input without manually configuring settings, and its noise robustness lets applications work in loud, unpredictable environments[1]. Google says it can be used for live interpretation in multilingual calls, meetings, lessons, broadcasts and more[1].

Using the Gemini Live API, developer platforms such as Agora, Fishjam, LiveKit, Pipecat and Vision Agents make it easier to build and deploy voice translation apps[1]. These integrations handle the complex real-time media streaming infrastructure so developers can focus on the user experience[1].

Real-world testing is also under way. The ride-hailing service Grab is trialing the model to enable near real-time multilingual communication between drivers and travelers at pickups[1]. Grab's users make more than 10 million voice calls per month through the service[1]. Companies including CJ ENM and LiveKit have also shared positive feedback, praising the translation quality, accuracy and low latency[1].

The experience in Google Meet and the Google Translate app

Speech translation in Google Meet will soon use 3.5 Live Translate[1]. Supported languages expand from the previous five to more than 70, and a single meeting can handle over 2,000 language combinations[1]. Previously translation was limited to and from English, but that restriction is being removed[1]. The interface is also updated to give instant access to speech translation[1]. The update is launching in private preview for select business Google Workspace customers this month, with a broader rollout planned later this year[1].

In the Google Translate app, the feature is rolling out globally on both Android and iOS[1]. When using the Live translate feature, simply connecting any pair of headphones lets you experience translation that mirrors the speaker's tone across 70+ languages[1].

On Android, Google is also starting to roll out a new "listening mode"[1]. It lets you hear translations directly through your phone's earpiece[1]. Holding the phone to your ear like a regular call streams the translated audio straight to you, which can help in situations where you want to hear translations quickly and discreetly without headphones on hand[1].

Generated audio watermarked with SynthID

All audio generated by the model is watermarked with SynthID[1]. This imperceptible watermark is woven directly into the audio output[1]. By keeping AI-generated content detectable, Google says it aims to help prevent the spread of misinformation[1]. The company outlines its approach to safety and responsibility in the model card[1].

Summary

Gemini 3.5 Live Translate is Google's new audio model that continuously translates across more than 70 languages while preserving the speaker's tone. It is being delivered in phases through a developer API, Google Meet and the Google Translate app, with practical use expected in meetings, calls and conversations abroad where language barriers get in the way. The generated audio carries a SynthID watermark, reflecting a design that also accounts for safety.

Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/