Gemini 3.5 Live Translate, simultaneous voice translation in over 70 languages
Google launches Gemini 3.5 Live Translate, a real-time speech-to-speech AI model translating over 70 languages across Google Meet, Translate, and APIs.
Twenty years after its first steps in machine translation, Google reaches a milestone with Gemini 3.5 Live Translate, a near real-time speech-to-speech audio translation model. Unlike turn-by-turn systems that wait for the end of a sentence, the model generates speech continuously, staying a few seconds behind the speaker, while preserving their intonation, pace, and pitch. It automatically detects over 70 languages within a single session and filters ambient noise to operate in challenging sound environments.
The rollout is taking place on three fronts: in public preview for developers via the Gemini Live API and AI Studio, in private preview in Google Meet for select Workspace customers (which, by the way, expands from 5 to 70 languages and from English as a pivot language to over 2,000 combinations per meeting), and in the Google Translate app on Android and iOS. Android additionally gets a listening mode that streams the translation directly into the phone's earpiece, held against the ear like a regular call. Grab is already testing the model for exchanges between drivers and passengers. All generated audio is marked by SynthID, Google's inaudible watermark.
On the API side, input is limited to audio, and Google acknowledges possible voice replication inconsistencies during long pauses or fast-paced multi-speaker conversations.