Gemini 3.8 turns text-to-speech into a voice direction studio
Gemini 3.8 Flash TTS and Flash-Lite TTS expand Google's voice generation stack with prompt-based voice design, a library of more than 2,000 profiles, and line-by-line performance control. Flash focuses on creative direction, while Flash-Lite targets high-volume production.
Describe a voice instead of selecting one
Choosing a preset is no longer the only starting point. With Gemini 3.8 Flash TTS, users can build a vocal identity through natural language by describing its role, accent, and performance characteristics.
Google says generative voice design covers more than 100 languages and dialects, alongside a library of over 2,000 production-ready voices. Regional varieties include Quebec French, Mexican Spanish, and Scots English.
Flash-Lite TTS carries over fine-grained control with a different focus. Google positions it for high-volume dubbing, audio production, and expressive voice agents, with greater emphasis on cost-efficient scaling. Thirty seconds to replicate a vocal profile
Gemini 3.8 Flash TTS also adds voice replication from a 30-second audio sample. Users must provide their own voice or one they have the rights to use.
Google requires consent verification: a verbal consent recording from the voice owner must match the reference speaker before the voice can be created. Audio generated by Gemini Audio models is also watermarked with SynthID, while Google cites C2PA credentials as another safeguard around voice replication.
The feature is geographically restricted. Voice replication through Google AI Studio is unavailable in the European Economic Area, the UK, Switzerland, India, Illinois, and Texas. Direct every line
Both models interpret performance instructions directly within a script. Tone, pacing, accent, whispers, and acting cues can change from one line to another without generating every variation separately.
Scripts can also include nonverbal cues such as laughter, sighs, and gasps, alongside short backchannel responses used in conversation. Native two-speaker staging generates dialogue from a single script while keeping vocal identities separated and preserving conversational turn-taking.
Google also says the models are designed to maintain voice quality, pacing, and character timbre across hours of long-form generation, targeting formats such as podcasts and audiobooks. External benchmarks for voice customization
On Hume AI's Voice Design Benchmark, Gemini 3.8 Flash TTS scores 71.4 in results reported by Google, including 60.8 for accent modeling. Flash and Flash-Lite also occupy the first two positions on the benchmark's Overall Quality Index.
Google additionally reports strong placements in Voice Arena's blind human preference evaluations across languages including Japanese, Brazilian Portuguese, Modern Standard Arabic, Mexican Spanish, and Hindi. These results remain dependent on the benchmark protocols and competing models included at the time of evaluation. A dedicated playground in AI Studio
Google AI Studio brings these capabilities into a voice design workspace. Users can describe or replicate a voice, save it, and bring it into a dual-speaker screenplay editor for line-by-line direction.
Both models are also available to developers through the Gemini API. Gemini 3.8 Flash TTS is being integrated into Gemini Notebook, while Flash-Lite TTS is being added to Google Vids. Gemini Enterprise API access is planned for a later stage.