Google has released two new text-to-speech models: Gemini 3.8 Flash TTS, which lets users build an entirely new voice from a written description, and Gemini 3.8 Flash-Lite TTS, optimized for fast, low-cost, high-volume processing. Earlier Google TTS models only let users pick from a fixed list of preset voices. The new models can generate a voice from scratch based on a role, accent, or vocal quality described in plain language, opening up new possibilities for games, audiobooks, and podcast production.
From Choosing a Voice to Designing One
Previous Google TTS models offered around 30 preset voices to choose from. With Flash TTS, users can still pick from those presets, but the headline change is the ability to describe a role, accent, or vocal texture in natural language and have an entirely new voice generated on the spot. Google also built a library of more than 2,000 voices covering regional speech variations, and plans to soon add a "voice remix" feature that lets users adjust tone, pitch, pace, and accent of a library voice through a prompt. Flash TTS supports 130 languages, while the high-volume Flash-Lite TTS supports 101, and both include Japanese.
The models also include a "voice replication" feature that can reproduce a speaker's voice from about 30 seconds of audio. It is limited to a user's own voice or a voice they have the rights to use, and requires the speaker to submit a spoken consent recording; the system checks that the voice in that recording matches the reference audio. Voice replication in Google AI Studio is not available in Illinois, Texas, the European Economic Area, the UK, Switzerland, or India.
Direction Cues and Long-Form Content
Both models let creators insert stage-direction-style cues line by line in a script to control pacing and emotion. They can generate hours of continuous audio while maintaining consistent voice quality and natural pauses, and a single script can be turned into a conversation between two speakers. Non-verbal sounds such as laughter, sighs, and gasps, as well as verbal backchannels, can be inserted through tags in the script. Every generated clip carries Google's SynthID audio watermark, and clips made with voice replication also carry C2PA provenance metadata.
Availability and Pricing
Both models began rolling out the same day. Developers can access them through the Gemini API and Google AI Studio. General users can use Flash TTS via the audio-overview feature in Gemini Notebook, while Flash-Lite TTS is available through Google Vids. Access via Gemini Enterprise's API is coming soon for business customers. Developer platforms including Agora, LiveKit, Pipecat, and Vercel are integrating the models through the Gemini API, and companies such as Figma and HeyGen are building them into their own products.
On paid Gemini API plans, Flash TTS costs $0.50 (about ¥79) per million input tokens and $9 (about ¥1,400) per million output tokens, while Flash-Lite TTS costs $0.50 (about ¥79) per million input tokens and $6 (about ¥940) per million output tokens. Both prices hold through December 31, 2026, and will roughly double starting January 1, 2027. A free tier is available, and asynchronous Batch and Flex processing are offered at half the standard rate.
On performance, Flash TTS took first place overall on voice-design benchmarks run by Hume AI, including the top score for accent expression, while Flash-Lite TTS placed second overall. On Voice Arena, where humans compare generated voices directly, both models ranked near the top across several languages, including Japanese. The two models join Google's existing Gemini Audio lineup, which already includes Gemini 3.8 Live for real-time voice conversation and Gemini 3.5 Transcribe for speech-to-text.
Summary
Google's newly released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS move beyond simply picking a preset voice, letting creators design an entirely new one through natural-language description. The models add fine-grained direction control and long-form stability while building in SynthID watermarking and C2PA provenance for authenticity, and they are available through a wide range of channels, from developer APIs to consumer tools like Gemini Notebook and Google Vids.
*1 USD = 157 JPY (as of September 25, 2026)
