Google has introduced two new text-to-speech AI models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, expanding the company’s Gemini Audio technology with more expressive and customizable voice generation.
According to Google, the new models are designed to go beyond traditional text-to-speech systems that simply read written text using fixed voices. Users can instead describe the type of voice they want through natural-language instructions, including characteristics such as accent, role and speaking style. The technology supports more than 100 languages and dialects.
Custom voices and realistic conversations
Gemini 3.8 Flash TTS is aimed at creative applications such as podcasts, audiobooks, games and interactive content. Google says developers can create new voices from scratch and control how individual lines are delivered, including pacing, emotion and acting style.
The system can also produce conversations involving two different speakers from a single script. Developers can add natural reactions such as laughter, sighs and other conversational sounds to make generated speech feel more realistic.
Google is also offering access to more than 2,000 production-ready voices. In addition, the company says a consistent voice profile can be created from a 30-second audio sample when the user owns the voice or has permission to use it.
Flash-Lite focuses on large-scale audio production
The second model, Gemini 3.8 Flash-Lite TTS, is designed for high-volume and cost-efficient applications. Google says it can be used for large-scale dubbing, audio content creation and expressive voice agents while still providing controls for tone, pacing and other voice characteristics.
The new models are being rolled out through the Gemini API and Google AI Studio, while Flash TTS is also available in Gemini Notebook. Flash-Lite TTS is being introduced in Google Vids, with wider enterprise availability planned.
Google adds safeguards for AI-generated voices
Google has also included safeguards designed to reduce misuse of generated and replicated voices. Voice replication requires a verbal consent recording from the voice owner, while AI-generated audio receives an invisible SynthID watermark intended to help identify AI-generated speech.
The launch could expand the use of AI-generated voices across news narration, podcasts, video dubbing, education, entertainment and voice-based applications. For publishers and content creators, the technology could make it easier to turn written scripts into natural-sounding audio without recording every piece of content manually.
