Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with prompt-designed voices and line-by-line direction
The models can create or replicate voices from natural-language prompts and ship with built-in watermarking.

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which it calls its most expressive audio generation models so far. The text-to-speech models are available in Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids. Users can create custom character voices from scratch or replicate existing ones with plain-language prompts. They can also direct scene dialogue line by line, controlling pacing, emotion and realistic conversational sounds. Google is pitching the models for audiobooks, podcasts, games and real-time voice agents at scale. The company says the models include built-in safety tools such as watermarking for generated audio. That matters because the ability to replicate an existing voice from a prompt raises the risk of impersonation. The release puts Google in direct competition with other providers of controllable speech synthesis for developers and enterprise media tools. The same day, Alibaba's Qwen team cut prices on its own speech models. Google's announcement does not include pricing or benchmark figures.