Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

blog+1blogblog+1Google Alphabet Inc. on Wednesday introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models it called its most expressive audio generation tools to date. Both models began rolling out immediately in the Gemini API and Google AI Studio.blog
The announcement, written by Group Product Manager Leland Rechis and Director of Research Science Alan Cowen on behalf of the Gemini Audio Team, positions the two models as complementary tools: Flash TTS for deep creative direction and Flash-Lite TTS for high-volume, cost-efficient applications.unite+1
Gemini 3.8 Flash TTS allows users to create original voices from scratch through natural language prompts, specifying role, accent, and vocal characteristics across more than 100 languages and dialects. The release expands Google's voice library from 30 original voices to more than 2,000 production-ready options, covering regional varieties such as Mexican Spanish, Quebec French, and Scots English.ndtvprofit+2
A voice replication feature can recreate a vocal profile from a 30-second audio sample, provided the user has rights to the voice. Google said the system requires a verbal consent recording from the voice owner, and all generated audio carries SynthID watermarking and C2PA credentials. Voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India.unite+1
Both models support line-by-line performance direction, long-form generation across hours of audio, and native two-speaker scene staging for podcasts and dramatic storytelling. Non-verbal cues such as laughter, sighs, and active-listening interjections can be scripted into dialogue.blog
Google reported that Gemini 3.8 Flash TTS took the top position on Hume AI's Voice Design Benchmark with a score of 71.4 and led accent modeling at 60.8. The two models hold the first and second spots on Hume AI's Overall Quality Index. In blind human preference evaluations on Voice Arena, the models ranked at the top in languages including Japanese, Brazilian Portuguese, Vietnamese, and Hindi.ndtvprofit+1
Developer platforms Agora, LiveKit, Pipecat, and Vercel support the models through the Gemini API, while companies including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang are integrating the TTS models for dubbing, media localization, and voice agent applications. Enterprise access via Gemini Enterprise is listed as coming soon.unite+2