Google Expands AI Voice Generation With New Gemini Models

Google has introduced two new text-to-speech models designed to make AI-generated audio more flexible, realistic, and suitable for large-scale production. The Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models are aimed at different use cases, ranging from entertainment and narration to automated dubbing and conversational applications.

Two Models For Different Uses

Gemini 3.8 Flash TTS is designed for interactive entertainment, game development, and long-form narration where creators need more control over AI-generated voices. Gemini 3.8 Flash-Lite TTS is focused on high-volume applications such as media dubbing, customer-facing conversational agents, and translation.

The models expand Google’s existing collection of audio systems and replace a fixed selection of 30 legacy voices with access to more than 2,000 pre-built vocal profiles. These profiles cover regional variations across more than 100 languages, including Quebec French, Scots English, and Mexican Spanish.

Google also plans to introduce voice remixing capabilities that will allow audio engineers to modify characteristics such as timbre, pitch, pace, and accent using text instructions.

Strong Voice Quality Results

Third-party testing has shown strong results for Gemini 3.8 Flash TTS. On the Hume AI Voice Design Benchmark, the model received an overall score of 71.4 and achieved a 60.8 score for accent modelling.

The model also received the highest position in Hume AI’s Overall Quality Index, while Gemini 3.8 Flash-Lite TTS ranked second. Both models also performed well in double-blind human testing involving languages such as Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

More Natural Conversations

Google has also focused on making the models useful for longer audio productions. A single script can direct conversations between two speakers while maintaining distinct voices and natural turn-taking.

The models are designed to maintain consistent audio quality and character timbre during multi-hour recordings. This could make them useful for audiobooks, podcasts, and other extended audio projects.

Creators can also add non-verbal cues directly into scripts. Tags such as <laughs>, <sigh>, and <gasp> can be used alongside conversational expressions such as mhm and yeah to create more natural-sounding speech.

Controls For Voice Cloning

Google has included safeguards for custom voice creation. Users seeking to recreate a person’s voice must provide a 30-second reference recording along with an explicit consent recording from the original voice owner.

Google checks the acoustic alignment between the recordings before allowing the custom voice profile to be processed.

Generated audio also includes invisible SynthID watermarks and cryptographic C2PA provenance metadata. These features are designed to help identify AI-generated audio and provide information about its origin.

Enterprise Availability

The new models are available through Google AI Studio and the Gemini API, with integrations spanning platforms and developer frameworks including Agora, LiveKit, Pipecat, and Vercel.

Early commercial applications include Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, covering areas such as regional media translation and customer service.

Flash TTS is also available through Gemini Notebook, while Flash-Lite TTS is being integrated into Google Vids. Google says Gemini Enterprise customers will receive administrative API access as part of a future rollout.

Source: https://www.artificialintelligence-news.com/news/google-gemini-3-8-flash-tts-voice-models/

Facebook
Twitter
LinkedIn

Related Posts

Leave a Reply

Your email address will not be published. Required fields are marked *