September 24, 2026
The-new-Gemini-TTS-can-now-replicate-voices-and-recite.jpg

Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models focused on improving the way AI-generated speech sounds. They can create voices from descriptions, follow detailed voice directions, manage two-way conversations, and keep those voices consistent across longer recordings.

For people who use Gemini for narration, podcasts, audiobooks, or voiceovers, the update offers much more control over the finished audio. Users can describe the type of voice they want and tell Gemini how individual lines should be delivered.

You can tell Gemini how you want the voice to sound

Gemini 3.8 Flash TTS can create an entry from a natural language description. Users can specify a character, accent and other voice qualities or choose from more than 2,000 existing voices. The templates support over 100 languages ​​and dialects.

You can also add statements within a script. For example, a sentence may be whispered, spoken more slowly, spoken with a different emotion, or include laughter, sighs, gasps, and conversational reactions such as “mhm” or “yes.” The update also allows two-party conversations to be generated from a single script, while voices and pacing can remain consistent even in longer recordings.

Custom voices can also be saved and reused between projects, which should help prevent the same character or narrator from sounding gradually different. Google plans to add voice remixing later, allowing you to adjust the pitch, tempo, timbre, and accent of existing voices.

Voice replication comes with some guarantees

Gemini can recreate a coherent voice profile from a 30-second recording, but it can’t simply feed it a voice sample. Google says the recording must belong to the user or be an entry they have the right to use. The owner of the voice must also provide a separate record of verbal consent that matches the reference speaker before the replica can be created.

The generated Gemini Audio clips also carry an imperceptible SynthID watermark, while Google says the vocal replication includes C2PA credentials to help identify AI-generated material. Gemini 3.8 Flash TTS is being rolled out via Gemini Notebook, Gemini API, and Google AI Studio. The cheaper Flash-Lite model will come to Google Vids and is designed for large-scale work such as dubbing, narration and voice agents.

Avatar photo
Written by

Hafizur Rahman

Hafizur is a writer and contributor covering breaking technology and science news, emerging innovations, digital trends, gadgets, artificial intelligence, space, and major scientific discoveries. He follows the latest developments across the technology and science industries and turns complex stories into clear, engaging, and easy-to-understand articles. His work aims to keep readers informed about the innovations, discoveries, and technological changes shaping the world.

Leave a Reply

Your email address will not be published. Required fields are marked *