Gemini 3.8 Flash TTS with Voice Cloning

Gemini 3.8 Flash TTS with Voice Cloning

More

Summary

Sam Witteveen examines Google’s two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash Lite TTS. Flash TTS targets expressive work such as character voices, audiobooks, podcasts and games, while Flash Lite TTS is built for high-volume jobs like dubbing, bulk audio generation and voice agents where cost matters most.

The headline feature is voice design, where users describe a voice in plain language, including role, accent and characteristics, across more than 100 languages. The models also offer a library of over 2,000 voices, line-by-line stage directions such as whispered or sarcastic delivery, non-verbal tags like sighs and laughs, and two-speaker scenes from a single script. Voice cloning from a 30-second sample is included, protected by a spoken consent check, SynthID watermarking and C2PA credentials, though availability is limited in some countries and US states.

The host compares Google’s claims with independent benchmark numbers, which look more mixed than the launch messaging, and notes that Google is mid-pack on price, with Flash Lite not much cheaper than Flash. He then tries the models in AI Studio, where voice generation failed during recording, and in a Colab notebook, comparing the two models on the same script.


📺 Source: Sam Witteveen · Published September 24, 2026
🏷️ Format: Review

1 Item

Channels

1 Item

Companies