Gemini 3.8 Flash Text-to-Speech (TTS), Google’s new technology for turning text into spoken audio, designs voices from ordinary descriptions. It affects you if you narrate podcasts, learning materials, audiobooks, or multilingual videos, as well as if you transcribe conversations or use voice assistants. At the same time, open audio models show that Artificial Intelligence (AI) can do more than speak: it can listen, distinguish speakers, and respond directly to spoken tasks.
Designing voices through descriptions
Google has introduced two models with different priorities. Gemini 3.8 Flash TTS is intended for expressive voices and creative productions, while Flash-Lite TTS targets lower-cost speech generation at scale. According to the initial report on the two Gemini models, they are available through Google AI Studio and the Gemini application programming interface and support more than 100 languages.
The main difference from a conventional voice library is prompt-based control, meaning instructions written in natural language. You can describe a role, accent, mood, and vocal character instead of selecting from only a few presets. Google also provides a library of more than 2,000 ready-made voices, including Mexican Spanish, Quebec French, and Scottish English.
For a podcast, for example, you could assign different voices to two characters and direct the pace and emotion of each line separately. In an audiobook, the narrator could remain calm while individual characters sound more energetic or restrained. The Gemini 3.8 TTS Playground demonstrates these multi-speaker conversations and lets users save and share their settings through a URL.
Google also says a voice can be replicated from a 30-second audio sample. The company requires spoken consent from the person whose voice is being copied, and users are supposed to provide their own voice or one they have the rights to use. According to Google DeepMind’s product description, the models also include watermarking as a built-in safety measure.
Current quality evidence mainly consists of a provider claim and one benchmark result. MarkTechPost reports that Flash TTS scored 71.4 and ranked first on Hume AI’s Voice Design Benchmark. The supplied sources do not independently verify that result, and a single score says little about technical terms, regional language variants, or the consistency of a long production. For practical work, listening to the complete output remains more useful than admiring a leaderboard.
Recognizing speech and organizing conversations
Speech AI now covers much more than synthetic voices. Alibaba’s Qwen team has introduced five Qwen-Audio-3.1 models for automatic speech recognition, speech synthesis, and real-time interaction. Automatic Speech Recognition (ASR) converts spoken language into text; Qwen’s model is also said to recognize multiple languages and dialects while automatically removing filler words and repetitions.
The more advanced ASR-Next assigns timestamps to multiple speakers and, according to the provider, recognizes emotions, environmental sounds, and machine noises. A recorded team meeting could therefore become a more readable transcript, with contributions separated by time and speaker. For a long dictation, automatic removal of “um” and repeated sentence openings could save editing time, although cleanup can also remove wording that was intentional.
Qwen’s TTS model also uses text instructions to control emotion, speed, and style, and it can transfer a voice across languages. TTS-Next generates speech, sound effects, and background audio in one pass. The real-time model is designed to speak and listen simultaneously, accept interruptions at any moment, and respond more slowly and sympathetically when it detects a subdued mood. These claims come from the description of the Qwen-Audio-3.1 family and have not been independently verified in the supplied sources.
Nvidia offers a more specialized open approach. Nemotron 3 Diarization is an open-weight model designed to determine who spoke and when. Speaker diarization is the process of assigning parts of an audio recording to individual speakers; the provider says the model can track up to eight people in real time, including overlapping voices, and process both recordings and live audio streams.
That makes it a useful addition to ordinary transcription. A verbatim record with no visible speaker changes is often only partly helpful. Open weights mean the trained model is available for independent deployment, offering more operational control but also requiring technical infrastructure. They do not automatically resolve privacy concerns or recognition errors.
Kyutai’s Voice of Reason takes a different route. Its two open-weight models process spoken mathematics directly as speech, without a transcription stage and without a separate text-based language model. According to the report on Voice of Reason, supervised fine-tuning and reinforcement learning raised accuracy on the spoken GSM8K mathematics test from 27.3 to 77.1 percent.
Reinforcement learning is a training method in which a model improves its behavior using feedback on its results. The reported increase represents substantial progress within that particular test, but it does not prove that the model can reliably solve arbitrary spoken problems. Running it on a single Nvidia H100 accelerator also places it closer to organizations with dedicated computing infrastructure than to an ordinary home laptop.
Pros and Cons of the new audio models
Pros:
- Flexible voices – You can describe roles, accents, pacing, and emotion in text without recording every variation separately.
- Multilingual content – Support for more than 100 languages can simplify narration for learning materials, podcasts, and videos aimed at different audiences.
- Better conversation records – Speech recognition, timestamps, and diarization can make dictations and meetings easier to search and review.
- More deployment choices – Open weights from Kyutai and Nvidia allow use outside an exclusively closed cloud service.
Cons:
- Potential misuse – Replicated voices can be used for deception or unauthorized imitation despite consent requirements and watermarking.
- Unverified claims – Many statements about emotion detection, quality, and real-time performance come from providers and lack independent confirmation.
- Context errors – Dialects, overlapping voices, technical terms, and background noise can distort transcripts or speaker assignments.
- Unclear total costs – Flash-Lite is positioned as the cheaper option, but the supplied sources do not provide a complete price list for typical projects.
What this means for your work
If you are getting started, use a short, non-sensitive text and a newly designed synthetic voice. Describe the role, approximate vocal range, speaking speed, mood, and language in a few clear sentences, then listen to every version from beginning to end. A short lesson or podcast introduction is a better test than an entire audiobook, because small pronunciation problems do not become more charming after a hundred pages.
Short experiments may be inexpensive, but one result cannot establish a general production cost. Simon Willison generated one minute and 18 seconds of audio with Flash TTS in about 20 seconds at a cost of 2.74 cents. That was one test with specific settings, not a universal price for podcasts, dubbing, or the less expensive Flash-Lite service.
More advanced users can get better results by defining each character separately, adding direction line by line, and systematically comparing different language versions. For meetings or interviews, combining transcription with diarization can preserve both the words and the speaker changes. Names, numbers, quotations, pronunciation, and speaker labels should still be checked manually before publication.
For users in Switzerland, broad multilingual support is potentially useful for content in German, French, or Italian. However, the sources do not confirm specific performance for Swiss German or provide details about regional availability, data storage, or processing in Switzerland. Language support alone therefore does not demonstrate suitable privacy conditions for meetings, customer conversations, or educational data.
Price reductions elsewhere in the market also deserve attention without being overinterpreted. Alibaba reports cuts of about 70 percent for TTS, roughly 85 percent for real-time models, and up to 95 percent for ASR. That adds competitive pressure, but without matching usage units, quality measurements, and availability information, it does not allow a direct cost comparison with Google.
What remains open about quality and control
An expressive voice is not automatically a reliable one. A system can sound convincing while mispronouncing names, changing numbers, or assigning the wrong emotion to a character. That is irritating in an audiobook; in training material, a customer interaction, or a meeting transcript, it can alter the substance.
Emotion recognition requires additional caution. A subdued voice may indicate fatigue, concentration, a poor connection, or simply someone’s normal way of speaking. When a voice system infers an emotional state and changes its behavior accordingly, the interpretation remains uncertain even if the response sounds considerate.
Voice replication also intensifies questions about consent and control. Google’s spoken-consent requirement and announced watermarking are concrete safeguards, but they cannot prevent every redistribution or later change of purpose. The sources also do not explain how reliably those watermarks can be detected outside Google’s own tools.
Open models shift part of the responsibility to whoever operates them. They can provide more control over deployment and audio data, but they require suitable hardware, maintenance, and independent checking of the results. In this context, “open” refers to access to model weights; it does not automatically mean simple operation, low cost, or risk-free use.
Gemini 3.8 Flash TTS makes multilingual voice design more accessible, while Qwen, Nemotron, and Voice of Reason show how broad speech AI has become. The clearest benefits currently lie in limited tasks such as short narrations, cleaned-up dictation, and structured meeting records. The unresolved risk is the combination of a convincing voice, incorrect interpretation, and uncertain control over sensitive audio data.
Sources
- Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design – MarkTechPost, 2026-09-23
- Gemini 3.8 Flash TTS: Google will mit KI-Stimmen Hörbücher, Spiele und Podcasts verändern – The Decoder, 2026-09-23
- Gemini 3.8 TTS Playground – Simon Willison’s Weblog, 2026-09-23
- Gemini 3.8 text-to-speech says hello – Google DeepMind, 2026-09-23
- Alibaba stellt Sprechmodelle für Erkennung, Synthese und Echtzeit-Interaktion vor – The Decoder, 2026-09-23
- Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning – MarkTechPost, 2026-09-23
- NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time – MarkTechPost, 2026-09-23


Image: Jeremy Enns via Pexels
