KO / EN

Gemini 3.8 Flash TTS Arrives: Google Trades 30 Preset Voices for Voices You Design by Prompt

Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, moving voice generation from a catalog of 30 original preset voices to generative voice design driven by natural-language prompts. Both models cover more than 100 languages and dialects, ship a library of over 2,000 production-ready voices, and require a verbal consent recording from the voice owner before replicating a voice from a 30-second sample. Google reports that 3.8 Flash TTS takes the number one overall spot on Hume AI's Voice Design Benchmark at 71.4 and leads accent modeling at 60.8, while 3.8 Flash TTS and Flash-Lite TTS hold first and second place on Hume AI's Overall Quality Index. ASAP works only from Google's own announcement to separate what changed from what is still undisclosed, and argues that the consequential shift here is not the scores but the move of the competitive axis from audio fidelity to directing authority and consent procedure.

The two models split by the character of the work, not by quality tier

The two models Google shipped on the same day are not a premium version and a budget version of one product. Gemini 3.8 Flash TTS is built for deep creative direction and character design, aimed at gaming, immersive audiobooks, podcasts, and interactive media where a new character voice has to be invented. Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale, aimed at large dubbing runs, audio content production, and expressive voice agents.

What both models share is performance control. A script can carry stage directions written line by line, or Gemini can steer delivery from natural script cues, with acting cues, pacing, dialect shifts, and backchanneling all inside the control surface. Backchanneling here means the short reaction sounds a listener drops into someone else's speech. Google demonstrates non-verbal cues such as laughs, sighs, and gasps, plus active-listening interjections, written directly into the script as tags to land comedic timing and reaction beats.

Both models extend a line Google now groups as the Gemini Audio family. Gemini 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking came earlier, and these two TTS models fill in the synthesis side of that family.

The route to inventing a voice matters more than the 2,000-voice library

The feature Google lists first is generative voice design. In 3.8 Flash TTS, role, accent, and voice characteristics can be specified by natural-language prompt to build a voice from scratch, across more than 100 languages and dialects. The examples Google gives are characters that no preset catalog would contain: a dramatic fire-breathing dragon, a high-energy DJ from Melbourne, a super-tinny monotone robot, and a dragon speaking Japanese.

The prepared catalog grew as well. More than 2,000 voices are production-ready, and the coverage includes regional varieties such as Mexican Spanish, Quebec French, and Scots English. Designed voices can be saved and managed so timbre does not drift across an ongoing project, and voice remixing, which picks a library voice and tunes timbre, pitch, pace, and accent, is listed as coming soon.

Voice replication runs on a 30-second audio sample. The stated condition is that the voice is the user's own or one they hold the rights to, and the flow carries consent verification, SynthID watermarking, and C2PA credentials.

Citing a number one benchmark placement means naming who did the scoring

Google's reported numbers fall into three groups, and every one of them comes from an external scorer rather than from Google's own internal evaluation. First, 3.8 Flash TTS takes first place overall on Hume AI's Voice Design Benchmark at 71.4 and first in accent modeling at 60.8. Second, on Hume AI's Overall Quality Index, 3.8 Flash TTS is first and 3.8 Flash-Lite TTS is second. Third, in blind human preference evaluations on Voice Arena, both models place near the top in key languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

The qualifier to attach is who scored the models. All three results come from external benchmarks and arenas rather than Google's internal evaluation, yet the announcement is Google selecting which of its own results to quote. The Overall Quality Index placements in particular arrive as ranks without scores, so the gap between first and second cannot be read from this announcement at all.

The comparison baseline is also worth naming. The major improvements Google cites, across use cases such as long-form content and dual-speaker screenplay control, are measured against Gemini 3.1 Flash TTS rather than against a competitor. Generational improvement and competitive advantage are different claims, and merging the two layers when citing this launch overstates it.

Pricing is absent. Flash-Lite TTS is positioned as the high-volume, cost-efficient option, but no per-token or per-second rate accompanies it, so a team planning a dubbing pipeline cannot build a cost model until the rate appears.

The competitive axis moved from audio fidelity to directing authority

What separates this launch from the recent run of TTS releases is the unit of work rather than a fidelity metric. Commercial TTS has largely been chosen on how human a voice sounds, and what a user controlled were global parameters: which voice, how fast, what pitch. What these models sell instead is per-line acting direction, two-speaker scene staging, and inserted non-verbal cues. The unit rose from the voice to the scene.

That shift changes which skill is scarce. Once selecting a voice from a dropdown becomes writing a script with direction in it, output quality is decided less by the model than by the directing judgment of whoever writes the cues. Google's partner list supports that reading. On the developer platform side are Agora, LiveKit, Pipecat, and Vercel, while the integrating companies named are Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang. The stated purposes are accelerating global dubbing, localizing media with nuanced regional accents, and powering conversational voice agents. All are areas where script and direction already sit at the center of the work.

For content teams outside the United States, two practical consequences separate out. Support for more than 100 languages and regional varieties lowers localization cost, and prompt-designed character voices add a draft stage ahead of casting human voice talent. One caveat: this announcement carries no separate result for Korean, and Korean is not among the languages listed as top-placing on Voice Arena, so Korean quality has to be tested directly.

Making the consent recording mandatory is the real product-design fork

For voice replication, Google chose prior verification over after-the-fact reporting: before a voice can be created, users must supply a verbal consent recording from the voice owner that matches the reference speaker, and without that match the replication does not proceed. On top of that, every audio clip from the Gemini Audio models carries a SynthID watermark, imperceptible to a listener, so that AI-generated speech stays detectable.

This matters because misuse of a cloned voice does not resemble misuse of a generated image or text. A voice has long served as a stand-in for identity verification over a phone call, and dropping the cloning threshold to 30 seconds destabilizes that assumption. Building consent verification into the product flow reads as a decision to handle the problem as a feature rather than as a line in the terms of service.

The gap that remains is scope. Consent recordings are required on the path that runs through Google, and synthetic speech produced by other routes carries no such step. SynthID likewise matters only where a detector exists, so whether the watermark is actually checked at distribution becomes the next question. Google itself points to the model card for the full safety and responsibility details.

What is left to verify is pricing, Korean, and real-world variance

Availability is currently split. Gemini 3.8 Flash TTS reaches developers through the Gemini API and Google AI Studio starting the day of launch, arrives for enterprises via API in Gemini Enterprise as coming soon, and reaches everyone in Gemini Notebook. Gemini 3.8 Flash-Lite TTS follows the same developer path, with Google Vids as the consumer surface. Google AI Studio also opened an audio playground where a voice can be prompted from scratch or replicated, then carried into a dual-speaker screenplay editor to direct line-by-line delivery.

To summarize what is established: the division of roles between the two models, the feature list, the external benchmark results Google chose to quote, the consent and watermark procedure, and the rollout paths. What is not established: pricing, Korean quality, and how consistently per-line direction reproduces in an actual production pipeline. The claim of minimal speaker drift across long-form generation also belongs in the column that has to be verified at hour-scale length. For anyone putting speech synthesis into a workflow, measuring variance on your own scripts is more useful than the benchmark ranking.

Source: Gemini 3.8 text-to-speech says hello (Google, September 23, 2026)

ASAP — AGI Soon As Possible

AI & tech,
read in depth

AGI Soon As Possible · asapai.co.kr

← All posts