Gemini 3.8 Live: A Voice Model That Reasons and Speaks at the Same Time Takes the Top Spot at 82.6
Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026. Extended Thinking captures the number one overall spot on Artificial Analysis' Speech to Speech Quality Index at 82.6, leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking, and scores 97.7% on Big Bench Audio. The genuinely new part is structural rather than numerical. The model reasons and speaks simultaneously, executes tool calls in the background without pausing the conversation, and switches automatically among 97 supported languages mid-dialogue. ASAP works only from Google's own announcement to separate what changed from what has not been disclosed.
The two models are different roles, not two tiers of the same one
Google splits them by the character of the load, not by capability grade. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, carrying increased intelligence and multi-step reasoning.
The phrase Google leans on is that Extended Thinking reasons and speaks simultaneously. It acknowledges a prompt with early verbal cues such as "Let me check that…" and narrates progress through multi-step background tasks while they run. Given that the most awkward stretch in any voice agent has been the silence while the model thinks, this design targets an experience problem rather than a score.
The 3.8 Live model was measured on a different axis. It took second place in the Speech Agent Arena, where users express preference, and Google presents it as the cost-effective option built for scale. On ServiceNow's EVA-Bench, a benchmark for evaluating voice agents, Google states that both models push the Pareto frontier for complex workflows by balancing accuracy against conversational quality, with a note that the run used the Live API on the Gemini Enterprise Agent Platform.
Put the four scores side by side and they read differently
The most revealing number in the release is not the 82.6 first-place score but the distance between 68.6% and 35.1%. Both measure agentic task completion, but the first is τ-Voice and the second is τ-Voice-banking, Sierra's banking-specific version. A model that finishes close to seven of ten general voice-agent tasks drops to roughly three of ten once the domain narrows to banking.
That gap is the number that matters in practice. What most teams actually want to hand a voice agent is exactly the tightly proceduralized work of account lookups and booking changes, and the published figure for that class of task is 35.1%. The 97.7% on Big Bench Audio is best read as evidence that reasoning is not the bottleneck. What remains is not how smart the model is but whether it can run a multi-step procedure to the end without drifting off it.
The first-place quality score needs its own qualifier. Artificial Analysis' Speech to Speech Quality Index is a composite of several components, and Google pairs the 82.6 with a claim of a highly competitive price point against other frontier models. The announcement contains neither a unit price nor a latency figure. Latency is listed among the things partner companies called impressive, alongside fluidity and tool calling, yet it is never given as a number.
Three mechanisms that keep the conversation from stopping
The capabilities Google describes are three distinct mechanisms, each aimed at a different reason a voice agent normally goes quiet. First, near real-time visual input: 3.8 Live processes what the camera sees as conversational context, with Google's examples being live employee onboarding guidance and playing chess. Second, language switching: the model detects and transitions between 97 supported languages mid-conversation. Third, background tool execution: the model runs tools and API calls behind the dialogue so it can acknowledge a request and keep talking while the work finishes.
The Extended Thinking demos are chosen to stack all three at once. Google's published examples are turning raw sketches plus near real-time voice feedback into functional React components, coordinating multi-step bookings and asynchronous function calls without interrupting live conversation, and building complete business plans and marketing toolkits through natural speech alone. None of these completes within a single utterance, and holding the conversation together while the model works for seconds at a time is the substantive claim of this release.
Competition in the voice stack is moving from models to infrastructure
The longest section of the announcement is an ecosystem list rather than a benchmark table. Google names seven developer platforms building on the Gemini Live API, Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents, and explains that they manage the real-time media streaming infrastructure so developers can focus entirely on the user experience. Salesforce, Genspark and Lumeris are named as companies excited about the two models.
That framing shows where the competitive axis has shifted. Shipping a voice product has always been harder at the infrastructure layer, real-time audio streaming, barge-in handling, session management and tool-call relaying, than at the model layer. Putting seven media-infrastructure companies at the front of the announcement is a decision to let partners own that layer, which means a developer now chooses a real-time stack before choosing a model.
The watermarking policy belongs in the same reading. All audio generated by Google's AI products carries an imperceptible SynthID watermark woven into the output. With synthetic speech now past the quality bar of a phone call, shipping machine-detectable provenance as a default matters, and it is the first item any enterprise deploying voice agents into customer contact will be asked about by its regulators.
What is open today and what is not
Availability splits three ways for both models, with different conditions per audience. Developers get both in the Gemini API and Google AI Studio starting today. Enterprises reach them through a private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience listed as coming soon. For everyone, 3.8 Live arrives in Search Live and Extended Thinking arrives in Gemini Live and in Workspace.
The Workspace conditions depend on subscription tier. Google AI Pro and Ultra subscribers get Extended Thinking in Docs, and all Google AI subscribers get it in Gmail and Keep. Workspace business customers are covered only by a coming-soon line. Individual subscribers therefore have partial access today while enterprises reach it through preview access alone.
For teams outside English-speaking markets the operative detail is the 97-language support with automatic mid-conversation switching. Google published the language count and nothing about per-language quality or the performance of any individual language. Any team evaluating this for a voice agent should build its own procedural task set in the spirit of τ-Voice-banking rather than copying the launch numbers, because a published banking figure of 35.1% next to a first-place index score is itself the evidence that a benchmark win and a completed workflow are not measuring the same thing.
Source: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (Google, September 15, 2026)

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr