Gemini 3.8 Live: What Google's Voice Models Change
Google released Gemini 3.8 Live and Extended Thinking on 15 September 2026. The headline is not the benchmark score, it is that the model runs tools without pausing the conversation.
Gemini 3.8 Live: What Google's Voice Models Change
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026, a pair of speech-to-speech models aimed at live dialogue rather than transcription. The headline capability is not the benchmark score. It is that the model can run tools and API calls in the background while the conversation keeps going, which is the thing that has made voice agents feel broken up to now.
If you are building anything that talks back, that single change matters more than the rest of the announcement combined.
What Google actually shipped
Two models, split by cost rather than by capability tier alone:
Model | Positioned for | Where it is available |
|---|---|---|
Gemini 3.8 Live | Scale and cost efficiency, fluid dialogue, visual grounding | Gemini API, Google AI Studio, Search Live |
Gemini 3.8 Live Extended Thinking | High-complexity, multi-step reasoning tasks | Gemini API, AI Studio, Gemini Live and Workspace for Pro and Ultra subscribers, private preview in Gemini Enterprise |
Both detect and switch between 97 languages mid-conversation, without being told the language in advance. Google did not publish pricing in the announcement post.
The benchmark numbers, read properly
Extended Thinking scores 82.6 on the Artificial Analysis Speech to Speech Quality Index and 97.7% on Big Bench Audio reasoning. Those are strong, and they are also the numbers a press release leads with.
The more useful figure is further down. On τ-Voice, a benchmark for completing agentic tasks by voice, Extended Thinking hits 68.6%. On Sierra's τ-Voice-banking, a narrower and more realistic domain, it hits 35.1%.
That gap is the story. A model that completes roughly a third of realistic banking tasks end to end is not a model you point at your customers' accounts unsupervised. It is a model you point at a scoped, low-stakes slice of the job, with a human path out. Google published the number rather than burying it, which is to its credit, but plenty of coverage this week quoted the 97.7% and skipped the 35.1%.
The same reading discipline applies to every release. A single composite score tells you how a model does on the average of many things, none of which is your thing.
Why background tool calling is the real change
Here is the failure mode every voice agent built before this has: the user asks a question, the model needs to look something up, and the conversation stops. Dead air for two or three seconds while an API call resolves. Users fill silence by repeating themselves, which corrupts the turn, which makes the agent ask them to repeat again.
Gemini 3.8 Live executes tools and API calls without interrupting the conversation, and the Extended Thinking variant can reason and speak at the same time, producing natural verbal acknowledgements while work happens underneath. In practice that means the agent can say something true and unremarkable while the lookup runs, rather than going silent.
This is a latency problem solved by architecture rather than by getting faster, and it is the difference between a demo and something you would put on a phone line.
What this means if you are building
Three practical notes.
Your prompt design changes. If the model can acknowledge while working, your instructions need to say what it should acknowledge with, or you get filler. This is a specific case of a general skill, and the same rules apply as when you prompt any AI voice agent: be explicit about the conversational contract, not just the task.
Language switching is now a default, not a feature. Ninety-seven languages detected mid-conversation means your agent may switch on you if a caller code-switches. Decide whether you want that and constrain it if you do not.
Scope by benchmark, not by vibe. The τ-Voice-banking number is a reasonable proxy for "regulated domain, real consequences, narrow tolerance for error". Use that class of number to decide what the agent owns and what it hands to a person.
For the mechanics of wiring speech into something you have already built, adding voice input to an AI-built app covers the plumbing side.
Is this release a big deal?
For most people building with AI, moderately. It is not a new frontier capability, and it does not change what is possible. It removes a specific, extremely annoying failure that made voice agents feel like toys.
Google shipped Gemini 3.8 Flash earlier this month, so this is the second 3.8-family release in two weeks. Cadence at that pace is worth noting on its own: if you are choosing a voice stack this quarter, the thing you pick will be superseded within one, and your integration should assume that. A useful habit here is judging whether a model release actually matters before you rearchitect anything around it, and keeping up with AI news in a way that does not cost you a day a week.
FAQ
How much does Gemini 3.8 Live cost?
Google did not publish pricing in the launch announcement. Check the Gemini API pricing page for current rates before you build a cost model around it.
Can Gemini 3.8 Live see as well as hear?
Yes. Google describes near real-time visual processing with contextual grounding, so the model can take video or image input alongside speech in the same session.
What is the difference between the two models?
Gemini 3.8 Live is positioned for cost efficiency and fluid dialogue at scale. Extended Thinking is for tasks needing multi-step reasoning, and it can reason while speaking rather than pausing to think first.
Does it work in languages other than English?
It supports 97 languages and switches between them automatically mid-conversation, without needing the language declared up front.
Should I switch my voice agent to it?
Only if silent pauses during tool calls are a real problem for you. That is the specific thing this release fixes. If your agent does not call tools mid-conversation, the case is much weaker.
When a provider changes the underlying model behind a feature you ship, your users notice a difference even if you did not change your own product. See how to tell users you changed the AI model for how to communicate that clearly.
Google followed this voice-only launch with a video layer on top of it, covered in Google adds Live Avatar to Gemini Enterprise.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


