Back to list
Sep 16, 2026
8
0
0
GeminiNEW

Gemini 3.8 Live and Extended Thinking Split Voice AI in Two

Google splits real-time voice AI into two tiers: Gemini 3.8 Live for scale, Extended Thinking for complex reasoning - both share one price table.

#Gemini#Google#Gemini Live#Voice AI#Extended Thinking
Gemini 3.8 Live and Extended Thinking Split Voice AI in Two
AI Summary

Google splits real-time voice AI into two tiers: Gemini 3.8 Live for scale, Extended Thinking for complex reasoning - both share one price table.

Introduction

On September 15, 2026, Google introduced two new voice AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The announcement, posted at 17:00 UTC and co-authored by Tom Ouyang, Principal Engineer, and Malini Jaganathan, Member of Technical Staff, speaks on behalf of the Gemini Audio Team. Google describes the pair as "our most advanced live dialogue models yet." Rather than shipping a single model, the launch splits real-time voice AI into two purpose-built tiers: Gemini 3.8 Live, built for scale and cost efficiency, and Gemini 3.8 Live Extended Thinking, built for high-complexity tasks that benefit from reasoning during a spoken response. Both begin rolling out the same day across developer tools, enterprise previews, and consumer surfaces.

Feature Overview

The two tiers are aimed at different workloads, and the benchmark numbers attach to specific models rather than the pair as a whole.

Gemini 3.8 Live: built for scale

Gemini 3.8 Live pairs conversational intelligence with fluid dialogue and visual grounding, aimed at handling high volumes of live conversations without the overhead of the reasoning tier. Google reports it placed second in the Speech Agent Arena tracked by Artificial Analysis. It processes visual input in near real time and can automatically detect and switch between any of 97 supported languages mid-conversation without a manual reset. It also executes tool calls and API requests in the background: rather than pausing the conversation to wait on a lookup or an action, the model can acknowledge the request, keep talking, and return with results once the task finishes.

Gemini 3.8 Live Extended Thinking: built for complexity

The Extended Thinking tier targets more demanding agentic workflows. Google reports it ranks #1 overall on Artificial Analysis' Speech to Speech Quality Index, with a score of 82.6. On agentic task completion, Google cites 68.6% on the tau-Voice benchmark and 35.1% on Sierra's tau-Voice-banking benchmark, alongside 97.7% on Big Bench Audio. Functionally, the model reasons and speaks at the same time: instead of a silent processing pause, it can respond with a verbal placeholder such as "Let me check that..." and narrate progress out loud while a multi-step background task runs. Google says it maintains "a highly competitive price point compared to other frontier models" despite the added reasoning step.

Shared ground: enterprise benchmark and audio watermarking

Both models were evaluated on ServiceNow's EVA-Bench, a voice-agent evaluation suite. Google says the pair "push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality" - though its own footnote clarifies that particular test ran the Live API on the Gemini Enterprise Agent Platform specifically, not a generic deployment. Every piece of audio the models generate carries a SynthID watermark, and Google points to a public model card for details on its safety and responsibility approach.

Usability Analysis

The pricing structure is one of the more interesting decisions in this launch. Google publishes a single Live API price table covering gemini-3.8-live, gemini-3.8-live-extended-thinking, and the older gemini-3.1-flash-live-preview together. Paid-tier input runs $0.75 per million text tokens, $3.00 per million audio tokens (or $0.005 per minute), and $1.00 per million image/video tokens (or $0.002 per minute); output, including thinking tokens, runs $4.50 per million text tokens and $12.00 per million audio tokens (or $0.018 per minute). A free tier remains available in Google AI Studio, and Grounding with Google Search is supported, with 5,000 free search requests a month shared across the Gemini 3.x family before a $14-per-1,000-request charge applies. The notable part is what Google did not do: it did not price Extended Thinking higher than the standard tier. Both currently sit on the same rate card, so the added reasoning capability is not gated behind a separate cost, at least at launch.

Access is staged by audience rather than opened all at once. Developers can reach both models today through the Gemini API and Google AI Studio. Enterprise customers get a private preview inside Gemini Enterprise, with Gemini Enterprise for Customer Experience listed as "coming soon" - and, for Extended Thinking specifically, a coming-soon slot for Google Workspace business customers as well. Everyday users see the split most directly: Gemini 3.8 Live surfaces through Search Live, while Extended Thinking becomes the model behind Gemini Live, reaches Google AI Pro and Ultra subscribers inside Workspace Docs, and reaches all Google AI subscribers in Gmail and Keep. On the developer ecosystem side, Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents are named as platforms integrating through the Gemini Live API, and Google names Salesforce, Genspark, and Lumeris as partners citing latency, conversational fluidity, and tool-calling as their reasons for adopting it.

What Google has not published matters just as much. There are no latency figures in milliseconds for either model, no parameter counts, no context-window size, and no committed general-availability date for the enterprise private preview - details that would normally accompany a benchmark-heavy launch like this one.

Pros and Cons

On the strength side, splitting the models by workload is a sensible response to a real tradeoff: cost-sensitive, high-volume conversations don't need the same reasoning overhead as a complex banking or multi-step agentic task, and Google now has a model tuned for each. Extended Thinking's benchmark position - first on Artificial Analysis' Speech to Speech Quality Index - combined with its ability to narrate progress verbally during background work, addresses one of live voice AI's persistent weaknesses: dead air during processing. Keeping both models on one shared price table is a genuinely user-friendly pricing decision, and background tool execution on the standard tier means simple lookups no longer interrupt a live conversation.

The limitations are concrete rather than speculative. Every hard performance number in the announcement is vendor-reported, sourced to Google's own citation of third-party leaderboards such as Artificial Analysis and Sierra, rather than independently reproduced. Enterprise availability is gated behind a private preview with no announced timeline for general release. Google also discloses nothing about latency, model size, or context window - figures that matter directly to anyone evaluating a live voice product for production use.

Outlook

The two-tier structure suggests Google intends to keep splitting future Live releases by workload rather than shipping one model that tries to do everything. If the shared pricing between Gemini 3.8 Live and Extended Thinking holds beyond launch, developers could prototype against the cheaper tier and move to the reasoning tier for harder tasks without renegotiating cost assumptions - a meaningfully different posture than pricing the more capable model at a premium from day one. The list of integrating platforms, spanning infrastructure players like LiveKit and Agora alongside named enterprise adopters like Salesforce, suggests Google is pushing the Live API beyond its own consumer surfaces. The coming-soon slots for Workspace and Customer Experience are the places to watch for how far that ambition extends.

Conclusion

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking give Google two distinct answers to the same problem - real-time voice AI - rather than one compromise. Developers building high-volume, latency-sensitive voice apps have a clear reason to look at Gemini 3.8 Live's background tool execution and mid-conversation language switching; teams building complex voice agents that need to reason through multi-step tasks out loud are the target audience for Extended Thinking. Both are rolling out now through the Gemini API and Google AI Studio, with enterprise and consumer access arriving on a staggered, and still partly unconfirmed, timeline.

Editor's Verdict

Gemini 3.8 Live and Extended Thinking Split Voice AI in Two earns a solid recommendation within the Gemini space.

The strongest case for paying attention: purpose-built two-tier design matches model capability to workload instead of forcing one model to serve both cheap, high-volume chat and complex reasoning tasks. That alone raises the bar for what readers should expect in this space. Reinforcing that, Extended Thinking ranks #1 overall on Artificial Analysis' Speech to Speech Quality Index and can narrate progress during background tasks instead of leaving dead air — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the launch splits voice AI by workload rather than shipping one model: Gemini 3.8 Live targets scale and cost efficiency, Extended Thinking targets complex, reasoning-heavy agentic tasks. On the other side of the ledger, one constraint is real rather than a marketing footnote: all headline benchmark figures are vendor-reported, citing third-party leaderboards like Artificial Analysis and Sierra rather than independently reproduced results. It should factor into any serious decision. Layered on top of that, enterprise access is limited to a private preview with no announced general-availability date for Gemini Enterprise or Gemini Enterprise for Customer Experience — which narrows the set of teams for whom this is an obvious yes.

For Google Cloud and Workspace integrators, multimodal-first teams, and Gemini API adopters, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Advertisement

Pros

  • Purpose-built two-tier design matches model capability to workload instead of forcing one model to serve both cheap, high-volume chat and complex reasoning tasks.
  • Extended Thinking ranks #1 overall on Artificial Analysis' Speech to Speech Quality Index and can narrate progress during background tasks instead of leaving dead air.
  • Both models currently share one Live API price table, so the more capable reasoning tier is not priced at a premium over the standard tier.
  • Gemini 3.8 Live's background tool and API execution lets simple lookups complete without interrupting the live conversation.
  • Broad ecosystem support at launch, with Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents integrating via the Gemini Live API.

Cons

  • All headline benchmark figures are vendor-reported, citing third-party leaderboards like Artificial Analysis and Sierra rather than independently reproduced results.
  • Enterprise access is limited to a private preview with no announced general-availability date for Gemini Enterprise or Gemini Enterprise for Customer Experience.
  • Google discloses no latency figures in milliseconds, no parameter counts, and no context-window size for either model.
  • The EVA-Bench result covering both models applies specifically to the Live API on the Gemini Enterprise Agent Platform, not deployments generally.
Advertisement

Comments0

Key Features

1. Two-tier launch - Gemini 3.8 Live (built for scale and cost efficiency) and Gemini 3.8 Live Extended Thinking (built for high-complexity, reasoning-heavy tasks) ship the same day, September 15, 2026. 2. Gemini 3.8 Live - fluid dialogue with visual grounding, near real-time visual processing, automatic mid-conversation switching across 97 supported languages, and background tool/API execution while the conversation continues. 3. Gemini 3.8 Live Extended Thinking - #1 overall on Artificial Analysis' Speech to Speech Quality Index (82.6); reasons and speaks simultaneously with verbal acknowledgment cues and live progress narration through multi-step background tasks. 4. Shared enterprise benchmark - both models were evaluated on ServiceNow's EVA-Bench (Live API on the Gemini Enterprise Agent Platform), which Google says balances accuracy with conversational quality for complex workflows. 5. Safety - every generated audio output carries a SynthID watermark; a public model card documents the underlying audio architecture. 6. Shared pricing - one Live API rate table currently covers gemini-3.8-live, gemini-3.8-live-extended-thinking, and gemini-3.1-flash-live-preview, so the reasoning tier is not priced at a premium over the standard tier.

Key Insights

  • The launch splits voice AI by workload rather than shipping one model: Gemini 3.8 Live targets scale and cost efficiency, Extended Thinking targets complex, reasoning-heavy agentic tasks.
  • Gemini 3.8 Live Extended Thinking, not the standard tier, holds the headline benchmark numbers: #1 on Artificial Analysis' Speech to Speech Quality Index (82.6), 68.6% on tau-Voice, 35.1% on Sierra's tau-Voice-banking, and 97.7% on Big Bench Audio.
  • Gemini 3.8 Live's distinguishing features are second place in the Speech Agent Arena, 97-language mid-conversation switching, and background tool execution that doesn't interrupt the conversation.
  • Both tiers currently share one Live API price table, meaning Extended Thinking's added reasoning is not gated behind a higher cost at launch.
  • Extended Thinking can narrate progress verbally during multi-step background tasks, directly addressing the dead-air problem common to live voice AI.
  • The EVA-Bench result Google cites for both models was specifically run on the Live API within the Gemini Enterprise Agent Platform, not a generic deployment - a distinction Google states in its own footnote.
  • Enterprise access is a private preview with no announced general-availability date, and Google discloses no latency, parameter count, or context-window figures for either model.
  • Consumer rollout differs by tier: Gemini 3.8 Live surfaces through Search Live, while Extended Thinking becomes the model behind Gemini Live and reaches Workspace Docs, Gmail, and Keep.

Was this review helpful?

Share

Twitter/X
Advertisement