Gemini 3.8 LIVE is INSANE: Google's New AI Can Think WHILE It Talks
Link to our newsletter: https://bitbiased.ai/ Gemini 3.8 Live Extended Thinking just hit 82.6 on Artificial Analysis’ Speech-to-Speech Quality Index — edging out OpenAI’s closest comparable realtime voice model. But the benchmark isn’t the most interesting part. Google says this model can actually “think while talking.” That sounds far more human than what’s really happening. Google launched two models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both are built on the Gemini 3 Pro architecture and use native audio-to-audio processing, taking audio, images, and video directly in while streaming speech back out — instead of relying on the traditional speech-to-text → language model → text-to-speech pipeline. That architecture brings some serious upgrades: roughly 1.18-second time to first spoken word for standard Live, barge-in support, 97+ languages, mid-conversation language switching, and a massive 131,072-token input context window. Vision is now part of the conversation too. Google demonstrated Gemini tracking a physical chess board through a live camera, turning a hand-drawn UI sketch into React code, and interacting with content inside products like Docs and Gmail. Then there’s Extended Thinking. Google describes it as reasoning and speaking simultaneously, but the underlying mechanism is more precise. The model can give you an interim spoken response while its interaction remains “IN_PROGRESS,” continue an asynchronous search or function call in the background, and then resume speaking when that work finishes. Developers can even control the reasoning level through `thinking_config`. So this isn’t a model literally maintaining two human-like trains of thought at once. It’s an asynchronous architecture designed to eliminate the awkward silence that normally appears while an AI works on a more complicated task. And the benchmark results reveal why that matters. Gemini 3.8 Live Extended Thinking scores 82.6 on Artificial Analysis’ Speech-to-Speech Quality Index, compared with 81.5 for GPT-Live-1 “Astra.” But on τ-Voice — a benchmark focused on completing multi-step tasks through voice — Extended Thinking reaches 68.6% versus just 30.1% for standard Gemini 3.8 Live. That jump may be much more important than the headline voice-quality score. There’s also a major price difference between the modes. Standard Gemini 3.8 Live works out to roughly $0.84 per hour of audio based on Artificial Analysis’ pricing methodology, while Extended Thinking sits around $3.50 per hour because of the additional reasoning compute. And there are catches. Extended Thinking is gated behind Google AI Pro and Ultra for Workspace use, everything remains cloud-based, tool calls are asynchronous, and Google’s polished vision demonstrations don’t establish how reliably the system performs under messy real-world conditions. So is Gemini 3.8 Live actually a major step forward for conversational AI, or is “thinking while talking” mostly smart asynchronous engineering wrapped in much more ambitious language? We break down what Google actually shipped, how the architecture works, the benchmarks against OpenAI’s current competition, where the system breaks, who gets access, and which numbers actually matter. CHAPTERS 00:00 Gemini’s New Voice AI Takes the Lead 00:49 What Actually Launched 02:15 The Architecture Behind the Voice 03:37 Vision Enters the Conversation 04:45 “Thinking While Talking” — What’s Actually Happening 05:58 The Benchmarks, Against the Actual Competition 08:02 Where It Breaks 09:52 Who Actually Gets Access 10:35 The Real Verdict #gemini #googleai #gemini38 #voiceai #artificialintelligence