
Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date, with a particular focus on making voice-based AI more intelligent, responsive, and capable of handling complex tasks while maintaining a natural conversation.
The two new Gemini models introduce near real-time reasoning, allowing AI agents to do more than simply listen to a question and provide a spoken answer. They can reason, use tools, process visual information, and work through multi-step tasks in the background while continuing to talk with you.
I think this is where voice AI needs to head if it is going to become better. Talking naturally to an assistant is great, but having to stop the conversation every time it needs to look something up or perform an action quickly breaks the illusion. Gemini 3.8 Live is designed as the more scalable and cost-efficient model, combining conversational intelligence with fluid dialogue and visual grounding. According to Google, the model can process visual inputs in near real-time, automatically detect and switch between 97 supported languages during a conversation, and execute tools or API calls asynchronously.
So, rather than falling silent while performing a task, Gemini can acknowledge what you’ve asked, continue the conversation, and complete tool calls in the background.
For more demanding workloads, Google has introduced Gemini 3.8 Live Extended Thinking. As its name suggests, this version is designed for situations requiring deeper reasoning and multi-step problem solving.
Extended Thinking can reason in the background while continuing to speak. Google says it can provide natural verbal acknowledgements and progress updates as it works through more complicated tasks, rather than forcing users to sit through an awkward silence while the AI figures everything out.
This could make voice interactions feel significantly different from the traditional voice-assistant experience where every request effectively becomes a separate command.
The Extended Thinking model supports text, image, audio, and video inputs, with text and audio output, along with asynchronous function calling and Search grounding. Google positions it specifically for complex real-time interactions where additional reasoning is required behind the scenes.
Google demonstrated some interesting possibilities, including Gemini using visual context to assist with employee onboarding and even playing chess in near real-time while combining visual understanding, reasoning, and conversation.
The technology isn’t limited to Google’s own services either. Developers can access Gemini 3.8 Live and Live Extended Thinking through the Gemini Live API and Google AI Studio, with platforms including LiveKit, Vercel, Agora, Fishjam and Pipecat supporting developers building voice-driven applications around the models.
Google lists pricing for developers at US$0.005 per minute for audio input and US$0.018 per minute for audio output, which suggests Gemini 3.8 Live isn’t simply an experimental demonstration but something Google expects developers to deploy at scale.
For regular users, Gemini 3.8 Live is beginning its rollout through Search Live. Gemini 3.8 Live Extended Thinking is rolling out through Gemini Live, while Google AI subscribers are also getting access across parts of Google Workspace. Google AI Pro and Ultra subscribers can use it in Docs, while Google AI subscribers can access the technology through Gmail and Keep.
All audio generated by Google’s AI products is also watermarked using SynthID, allowing AI-generated audio to remain detectable.
Voice has always felt like one of the most natural ways to interact with an AI assistant, but latency and limited reasoning have also made it obvious that you’re talking to software. Gemini 3.8 Live and particularly Live Extended Thinking look like Google’s attempt to close that gap — not simply by generating more natural speech, but by letting the AI actually work, reason and use tools without bringing the conversation to a halt.
And that may end up being the more important advancement here. The future of voice assistants probably isn’t about getting a better spoken answer to a question. It’s being able to say what you want done, continue talking naturally, and have the AI actually get on with doing it.
Check out the official blog post for more details.




