metamorworks - stock.adobe.com
Gemini 3.8 Live Transforms Conversational AI
The models support real-time conversation by enabling users to interact without the traditional prompt-response structure.
Google introduced two new audio models that enable enterprises to deploy intelligent AI agents with reasoning depth and more interactivity and conversational speed.
Launched on September 15, Gemini 3.8 Live and Live Extended Thinking let developers build voice agents that can reason and execute tasks while maintaining the flow of the conversation, Google said. The models’ key capabilities include performing API and tool calls while continuing to stream audio responses to users, providing live visual inputs that help agents understand what users say and see, and supporting more than 97 languages. The 3.8 Live Extended Thinking version supports configurable thinking, a feature that lets users turn an AI model’s internal reasoning on or off.
The release of the audio models comes two weeks after Google initially introduced the Gemini 3.8 Flash and 3.8 Cyber foundation models. With 3.8 Live’s capabilities, Google appears to be catching up to vendors such as Apple, which already offers real-time transcription in productivity apps such as Notes and Voice Memos, said Bradley Shimmin, an analyst at Futurum Group. Still, Google tech takes it to the next level by changing how users interact with AI, he said.
Real conversations
“The way it's architected … you can customize this to behave in a lot of different ways,” Shimmin said. He noted that while enterprises can use the models for traditional translation use cases, Google has also designed them so users don’t interact with them in a prompt-response way; rather, they can have something of a normal conversation.
“It's just a natural conversation with the ability to interrupt it in real time to inject,” Shimmin said. “You can inject context into the conversation without interrupting what it's doing.”
With Google’s advanced speech technology capability, the models don’t just answer a question at face value, said Sid Nag, founder of Tekonyx.
“The model is designed to do background reasoning,” Nag said. “It’s doing it during the conversation so you can keep the conversation alive while it’s reasoning.”
This technology is a development in which interactive latency is separated from reasoning latency. Interactive latency is the delay between a user performing an action and the system responding. Reasoning latency is the total time it takes for an AI model to process information.
“It’s doing it in real time, rather than halting and then coming back with another answer,” Nag said.
The speech models also change how an AI agent responds and interacts because “an AI agent doesn’t necessarily have to choose between being fast and being thoughtful,” Nag continued. “It can actually maintain a real-time interaction.”
Enterprise applications for the new models include customer support chatbots, sales processes and multimodal agentic applications that combine text, voice and video, Nag said.
Esther Shittu is an Informa TechTarget news writer and podcast host covering AI technology.