Summary
Valeria Wu Fon (product lead) and Tom Ouyang (engineer) from Google DeepMind present the research direction and architectural thinking behind the speech-to-speech model powering Gemini Live. The talk covers two decades of progress in speech AI — from multi-component acoustic-plus-language-model pipelines pre-2018, through end-to-end ASR systems, to today’s natively multimodal LLMs — framing Gemini’s approach as a unified token embedding space that learns audio, video, and text jointly during pre-training.
The central thesis is that voice is the most natural human interface, and that cascaded STT-LLM-TTS pipelines cannot deliver the latency, naturalness, or agentic capability required for robust universal voice agents. The DeepMind team describes its research as pulling on three simultaneous axes: low-latency conversational naturalness, strong instruction-following and reasoning intelligence, and broad multimodal input-output capability — with the added constraint that all three must work across non-English languages, which represent the majority of Gemini’s user base.
The talk highlights a core tension the team is actively researching: increasing model intelligence (e.g., extended thinking before tool calls) degrades latency and naturalness. Demos shown include live translation and embodied agent interfaces. The presentation gives a rare inside look at how Google DeepMind is positioning speech-to-speech models as the foundation for the agentic future across both consumer products and enterprise cloud APIs.
📺 Source: AI Engineer · Published September 15, 2026
🏷️ Format: Keynote Launch







