YouTube15 Sept 2026
16m

Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

Podcast cover

AI Engineer

Speech-to-speech models represent the future of human-computer interaction, moving beyond traditional cascaded systems toward natively multimodal, end-to-end architectures. By integrating audio, video, and text into a unified token embedding space, these models enable sophisticated agentic capabilities, including real-time multilingual translation, proactive noise management, and visual presence through avatars. The core development challenge lies in balancing conversational latency with high-level reasoning and instruction-following, ensuring models remain responsive while maintaining deep intelligence. These advancements facilitate versatile applications, from roadside assistance agents that handle alphanumeric data under pressure to interactive search tools that interpret visual environments. Ultimately, the transition toward spoken-first interfaces promises to make artificial intelligence more accessible and natural, allowing for seamless, fluid communication across diverse languages and complex, real-world scenarios.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise