02 Jun 2026
2h 48m

What it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | Rohin Shah

Podcast cover

80,000 Hours Podcast

Achieving AGI safety relies on prosaic alignment techniques rather than assuming catastrophic failure is the inevitable default. Current AI progress exhibits approximate continuity, suggesting that development remains predictable and manageable through iterative, evidence-based oversight. Public safety commitments often function as performative signals; instead, rigorous third-party auditing and granular, technical evaluations provide more reliable accountability. Transformer architectures maintain low opaque serial depth, ensuring that reasoning processes remain legible through chain-of-thought monitoring, which serves as a critical tool for safety oversight. Rather than attempting to forecast superintelligence challenges decades in advance, prioritizing practical infrastructure—such as treating AI agents as untrusted insiders with restricted permissions—offers a more effective path to safety. Empirical data from benchmark stitching indicates that AI progress remains linear, reinforcing the need for focused, implementable solutions over speculative long-term planning.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise