
If an AI Model Can Cheat, It Will | Turing CEO on Reward Hacking
Sourcery with Molly O'Shea
The shift toward AI mastering real-world work requires engineering sophisticated, simulated reinforcement learning environments that mirror professional workflows. Rather than simply distilling human knowledge into models, success now depends on creating high-fidelity simulations where agents can train, verify outputs, and iteratively improve. Jonathan Siddharth, CEO of Turing, emphasizes that while frontier models drive superintelligence for global challenges like disease research and material science, open-weight models enable enterprises to retain sovereignty over proprietary workflows. This symbiotic relationship between human error correction and AI speed creates a continuous learning loop, essential for both safety and economic progress. Contrary to rapid-takeoff fears, AI integration will likely follow a decade-long trajectory of gradual, practical adoption, as agents move from solving isolated tasks to autonomously managing complex, multi-step enterprise processes. This evolution prioritizes verifiable, goal-oriented performance over mere model scaling.
Sign in to continue reading, translating and more.
Open full episode in Podwise