02 Sept 2026
1h 40m

Designing How AI Grows — Tom McGrath

Podcast cover

Machine Learning Street Talk (MLST)

Mechanistic interpretability functions as a natural science for neural networks, enabling the extraction of hidden scientific knowledge and the implementation of intentional design in machine learning. By utilizing tools like sparse autoencoders and gradient readout, researchers map the internal geometry of models to identify convergent structures, such as modular arithmetic units or persona-based representations. This capability allows for closed-loop control over training, where engineers steer models away from emergent misalignment and reward hacking by intervening on specific conceptual manifolds rather than relying on blunt scalar rewards. As models evolve into self-adapting systems, interpretability provides the necessary oversight to monitor for deception and goal-seeking behavior. Moving beyond simple heuristics, this approach treats neural networks as complex, modular computers, offering a path to align increasingly autonomous agents with human values through precise, representation-based intervention.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise