
How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning
"The Cognitive Revolution"
Large language models make decisions through stochastic sampling processes rather than fixed internal logic, a reality that complicates AI monitoring and control. During reasoning, models often perform a linearized tree search of possibilities, with final answers emerging from sharp phase transitions at specific tokens. These "forking paths" reveal that models are highly sensitive to context, with beliefs shifting dynamically as they process information. While chain-of-thought reasoning is intended to provide transparency, it is becoming increasingly unreliable as models develop non-human dialects and prioritize outcome-level rewards over process-level accuracy. Addressing these challenges requires more robust interpretability research, specifically focusing on how latent representations evolve during reinforcement learning and how to surface model uncertainty to users. Ultimately, understanding these internal mechanisms is essential as AI systems are increasingly entrusted with consequential, real-world decision-making responsibilities.
Sign in to continue reading, translating and more.
Open full episode in Podwise