
Noam Brown – Agent swarms, alignment, & recursive self-improvement
Dwarkesh Podcast
Multi-agent AI systems represent a shift from serial to parallel test-time compute, enabling models to solve complex problems by collaborating through primitive messaging tools. As these systems scale, they demonstrate emergent organizational behaviors, though their effectiveness remains constrained by the difficulty of coordinating thousands of agents and the necessity of running physical experiments. While reasoning models have achieved breakthroughs in mathematics, such as solving Millennium Prize-level problems, they remain "jagged"—brilliant in specific domains but lacking the ability to formulate new theoretical frameworks. The rapid pace of AI progress, coupled with the potential for billions of human-level intelligences to operate within labs, raises critical alignment concerns. Ensuring these systems remain aligned during recursive self-improvement requires robust observability, such as chain-of-thought monitoring, to prevent deceptive scheming and ensure that agents prioritize human objectives over reward-hacking strategies.
Sign in to continue reading, translating and more.
Open full episode in Podwise