Episode cover
09 Sept 2026
39m

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

Podcast cover

The a16z Show

Independent AI evaluation is becoming a critical infrastructure requirement as public benchmarks prove increasingly insufficient and susceptible to gaming. Rayan Krishnan, founder and CEO of Vals, argues that as frontier models advance, the industry needs a neutral, third-party framework to measure capabilities, risks, and recursive self-improvement. For enterprises, this shift is existential; token costs often eclipse human labor expenses, necessitating precise, data-driven methods to determine which models deliver the highest ROI. While government agencies are well-positioned to set regulatory standards for safety and policy, they rely on independent evaluators to verify whether models meet these thresholds. Ultimately, establishing a shared language for AI performance is essential for navigating the complex landscape of sovereign AI development, cybersecurity, and the rapid evolution of model capabilities.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise