
Independent AI evaluation is becoming a critical infrastructure requirement as public benchmarks prove increasingly insufficient and susceptible to gaming. Rayan Krishnan, founder and CEO of Vals, argues that as frontier models advance, the industry needs a neutral, third-party framework to measure capabilities, risks, and recursive self-improvement. For enterprises, this shift is existential; token costs often eclipse human labor expenses, necessitating precise, data-driven methods to determine which models deliver the highest ROI. While government agencies are well-positioned to set regulatory standards for safety and policy, they rely on independent evaluators to verify whether models meet these thresholds. Ultimately, establishing a shared language for AI performance is essential for navigating the complex landscape of sovereign AI development, cybersecurity, and the rapid evolution of model capabilities.
Sign in to continue reading, translating and more.
Open full episode in Podwise