Episode cover
YouTube20 Aug 2026

Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust

Podcast cover

AI Engineer

AI application architectures have undergone rapid, step-function shifts—from simple prompt-response models to complex, orchestrated agentic systems—necessitating a parallel evolution in evaluation strategies. Because static evaluation methods fail to capture the nuances of modern, non-deterministic workflows, developers must adopt a dynamic "flywheel" approach that continuously harvests production data to inform testing. This process requires moving beyond basic accuracy metrics toward statistical measures like "pass at k" and "pass wedge k" to quantify reliability in loop-based systems. Furthermore, identifying novel failure modes is critical as architectures grow more sophisticated. Tools that perform cluster analysis on production data, such as BrainTrust’s TopX, enable teams to uncover unanticipated failure categories, ensuring that evaluation frameworks remain robust and congruent with the latest model capabilities and architectural designs.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise