YouTube14 Jul 2026
21m

Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind

Podcast cover

AI Engineer

AI agents require rigorous evaluation to ensure performance and reliability, as shipping skills without testing leads to unpredictable behavior and negative outcomes. Skills—defined as modular instructions for models—should be treated as code, requiring clear directives rather than verbose descriptions. Developers should implement evaluation harnesses using JSON or YAML test cases to validate whether a skill triggers correctly and achieves the desired task. Key best practices include keeping skills concise, removing "no-ops" that consume tokens without changing behavior, and running ablation tests to determine when a skill has become redundant due to model improvements. By maintaining a suite of regression tests, teams can ensure consistent agent performance, optimize token usage, and confidently retire outdated skills that foundation models can now handle natively.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise