Episode cover
YouTube17 Jul 2026

How To Build AI Evals

Podcast cover

Hamel Husain

Implementing robust evaluation systems for AI agents requires a systematic approach that prioritizes observability and high-quality, human-labeled data. Lucas, a developer at Nova Escola, details his transition from manual, spreadsheet-based assessments to an automated pipeline integrated with Claude and Langfuse. By instrumenting production systems to capture traces and utilizing LLM-as-a-judge, the team successfully identified critical failure modes and improved the accuracy of generated lesson plans. The process emphasizes that effective evaluation focuses on persistent, high-impact issues rather than exhaustive metrics. Establishing a "benevolent dictator" for labeling and maintaining clear, specific rubrics minimizes inter-annotator disagreement, ensuring the judge remains reliably calibrated. This workflow transforms evaluation from a significant technical barrier into a strategic asset, enabling faster development cycles and higher confidence in the quality of AI-generated educational content.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise