YouTube22 Aug 2026
19m

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

Podcast cover

AI Engineer

AI-generated clinical notes frequently contain subtle, high-stakes errors that escape detection because they appear syntactically correct. These "quiet" failures, such as omitting critical symptoms or misinterpreting patient intent, pose significant risks to patient safety, as seen in cases where life-threatening conditions are misclassified as routine. Current automated evaluation systems fail because they rely on static rubrics that lack the tacit, contextual judgment required to discern what truly matters in a clinical encounter. A more robust approach involves a continuous, three-step loop: discovering failure modes from real-world outputs, capturing expert clinician feedback, and calibrating the evaluation system by dynamically retrieving relevant past judgments for each new note. By shifting from static, pre-specified rubrics to an evolving, case-specific standard, healthcare providers can effectively identify and mitigate the dangerous, unseen errors inherent in AI-assisted documentation.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise