Episode cover
YouTube03 Aug 2026

How Uber Evaluates AI Agents at Production Scale | Arize Observe 2026

Podcast cover

Arize AI

Building reliable AI agents at scale requires moving beyond treating evaluations as a mere launch gate. The transition from prototype to production hinges on making observability and tracing a default, rather than an afterthought. By automating evaluator creation and integrating signals directly into team workflows, developers can overcome the cold start problem and maintain data set relevance. Crucially, shifting the responsibility of managing these evaluations from engineers to product owners and designers—who possess deeper customer context—significantly improves agent quality. The Rider Voicebooking team, for instance, identified critical failures through turn-count spikes rather than static metrics, demonstrating that effective evaluations must be dynamic, trust-based, and deeply integrated into the product development lifecycle. Ultimately, evaluations serve as the engine for product insight, transforming how teams iterate and improve agent performance in real-time.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise