YouTube03 Oct 2026
18m

Your LLM App Returned 200 OK. It Was Still Wrong. — Marina Petzel, Datadog

Podcast cover

AI Engineer

Traditional "golden signals"—latency, error, traffic, and saturation—are insufficient for evaluating the health of generative AI applications. Unlike deterministic software, GenAI requires a multi-dimensional monitoring strategy focused on cost, safety, and output quality. Financial stability depends on granular tracking of token usage, model selection, and caching, as unoptimized queries can lead to significant budget overruns. Security demands robust defenses against prompt injection, PII leakage, and jailbreaking, which standard server error codes fail to detect. Furthermore, quality assessment must move beyond binary success metrics to include hallucination rates, relevance scoring, and RAG performance. By integrating these specialized metrics into the observability stack, developers gain visibility into the subjective and dynamic nature of GenAI, ensuring applications remain secure, cost-effective, and truly helpful for end users.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise