YouTube07 Jul 2026
14m

How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI

Podcast cover

AI Engineer

The "knowledge gap" between rapidly evolving LLM reasoning capabilities and stagnant search retrieval systems limits the effectiveness of AI agents in complex tasks. While models are highly capable, they often rely on inefficient, keyword-heavy queries optimized for legacy benchmarks, leading to significant performance drops in noisy environments. By implementing a specialized search agent that utilizes a multi-tool harness—including overview, semantic, filter, and grep functions—it is possible to recover near-theoretical "oracle" performance. Training these agents through supervised fine-tuning and on-policy reinforcement learning encourages natural language query formulation and precise tool selection. This approach significantly improves retrieval accuracy on benchmarks like Oply Congress and MetQA, demonstrating that optimizing the interaction between reasoning layers and retrieval tools is essential for building fast, precise, and cost-effective knowledge agents.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise