YouTube22 Aug 2026
21m

FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft

Podcast cover

AI Engineer

The shift from "token maxing" to "value maxing" is essential for managing the high costs of agentic AI workflows. Current industry practices often rely on request-level gateways that fail to prevent runaway loops or excessive consumption. TokenOps introduces an out-of-bound control plane that operates at the agent run layer, providing granular attribution and real-time governance. By implementing policies like context compaction and tool output reduction, systems can steer agent behavior to stay within budget rather than simply halting execution. Benchmarks on frameworks like MetaGPT demonstrate that this approach reduces average spend by 78% while simultaneously boosting task completion rates from 67% to 96%. This architecture enables developers to maintain cost-effective, reliable agentic systems through proactive, in-place policy enforcement that balances performance with strict fiscal constraints.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise