YouTube08 Jul 2026
9m

I Run a Fleet of AI Agents Across Three Machines. Here's What Broke. - Kyle Jaejun Lee, KRAFTON

Podcast cover

AI Engineer

Scaling AI coding agents requires moving beyond simple terminal interactions to a structured, hierarchical organization. Managing multiple agents manually creates a bottleneck where human attention becomes the limiting factor. By implementing a CEO-to-worker hierarchy, agents operate within scoped contexts and approval boundaries, significantly reducing cognitive load. State persistence is achieved by moving data out of model context windows and into file-based workspaces, allowing for system resets without losing progress. Hardware failures and resource constraints necessitate a distributed approach, where tasks are offloaded to headless Linux machines and controlled via a centralized gateway. Ultimately, the transition from manual management to a robust, automated infrastructure—leveraging concepts like Kubernetes for scheduling and resource allocation—is essential for maintaining reliable, scalable agent operations across diverse computing environments.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise