YouTube19 Sept 2026
15m

Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave

Podcast cover

AI Engineer

CoreWeave’s inference platform architecture optimizes for diverse AI workloads by balancing serverless and dedicated consumption models. The platform utilizes KVCache-aware routing to efficiently manage agentic and chat traffic, significantly reducing costs by minimizing redundant prefill computations. For latency-sensitive tasks like real-time voice and video, the system employs specialized hardware distribution and performance levers such as quantization and speculative decoding. By offloading KVCache to high-bandwidth storage, the infrastructure maintains low latency across multi-turn conversations without requiring full re-computation. Furthermore, the platform supports flexible scheduling, allowing customers to transition capacity between real-time inference and batch processing based on demand. These design choices ensure high price-performance efficiency while accommodating varying workload shapes and SLA requirements across different GPU generations.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise