Episode cover
03 Aug 2026
1h 41m

The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

Podcast cover

Latent Space: The AI Engineer Podcast

Inference engineering has evolved from simple model deployment into a complex, integrated discipline where training and inference are increasingly unified. Optimizing large language model performance requires a multi-layered approach, including cache-aware routing, speculative decoding, and precise quantization strategies that preserve model fidelity while maximizing throughput. Hardware constraints, particularly memory bandwidth and interconnect speeds, dictate the effectiveness of techniques like tensor and expert parallelism. As models grow in size and complexity, the industry is shifting toward infrastructure-centric solutions, where GPUs function more like specialized ASICs to handle massive token volumes. Future advancements in video generation and continual learning will likely rely on hybrid architectures that combine diffusion and autoregressive methods, alongside improvements in inter-node communication to overcome current bottlenecks in KV cache management and data transfer.

Outlines

Part 1: Inference, Structured Output

Part 2: Engineering, Quantization, Scaling

Part 3: Parallelism, GPU Architecture

Part 4: Modalities, Future Convergence

Sign in to continue reading, translating and more.

Open full episode in Podwise