
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Latent Space: The AI Engineer Podcast
Inference engineering has evolved from simple model deployment into a complex, integrated discipline where training and inference are increasingly unified. Optimizing large language model performance requires a multi-layered approach, including cache-aware routing, speculative decoding, and precise quantization strategies that preserve model fidelity while maximizing throughput. Hardware constraints, particularly memory bandwidth and interconnect speeds, dictate the effectiveness of techniques like tensor and expert parallelism. As models grow in size and complexity, the industry is shifting toward infrastructure-centric solutions, where GPUs function more like specialized ASICs to handle massive token volumes. Future advancements in video generation and continual learning will likely rely on hybrid architectures that combine diffusion and autoregressive methods, alongside improvements in inter-node communication to overcome current bottlenecks in KV cache management and data transfer.
Part 1: Inference, Structured Output
Part 2: Engineering, Quantization, Scaling
Part 3: Parallelism, GPU Architecture
Part 4: Modalities, Future Convergence
Sign in to continue reading, translating and more.
Open full episode in Podwise