
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Latent Space: The AI Engineer Podcast
Inference engineering has evolved from simple GPU optimization into a complex discipline integrating training and inference loops. Modern production-ready APIs rely on techniques like cache-aware routing, disaggregated prefill and decode, and speculative decoding to maximize throughput and minimize latency. Quantization remains a critical lossy optimization, where strategic layer selection and calibration preserve model fidelity while significantly reducing resource requirements. As hardware architectures like NVIDIA’s Rubin emerge, the bottleneck is shifting toward interconnect bandwidth and memory management, necessitating systems-level solutions like KV-cache offloading. The future of the field lies in the unification of training and inference, where models continuously learn from live traces and optimize their own kernels, effectively bridging the gap between general-purpose GPU compute and specialized ASIC performance.
Sign in to continue reading, translating and more.
Open full episode in Podwise