Episode cover
YouTube01 Oct 2026

Inference Engineering: Frontier Open Models on the Hardware You Already Own with Prince Caluma

Podcast cover

Lisbon AI

Apple Silicon serves as a robust distributed compute base for running large language models locally, effectively bridging the gap between consumer hardware and cloud-based performance. Achieving this requires mastering inference engineering through techniques like weight and KV cache quantization, which minimize memory footprints without significant accuracy loss. Speculative decoding, particularly methods like MTP and D-Flash, provides substantial speed improvements for dense models, while token eviction strategies further optimize context management. These advancements enable resource-efficient, real-time applications, such as voice-integrated coding agents and multi-modal processing, directly on devices. By leveraging unified memory and optimized engines like MLX VLM, developers can handle 70% to 90% of standard AI workloads locally, drastically reducing operational costs and latency while maintaining high intelligence per watt.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise