
Ep. 034 - Engrams: How DeepSeek Offloads KV Cache to DRAM and SSD (Core Research) | Jordan Nanos, Cam Quilici, Alec Ibarra, Bryan Shan
SemiAnalysis Weekly
Architectural co-design and hardware optimization drive significant gains in AI inference efficiency. Engrams enable the offloading of model components to DRAM or SSD, effectively bypassing HBM bandwidth constraints by leveraging token ID-based retrieval. The AgentX benchmark demonstrates that agentic coding workloads exhibit high cache hit rates, making optimized serving systems highly profitable. Recent hardware advancements, including TPU v7 and Vera Rubin, provide substantial performance boosts, further enhanced by specialized software techniques like lookup table quantization and mega-kernels. While general-purpose GPUs remain central to the ecosystem, the development of latency-optimized engines like TileRT allows for massive throughput improvements. These innovations collectively address the growing demand for high-concurrency, interactive AI services, proving that strategic hardware-software integration is essential for scaling modern large language model inference.
Sign in to continue reading, translating and more.
Open full episode in Podwise