Episode cover
23 Jul 2026
1h 12m

The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

Podcast cover

The MAD Podcast with Matt Turck

AI inference speed serves as the primary driver for modern computing, as productive AI applications demand rapid token delivery to maintain real-time engagement. Cerebras CEO Andrew Feldman explains that traditional GPU architectures struggle with inference because they rely on slow memory movement between DRAM and compute units. By utilizing wafer-scale technology, Cerebras integrates massive amounts of fast SRAM directly onto a single chip, significantly reducing data movement bottlenecks. This architectural shift enables faster processing for complex AI workloads, including reasoning and agentic tasks. While the industry faces constraints in HBM memory, packaging, and factory capacity, the shift toward multi-silicon environments reflects a maturing ecosystem. OpenAI’s massive investment in Cerebras’s infrastructure highlights the growing necessity for specialized, high-performance silicon to meet the exponential demand for AI compute capacity.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise