
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
The MAD Podcast with Matt Turck
AI inference speed serves as the primary driver for modern computing, as productive AI applications demand rapid token delivery to maintain real-time engagement. Cerebras CEO Andrew Feldman explains that traditional GPU architectures struggle with inference because they rely on slow memory movement between DRAM and compute units. By utilizing wafer-scale technology, Cerebras integrates massive amounts of fast SRAM directly onto a single chip, significantly reducing data movement bottlenecks. This architectural shift enables faster processing for complex AI workloads, including reasoning and agentic tasks. While the industry faces constraints in HBM memory, packaging, and factory capacity, the shift toward multi-silicon environments reflects a maturing ecosystem. OpenAI’s massive investment in Cerebras’s infrastructure highlights the growing necessity for specialized, high-performance silicon to meet the exponential demand for AI compute capacity.
Sign in to continue reading, translating and more.
Open full episode in Podwise