
Ep. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos
SemiAnalysis Weekly
The AI hardware industry is pivoting from maximizing HBM capacity to prioritizing bandwidth and supply chain efficiency, as evidenced by the redesign of NVIDIA’s Rubin Ultra from 1TB to 192GB of HBM. This shift toward 4-high and 8-high HBM stacks addresses critical manufacturing yield losses and severe memory supply shortages. While previous generations favored increasing stack height to accommodate larger model parameters, current research indicates that parameter scaling has slowed due to advanced techniques like loop transformers and efficient post-training. Consequently, the industry is optimizing for tokens-per-watt and tokens-per-dollar rather than raw capacity. Because memory production remains constrained by physical limitations—such as clean room availability and EUV tool scarcity—this move toward lower-height stacks serves as a strategic necessity to maximize accelerator output and alleviate bottlenecks across the broader semiconductor ecosystem.
Sign in to continue reading, translating and more.
Open full episode in Podwise