
Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon
No Priors: AI, Machine Learning, Tech, & Startups
Generative AI development is shifting from autoregressive models toward diffusion-based architectures to prioritize inference speed and efficiency. Stefano Ermon, a Stanford professor and co-founder of Inception, explains that while autoregressive models process tokens sequentially—creating memory-bound bottlenecks—diffusion models enable parallel, coarse-to-fine generation that maps more effectively to GPU hardware. This shift offers significant performance gains, particularly for latency-sensitive applications like voice agents. Beyond speed, diffusion models provide superior controllability and data efficiency, potentially offering a more scalable path for future multimodal AI systems. By building custom serving engines and optimizing inference, Inception aims to challenge the dominant transformer-based paradigm, demonstrating that architectural innovation can yield substantial competitive advantages even in a compute-constrained industry.
Sign in to continue reading, translating and more.
Open full episode in Podwise