Episode cover
YouTube25 Aug 2026

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

Podcast cover

Invest Like The Best

Making AI token inference as cheap as possible requires a full-stack approach that optimizes software, hardware, and energy usage. Neil Movva, founder of SAIL Research, argues that the future of intelligence lies in background, long-horizon agents rather than latency-sensitive chatbots. By adopting a "scavenger strategy," the company achieves unbeatable economics by utilizing diverse, non-bleeding-edge hardware and distributed, lower-uptime data centers that others overlook. This approach involves writing custom GPU kernels, hybridizing compute and memory architectures, and leveraging intermittent power sources like solar and wind. Ultimately, the goal is to commoditize intelligence, enabling agents to operate on human time scales and tackle complex, verifiable tasks—such as deep research and cybersecurity—without the constraints of traditional, expensive inference models. This shift toward abundant, cheap tokens will unlock new product categories and fundamentally change how organizations and individuals harness machine intelligence.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise