YouTube21 Jun 2026
33m

⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai

Podcast cover

Latent Space

Continual learning represents the next major evolution in AI, shifting from static, pre-trained models to systems that dynamically improve through real-world user interactions. Ronak Malde, co-founder of Trajectory and former researcher at DeepMind and Windsurf, argues that capturing expert traces and user corrections provides a critical flywheel for model optimization. By utilizing self-distillation policy optimization (SDPO), Trajectory enables models to learn from nuanced, non-binary feedback, significantly outperforming traditional reinforcement learning approaches that rely on sparse reward signals. This infrastructure allows companies in regulated sectors like legal and finance to deploy faster, cheaper, and more accurate models by integrating production data directly into the training loop. Moving beyond coding-specific applications, this paradigm shift empowers enterprises to build intelligence that evolves alongside their specific workflows, effectively closing the gap between offline training and real-world performance.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise