YouTube12 Aug 2026
23m

Scaling up Continual Learning — Ronak Malde, Trajectory

Podcast cover

AI Engineer

Continual learning represents the next frontier for AI, shifting from static benchmark training to systems that improve through real-world interaction. Current reinforcement learning methods, such as GRPO, suffer from high infrastructure demands, off-policy task distributions, and sequence-level reward limitations. On-Policy Self-Distillation (OPSD) addresses these bottlenecks by utilizing privileged information as a "hint" to guide student models, enabling token-level feedback and eliminating the need for massive parallel rollouts. While scaling OPSD to long-horizon tasks introduces challenges like reward hacking through hint leakage and divergence in tool calling, techniques such as step-level divergence weighting and residual guidance effectively mitigate these issues. By transforming production agent traces into continuous improvement loops, this approach allows AI systems to evolve dynamically, turning every interaction into a source of model refinement.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise