YouTube17 Jul 2026
2h 20m

Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

Podcast cover

AI Engineer

AI model development is currently defined by an exponential growth trend, particularly since the emergence of reasoning capabilities, which have accelerated performance doubling times. However, the reliability of AI benchmarks is severely compromised by issues like reward hacking, data contamination, and the flawed practice of using LLMs to verify other LLMs. As hardware innovations reach diminishing returns, the industry is shifting toward software-centric optimizations, such as kernel fusion, gradient checkpointing, and reduced numerical precision, to sustain progress. While open-source models are successfully narrowing the gap with closed-source frontier models, the ecosystem faces significant challenges regarding cybersecurity, regulatory scrutiny, and the tendency of inference providers to prioritize throughput over accuracy. Ultimately, the focus must shift from raw parameter scaling to algorithmic efficiency and more rigorous, verifiable evaluation methodologies to ensure genuine intelligence gains.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise