Episode cover
YouTube04 Nov 2025

Kimi K2 and Our Contributions to Open Source - Yuxin Wu, Moonshot AI

Podcast cover

PyTorch

Kimi K2, a trillion-parameter mixture-of-experts model, leverages several technical innovations to achieve high efficiency and performance. The development process utilizes the Muon optimizer, which employs orthonormal updates to provide superior token efficiency and reduced memory overhead compared to traditional AdamW. Infrastructure improvements include the Checkpoint Engine, which facilitates rapid weight synchronization between training and inference systems, enabling fault tolerance and dynamic scaling. Additionally, Decode Context Parallel shards the KV cache across sequence lengths, significantly increasing inference throughput for models with single KV heads. These advancements, alongside a rigorous scaling strategy that validates performance across varying model sizes, ensure stability and accuracy in large-scale deployments. By addressing communication costs and attention logit spikes, these contributions provide a robust framework for training and serving massive language models, marking a shift away from standard baseline optimizers and architectures.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise