Episode cover
YouTube16 Jul 2026
33m

But what is cross-entropy? | Compression is Intelligence Part 2

Podcast cover

3Blue1Brown

Cross-entropy serves as a fundamental link between data compression and the training of large language models. By measuring the efficiency of a code optimized for one context when applied to another, this metric quantifies the divergence between predicted and actual data distributions. In machine learning, minimizing cross-entropy loss forces a model to align its probability outputs with the statistical patterns found in training data, effectively turning the model into a sophisticated text compressor. The choice of the negative log function is mathematically necessary to ensure that loss is minimized only when the model’s predictions match the underlying data statistics. Beyond pre-training, this framework enables model distillation, where smaller, efficient models learn to replicate the performance of larger systems by minimizing the cross-entropy between their respective output distributions.

Outlines

Sign in to continue reading, translating and more.

Open full episode in Podwise