Skip to content
Kudos AI

Kullback-Leibler Divergence

The number of extra bits per symbol paid for describing one distribution with a code built for another. It is zero only when the two agree, it is never negative, and it is not symmetric, so it is a cost rather than a distance.

Also known as: KL divergence, Relative entropy

Understanding Kullback-Leibler Divergence

Suppose a source emits symbols according to p and you build the shortest possible code for a different distribution q. Every symbol then costs log2(1/q) bits instead of log2(1/p), and the average bill is the cross-entropy. The Kullback-Leibler divergence is the part of that bill you did not have to pay: the difference between what you spend and the entropy of the source.

The decomposition is exact and it is the whole reason cross-entropy is the standard loss. For a source with probabilities 0.6, 0.25, 0.1 and 0.05, the entropy is 1.4905 bits. Coding it with a scheme built for the uniform distribution costs exactly 2.0000 bits per symbol, and the surcharge is 0.5095. The first number is fixed by the world; the second is the only part a model can improve, so training on cross-entropy is training on divergence from the truth.

It is not a distance, and the asymmetry is not a technicality. For a source that emits one symbol 98% of the time, the divergence from that source to the uniform distribution is 1.4235 bits, while the divergence in the other direction is 2.8540 - just over twice as large. KL(p||q) is dominated by outcomes that p produces and q calls unlikely, so it punishes a model for failing to cover what actually happens; the reverse direction punishes a model for spreading mass where nothing occurs. Which direction a method minimises decides whether it hedges or commits.

The penalty for a confident mistake grows without bound. Giving the truth a probability of 0.5 costs 1 bit, 0.1 costs 3.32, 0.01 costs 6.64 and 0.001 costs 9.97. This is exactly why cross-entropy is sensitive to calibration in a way accuracy is not, and why a handful of confidently wrong examples can dominate a gradient.

How to Calculate

KL(p||q) = Σ p(x) log2( p(x) / q(x) ) and H(p,q) = H(p) + KL(p||q)

where

p
the distribution that actually generates the data
q
the model: the distribution whose code you are using
H(p,q)
cross-entropy, the average cost of coding p with q
KL(p||q) ≥ 0
zero exactly when p and q agree everywhere

Example of Kullback-Leibler Divergence

Source (0.6, 0.25, 0.1, 0.05) against a uniform model: entropy 1.4905 bits, cross-entropy 2.0000 bits, divergence 0.5095 bits per symbol.

A source that is 98% one symbol: KL(source||uniform) = 1.4235 bits, KL(uniform||source) = 2.8540 bits. Same pair of distributions, twice the cost in one direction.

Per-example loss when the model gives the true label probability 0.5, 0.1, 0.01 and 0.001: 1.00, 3.32, 6.64 and 9.97 bits.

Frequently Asked Questions

Why not use a symmetric measure instead?

Symmetric alternatives exist, and they are useful when you genuinely want to compare two distributions as objects. But the asymmetric version is the one that answers "what does this model cost me", which is the question a loss function has to answer.

What happens if the model assigns zero to something that occurs?

The divergence is infinite, and this is not a numerical accident: the code has no word for that symbol. In practice it is why probabilities are smoothed or clipped, and why a zero in a model should always be a deliberate claim rather than an artefact of a small sample.

Is minimising cross-entropy the same as maximum likelihood?

Yes, up to a constant and a change of units. The average negative log-likelihood of the data under the model is the empirical cross-entropy, so the two procedures select the same model and differ only in how the number is reported.

The Bottom Line

The divergence is the price of being wrong about the distribution, measured in bits and payable per observation. It splits cross-entropy into a part you cannot change and a part that is entirely your model, which is what makes it the natural thing to minimise - as long as you remember which direction you are minimising.