Understanding Cross-Entropy
Cross-entropy asks how costly it is to describe outcomes drawn from a true distribution using a code built for a predicted one. If the prediction matches reality the cost is minimal and equals the true distribution’s own entropy. The worse the mismatch, the more wasted bits, so the quantity works naturally as a measure of how wrong a predicted distribution is.
For classification the true distribution is usually a one-hot vector: the correct class has probability 1 and all others 0. The sum then collapses to a single term, the negative logarithm of the probability the model assigned to the correct class. Predicting 0.9 for the right class costs −log(0.9) ≈ 0.105; predicting 0.1 costs −log(0.1) ≈ 2.303, more than twenty times as much.
That asymmetry is deliberate and important. As the predicted probability of the true class approaches zero, the loss grows without bound. A model that is confidently wrong is punished severely, which discourages the overconfidence that a symmetric loss such as squared error would tolerate.
The pairing with softmax or sigmoid outputs is what makes it practical. On its own the logarithm has an awkward derivative, but combined with these activations the gradient of the loss with respect to the pre-activation output simplifies to the difference between the predicted probability and the true label. The gradient is proportional to the error, does not saturate, and is numerically stable, which is precisely why this pairing is the default.
How to Calculate
H(p, q) = − Σᵢ p(xᵢ) log q(xᵢ); for one-hot labels this reduces to −log q(correct class)
where
- p
- the true distribution, typically one-hot over classes
- q
- the distribution predicted by the model
- H(p, q)
- the cross-entropy, the average cost of coding p using q
Example of Cross-Entropy
A three-class problem with true class B gives the one-hot target (0, 1, 0). If the model predicts (0.2, 0.7, 0.1), the loss is −log(0.7) ≈ 0.357.
If instead it predicts (0.2, 0.1, 0.7), placing most of its confidence on the wrong class, the loss is −log(0.1) ≈ 2.303. The probability assigned to the correct class fell by a factor of seven, and the loss rose by a factor of about six and a half.
A near-certain correct prediction of 0.99 costs only −log(0.99) ≈ 0.01, while a near-certain wrong one of 0.01 costs ≈ 4.61. The steepness at the wrong end is what drives the model away from confident errors during training.
Frequently Asked Questions
Why not use squared error for classification?
Squared error paired with a sigmoid output produces a non-convex objective with gradients that vanish exactly where the model is most wrong, so learning stalls on the examples that matter most. Cross-entropy keeps the gradient proportional to the error and remains convex for logistic regression.
How does cross-entropy relate to maximum likelihood?
They are the same optimization. The likelihood of the observed labels is a product of predicted probabilities; taking its negative logarithm turns that product into the sum that defines cross-entropy. Minimizing one maximizes the other.
What is the difference between cross-entropy and KL divergence?
Cross-entropy equals the entropy of the true distribution plus the KL divergence from the true distribution to the predicted one. Since the true distribution is fixed during training, its entropy is a constant, so minimizing cross-entropy and minimizing KL divergence are equivalent.
The Bottom Line
Cross-entropy scores a predicted distribution by the log-probability it assigned to what actually happened, punishing confident mistakes without bound. Combined with softmax it yields the clean, non-saturating gradient that makes it the default classification loss.