Skip to content
Kudos AI

Entropy

A measure of the uncertainty in a random variable, equal to the average number of bits needed to encode its outcome.

Also known as: Shannon entropy, Information entropy

A distribution reshaped from uniform to near-certain, with the entropy in bits tracking it down to almost nothing.

Understanding Entropy

Entropy quantifies how much you do not know about the outcome of a random variable. If a variable always takes the same value, observing it tells you nothing and its entropy is zero. The more evenly probability is spread across possible outcomes, the more uncertain the result and the higher the entropy.

The units are made concrete by the base-2 logarithm: entropy counts bits. Russell and Norvig make the calibration explicit. A fair coin is equally likely to land either way, and that counts as exactly 1 bit. A fair four-sided die has 2 bits, because two bits are needed to describe one of four equally probable choices. An unfair coin that lands heads 99% of the time carries much less uncertainty, and its entropy should be close to zero while remaining positive.

The definition weights each outcome’s surprise, −log₂ p, by how often that outcome actually occurs. Rare events are individually very surprising but contribute little because they seldom happen; common events are unsurprising but frequent. Entropy is the average of this surprise, which is why it peaks at the uniform distribution, where no outcome can be anticipated.

This is not merely a metaphor about information. Shannon’s source coding theorem establishes entropy as a hard limit: no lossless encoding of a source can use fewer bits per symbol on average than the source’s entropy. The same quantity reappears throughout machine learning, in the information gain used to split decision trees, in cross-entropy loss, and in the KL divergence between distributions.

How to Calculate

H(X) = − Σᵢ p(xᵢ) log₂ p(xᵢ)

where

H(X)
the entropy of the random variable X, in bits
p(xᵢ)
the probability of outcome xᵢ
log₂
base-2 logarithm, which makes the unit the bit
−
makes the result positive, since log of a probability is negative

Example of Entropy

A fair coin has p = 0.5 for each face, giving H = −(0.5 log₂ 0.5 + 0.5 log₂ 0.5) = 1 bit exactly. A fair four-sided die has four outcomes at p = 0.25, giving H = 2 bits, matching the intuition that two binary digits identify one of four options.

Now take the biased coin that lands heads 99% of the time. Computing −(0.99 log₂ 0.99 + 0.01 log₂ 0.01) gives approximately 0.0808 bits: close to zero, as expected, but strictly positive, because the rare tail still carries genuine surprise when it occurs.

The compression reading is direct. A long sequence of flips from the biased coin can be encoded in about 0.081 bits per flip on average, more than a tenfold saving over the one bit per flip a naive encoding would spend, because the sequence is overwhelmingly heads and that regularity can be exploited.

Frequently Asked Questions

Why is there a minus sign in the formula?

Probabilities lie between 0 and 1, so their logarithms are negative or zero. The minus sign flips the sum so entropy is reported as a non-negative quantity.

What is the difference between entropy and cross-entropy?

Entropy measures the uncertainty of a single distribution. Cross-entropy measures the average cost of encoding outcomes drawn from one distribution using a code optimized for a different one, which is why it works as a loss function comparing predictions against truth.

Why do decision trees use entropy?

A good split makes the resulting groups purer, that is less uncertain about the class. Information gain measures exactly that: the entropy before the split minus the weighted average entropy after it. The split that removes the most uncertainty is chosen.

The Bottom Line

Entropy measures uncertainty in bits, is zero for a certain outcome and maximal for a uniform one, and sets the hard floor on lossless compression. It reappears throughout machine learning wherever purity, surprise, or the distance between distributions has to be quantified.