Skip to content
Kudos AI

Perplexity

The exponential of a model’s average cross entropy, read as the number of equally likely options it is effectively choosing between at each step.

Also known as: Effective branching factor

Understanding Perplexity

Cross entropy charges a model the negative logarithm of the probability it assigned to what actually happened, averaged over the sequence. The number is in nats or bits and is hard to feel. Exponentiating it produces a count: a model with the same cross entropy as a uniform choice among 50 options has a perplexity of 50, whatever the real vocabulary is. Lower is better, and 1 is perfect prediction.

The uniform case gives the scale a fixed point. A model with no preference assigns 1/V to every option, pays log V, and has a perplexity of exactly V. That makes an untrained model’s expected loss computable in advance, and it makes the number a free diagnostic: a fresh model reporting far below log V has seen something it should not have, and one reporting far above has a bug in its loss or its tokenisation.

Comparisons need care. Perplexity depends on the tokenisation as much as on the model, because a vocabulary that splits words into more pieces spreads the same text over more, easier predictions. Two models are comparable on perplexity only if they are scored on the same text with the same tokeniser, which is why published numbers travel with a dataset name.

What it does not measure is worth stating plainly. Perplexity rewards assigning probability to the observed continuation. A model can be fluent, confident and wrong, and nothing about a low perplexity distinguishes those cases from a model that is right.

How to Calculate

\mathrm{perplexity} = e^{L}, \qquad L = -\frac{1}{N}\sum_{t=1}^{N} \log p_\theta(x_t \mid x_{<t})

where

L
average cross entropy, in nats
p_\theta(x_t \mid x_{<t})
the probability the model gave the token that actually came next
N
the number of tokens scored

Example of Perplexity

For a vocabulary of 50,257 tokens the uniform reference is log 50,257 = 10.8249 nats. An untrained model measured at 10.7940 over six tokens sits 0.03 below it, a perplexity of 48,725.8, or 96.95 per cent of the vocabulary still in play. Over the full loaders the same untrained model reports 10.9876, a little above the reference, because random logits are not exactly uniform - log V is a reference point, not a bound.

That share is the reading worth keeping. It says the model is choosing almost uniformly, which is what an untrained model should do, and it says so in a form that transfers: half the vocabulary in play reads the same at a thousand tokens as at fifty thousand, where the two raw losses differ by nearly four nats and neither means anything alone.

At the other end, a perplexity of 1 corresponds to a loss of 0 and means the model gave probability 1 to every token that occurred. On held-out text that is a sign of a leak rather than of a good model.

Frequently Asked Questions

Is a perplexity of 20 good?

Unanswerable without the vocabulary and the test set. Twenty out of a vocabulary of 50,257 is a strong model; twenty out of a vocabulary of 32 is barely better than guessing.

Can perplexity be below one?

No. It is the exponential of an average of non-negative terms, so it is at least 1, and equals 1 only if every observed token was predicted with probability 1.

The Bottom Line

Perplexity is cross entropy on a countable scale: how many options the model is effectively weighing. Read it against log V rather than on its own, and remember it scores probability assigned to the observed text, not correctness.