Perplexity
One number for how surprised a model is by a piece of text. It travels much less well than people assume.
Part of the Turning the dials track on lAItest.
There is a single number for how surprised a model is by a piece of writing.
Lower means less surprised. That is where the simplicity ends.
Perplexity asks: on average, how many options was the model choosing between?
Take text the model did not write. At each position, look at the probability it would have assigned to the token that actually came next. Perplexity boils those probabilities down to one number. A perplexity of 1 means it called every token with certainty. A perplexity of 20 means it was, on average, about as unsure as someone picking between twenty equally likely options.
A common misconception
Commonly believed: Lower perplexity means a better model.
Actually: It means a better fit to that particular text, measured with that particular tokenizer. Change the tokenizer and the number moves while the model stands still. Compare two models on different corpora and the comparison means nothing. Perplexity is also silent on whether the text is true or useful: a model can be fluently, confidently wrong at very low perplexity.
Model A scores lower perplexity than model B on some text. What follows?
Answer: A found that text less surprising, on that corpus with that tokenizer. Perplexity is scoped to one corpus and one tokenizer. Inside a single training run it is a useful signal and a fair sanity check. Carried across models or datasets it stops meaning anything, which is why capability is argued over with task benchmarks instead.
In one sentence
Perplexity scores how well a model predicts a text. It says nothing about whether that text was worth predicting.