Top-p and top-k

Temperature reshapes the whole list. These two just delete the bottom of it.

Part of the Turning the dials track on lAItest.

Temperature reshapes every candidate. Often you would rather just throw the bad ones away.

Top-k keeps a fixed number of candidates. Top-p keeps a fixed amount of probability.

Top-k says: discard everything except the k highest-scoring tokens, then sample from what survives. Top-p, also called nucleus sampling, says: take tokens from the top until their probabilities add up to p, then discard the rest. The difference is adaptivity. Top-k keeps k candidates whether the model is certain or wide open. Top-p keeps one candidate when the model is sure and forty when it is not.

Try it

Pull top-p down and watch candidates get cut. Compare a step where the model is certain with one where it is genuinely undecided. This step is an interactive widget; open the lesson to use it.

The model is almost certain of the next token. What does top-p at 0.9 do?

Answer: Keeps almost nothing beyond the top token. Top-p accumulates probability until it reaches 0.9. If the top token alone holds 0.95, the cut lands after a single candidate. That is the point of it: one setting is strict when the model is confident and generous when it is not.

A common misconception

Commonly believed: Set temperature, top-p and top-k together for really fine-grained control.

Actually: They stack, and stacked cuts are hard to reason about. Common practice is to move one and leave the others at their defaults. Raising temperature while top-p is tight does almost nothing, because the tokens the heat was meant to reach had already been deleted from the list.

In one sentence

Temperature changes the odds. Top-p and top-k change who is still on the list when the draw happens.