Sampling and decoding

Same prompt, same model, different answer. The model did not change its mind; the draw did.

Part of the Turning the dials track on lAItest.

Ask the same model the same question twice and you can get two different answers.

Nothing about the model changed between the two.

Something has to pick one token from the odds. How it picks is the decoding strategy.

There are two basic moves. Greedy decoding takes the highest-probability token every time and never varies. Sampling draws one token at random, weighted by the odds, so a token sitting at 60% comes up roughly six times in ten. Everything else in this track is either a variation on those two or a way of reshaping the odds before the draw.

A common misconception

Commonly believed: Greedy decoding gives the best answer, because it takes the best token every time.

Actually: It takes the best next token, which is not the same as the best sentence. A strong opening word can lead into a corner where every continuation is poor, and greedy decoding cannot look ahead or back out. In practice it tends to produce flat text that repeats itself and sometimes loops outright.

You are sampling, and one token sits at 20%. What happens to it?

Answer: It is chosen roughly one run in five. Sampling draws in proportion to the odds, so 20% means about one time in five. The same arithmetic explains rare bad outputs: a wrong token at 2% is not impossible, it is two times in a hundred, which you will meet on day one of production traffic.

In one sentence

The model hands you odds. The decoding strategy is what turns odds into words, and that choice is yours.