Speculative decoding

A small model guesses the next few tokens and the big model grades them in one pass. Accepted guesses are exactly what the big model would have written.

Part of the Cheap and fast track on lAItest.

A small model guesses the next few words. The big model only marks them right or wrong.

Checking is cheaper than writing.

Writing is serial, checking is not

Generation is strictly one token at a time: the fifth token cannot be written before the fourth exists. Checking has no such rule. If a small draft model proposes four tokens, the large model can score all four in a single pass, because the whole proposed run is already there to look at. Every guess that survives was produced without its own slow step.

The output does not get worse

Speculative decoding (arXiv:2211.17192) is arranged so that an accepted token is exactly the token the large model would have chosen on its own. When the draft is wrong, the guess is discarded and the large model supplies its own. You get the same answers, sooner. That is rare. Nearly every other speed trick trades quality away.

Where does speculative decoding show up on a hosted API bill?

Answer: Nowhere directly — you are billed per token, not per unit of work. Providers bill for tokens produced. Speculative decoding changes how many chips are busy and for how long, not how many tokens you asked for. You feel it as lower latency; the savings sit with whoever owns the hardware.

In one sentence

Draft, verify, discard. It is the only common way to make generation faster without making it worse.