Paying for the pause
Letting a model produce more tokens before it answers really does make it more accurate — and it is the most expensive sentence in this course.
Part of the Fast, slow, and neither track on lAItest.
The model gets better if you let it ramble first.
This is not folklore. It is measured, repeatedly.
Chain of thought
In 2022 Wei and colleagues showed that putting worked examples in the prompt, each one reasoning its way to the answer rather than just stating it, lifted accuracy on problems the model otherwise failed. It only worked once the model was large enough. Below roughly a hundred billion parameters it made things worse, because smaller models produced fluent reasoning that did not follow.
Then: buy the accuracy with compute
The idea generalised into spending more compute at answer time rather than at training time. Snell and colleagues found that on an equal compute budget, spending it at answer time can beat a model roughly fourteen times larger. That comes with two conditions worth remembering, because the second one is about you: the smaller model must not be hopeless at the task, and you must not be serving much traffic. At high request volume the same paper finds the trade reverses and the parameters win. DeepSeek-R1 then shipped the technique with open weights.
A common misconception
Commonly believed: Thinking tokens are a free upgrade — you get the better answer at the usual price.
Actually: Thinking tokens are billed as output, and output is the expensive side of the bill. A model reasoning by default can spend a substantial budget before it emits a single visible character. This is why setting a low output ceiling on a reasoning model produces an empty response rather than a short one: the ceiling is consumed by thinking the reader never sees.
Output is the side that costs
| What | Value | Provenance |
|---|---|---|
| Claude Fable 5.1 input | $10 per 1M | source, verified . |
| Claude Fable 5.1 output | $50 per 1M | source, verified . Thinking tokens bill here, not on the input line — and you do not see them. |
| GPT-6 Astra output | $50 per 1M | source, verified . |
| Gemini 3.8 Flash output | $3.75 per 1M | source, verified . The cheap tier exists because not every task needs the pause. |
Try it
Raise the output tokens and watch which line of the bill moves. This step is an interactive widget; open the lesson to use it.
In one sentence
Inference-time compute buys accuracy, and you pay for it on the output line whether or not anyone reads the output.