Paying for the pause

Letting a model produce more tokens before it answers really does make it more accurate — and it is the most expensive sentence in this course.

Part of the Fast, slow, and neither track on lAItest.

The model gets better if you let it ramble first.

This is not folklore. It is measured, repeatedly.

Chain of thought

In 2022 Wei and colleagues showed that putting worked examples in the prompt, each one reasoning its way to the answer rather than just stating it, lifted accuracy on problems the model otherwise failed. It only worked once the model was large enough. Below roughly a hundred billion parameters it made things worse, because smaller models produced fluent reasoning that did not follow.

Then: buy the accuracy with compute

The idea generalised into spending more compute at answer time rather than at training time. Snell and colleagues found that on an equal compute budget, spending it at answer time can beat a model roughly fourteen times larger. That comes with two conditions worth remembering, because the second one is about you: the smaller model must not be hopeless at the task, and you must not be serving much traffic. At high request volume the same paper finds the trade reverses and the parameters win. DeepSeek-R1 then shipped the technique with open weights.

A common misconception

Commonly believed: Thinking tokens are a free upgrade — you get the better answer at the usual price.

Actually: Thinking tokens are billed as output, and output is the expensive side of the bill. A model reasoning by default can spend a substantial budget before it emits a single visible character. This is why setting a low output ceiling on a reasoning model produces an empty response rather than a short one: the ceiling is consumed by thinking the reader never sees.

Output is the side that costs

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Claude Fable 5.1 input$10 per 1Msource, verified .
Claude Fable 5.1 output$50 per 1Msource, verified . Thinking tokens bill here, not on the input line — and you do not see them.
GPT-6 Astra output$50 per 1Msource, verified .
Gemini 3.8 Flash output$3.75 per 1Msource, verified . The cheap tier exists because not every task needs the pause.

Try it

Raise the output tokens and watch which line of the bill moves. This step is an interactive widget; open the lesson to use it.

In one sentence

Inference-time compute buys accuracy, and you pay for it on the output line whether or not anyone reads the output.