Chain of thought

Making a model write out its steps changes the answer, because the steps are where the computation actually happens.

Part of the Getting good answers track on lAItest.

A model that answers a hard question instantly is often wrong. The same model, asked to work through it, is often right.

The written words are the workspace.

A model does a fixed amount of work per token.

Producing one token is one pass through the network, and that pass is the same size whether the question is "what is 2 plus 2" or a five-step logic puzzle. If the answer needs more work than one pass provides, the only way to buy more is to produce more tokens. Writing out the steps is not a performance. It is the extra passes.

A common misconception

Commonly believed: The chain of thought shows you how the model reached its answer.

Actually: It shows a plausible account of how it could have. The written steps genuinely feed into every token after them, so they do shape the answer, but a model can also write correct-looking steps and then state a conclusion that does not follow from them. Treat the steps as work in progress, not as an audit trail.

Two ways to get the steps.

The old way is a prompt: ask for the reasoning before the answer, or show an example of step-by-step working. The newer way is a model setting, a dial for how hard to think before answering. Both produce the same thing: tokens generated before the answer, and billed like any other token you generate.

Why does asking for step-by-step working help on a hard arithmetic problem?

Answer: It buys more forward passes, and the written steps become input for the tokens that follow. There is no careful mode. Each token gets one pass through the network, so the only way to spend more computation on a problem is to emit more tokens. Written intermediate results also become part of the input for everything after them, which is how the model can use step two while producing step three.

In one sentence

For a model, thinking is literally writing. Fewer tokens means less computation, whatever the question was.