Test-time compute

Spending more computation when the question arrives, rather than only back when the model was trained.

Part of the Getting good answers track on lAItest.

For years a model got better by being trained bigger. Now it also gets better by being given longer to answer.

Same weights. More compute.

Two places to spend compute.

Training compute is spent once, before anyone asks anything. Test-time compute is spent per question, while the user waits. The second kind is what turned into a product control, and it changes the shape of the bill: on the same model, a hard question can now cost far more than an easy one, because you chose to let it work longer.

Three families, and they are not interchangeable.

Search: build a tree of partial solutions and explore the promising branches. Sequential: produce an answer, criticise it, revise, repeat, each pass conditioned on the last. Parallel: generate several independent attempts and pick one using a verifier or a vote. Self-consistency is the plainest version of parallel.

A common misconception

Commonly believed: Thinking for longer always produces a better answer.

Actually: It helps where there is something to check against: arithmetic, code that runs, a puzzle with a verifiable answer, a plan with hard constraints. On a matter of taste, or a question whose answer is simply not available, extra passes mostly yield a longer and more confident version of the same answer. Compute cannot manufacture information the model does not have.

Which of these is a parallel test-time-compute method?

Answer: Generating five independent answers and keeping the one a verifier accepts. Parallel means independent attempts, scored afterwards. Revision is sequential, because each pass reads the previous one. Tree building is search. Training is the other budget entirely, spent once, long before the question existed.

In one sentence

Test-time compute is a dial between speed and quality that lives in the request rather than in the model, which is why analysts expect inference, not training, to dominate AI compute.