Thinking tokens you pay for and never see
Hidden reasoning is billed at output rates, so a short answer can carry a bill many times larger than it looks.
Part of the Getting good answers track on lAItest.
The answer was 500 tokens. The bill was for 4,500. Nothing had gone wrong.
You paid for the thinking.
Reasoning happens in tokens, and tokens cost money.
When a model thinks before answering, it generates thinking tokens first and the visible answer afterwards. On the top models the raw thinking is not returned to you at all: it is generated, used, billed and dropped. Those tokens are charged at the output rate, the expensive one, because producing them costs exactly as much as producing an answer.
Posted rates, both sides of the bill
| What | Value | Provenance |
|---|---|---|
| Claude Opus 5 — input | $5 per 1M | source, verified . |
| Claude Opus 5 — output | $25 per 1M | source, verified . thinking tokens are billed at this rate |
| Claude Sonnet 5 — output | $10 per 1M | source, verified . |
| DeepSeek V4-Pro — output | $3.96 per 1M | source, verified . |
Try it
Set the visible answer to 500 tokens. Now add 4,000 thinking tokens you never get to read. Watch which number moves. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: Hidden reasoning is free. You only pay for what you can read.
Actually: You pay for every token generated, read or not. Four thousand tokens of thinking in front of a five-hundred-token answer is nine times the billed output, for a reply that looks short. This is why lowering effort saves real money, and why a brief answer is no evidence of a small bill.
A request returns a 200-token answer after 3,000 tokens of hidden thinking. How is it billed?
Answer: All 3,200 at output rates. Thinking tokens are generated, which makes them output tokens, which means they are priced on the expensive side of the sheet. Nothing about being hidden makes them cheaper. Input rates apply to what you send, and you did not send the thinking.
In one sentence
The part of the answer you cannot see is usually the part that costs the most.