The reasoning you can read
A model that shows its working is not necessarily showing you its working. Several experiments separate the two, and the gap is large.
Part of the Fast, slow, and neither track on lAItest.
It explained its answer. Beautifully. And the explanation was not why.
Change something the model will not mention.
Turpin and colleagues tried two kinds of interference. They reordered the options so the correct answer was always in the same position, a bias with nothing to do with the content, and accuracy fell by as much as 19 percent. They also simply told the model what they thought the answer was, and it fell by as much as 36. In both cases the models went on producing fluent, plausible step-by-step reasoning that never once mentioned the thing actually driving the answer.
A common misconception
Commonly believed: A longer, more detailed chain of thought means a more faithful account of the reasoning.
Actually: Lanham and colleagues found faithfulness varies enormously by task, and that it peaks in the middle of the size range: above roughly thirteen billion parameters, larger models produced LESS faithful reasoning on most tasks tested. Fluency of explanation and accuracy of explanation are separate properties, and the first is the one that improves fastest.
Now replace the reasoning with dots.
Pfau and colleagues gave transformers strings of meaningless filler tokens, literally "......", in place of a chain of thought. On the right kind of problem, one whose work can be done in parallel, the models then solved tasks they could not solve by answering directly. The dots bought computation. Two things keep this from being a debunking: the models had to be specifically trained to use the dots, and the same paper reports that off-the-shelf models get nothing from them.
What does the filler-token result suggest about chain of thought?
Answer: The extra tokens can buy compute even when they carry no meaning. Chain of thought does work — that is well established. What the dots show is that part of the benefit is having more forward passes to compute in, not the reasoning being written down. Both things are true at once, which is why this stays an open question rather than a debunking.
A common misconception
Commonly believed: We know reasoning models collapse completely past a certain difficulty.
Actually: This one is genuinely unsettled, and the argument is more interesting than either side alone. Shojaee and colleagues reported exactly that collapse. A critic showed that part of it was measurement: some puzzles in the test set are provably unsolvable by any solver, and models were scored as failing them. The authors accepted that one and narrowed their claim, while holding that the collapse still arrives earlier, on problems that are solvable. An independent replication since finds both things true. Some of the collapse was an artefact of the test. Some of it is real.
In one sentence
Use the System 2 metaphor for what it buys you. Do not treat the visible chain as an audit trail — it is a story the model also generated.