Needle in a haystack
The test that made long context look solved, and the result that showed it is not.
Part of the Memory and truth track on lAItest.
Hide one odd sentence in a very long document, ask the model to find it, and almost every model scores near perfect.
This is roughly where the industry stopped testing.
Needle in a haystack is an easy test in a hard costume.
The needle is normally a sentence with nothing else like it anywhere in the haystack. Spotting the one thing that does not belong is a far simpler job than using one of many things that do. A model can ace that test and still lose track of the fourth of nine near-identical clauses in a contract, which is the thing you actually needed.
A common misconception
Commonly believed: If a model advertises a million-token window, the last token gets the same attention as the first.
Actually: Chroma tested 18 models and found performance changing with input length even on tasks that are trivial when the input is short. It gets worse when the question is worded unlike the target text, when the haystack contains topically related distractors, and depending on how the haystack is structured. This is what people mean by context rot. The advertised window is what the model will accept, not what it will use evenly.
What these models will accept
| What | Value | Provenance |
|---|---|---|
| Gemini 3.1 Pro | 1M | source, verified . |
| Kimi K3 | 1M | source, verified . |
| DeepSeek V4-Pro | 1M | source, verified . Accepting this much is not a claim about using it evenly. |
In one sentence
A big window is a promise about what fits, not a promise about what gets understood.