Long context vs RAG
Windows got enormous. That moved the line where retrieval is worth it, and settled nothing.
Part of the Memory and truth track on lAItest.
If a model will read a million tokens in one request, why not skip retrieval and paste in the entire handbook?
How much text a model will take in one request
| What | Value | Provenance |
|---|---|---|
| Largest frontier context window | 1.1M | source, verified . |
| Claude Opus 5 | 1M | source, verified . |
| Grok 4.5 | 500K | source, verified . Smaller than the Grok model it replaced. Windows do not only grow. |
| Claude Haiku 4.5 | 200K | source, verified . Huge windows are a frontier feature, not a universal one. |
A common misconception
Commonly believed: Long context killed RAG.
Actually: This is a live argument, not a finished result. Three reasons retrieval keeps earning its place: the weights have a cutoff and your data does not; attention cost climbs sharply with length, so long prompts are slower and dearer; and a bigger window drags in more irrelevant text alongside the relevant text. Long context wins on small, tidy, stable material. Retrieval wins as the pile gets bigger, fresher and messier.
Most serious systems do both, behind a router.
The decision is made per question, not once for the product. A document or two of stable material: load it all, and let prompt caching make re-sending it cheap. A large, private or fast-moving corpus: retrieve. Where exactly the line falls depends on the model, and it moves with every release — which is the honest reason the argument has not been settled.
Try it
Pour a repository into a window, then switch to a smaller model. A chat product drops the overflow silently; an API call over the limit is rejected outright. This step is an interactive widget; open the lesson to use it.
In one sentence
Long context and retrieval are not rivals. They are two ways to get text in front of a model, and grown-up systems keep both.