The context window
There is a hard limit on how much text can sit in front of the model at once, and what happens when you exceed it depends on who is asking.
Part of the What is this thing? track on lAItest.
There is a maximum amount of text the model can have in front of it. Go past it in a chat app and the oldest part quietly stops existing.
Call the API directly and it refuses instead.
A window, measured in tokens
The context window is the total number of tokens allowed in one run: system prompt, the conversation so far, every attached document, and the answer being written, all sharing one budget. It is a size limit on the page, not a memory. When the page is full, something has to come off it.
Windows differ enormously, even inside one maker
| What | Value | Provenance |
|---|---|---|
| Claude Haiku 4.5 — tokens that fit | 200K | source, verified . |
| Claude Opus 5 — tokens that fit | 1M | source, verified . |
| DeepSeek V4-Pro — tokens that fit | 1M | source, verified . |
| Largest window on the current frontier | 1.1M | source, verified . Across every model this index treats as frontier today. |
Try it
Pour things in until it overflows. Then switch to a bigger window and pour the same things again. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: A bigger window means the model remembers better.
Actually: It means more text fits. Whether the model uses all of it is a separate question, and a measured one: testing across many models found performance changing with how much text is in the window even on trivial tasks, and getting worse when the window is packed with material that is related but not the answer. Room on the page is not attention on the page.
In one sentence
The context window is the size of the page, not the size of the memory — and when it overflows, it overflows quietly.