Long-horizon autonomy
Long runs fail by drift and repetition rather than by one dramatic wrong decision.
Part of the Agents track on lAItest.
Agents look excellent for twenty minutes. The interesting failures start around hour four, when nobody is watching.
Window size stopped being the bottleneck
| What | Value | Provenance |
|---|---|---|
| Largest context window on the board | 1.1M | source, verified . Capacity, not coherence |
| Frontier models tracked | 18 | source, verified . |
How long runs actually fail
Not with one dramatic wrong decision. Errors compound: a slightly wrong observation early becomes a confident premise later. The transcript fills with dead ends that are re-read on every step. A failing action gets retried because nothing in the system notices repetition. And a model that has been agreeable for two hundred steps stays agreeable about its own progress.
A common misconception
Commonly believed: Let the agent curate its own knowledge and memory, and it gets better the longer it runs.
Actually: That claim has industry traction and thin empirical backing. One study of agentic retrieval found an advantage on direct retrieval over long documents while overall accuracy stayed close to simply using the full context. Treat self-managed agent memory as an open question being argued about, not a settled feature.
Try it
Grow a transcript step by step and watch how much of the window becomes dead ends rather than facts the next step needs. This step is an interactive widget; open the lesson to use it.
In one sentence
Long-horizon autonomy fails by drift and repetition, so the fixes are checkpoints, fresh contexts and hard budgets rather than a bigger window.