Long-horizon autonomy

Long runs fail by drift and repetition rather than by one dramatic wrong decision.

Part of the Agents track on lAItest.

Agents look excellent for twenty minutes. The interesting failures start around hour four, when nobody is watching.

Window size stopped being the bottleneck

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Largest context window on the board1.1Msource, verified . Capacity, not coherence
Frontier models tracked18source, verified .

How long runs actually fail

Not with one dramatic wrong decision. Errors compound: a slightly wrong observation early becomes a confident premise later. The transcript fills with dead ends that are re-read on every step. A failing action gets retried because nothing in the system notices repetition. And a model that has been agreeable for two hundred steps stays agreeable about its own progress.

A common misconception

Commonly believed: Let the agent curate its own knowledge and memory, and it gets better the longer it runs.

Actually: That claim has industry traction and thin empirical backing. One study of agentic retrieval found an advantage on direct retrieval over long documents while overall accuracy stayed close to simply using the full context. Treat self-managed agent memory as an open question being argued about, not a settled feature.

Try it

Grow a transcript step by step and watch how much of the window becomes dead ends rather than facts the next step needs. This step is an interactive widget; open the lesson to use it.

In one sentence

Long-horizon autonomy fails by drift and repetition, so the fixes are checkpoints, fresh contexts and hard budgets rather than a bigger window.