Sub-agents and the clean context

A fresh, short context beats a long one. Sub-agents exist to protect the parent from the mess.

Part of the Agents track on lAItest.

An agent forty steps in is mostly reading thirty-nine steps of its own mess.

Bigger windows did not fix this.

A common misconception

Commonly believed: A million-token window means you can keep appending everything and stop thinking about it.

Actually: Chroma tested this across eighteen models: performance changes with input length even on trivial tasks. It degrades further when the answer is worded differently from the question, when topically related but wrong material sits in the window, and depending on how the surrounding text is structured. Length is capacity, not attention.

Context engineering, as distinct from prompt engineering

Prompt engineering writes the instructions. Context engineering curates the whole set of tokens a running agent sees: compacting old turns into summaries, keeping structured notes outside the transcript, fetching detail just in time by identifier instead of pasting it in up front, and handing narrow jobs to sub-agents that start clean.

Why a sub-agent helps

A sub-agent is a fresh loop with an empty transcript, one job, and a small tool set. It does the searching or the reading and returns a short result. The parent never sees the thirty dead ends it walked through. The gain is not parallelism. The gain is that the parent context stays short and stays relevant.

Try it

Fill the window with a long agent transcript, then ask for a detail buried in the middle. Find where it starts to slip. This step is an interactive widget; open the lesson to use it.

In one sentence

The scarce resource in an agent is not context length, it is the share of the context that still matters.