Chunking
Cutting documents into retrievable pieces — the dullest decision in retrieval and the one that decides most.
Part of the Memory and truth track on lAItest.
Cut a document every few hundred tokens and sooner or later you slice a sentence in half and file the two halves as unrelated ideas.
One embedding has to speak for one whole chunk.
Too big, and a single vector averages several topics, so it ends up close to nothing in particular. Too small, and each piece loses the context that gave it meaning: "it costs forty euros" is useless without the sentence naming the product. The usual starting point is a few hundred tokens, cut on structure such as headings and paragraphs rather than on raw length, with a little overlap so a sentence on a boundary appears in both pieces.
Try it
Change how many chunks come back for the same question. Watch a needed fact drop out of the answer. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: Chunking is a settled detail. Pick a size and move on.
Actually: It is one of the most actively attacked problems in retrieval, and every attack tries to give a chunk back the context it lost. Anthropic prepends a short generated description of where the chunk sits in its document before embedding it. Jina runs the whole document through the model first and pools per chunk afterwards. Voyage trains the contextualisation into the embedding model itself. Same complaint, three routes.
In one sentence
A chunk is the smallest thing your system can find. If an answer spans two chunks, your system cannot find it whole.