Reranking
A second, slower model re-reads the shortlist alongside the query and puts it in the right order.
Part of the Memory and truth track on lAItest.
Your retriever returns 150 candidates. Your prompt has room for 20. Choosing the right 20 is a different job from finding the 150.
The first pass never compares the query and the document directly.
For search to be fast, documents must be turned into vectors in advance, long before your question existed, so the document embedding could not take your question into account. A reranker drops that constraint. It reads the query and one candidate together in a single pass and scores the pair. Much more accurate, and far too slow to run across a million documents — so you run it across the survivors.
Retrieve wide, then rerank narrow.
Cast a cheap wide net, then spend the expensive model only on what came back. Anthropic's published pipeline retrieves 150 candidates and reranks down to the top 20. The wide net protects recall; the reranker protects precision. Neither does the other job well.
A common misconception
Commonly believed: With good enough embeddings, reranking is optional polish.
Actually: Anthropic measured the whole stack on their own retrieval benchmark. The right chunk was missing from the top results 5.7% of the time as a baseline, 3.7% with contextual embeddings, 2.9% after adding contextual keyword search, and 1.9% once a reranker was added. That last step removed about a third of what was still going wrong, on candidates the retriever had already selected.
In one sentence
Search finds candidates. Reranking decides which of them actually answers the question.