RAG, end to end
Retrieval-augmented generation: fetch the relevant text, paste it into the prompt, then answer from it.
Part of the Memory and truth track on lAItest.
RAG is not a clever mechanism inside the model. It is copy and paste, performed by a program, moments before the model reads your question.
Two phases: one you run in advance, one you run per question.
In advance: cut the documents into chunks, embed each chunk, store the vectors. Per question: embed the question, find the nearest chunks, paste them above the question with an instruction to answer from them, and send the lot. The model then does the only thing it ever does, which is continue text — except now the relevant paragraph is sitting right there in front of it.
Try it
Run a question through the pipeline, then compare the answer with and without the corpus attached. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: RAG teaches the model my documents.
Actually: It teaches it nothing. No weights change, and the documents are visible for exactly one request before they are gone again. That is also the good news: remove a document from the store and the next answer cannot cite it, with no retraining and no waiting. Access control stays where it belongs, in the retrieval step.
Which is why most RAG failures are retrieval failures.
If the right chunk never came back, no rewording of the prompt saves the answer. You are asking about text the model cannot see, and a fluent model will fill the gap rather than stop. When a retrieval system is wrong, look at what it retrieved before you touch the wording.
In one sentence
RAG does not make a model smarter. It makes sure the right paragraph is on the page.