Memory and truth
How a model reaches things it was never trained on — embeddings, retrieval, context and memory.
13 concepts, about 78 minutes of reading at roughly six minutes each. 2 of them are free to read with no account; the rest need a paid plan.
What is in this track
- Why it cannot know your files — The weights were finished on a date, and your documents were never in them. (free)
- Meaning as coordinates — An embedding turns text into a list of numbers, positioned so similar meanings land near each other. (free)
- Cosine similarity — Comparing two embeddings by the angle between them, and why the number it gives back is not a percentage.
- Meaning search vs keyword search — One finds what you meant, the other finds what you typed, and each fails exactly where the other works.
- Vector databases — What stores millions of embeddings, and why "find the nearest" is deliberately an approximation.
- RAG, end to end — Retrieval-augmented generation: fetch the relevant text, paste it into the prompt, then answer from it.
- Chunking — Cutting documents into retrievable pieces — the dullest decision in retrieval and the one that decides most.
- Reranking — A second, slower model re-reads the shortlist alongside the query and puts it in the right order.
- Hybrid search — Running keyword and meaning search side by side, then fusing two rankings that share no common scale.
- Long context vs RAG — Windows got enormous. That moved the line where retrieval is worth it, and settled nothing.
- Needle in a haystack — The test that made long context look solved, and the result that showed it is not.
- Context engineering — Prompt engineering writes the instruction. Context engineering decides what else is in the window at all.
- Agent memory — There is no memory inside the model. Anything that persists is a file some program wrote and read back.
Before this: Getting good answers
After this: What's inside
Every track · Pricing · Claims we checked and could not stand behind