Memory and truth

How a model reaches things it was never trained on — embeddings, retrieval, context and memory.

13 concepts, about 78 minutes of reading at roughly six minutes each. 2 of them are free to read with no account; the rest need a paid plan.

What is in this track

  1. Why it cannot know your files — The weights were finished on a date, and your documents were never in them. (free)
  2. Meaning as coordinates — An embedding turns text into a list of numbers, positioned so similar meanings land near each other. (free)
  3. Cosine similarity — Comparing two embeddings by the angle between them, and why the number it gives back is not a percentage.
  4. Meaning search vs keyword search — One finds what you meant, the other finds what you typed, and each fails exactly where the other works.
  5. Vector databases — What stores millions of embeddings, and why "find the nearest" is deliberately an approximation.
  6. RAG, end to end — Retrieval-augmented generation: fetch the relevant text, paste it into the prompt, then answer from it.
  7. Chunking — Cutting documents into retrievable pieces — the dullest decision in retrieval and the one that decides most.
  8. Reranking — A second, slower model re-reads the shortlist alongside the query and puts it in the right order.
  9. Hybrid search — Running keyword and meaning search side by side, then fusing two rankings that share no common scale.
  10. Long context vs RAG — Windows got enormous. That moved the line where retrieval is worth it, and settled nothing.
  11. Needle in a haystack — The test that made long context look solved, and the result that showed it is not.
  12. Context engineering — Prompt engineering writes the instruction. Context engineering decides what else is in the window at all.
  13. Agent memory — There is no memory inside the model. Anything that persists is a file some program wrote and read back.

Before this: Getting good answers

After this: What's inside

Every track · Pricing · Claims we checked and could not stand behind