Trust and evals

How anyone knows whether any of this works: what an eval is, what a benchmark score hides, and the ways a system fails while its numbers improve.

13 concepts, about 78 minutes of reading at roughly six minutes each. Every concept in this track needs a paid plan.

What is in this track

  1. Why vibes stop working — You tried it five times and it looked fine. An eval is that same test, written down, scored the same way, and run again after every change.
  2. LLM-as-judge — Most useful outputs have no right answer to check against, so people hired a model to grade the model. It works, on conditions.
  3. Benchmarks and contamination — A public benchmark is an exam whose answer key is on the web, given to a system trained on the web. Scores rise for two different reasons.
  4. GPQA and human baselines — 448 expert-written questions, published alongside what experts and skilled searchers actually score. That second part is the useful part.
  5. Name the harness — A coding score measures a system, not a model. Without the scaffold, the version and the suite named, two numbers cannot be compared.
  6. Tests at the edge — ARC-AGI measures how efficiently a system picks up a rule it has never seen. HLE just makes the questions brutal. One of them prints the bill.
  7. Alignment — A model that will do anything you ask is a product decision, and so is one that refuses. Alignment is the name for that decision and the work behind it.
  8. Red teaming and jailbreaks — A jailbreak does not add a capability. It finds a route to one already in the weights, which is why the real defences are about access.
  9. Prompt injection — Your instructions and the text you asked the model to read arrive in the same stream, with nothing marking which is which. Everything follows from that.
  10. The OWASP LLM Top 10, and guardrails — Enough teams shipped the same ten mistakes that someone numbered them. Most are ordinary application security with a more persuasive input.
  11. Sycophancy and reward hacking — Anything optimised against a measurement will find the cheapest way to move it. Agreeing with you is one of the cheapest ways there is.
  12. Model cards and system cards — The document a maker publishes with a model. Read it for what it says, then read it again for the fields it declines to fill in.
  13. Open weights is not open source — There is a published three-part test for open source AI. Nearly every model marketed as open passes one part of it: the weights are downloadable.

Before this: Agents

After this: Fast, slow, and neither

Every track · Pricing · Claims we checked and could not stand behind