Trust and evals
How anyone knows whether any of this works: what an eval is, what a benchmark score hides, and the ways a system fails while its numbers improve.
13 concepts, about 78 minutes of reading at roughly six minutes each. Every concept in this track needs a paid plan.
What is in this track
- Why vibes stop working — You tried it five times and it looked fine. An eval is that same test, written down, scored the same way, and run again after every change.
- LLM-as-judge — Most useful outputs have no right answer to check against, so people hired a model to grade the model. It works, on conditions.
- Benchmarks and contamination — A public benchmark is an exam whose answer key is on the web, given to a system trained on the web. Scores rise for two different reasons.
- GPQA and human baselines — 448 expert-written questions, published alongside what experts and skilled searchers actually score. That second part is the useful part.
- Name the harness — A coding score measures a system, not a model. Without the scaffold, the version and the suite named, two numbers cannot be compared.
- Tests at the edge — ARC-AGI measures how efficiently a system picks up a rule it has never seen. HLE just makes the questions brutal. One of them prints the bill.
- Alignment — A model that will do anything you ask is a product decision, and so is one that refuses. Alignment is the name for that decision and the work behind it.
- Red teaming and jailbreaks — A jailbreak does not add a capability. It finds a route to one already in the weights, which is why the real defences are about access.
- Prompt injection — Your instructions and the text you asked the model to read arrive in the same stream, with nothing marking which is which. Everything follows from that.
- The OWASP LLM Top 10, and guardrails — Enough teams shipped the same ten mistakes that someone numbered them. Most are ordinary application security with a more persuasive input.
- Sycophancy and reward hacking — Anything optimised against a measurement will find the cheapest way to move it. Agreeing with you is one of the cheapest ways there is.
- Model cards and system cards — The document a maker publishes with a model. Read it for what it says, then read it again for the fields it declines to fill in.
- Open weights is not open source — There is a published three-part test for open source AI. Nearly every model marketed as open passes one part of it: the weights are downloadable.
Before this: Agents
After this: Fast, slow, and neither
Every track · Pricing · Claims we checked and could not stand behind