Why vibes stop working
You tried it five times and it looked fine. An eval is that same test, written down, scored the same way, and run again after every change.
Part of the Trust and evals track on lAItest.
You tried the prompt five times and it looked good. Nobody knows what it does on the ten thousandth run, including you.
That gap is what an eval is for.
A common misconception
Commonly believed: I tested it. I ran my prompt a few times, read the answers, and they were fine. If something breaks I will notice.
Actually: You sampled five outputs from a system that does not repeat itself. Even at temperature zero the same prompt is not guaranteed to give the same answer — Anthropic says so in its own documentation. Five good samples tell you the failure rate is probably not enormous. They cannot tell you whether it is one in ten or one in a thousand, and they say nothing at all about the version you ship next week.
An eval is three boring things
A fixed set of inputs. A rule for scoring an output, applied the same way every time. A number you write down and compare against the last run. That is the whole idea. The work is not clever: it is choosing inputs that look like real traffic, including the ugly ones, and writing a scoring rule you would still accept on the day it tells you your new prompt is worse. The published advice for picking a retrieval model puts a size on it — use leaderboards to make a shortlist, then run about a hundred labelled examples of your own.
Try it
Send the same prompt twice at the same settings, then raise the dial. Ask how many samples you would need before you would bet on the answer. This step is an interactive widget; open the lesson to use it.
A new prompt scores better on your eval set. What does that actually tell you?
Answer: The new prompt is better on those inputs, scored that way. An eval measures what is in it. A gain on your set is real evidence about your set, which is only useful if your set resembles your traffic. It says nothing about inputs you left out, and nothing about the model, which did not change. It is also why a set you never revisit slowly stops describing the thing you are running.
In one sentence
An eval is a fixed set of inputs and a fixed way of scoring them. Without one you are not measuring your change, you are remembering the last five answers you happened to read.