Benchmarks and contamination
A public benchmark is an exam whose answer key is on the web, given to a system trained on the web. Scores rise for two different reasons.
Part of the Trust and evals track on lAItest.
An exam whose answer key is published on the web, sat by a system trained on the web.
Scores go up. Two very different things could be causing it.
What a benchmark is, and why it was a good idea
A shared test: a fixed set of questions, one scoring rule, published so different systems can be compared on the same thing rather than on anecdotes. MMLU is the workhorse example, a broad multiple-choice exam across many subjects (arXiv:2009.03300). It gave the field a common yardstick. A 70-billion-parameter model called Chinchilla scored 67.5% on it while beating models several times its size, which is how the field learned that how much data you train on matters as much as how large the model is (arXiv:2203.15556).
Contamination is the default, not the scandal
A benchmark is published so everyone can use it. Publishing means it is on the web. Training data is scraped from the web. So the questions, and often the answers, end up inside the model that is later tested on them. A model that has seen the answer key scores well without being able to do the task. Nobody has to cheat for this to happen. It is what you should expect from publishing a test and then training on everything.
A common misconception
Commonly believed: Model A scores higher than model B on a benchmark, so model A is the better model.
Actually: It is the better model at that benchmark. The gap can come from capability, from having seen the test, from a different scoring setup, or from a maker choosing the board that flatters it. Scores are evidence, not verdicts. A fast habit that pays: notice which tests a maker chose to report, and which obvious comparable ones it did not.
A two-year-old benchmark shows every new model scoring above 90%. What is the most likely reading?
Answer: The test is saturated and probably contaminated, so it no longer separates models. Two things happen to an old public benchmark. It leaks into training data, and makers tune against it because it is the number that gets reported. Both push scores up faster than the underlying skill moves. Once everything lands in a narrow band near the ceiling the test has stopped doing its job, which is exactly why harder tests keep being built.
In one sentence
A benchmark stops measuring a model the moment the model has read it, which is why the interesting tests are always the new ones.