GPQA and human baselines
448 expert-written questions, published alongside what experts and skilled searchers actually score. That second part is the useful part.
Part of the Trust and evals track on lAItest.
Give a skilled person these questions, thirty minutes, and unrestricted web access. They get about a third of them right.
That is the design working, not failing.
The test, and the two humans it was measured against
GPQA is 448 expert-written questions (arXiv:2311.12022), published with two human numbers beside them. PhD-level experts working inside their own domain score 65%, or 74% if you drop the questions they later flagged as their own slips. Skilled non-experts, given more than thirty minutes and unrestricted web access, score 34%. The distance between those two figures is the thing the test exists to measure: knowing where to look is not the same as knowing.
What a model score means against those anchors
GPT-4 scored 39% when the benchmark was published. On its own that number says nothing. Read against 34% for a searching non-expert and 65% for an in-domain PhD, it becomes a position on a scale with two ends you can picture. That is why the human baselines were published at all, and their absence is the single fastest way to spot a benchmark claim that is not telling you much.
When a vendor says GPQA, check which GPQA.
The figures above describe the full 448-question set. Vendors almost always report GPQA Diamond instead — a 198-question subset kept precisely because in-domain experts got them right and skilled non-experts did not. It is the harder half, so a Diamond score and a full-set score are not comparable, and the 65% and 34% baselines here belong to the full set. Two numbers under one name is a recurring habit in benchmark reporting, not an accident.
A common misconception
Commonly believed: A model beats PhDs on a science exam, so it is better at science than a scientist.
Actually: The expert figure is 65%, from in-domain PhDs answering a fixed written test under conditions. Beating it means answering those 448 questions better than those people did on that day. Science is choosing the question, building the apparatus, and noticing when the result is wrong. A fixed question set cannot measure that, and a score above a human baseline is a claim about the test, not about the job.
Why does GPQA publish what humans score, when most benchmark reporting only publishes model scores?
Answer: Because a percentage is unreadable without knowing what a person gets. A raw 39% could mean anything. Placed between 34% for a skilled searcher and 65% for an in-domain PhD it becomes informative immediately. Human baselines are what turn a benchmark number into a comparison, and a benchmark that reports none is asking you to supply the scale yourself.
In one sentence
A benchmark score is only readable next to what a person scores on the same questions.