Name the harness

A coding score measures a system, not a model. Without the scaffold, the version and the suite named, two numbers cannot be compared.

Part of the Trust and evals track on lAItest.

Two labs report a score on SWE-bench. The numbers are far apart. Both are honest.

They did not run the same test.

The harness is part of what got measured

A coding benchmark scores a system, not a model. The model sits inside a scaffold: how many attempts it gets, which tools it can call, how the repository is prepared, what happens on a failed run. SWE-bench’s own leaderboard filters results by scaffold for exactly this reason. A score with no harness named is a number with its conditions stripped off, and comparing two of them across different scaffolds is meaningless rather than merely imprecise.

Same name, different test

It goes further than scaffolds. SWE-bench Verified, SWE-Bench Pro and FrontierSWE are different tests with confusingly similar names. In embeddings, MTEB v1 scores are not comparable to v2, MMTEB or RTEB, and vendors quote whichever board flatters them. In speech, one set of weights scores 7.83 average word error rate on one suite and 16.13 on a meeting-room suite. Same weights, same day, two numbers that look like a capability gap and are not.

A common misconception

Commonly believed: Benchmark numbers are comparable. Being comparable is the entire point of a benchmark.

Actually: They are comparable inside one suite, one version of it, and one harness. Cross any of those boundaries and a difference can be entirely procedural. Before comparing two scores, check three things: the same test, the same version of that test, the same scaffold around the model. If any answer is no, or unstated, you are holding two facts rather than a comparison.

A vendor reports 80% on a coding benchmark and names no harness. What can you conclude?

Answer: Only that their system scored 80% under conditions they did not describe. An unnamed harness is missing information, not proof of dishonesty. The score is real for whatever setup produced it. What you cannot do is line it up against someone else’s number, because the scaffold is part of what was measured and you do not know theirs.

In one sentence

A benchmark score without its suite, its version and its harness is a number, not a measurement.