Tests at the edge
ARC-AGI measures how efficiently a system picks up a rule it has never seen. HLE just makes the questions brutal. One of them prints the bill.
Part of the Trust and evals track on lAItest.
One leaderboard prints what each answer cost to produce, right next to how often it was right.
Almost none of the others do.
ARC-AGI measures learning, not knowing
Most benchmarks ask what a system knows. ARC-AGI asks how efficiently it can acquire a rule it has never seen, from a handful of examples. That property is called skill-acquisition efficiency, and it is why having read the internet does not obviously help. On ARC-AGI-2 the average individual human scores about 66%. ARC-AGI-3 moves again, from static puzzles into novel interactive environments, because a test that stays still eventually gets solved by fitting to it.
Cost per task, printed beside accuracy
ARC-AGI publishes what each attempt cost to run. That is unusual, and it exposes what other boards hide: a system can spend more thinking to score higher. Private reasoning tokens bill at output rates — $25 per 1M on Claude Opus 5 — so a long thinking pass on a short answer can cost several times what the visible answer costs. Once money is on the board, "it scored higher" becomes "it scored higher, at this much per task".
Humanity’s Last Exam went the other way
Rather than testing a different ability, HLE makes the questions themselves brutal: 2,500 questions across more than a hundred subjects, written by roughly a thousand expert contributors from over five hundred institutions in fifty countries, published in Nature in January 2026. Its scores move month to month. That is not a flaw in the test; it is what a test deliberately built at the edge of what anyone can answer looks like while the frontier moves under it.
Try it
Put a long private reasoning pass against a short visible answer. Thinking bills at the output rate, so watch the total move while the answer does not. This step is an interactive widget; open the lesson to use it.
What a leaderboard screenshot does not carry
| What | Value | Provenance |
|---|---|---|
| Cheapest frontier input price | $0.75 per 1M | source, verified . Cost per task depends on which model you ran it with |
| Frontier models tracked here | 18 | source, verified . |
| Newest frontier release | Sep 2026 | source, verified . |
| This index last verified | Sep 2026 | source, verified . |
In one sentence
A score says what a system got right. A cost per task says what it spent getting there, and only one of those appears on most leaderboards.