GRPO and verifiable rewards
When a program can check the answer, training no longer needs a human rater or a reward model.
Part of the How models are made track on lAItest.
For maths and code, nobody has to judge the answer. You can just run it.
That single fact reshaped how reasoning models are trained.
RLVR: reinforcement learning from verifiable rewards.
If a task has a checkable answer, the unit tests pass, the final number matches, the program compiles, then the reward can come from running a checker instead of from a learned model of human taste. That reward cannot be flattered. You do not win by sounding confident, only by being right. Which is why the recipe is concentrated in maths, code and other domains that have a ground truth.
GRPO scores a group against itself.
Reinforcement learning needs a baseline: was this answer better or worse than expected? Usually that means training yet another network to estimate the expected value. Group Relative Policy Optimization drops it. Sample a group of answers to the same question, score them all, and use the group average as the baseline. Better than your siblings, push up. Worse, push down. One fewer network to train and serve.
The dates, precisely.
The work that put this recipe in front of the whole field appeared as an arXiv preprint in January 2025 and was published, peer-reviewed, in Nature on 17 September 2025. Those are two milestones eight months apart, and they get collapsed into one wrong date constantly.
A common misconception
Commonly believed: Verifiable rewards mean the model is finally trained on truth, so hallucination is solved.
Actually: It is trained on checkability, which is narrower. A grader exists for an arithmetic answer and for a failing test suite. No grader exists for "is this summary fair", "is this advice safe", "is this the right architecture". Whether skills learned under a verifier transfer to work nobody can verify is an open empirical question, not a settled result.
In one sentence
When a program can grade the answer, training stops needing anyone at all, and only some of the work we care about can be graded that way.