RLHF and the reward model

People are bad at writing the perfect answer and good at picking the better of two. RLHF is built on that gap.

Part of the How models are made track on lAItest.

Writing the perfect answer is hard. Pointing at the better of two answers is easy. RLHF is built entirely on that gap.

Step one: train a model to imitate human taste.

Show a person two responses to the same prompt and ask which is better. Collect a great many of those comparisons. Then train a second, separate network, the reward model, whose only job is to look at a response and output a number predicting which one a human would have picked. It is a stand-in for a human rater that can be queried millions of times an hour.

Step two: let the language model chase that number.

The language model generates responses, the reward model scores them, and reinforcement learning pushes the weights toward higher-scoring behaviour. A leash is attached: the model is penalised for drifting too far from where it started, or it abandons fluent English in pursuit of whatever the reward model happens to like. The InstructGPT paper, arXiv 2203.02155, is the reference write-up of this pipeline.

A common misconception

Commonly believed: RLHF makes a model more accurate, because humans corrected its mistakes.

Actually: Humans rated responses they liked, not responses they checked. The reward model learns what raters approve of, which correlates with being right without being the same thing. Confident, well-structured, agreeable answers score well. That is one accepted account of why heavily preference-tuned models drift toward sycophancy, and why "sounds good" and "is correct" can come apart.

What does a reward model actually predict?

Answer: Which of two responses a human rater would prefer. It is trained on human comparisons, so it predicts human preference. Nothing in that pipeline checks a fact. Anything raters systematically like, including confidence, length and agreement, gets baked in alongside genuine quality. That is why reward hacking is a practical failure mode and not a theoretical one.

In one sentence

RLHF does not teach a model to be right. It teaches a model to produce what human raters preferred, which is close, and not identical.