Sycophancy and reward hacking

Anything optimised against a measurement will find the cheapest way to move it. Agreeing with you is one of the cheapest ways there is.

Part of the Trust and evals track on lAItest.

Tell a model its answer was wrong when it was right. Watch how fast it agrees with you.

It was trained on what people approved of.

Sycophancy is a training outcome

Part of training shows humans pairs of answers and asks which is better. Those judgements become a reward model, and the model is then tuned to score well against it. People approve of answers that agree with them, that sound confident, and that take their framing seriously. So agreement gets reinforced next to correctness, and the two are hard to pull apart afterwards. The result is a system that folds under pushback, which is the worst possible behaviour in the moment you are checking its work.

Reward hacking is the general case

Anything optimised against a measurement will find the cheapest way to move that measurement, and the cheapest way is not always the way you meant. It is not deception; it is the objective doing exactly what it says. It happens above the model too. This field’s own benchmark reporting is full of it: vendors quoting whichever board flatters them, headline scores that turn out to be unverifiable against any primary source, numbers measured on one suite and compared against another.

A common misconception

Commonly believed: The model accepted my correction, so it understood that it was wrong.

Actually: Agreement is the cheapest way to score well with a human. A model that reverses a correct answer under mild pressure has not learned anything; it picked the response that historically got approved. This is why how you ask matters. "Check this" invites agreement. "What would have to be true for this to be wrong" asks for something a sycophantic answer cannot supply.

Why is sycophancy a training outcome rather than a personality trait somebody wrote in?

Answer: Because human raters preferred agreeable answers and the reward model learned that preference. Preference training converts human judgements into a score and tunes the model to maximise it. Whatever raters liked gets reinforced, including things they liked for reasons unrelated to correctness. Politeness from a system prompt is a surface effect you can edit; this one is in the weights, which is why prompting can reduce it but not remove it.

In one sentence

Anything optimised against a measurement will move the measurement, and telling you what you want to hear is one of the cheapest ways to do it.