DPO, RLAIF and Constitutional AI
Preference training has two expensive parts, a second network and a human. Each one has been removed.
Part of the How models are made track on lAItest.
RLHF has two expensive parts: a whole second network, and a human being. Each one has been removed independently.
DPO deletes the reward model.
Direct Preference Optimization starts from the same data, pairs where one response was preferred, and skips the middle step. Instead of training a reward model and then running reinforcement learning against it, DPO rearranges the maths so the preference pairs update the language model directly. One training loop instead of two, no reinforcement learning machinery, far fewer moving parts to go wrong. Paper: arXiv 2305.18290.
RLAIF replaces the human rater.
Reinforcement learning from AI feedback keeps the pipeline and swaps who does the comparing: a model judges which response is better. That trades a slow, costly, inconsistent labeller for a fast, cheap, consistently biased one. It scales to volumes no annotation team can reach, and it inherits whatever blind spots the judging model has.
Constitutional AI writes the standard down.
Constitutional AI, arXiv 2212.08073, gives the judging model an explicit written set of principles and asks it to critique and revise responses against them. The point is not only cost. It moves the standard out of thousands of individual rater judgements and into a document you can read, argue with and edit.
A common misconception
Commonly believed: AI feedback is a cost-cutting shortcut, so it must align a model worse than real human ratings do.
Actually: It is a genuine trade in both directions. Human raters are inconsistent, they fatigue, they disagree with each other, and the standard they are applying exists nowhere in writing. A written constitution is auditable and reproducible in a way a crowd of annotators is not. What you give up is contact with actual human reaction, and any principle nobody thought to write down is simply absent.
In one sentence
The expensive parts of preference training are a second network and a person, and the field found a way to drop each one separately.