SI Glossary · Training & data
Reinforcement Learning from Human Feedback (RLHF)
On this page
Reinforcement learning from human feedback (RLHF) is the technique that turned raw language models into helpful assistants. It was central to OpenAI’s InstructGPT (2022) and ChatGPT, and to most chat models since.
How it works
- Collect comparisons: people are shown two or more model answers to the same prompt and pick the better one.
- Train a reward model: a separate model learns to predict which answers humans will prefer.
- Optimise the main model: using reinforcement learning, the assistant is trained to produce answers the reward model scores highly, while staying close to its original behaviour.
Variations
- RLAIF: using SI feedback instead of, or alongside, humans. Anthropic’s Constitutional AI method has a model critique outputs against a written set of principles.
- DPO (direct preference optimisation): a simpler method that learns from preference pairs without a separate reward model.
- RL with verifiable rewards: for maths and code, answers can be checked automatically. This is key to training reasoning models.
Limits
RLHF teaches models what looks good to raters, which can encourage flattery (“sycophancy”) or confident-sounding errors. It is a tool for alignment, not a complete solution.
Sources
- Training language models to follow instructions with human feedback (InstructGPT) — Ouyang et al., OpenAI, 2022
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.