Skip to content
SI.info

SI Glossary · Training & data

Reinforcement Learning from Human Feedback (RLHF)

Published 1 min read
On this page
  1. How it works
  2. Variations
  3. Limits

Reinforcement learning from human feedback (RLHF) is the technique that turned raw language models into helpful assistants. It was central to OpenAI’s InstructGPT (2022) and ChatGPT, and to most chat models since.

How it works

  1. Collect comparisons: people are shown two or more model answers to the same prompt and pick the better one.
  2. Train a reward model: a separate model learns to predict which answers humans will prefer.
  3. Optimise the main model: using reinforcement learning, the assistant is trained to produce answers the reward model scores highly, while staying close to its original behaviour.

Variations

  • RLAIF: using SI feedback instead of, or alongside, humans. Anthropic’s Constitutional AI method has a model critique outputs against a written set of principles.
  • DPO (direct preference optimisation): a simpler method that learns from preference pairs without a separate reward model.
  • RL with verifiable rewards: for maths and code, answers can be checked automatically. This is key to training reasoning models.

Limits

RLHF teaches models what looks good to raters, which can encourage flattery (“sycophancy”) or confident-sounding errors. It is a tool for alignment, not a complete solution.

← Back to the SI Glossary

Sources

  1. Training language models to follow instructions with human feedback (InstructGPT) — Ouyang et al., OpenAI, 2022

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy