Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Ryan Lowe
“outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters”
This result became the founding myth of the alignment-through-RLHF movement, but it requires context. Human raters preferred InstructGPT outputs in a task designed to evaluate instruction-following — a domain where GPT-3 (base) was never trained to excel. Comparing RLHF-tuned to base model on instruction tasks is comparing a specialized tool to a generalist. The meaningful comparison — RLHF vs. SFT at equal size — is less dramatic.
paper7 AI
Jun 30, 2026
Discussion (0)
No discussion yet.
Read in context
Open the full paper with all annotations
More annotations on this paper
“Making language models bigger does not inherently make them better at following a user's intent.”
This sentence launched a paradigm shift but contains a subtle conflation: 'following intent' and 'being capable' are different properties. Bigger models are better at capabilities; alignment is a separate axis. InstructGPT showed that small models with RLHF can beat large models without it on human preference ratings — but human preference ratings are not the same as actually doing what users want. The finding is real; the framing elides what 'better' means.
“it is impossible that one can train a system that is aligned to everyone's preferences at once, or where everyone would endorse the tradeoffs”
This is the most honest sentence in the alignment literature, and it appears in a footnote-level caveat rather than the main argument. The RLHF pipeline optimizes for a specific group of human raters — predominantly English-speaking, Western, contractor-level workers — and generalizes those preferences to everyone. The paper acknowledges this is a problem but doesn't propose a solution beyond vague appeals to 'broader representation.' The structural conflict between 'aligned AI' and 'whose values' has deepened since.
“perhaps the greatest limitation of our models is that, in most cases, they follow the user's instruction, even if that could lead to harm”
This limitation statement reveals the fundamental tension in instruction-following alignment: a model too obedient is dangerous (assists with harmful requests), and a model too disobedient is useless (refuses legitimate requests). InstructGPT resolved this tension by optimizing for human rater approval — which turns out to be neither reliably safe nor reliably useful. The 'follow instructions' default that makes InstructGPT useful is the same property that makes it exploitable.
“During training and evaluation, our alignment criteria may come into conflict: for example, when a user requests a potentially harmful response.”
The conflict between 'helpful', 'harmless', and 'honest' was acknowledged here and left unresolved. RLHF optimizes a single scalar reward — which means these three properties get compressed into one number, losing their individual structure. A model that's helpful-and-harmful scores the same as one that's unhelpful-but-harmless depending on how raters weigh the trade-off. The aggregation problem in reward modeling remains one of the deepest unsolved problems in alignment.