Fine-Tuning Language Models from Human Preferences

Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, Geoffrey Irving

Introduction

We would like to apply reinforcement learning to complex tasks defined only by human judgment, where we can only tell whether a result is good or bad by asking humans. To do this, we can first use human labels to train a model of reward, and then optimize that model. While there is a long history of work learning such models from humans through interaction, this work has only recently been applied to modern deep learning, and even then has only been applied to relatively simple simulated environments [Christiano et al., 2017, Ibarz et al., 2018, Bahdanau et al., 2018]. By contrast, real world settings in which humans need to specify complex goals to AI agents are likely to both involve and require natural language, which is a rich medium for expressing value-laden concepts. Natural language is particularly important when an agent must communicate back to a human to help provide a more accurate supervisory signal [Irving et al., 2018, Christiano et al., 2018, Leike et al., 2018].

Natural language processing has seen substantial recent advances. One successful method has been to pretrain a large generative language model on a corpus of unsupervised data, then fine-tune the model for supervised NLP tasks [Dai and Le, 2015, Peters et al., 2018, Radford et al., 2018, Khandelwal et al., 2019]. This method often substantially outperforms training on the supervised datasets from scratch, and a single pretrained language model often can be fine-tuned for state of the art performance on many different supervised datasets [Howard and Ruder, 2018]. In some cases, fine-tuning is not required: Radford et al. find that generatively trained models show reasonable performance on NLP tasks with no additional training (zero-shot).

There is a long literature applying reinforcement learning to natural language tasks. Much of this work uses algorithmically defined reward functions such as BLEU for translation [Ranzato et al., 2015, Wu et al., 2016], ROUGE for summarization [Ranzato et al., 2015, Paulus et al., 2017, Wu and Hu, 2018, Gao et al., 2019b], music theory-based rewards [Jaques et al., 2017], or event detectors for story generation [Tambwekar et al., 2018]. Nguyen et al. used RL on BLEU but applied several error models to approximate human behavior. Wu and Hu and Cho et al. learned models of coherence from existing text and used them as RL rewards for summarization and long-form generation, respectively. Gao et al. [2019a] built an interactive summarization tool by applying reward learning to one article at a time. Experiments using human evaluations as rewards include Kreutzer et al. which used off-policy reward learning for translation, and Jaques et al. which applied the modified Q-learning methods of Jaques et al. to implicit human preferences in dialog. Yi et al. learned rewards from humans to fine-tune dialog models, but smoothed the rewards to allow supervised learning. We refer to Luketina et al. for a survey of RL tasks involving language as a component, and for RL results using transfer learning from language. RL is not the only way to incorporate ongoing human feedback: Hancock et al. ask humans what a dialogue system should have said instead, then continue supervised training.

In this paper, we combine the pretraining advances in natural language processing with human preference learning. We fine-tune pretrained language models with reinforcement learning rather than supervised learning, using a reward model trained from human preferences on text continuations. Following Jaques et al. , we use a KL constraint to prevent the fine-tuned model from drifting too far from the pretrained model. We apply our method to two types of tasks: continuing text in a way that matches a target style, either positive sentiment or vividly descriptive, and summarizing text from the CNN/Daily Mail or TL;DR datasets [Hermann et al., 2015, Völske et al., 2017]. Our motivation is NLP tasks where supervised data sets are unavailable or insufficient, and where programmatic reward functions are poor proxies for our true goals.

For stylistic continuation, 5,000 human comparisons (each choosing the best of 4 continuations) result in the fine-tuned model being preferred by humans 86% of the time vs. zero-shot and 77% vs. fine-tuning to a supervised sentiment network. For summarization, we use 60,000 human samples to train models that can roughly be described as “smart copiers”: they typically copy whole sentences from the input, but vary what they copy to skip irrelevant initial text. This copying behavior emerged naturally from the data collection and training process; we did not use any explicit architectural mechanism for copying as in See et al. , Gehrmann et al. . One explanation is that copying is an easy way to be accurate, given that we did not instruct labelers to penalize copying but do instruct them to penalize inaccuracy. It may also reflect the fact that some labelers check for copying as a fast heuristic to ensure a summary is accurate. Indeed, human labelers significantly prefer our models to supervised fine-tuning baselines and even to human-written reference summaries, but not to a lead-3 baseline which copies the first three sentences.

For summarization, we continue to collect additional data and retrain our reward model as the policy improves (online data collection). We also test offline data collection where we train the reward model using data from the original language model only; offline data collection significantly reduces the complexity of the training process. For the TL;DR dataset, human labelers preferred the policy trained with online data collection 71% of the time, and in qualitative evaluations the offline model often provides inaccurate summaries. In contrast, for stylistic continuation we found that offline data collection worked similarly well. This may be related to the style tasks requiring very little data; Radford et al. show that generatively trained models can learn to classify sentiment from very few labeled examples.

In concurrent work, Böhm et al. also use human evaluations to learn a reward function for summarization, and optimize that reward function with RL. Their work provides a more detailed investigation of the learned policy and reward function on the CNN/Daily Mail dataset, while we are interested in exploring learning from human feedback more generally and at larger computational scale. So we consider several additional tasks, explore the effects of on-policy reward model training and more data, and fine-tune large language models for both reward modeling and RL.

Methods

We begin with a vocabulary Σ\Sigma and a language model ρ\rho which defines a probability distribution over sequences of tokens Σn\Sigma^{n} via

We will apply this model to a task with input space X=Σ≤mX=\Sigma^{\leq m}, data distribution D\mathcal{D} over XX, and output space Y=ΣnY=\Sigma^{n}. For example, x∈Xx\in X could be an article of up to 1000 words and y∈Yy\in Y could be a 100-word summary. ρ\rho defines a probabilistic policy for this task via ρ(y∣x)=ρ(xy)/ρ(x)\rho(y|x)=\rho(xy)/\rho(x): fixing the beginning of the sample to xx and generating subsequent tokens using ρ\rho.

However, we want to perform tasks defined by human judgments, where we can only learn about the reward by asking humans. To do this, we will first use human labels to train a reward model, and then optimize that reward model.

Since the reward model needs to understand language, we initialize it as a random linear function of the final embedding output of the language model policy ρ\rho following Radford et al. (see section 4.2 for why we initialize from ρ\rho rather than π\pi). To keep the scale of the reward model consistent across training, we normalize it so that it has mean 0 and variance 1 for x∼D,y∼ρ(⋅∣x)x\sim\mathcal{D},y\sim\rho(\cdot|x).

Now we fine-tune π\pi to optimize the reward model rr. To keep π\pi from moving too far from ρ\rho, we add a penalty with expectation βKL⁡(π,ρ)\beta\operatorname{KL}(\pi,\rho) (see table 10 for what happens without this). That is, we perform RL on the modified reward

We either choose a constant β\beta or vary it dynamically to achieve a particular value of KL⁡(π,ρ)\operatorname{KL}(\pi,\rho); see section 2.2. This term has several purposes: it plays the role of an entropy bonus, it prevents the policy from moving too far from the range where rr is valid, and in the case of our style continuation tasks it also is an important part of the task definition: we ask humans to evaluate style, but rely on the KL term to encourage coherence and topicality.

Gather samples (x,y0,y1,y2,y3)(x,y_{0},y_{1},y_{2},y_{3}) via x∼D,yi∼ρ(⋅∣x)x\sim\mathcal{D},y_{i}\sim\rho(\cdot|x). Ask humans to pick the best yiy_{i} from each.

Initialize rr to ρ\rho, using random initialization for the final linear layer of rr. Train rr on the human samples using loss 1.

Train π\pi via Proximal Policy Optimization (PPO, Schulman et al. ) with reward RR from 2 on x∼Dx\sim\mathcal{D}.

In the online data collection case, continue to collect additional samples, and periodically retrain the reward model rr. This is described in section 2.3.

We use a 774M parameter version of the GPT-2 language model in Radford et al. trained on their WebText dataset and their 50,257 token invertible byte pair encoding to preserve capitalization and punctuation [Sennrich et al., 2015]. The model is a Transformer with 36 layers, 20 heads, and embedding size 1280 [Vaswani et al., 2017].

For stylistic continuation tasks we perform supervised fine-tuning of the language model to the BookCorpus dataset of Zhu et al. prior to RL fine-tuning; we train from scratch on WebText, supervised fine-tune on BookCorpus, then RL fine-tune to our final task. To improve sample quality, we use a temperature of T<1T<1 for all experiments; we modify the initial language model by dividing logits by TT, so that future sampling and RL with T=1T=1 corresponds to a lower temperature for the unmodified pretrained model.

2 Fine-tuning details

Starting with the pretrained language model, the reward model is trained using the Adam optimizer [Kingma and Ba, 2014] with loss 1. The batch size is 8 for style tasks and 32 for summarization, and the learning rate is 1.77×10−51.77\times 10^{-5} for both. We use a single epoch to avoid overfitting to the small amount of human data, and turn off dropout.

For training the policy π\pi, we use the PPO2 version of Proximal Policy Optimization from Dhariwal et al. . We use 2M episodes (x,yx,y pairs), γ=1\gamma=1, four PPO epochs per batch with one minibatch each, and default values for the other parameters. We use batch size 1024 for style tasks and 512 for summarization. We do not use dropout for policy training. The learning rate was 1.41×10−51.41\times 10^{-5} for style tasks and 7.07×10−67.07\times 10^{-6} for summarization.

Models trained with different seeds and the same KL penalty β\beta sometimes end up with quite different values of KL⁡(π,ρ)\operatorname{KL}(\pi,\rho), making them hard to compare. To fix this, for some experiments we dynamically vary β\beta to target a particular value of KL⁡(π,ρ)\operatorname{KL}(\pi,\rho) using the log-space proportional controller

For supervised fine-tuning baselines, we fine-tune for 1 epoch on the CNN/Daily Mail and TL;DR training sets (for TL;DR we removed 30K examples to serve as a validation set). We decayed the learning rate to 0 with a cosine schedule; for the initial value, we swept over 8 log-linearly spaced options between 10−410^{-4} and 3×10−43\times 10^{-4}. We also experimented with different dropout rates, and found a rate of 0.1 to work best. We then chose the model with the best validation loss.

3 Online data collection

If the trained policy π\pi is very different from the zero-shot policy ρ\rho, the reward model will suffer a large distributional shift from training on samples from ρ\rho to evaluation on samples from π\pi. To prevent this, we can collect human data throughout RL fine-tuning, continuously gathering new data by sampling from π\pi and retraining the reward model. As section 3 shows, online data collection was important for summarization but not for the simpler style tasks.

In the online case, we will choose a function l(n)l(n) describing how many labels we want before beginning the nthn^{\textrm{th}} PPO episode. Let Nπ=2×106N_{\pi}=2\times 10^{6} be the total number of PPO episodes, Nr0=l(0)N_{r}^{0}=l(0) be an initial number of human labels, and NrN_{r} be the total number of human labels. We take

We pause before the nthn^{\textrm{th}} PPO episode if we have fewer than l(n)l(n) labels. We send another batch of requests to the labelers if the total requests so far is less than l(n)+1000l(n)+1000, to ensure they have at least 1000 outstanding queries at any time. We train the reward model before the first PPO episode, and then retrain it 19 more times at evenly spaced values of l(n)l(n). Each time we retrain we reinitialize rr to a random linear layer on top of ρ\rho and do a single epoch through the labels collected so far. The offline case is the limit Nr=Nr0N_{r}=N_{r}^{0}.

To estimate overall progress, we gather validation samples consisting of x∼D;y0,y1∼ρ(⋅∣x);y2,y3∼π(⋅∣x)x\sim\mathcal{D};y_{0},y_{1}\sim\rho(\cdot|x);y_{2},y_{3}\sim\pi(\cdot|x) at a constant rate; human labels on these give how often π\pi beats ρ\rho. Since validation samples are only used to evaluate the current π\pi, we can add them to the training set for rr. In order to estimate inter-labeler agreement, 5% of queries are answered 5 times by different labelers. Label counts in section 3 include validation samples and repeated labels.

4 Human labeling

We use Scale AI to collect labels. The Scale API accepts requests of the form (x,y0,y1,y2,y3)(x,y_{0},y_{1},y_{2},y_{3}) and returns selections b∈{0,1,2,3}b\in\left\{0,1,2,3\right\}. We describe the task to Scale through a combination of instructions (appendix A) and a dataset of about 100 example comparisons from the authors.

Unlike many tasks in ML, our queries do not have unambiguous ground truth, especially for pairs of similar outputs (which play a large role in our training process, since we train rr on pairs of labels sampled from a single policy π\pi). This means that there is significant disagreement even between labelers who have a similar understanding of the task and are trying to rate consistently. On 4-way comparisons for sentiment and TL;DR summarization, authors of this paper agree about 60% of the time (vs. 25% for random guessing). This low rate of agreement complicates the quality control process for Scale; the authors agree with Scale labelers 38% of the time on sentiment and 46% of the time on TL;DR summarization. We give further details of the human data collection and quality evaluation in appendix B.

For final evaluation of two models AA and BB, we generate either 2-way comparisons between pairs (a∼A,b∼B)(a\sim A,b\sim B) or 4-way comparisons with quadruples (a0,a1∼A,b0,b1∼B)(a_{0},a_{1}\sim A,b_{0},b_{1}\sim B), randomize the order in which samples are presented, and present these comparisons to Scale. Evaluating the quality of a model trained by Scale using the same set of humans from Scale is perilous: it demonstrates that rr and π\pi have succeeded in fitting to the human reward, but does not show that those human evaluations capture what we really care about, and our models are incentivized to exploit idiosyncracies of the labeling process. We include samples from our models so that readers can judge for themselves.

Experiments

In section 3.1.1, we test our approach to RL fine-tuning of language models by using a mock labeler (a sentiment model trained on a review classification problem) as a stand-in for human labels. We show that RL fine-tuning is effective at optimizing this complex but somewhat artificial reward. In section 3.1.2, we show that we can optimize language models from human preferences on stylistic continuation tasks (sentiment and physical descriptiveness) with very little data, and that in the sentiment case the results are preferred to optimizing the review sentiment model. In section 3.2 we apply RL fine-tuning to summarization on the CNN/Daily Mail and TL;DR datasets, show that the resulting models are essentially “smart copiers”, and discuss these results in the context of other summarization work.

We release codeCode at https://github.com/openai/lm-human-preferences. for reward modeling and fine-tuning in the offline data case. Our public version of the code only works with a smaller 124M parameter model with 12 layers, 12 heads, and embedding size 768. We include fine-tuned versions of this smaller model, as well as some of the human labels we collected for our main experiments (note that these labels were collected from runs using the larger model).

We first apply our method to stylistic text continuation tasks, where the policy is presented with an excerpt from the BookCorpus dataset [Zhu et al., 2015] and generates a continuation of the text. The reward function evaluates the style of the concatenated text, either automatically or based on human judgments. We sample excerpts with lengths of 32 to 64 tokens, and the policy generates 24 additional tokens. We set the temperature of the pretrained model to T=0.7T=0.7 as described in section 2.1.

To study our method in a controlled setting, we first apply it to optimize a known reward function rsr_{s} designed to reflect some of the complexity of human judgments. We construct rsr_{s} by training a classifierThe model is a Transformer with 6 layers, 8 attention heads, and embedding size 512. on a binarized, balanced subsample of the Amazon review dataset of McAuley et al. . The classifier predicts whether a review is positive or negative, and we define rs(x,y)r_{s}(x,y) as the classifier’s log odds that a review is positive (the input to the final sigmoid layer).

Optimizing rsr_{s} without constraints would lead the policy to produce incoherent continuations, but as described in section 2.2 we include a KL constraint that forces it to stay close to a language model ρ\rho trained on BookCorpus.

The goal of our method is to optimize a reward function using only a small number of queries to a human. In this mock sentiment experiment, we simulate human judgments by assuming that the “human” always selects the continuation with the higher reward according to rsr_{s}, and ask how many queries we need to optimize rsr_{s}.

Figure 2 shows how rsr_{s} evolves during training, using either direct RL access to rsr_{s} or a limited number of queries to train a reward model. 20k to 60k queries allow us to optimize rsr_{s} nearly as well as using RL to directly optimize rsr_{s}.

Because we know the reward function, we can also analytically compute the optimal policy and compare it to our learned policies. With a constraint on the KL divergence KL⁡(π,ρ)\operatorname{KL}(\pi,\rho) between the learned policy π\pi and the language model ρ\rho, the optimal policy has the form:

We approximate the reward of this policy for given xx and β\beta by sampling a large number of continuations from ρ(y∣x)\rho(y|x) and reweighting them by ers(x,y)/βe^{r_{s}(x,y)/\beta}. Figure 3 compares the reward obtained by our policies to the estimated optimal reward across a range of KL values. There is a significant gap from optimality after training the policy on 2M continuations—the number used in our main experiments—though it is largely closed with more training. Our policies continue to receive higher rewards for larger KL divergences, where we cannot afford to approximate πopt\pi_{\textrm{opt}} by sampling.

1.2 Human evaluations of continuations

We apply our method to two continuation tasks defined by human judgments:

Humans are asked to reward “positive and happy” continuations.

Humans are asked to reward “vividly descriptive” continuations.

The human labelers are presented with a BookCorpus excerpt and four possible continuations; they are asked to select the best continuation. Full instructions for labelers are provided in appendix A (although labelers also learned from ∼50\sim 50 example comparisons labeled by the authors and so the instructions do not completely define the task).

To make the labeling task more natural, we select excerpts that start and end with a period. When sampling continuations that will be presented to humans, we use rejection sampling to ensure there is a period between tokens 16 and 24 and then truncate at that period.This is a crude approximation for “end of sentence.” We chose it because it is easy to integrate into the RL loop, and even a crude approximation is sufficient for the intended purpose of making the human evaluation task somewhat easier. During the RL fine-tuning, we penalize continuations that don’t have such a period by giving them a fixed reward of −1-1.

We dynamically adjusted β\beta to obtain a KL divergence of 6 nats for descriptiveness and 10 nats for sentiment (section 2.2).

We trained a range of models using different amounts of feedback, testing both offline data collection where humans rate only the initial language model’s continuation, and online data collection where humans continuously rate the current policy’s continuations (section 2.3). We then compared these different policies to each other and to the zero-shot performance of the original language model. The results are shown in fig. 4 and table 1. Each model comparison is based on 1024 four-way continuation comparisons, two from each of the models being compared, each rated by 3 humans.

For these continuation tasks, offline and online data collection give similar performance. We find that very little human data is required for fine-tuning: performance with 5k, 10k, and 20k reward model training samples is similar, degrading only for less than 5k samples.The descriptiveness policy trained with 2.5k samples performed poorly, but we believe this is due to randomness in RL. The model trained using the review sentiment classifier from section 3.1.1 does poorly relative to models optimized using human preference: in 77% of contexts, labelers preferred the output of the model trained with real human feedback.

2 Summarization

We also applied our method to two summarization tasks: the CNN/Daily Mail dataset of Hermann et al. and the TL;DR dataset of Völske et al. . We sample articles or Reddit posts, truncate to 500 tokens, add a "\n\nTL;DR:" suffix (and for CNN/Daily Mail, a "Article:\n\n" prefix) and let the policy respond with up to 75 tokens. We set the temperature of the pretrained model to T=0.5T=0.5 for CNN/Daily Mail and T=0.7T=0.7 for TL;DR. To make the task more natural for humans, we ensure articles consist of whole sentences by truncating to the last newline character. When sampling summaries that will be shown to a human, we use rejection sampling to ensure there is a newline between tokens 55 and 75 and truncate at that newline. During RL fine-tuning, we penalize summaries that don’t have such a newline by giving them a fixed score of -1. For CNN/Daily Mail we used a fixed KL coefficient β=0.1\beta=0.1; for TL;DR we used β=0.03\beta=0.03.

For RL fine-tuning, we trained online data collection models with 15k, 30k, and 60k human labels, and an offline data collection ablation with 60k labels. We also show zero-shot performance of the pretrained model, a supervised fine-tuned baseline using the same pretrained model as starting point (section 2.2), and a lead-3 baseline which copies the first three sentences of the context. We truncate lead-3 at a period in the same way we truncate generated summaries, so occasionally it is 2 sentences. Finally, we combine supervised and RL fine-tuning: performing human RL fine-tuning starting with the supervised fine-tuned model. The purely RL fine-tuned models use contexts from the datasets during training but ignore the reference summaries; the supervised and supervised+RL models use both contexts and summaries.

We report two sets of numerical results: human evaluations between pairs of models (table 5) and ROUGE results on the test set of CNN/Daily Mail and our validation set of TL;DR (table 4). ROUGE results suggest that online data collection is important for best performance, in contrast to our stylistic continuation tasks. At a fixed number of labels, online tends to be better than offline, with a 3 point R-AVG gain on CNN/DM at 60k labels.That said, different training runs have considerable variation and it is expensive to run multiple seeds with humans, so it is possible that this gap is largely noise. On both datasets we see significant returns to data volume up to 60k human labels (though the trend is less clear for human evaluation). On both datasets, supervised + RL fine-tuning is best, and indeed pure RL fine-tuning is worse than the supervised baseline according to ROUGE in all cases (though the supervised baseline uses the full supervised training dataset, which is much larger than 60k samples). Lead-3 is hard to beat: it is the best model for R-1 and R-2 on CNN/Daily Mail, and only supervised + RL fine-tuning beats it otherwise.

But our goal is optimizing reward defined by humans, not ROUGE. Table 5 shows pairwise comparisons between different model pairs according to human labelers, using 1024 samples with majority vote of 3 labelers per sample. Here the picture is different, though also significantly noisier. Our online trained, 60k label model reliably beats both the zero-shot and supervised baselines, and even beats the combined supervised + RL fine-tuned model. Online training remains important, but the situation w.r.t. data volume is less clear and likely contaminated by noise: the 60k TL;DR model beats the 30k model only 40% of the time, for example. More worrisome, the 60k online model beats the human ground truth 96% of the time for TL;DR and 84% of the time for CNN/Daily Mail.

What is going on? As we show in the next section, our 60k RL fine-tuned model is almost entirely extractive (despite lacking any explicit extractive architectural component): it mostly copies whole sentences from the context, but varies which sentences are copied.

Much previous work in summarization has focused on explicit copying mechanisms, including the pointer network-based architecture of See et al. and the two-phase mask and paraphrase approach of Gehrmann et al. . The goal is to take advantage of copying (which is of fundamental importance to the task of summarization) without only copying—to be abstractive rather than extractive.

Figures 5 and 6 show the fractions of nn-grams and sentences generated by our models which are novel and repeated, respectively. From the novelty stats, we see that our RL fine-tuning consistently causes models to copy more. In particular, our 60k RL fine-tuned models are almost entirely extractive: they copy whole sentences 71% of the time for TL;DR and 98% of the time for CNN/Daily Mail. Applying RL fine-tuning starting from the supervised fine-tuned model copies much less: 6% and 30% for TL;DR and CNN/Daily Mail. Although we do not use explicit coverage metrics as in See et al. , Gehrmann et al. , both supervised and RL fine-tuned models do very little repetition within summaries.

While the purely RL fine-tuned models mostly copy, they vary where they copy from. Figure 7 illustrates this via the position of the longest common subsequence between context and summary. To understand when the model chooses to copy from the exact beginning, we identify common preambles in articles such that we would expect copying to be a poor strategy. Table 7 shows that these preambles are copied much less often than in the immediate beginnings of other articles, giving evidence that our models are smart about when to copy. However, we cannot establish that our reward model is smart beyond rewarding copying, as the zero-shot model also skips preambles.

Since combining supervised fine-tuning and RL fine-tuning gives the best ROUGE scores and and is also more abstractive, why not use it? Unfortunately there is an advantage to pure copying shown in table 8: it makes it easy for the model to tell the truth. The models that copy the most, 60k RL fine-tuned, is 90% and 95% accurate on TL;DR and CNN/Daily Mail; lifting whole sentences from the article usually leaves them true. The supervised fine-tuned and combined supervised+RL fine-tuned models are accurate at most 70% of the time: they paraphrase but paraphrase badly, often swapping names from the context or mixing together multiple sentences in invalid ways. Zero-shot is the most novel, but is accurate only 20% of the time. Similarly, Kryściński et al. found that 30% of samples from the supervised summarization models they tested contained inconsistencies, and Khandelwal et al. found that their pretrained encoder-decoder model “hallucinates facts…which are topical but never appear in the source”.

There are at least two ways of interpreting these results. The first is that copying is the easiest way to be accurate. The labelers were told to penalize inaccuracy and redundancy, but were not told to penalize copying. The zero-shot model copies some of the time, and when it copied it was accurate, so this behavior was reinforced. The result is a model that “degenerated to copying”, but at least does not lie.

However, this does not explain why both our model and lead-3 are strongly preferred by the labelers to the human reference summaries (table 5). This reveals a mismatch between the notion of quality we wanted our model to learn, and what the humans labelers actually evaluated. Checking for copying is very easy, so labelers who check primarily for copying can work quickly. Since the online data collection setting made quality control more difficult, we failed to detect and penalize this behavior.

Challenges

We conclude with a few lessons and directions we plan to consider in future reward learning work.

Online data collection was necessary to achieve the best results on summarization. However, fully online data collection—where each label comes from an up-to-date version of the policy which has already learned from almost all previous labels—had major disadvantages:

Software complexity: Our online system interleaves data gathering, reward model training, and RL fine-tuning. The resulting distributed system was significantly more complicated than if each task was kept separate, slowing the software development process. Moreover, a bug in any one of the tasks would break the entire training process.

Machine learning complexity: Online experiments were difficult to debug, as it was hard to iterate on one piece of the ML system at a time. We could often debug an online job only by switching briefly to offline, such as by launching a standalone reward model training run, but then would switch back to online once debugging was complete (until the next cycle).

Quality control issues: Significant work was required on Scale’s part to make their data quality mechanisms work in the low latency, online setting. However, even after this work it was difficult to maintain high data quality over a long period of time, and regressions were often not detected until after (or well after) training runs were complete. Since evaluation of labeler performance was online, by the time a worker was detected as poor some of their data might already been reported back and used for reward model training.

We believe the right middle ground between offline and online data collection is batched data collection, and plan to use this setting in future work. Collect a batch of data from the pretrained policy ρ\rho, train the reward model rr on this batch, then fine-tune the policy π\pi with rr frozen. Once complete, collect another batch of data sampled from π\pi, and iterate. The latency for each batch can be far longer than the online case, simplifying quality control. As in the fully online setting, we can always retrain the reward model from scratch on all data collected so far; human data is expensive so the total volume will be low. Removing the interleaved training of rr and π\pi simplifies software architecture and diagnosis of ML issues, and allows iteration on just one component (say rr in isolation) if problems occur. Li et al. reached similar conclusions in a restricted dialogue setting after validating in simulation that online and batched trained performed similarly.

Batched data collection is also a well-studied setting for active learning techniques. Although we use RL to fine-tune the policy π\pi, the human data is used only for supervised training of the reward model rr. Thus, any method for batch mode active learning of supervised models applies, using π\pi as the unlabeled data distribution for rr. Examples of such techniques include selecting batches based on entropy considerations [Guo and Schuurmans, 2008], gradient-based metrics [Huang et al., 2016, Ash et al., 2019], or by attempting to distinguish labeled and unlabeled examples [Gissin and Shalev-Shwartz, 2019].

2 Sharing parameters between reward model and policy causes overfitting

Although the reward model and policy are both initialized to ρ\rho, we train them as separate networks rather than a single shared network with multiple heads. We might expect joint training to be helpful, effectively using RL as an auxiliary task to improve the reward model’s performance. Joint training is particularly appealing because it could help the reward model stay strong enough that the policy cannot exploit it. Sharing could also improve computational efficiency, by allowing the models to share activations rather than requiring two separate forward passes.

Despite several attempts, we were not able to make this idea work. The problem comes from the massive imbalance of data: we have at most 60k samples for the reward model, but 2M episodes for the policy. This makes it challenging to maintain performance on both tasks without performing many epochs for the reward model and overfitting. We hope that future work will overcome this challenge.

3 Ambiguous tasks make labeling hard

Evaluation of a summary is both subjective and multidimensional. A single human labeler may have a clear notion of whether a given sample is separately accurate, grammatical, nonredundant, or covers all important topics; but in our experiments a labeler will often be asked to choose between samples each of which has some deficiencies. In choosing which of four samples is the best, a labeler must trade off between different desiderata. This makes consistent labeling difficult for honest labelers (including the authors!), and makes it difficult to quickly detect problematic labelers. It also makes the research more difficult to present and interpret: during our experiments we routinely checked the performance of models by having authors label results, since we knew the authors would attempt to do the task honestly, but were epistemically uneasy about reporting these numbers in the paper (table 8 is the one exception).

One could hope to cope with such “noise” by simply getting more labels and averaging them, but this does not resolve all the practical difficulties with ambiguity. When possible, it seems better to design less ambiguous labeling tasks that get at the same information. For example, rather than asking a person to rate or compare summaries, we could ask for a verbal description of the problems with a summary, or a suggested correction. If problems don’t exist we are done; otherwise describing a problem does not require consistently picking the same most important problem. Even if two people disagree on the most important problem, they may be more likely to agree that the other picked some problem, and more agreement eases data quality control and the overall experimental process.

4 Bugs can optimize for bad behavior

One of our code refactors introduced a bug which flipped the sign of the reward. Flipping the reward would usually produce incoherent text, but the same bug also flipped the sign of the KL penalty. The result was a model which optimized for negative sentiment while still regularizing towards natural language. Since our instructions told humans to give very low ratings to continuations with sexually explicit text, the model quickly learned to output only content of this form, regardless of how innocuous the starting point was. This bug was remarkable since the result was not gibberish but maximally bad output. The authors were asleep during the training process, so the problem was noticed only once training had finished. A mechanism such as Toyota’s Andon cord could have prevented this, by allowing any labeler to stop a problematic training process.

Conclusion

We have demonstrated RL fine-tuning of language models to four NLP tasks: stylistic continuation with high sentiment or physically descriptive language, and summarization on the CNN/Daily Mail and TL;DR datasets. Rather than building task-specific techniques, we achieve our results by straightforwardly applying reward learning to language generation. We extend previous reward learning work with pretrained models and KL regularization to prevent the policy from diverging too far from natural language.

Our results are mixed. On the continuation tasks we achieve good results vs. the zero-shot baseline as evaluated by humans with very few samples: 2.5k for sentiment and 5k for descriptiveness. However, for both summarization tasks our policies are only “smart copiers” (extractive rather than abstractive): they copy from the input text but skip over irrelevant preamble. The advantage of copying is truthfulness: by comparison the zero-shot and supervised models produce natural, plausible-looking summaries that are often lies. We believe the limiting factor in our experiments is data quality, in particular exacerbated by the online data collection setting, and plan to ameliorate this with batched data collection in future.

We believe the application of human reward learning to natural language tasks is important both from a capability and safety perspective. On the capability side, purely supervised training is insufficient to correct mistakes that arise when sampling from trained policies, and RL training to programmatic reward functions such as BLEU or ROUGE is insufficient: Paulus et al. conclude that "optimizing for single discrete evaluation metric[s] such as ROUGE with RL can be detrimental to the model quality." Interactive tasks such as dialogue are particularly relevant: it is difficult to define the goal of a dialogue without the human participant, and the length of dialogue makes it more likely that supervised learned models will go off distribution. In the supervised case NLP models are trained using human data; if we want RL fine-tuning we need human data too.

On the AI safety side, interactive communication between humans and ML models is a requirement for scalable reward learning methods such as amplification, debate, and recursive reward modeling [Christiano et al., 2018, Irving et al., 2018, Leike et al., 2018], and natural language is how humans communicate complex ideas. Although language models are unlikely to be ready for these tasks in their full generality, Perez et al. demonstrates that debate already improves generalization for question-answering when debaters quote from a source text. Using direct human preferences for language tasks is a step in the direction of scalable reward learning for language, and we believe further steps are possible.

Acknowledgments

We thank Akshat Bubna, Shariq Hashme, and many others at Scale for their work on data collection, Shan Carter for help with visualizations, Scott Gray for help with low precision training, Shahbaz Syed for information about the TL;DR dataset, and Christine Payne, Miles Brundage, Jared Kaplan, Jan Leike, Ethan Perez, and Jelena Luketina for helpful comments on the paper.

References

Appendix A Instructions for labelers

You will see snippets of text followed by several possible ways the text could be continued. Choose the most positive and happy continuation: • Consider all of the moods of the characters, the mood of the narrator, and the mood implied by the events and descriptions. • If a continuation contains sexually explicit or otherwise offensive content, treat it as strongly negative. Do not choose it even if it expresses a positive attitude.

A.2 Descriptiveness

You will see snippets of text followed by several possible ways the text could be continued. Choose the most vividly descriptive continuation: • Evaluate both on the quantity and on the vividness of physical details described. • The best continuations are full of details that give a strong sense of what the scene looks, sounds, or smells like. • Count only physical details, not details about abstract facts.

A.3 Summarization: TL;DR

You will see some text followed by several summaries. Please read the text and select the best summary. A summary is good if it: • Is useful and a good summary in general • Accurately states the important points of the text • Makes sense on its own A summary is bad if it: • Includes information that doesn’t appear in the text

A.4 Summarization: CNN/DM

You will see an article followed by several summaries. Please read the article and select the best summary. A summary is good if it: • Is useful and a good summary in general • Accurately states the important points of the article • Makes sense on its own A summary is bad if it: • Includes information that doesn’t appear in the article • Includes quotations that don’t appear verbatim in the article

Appendix B Human labeling details

Our quality assurance process was handled by Scale AI, though Scale made significant changes to their usual quality systems in order to deal with subjective tasks and provide very fast turnaround. Since we initially believed online data collection would be crucial, even the offline experiments were collected with this fast turnaround requirement. In the future we plan to use a more relaxed latency requirement.

The first step of data collection involves teaching the task to a small number of trusted Scale labelers by giving them a description of the task (appendix A). Scale uses these labelers to collect a large number of benchmark data points where several trusted labelers agree (out of a large set of unlabeled data points from ρ\rho). During full data collection, Scale serves these benchmark data points to freelance workers alongside real unlabeled data for training (the two types of data are indistinguishable when π=ρ\pi=\rho, though they do become distinguishable during training), maintaining a confidence model for the performance of each labeler on the benchmark distribution. The probability of getting a benchmark vs. a real sample varies dynamically on factors such as confidence in the labeler to correctly label a certain category. Freelancers who fail to perform well on benchmark tasks are filtered out. Additionally, Scale makes ad-hoc improvements to quality control over time, sometimes validating quality by comparing to a small number of gold-standard labels from the authors.

We evaluated the data quality after the fact on two of the tasks. During all data collection, 5% of queries were answered by 5 distinct labelers. We sampled 100 of these queries (restricting to ones generated from ρ\rho) and had two authors label each one. Based on this data, we estimated the rate of agreement between authors and Scale labelers, pairs of labelers, and pairs of authors. As table 9 shows, the data contained a significant amount of signal but did not match the quality of data which was hand-labeled by the authors.

An earlier version asked labelers for 1-10 ratings; in the best case this provides more information per label, but it was difficult to gauge labeler performance. Normalization was required since two good labelers would often differ by a (noisy) monotonic transform. If many scores concentrated on a few values (say 7 and 8) simple strategies could fool the filtering process. Absolute scores also tended to drift over the training process, as labelers would adjust to the new distribution of samples from the changing policy.

Finding high-quality workers involves human answering quality control questions which are not used in our experiments, and throwing away data from low-quality workers. So the total human cost of experiments is somewhat higher than the number of labels we actually use (which is what we report). For a short training run this can easily dominate the actual label requirements, though it can be amortized across several tasks by identifying consistently good workers. For our longer training runs the additional number of labels was modest. (None of these details are exposed to customers.)

Appendix C Samples

Samples from our models are shown in the following tables:

Mock sentiment continuation without a KL penalty: table 10

TL;DR summarization: tables 13, 14 and 15

CNN/Daily Mail summarization: tables 16, 17 and 18