Pretraining Language Models with Human Preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, Ethan Perez
Introduction
Language models (LMs) are trained to imitate text from large and diverse datasets. These datasets often contain content that violates human preferences, e.g., falsehoods (Lin et al., 2022), offensive comments (Gehman et al., 2020), personally identifiable information (PII; Carlini et al., 2020) or low-quality code (Chen et al., 2021b). Imitating such data stands in stark contrast with the behavior people desire from language models, e.g., to generate text that is helpful, honest and harmless (Askell et al., 2021). In this paper, we explore alternative objectives for pretraining LMs on large amounts of diverse data that guide them to generate text aligned with human preferences.
Prior work on aligning LMs with human preferences almost exclusively focused on making adjustments to pretrained LMs. A widely adopted strategy of adding safety filters on top of pretrained LMs (Xu et al., 2020) works only to an extent: even the most effective safety filters fail to catch a large amount of undesirable content (Gehman et al., 2020; Welbl et al., 2021; Ziegler et al., 2022). Another approach involves finetuning LMs using either supervised learning on curated data (Solaiman & Dennison, 2021; Scheurer et al., 2023) or reinforcement learning from human feedback (RLHF; Ziegler et al., 2019; Ouyang et al., 2022; Bai et al., 2022; Menick et al., 2022), but this strategy is also limited by the fact that large LMs are quite resistant to forgetting their training data (an effect that increases with model size; Carlini et al., 2022; Vu et al., 2022; Ramasesh et al., 2022). While filtering out all undesirable content from pretraining data could seem to be a simple solution, it severely handicaps the capabilities of LMs (Welbl et al., 2021) which are already bottlenecked by high-quality data (Hoffmann et al., 2022; Villalobos et al., 2022). Moreover, reducing the diversity of training data can negatively impact alignment with human preferences by decreasing robustness (Hendrycks et al., 2019, 2020) and amplifying existing social biases (Xu et al., 2021; Welbl et al., 2021). These limitations suggest that while human preferences should be imposed in pretraining itself, content violating those preferences should still be present in the training data.
In this paper, we explore objectives for aligning LMs with human preferences during pretraining. Instead of filtering the training data, we propose pretraining with human feedback (PHF), where we estimate human preference judgments using a reward function (e.g. a toxic text classifier). In this way, we allow the LM to learn from undesirable content while guiding the LM not to imitate it at inference time. We experiment with four PHF objectives: conditional training (Keskar et al., 2019), dataset filtering, unlikelihood loss (Welleck et al., 2020) and two offline RL algorithms, reward-weighted regression (RWR; Peters & Schaal, 2007) and advantage-weighted regression (AWR; Peng et al., 2019). We compare them to maximum likelihood estimation (MLE), the standard pretraining objective.
We evaluate PHF objectives on three tasks: generating non-toxic text, text without personally identifiable information (PII), and PEP8-compliant Python (van Rossum et al., 2001). We compare LMs pretrained with feedback in terms of alignment (how well they satisfy preferences) and capabilities (how well they perform on downstream tasks). While different objectives offer different alignment–capabilities trade-offs for different tasks, we find that conditional training is on the Pareto frontier across all three tasks. Conditional training is a simple algorithm that learns a distribution over tokens conditional on their human preference score, reminiscent of decision transformer in reinforcement learning (Chen et al., 2021a). Conditional training decreases the frequency of undesirable content in LM samples up to an order of magnitude, reaping continued improvements with increasing training data (§4.1). Superior alignment persists when the LM is faced with an adversary prompting it to elicit undesirable behavior, as evaluated using the automated red-teaming approach from Perez et al. (2022) (§4.2). At the same time, conditional training achieves comparable performance to MLE-trained LMs on zero-shot benchmarks (Paperno et al., 2016; Chen et al., 2021b) and after finetuning on GLUE tasks (Wang et al., 2018) (§4.3); conditional training is able to learn representations from the entire training distribution, without learning to regurgitate undesirable content as MLE-trained LMs do.
Finally, in §5 we examine whether PHF improves over the standard practice of MLE pretraining followed by finetuning with human feedback. We find that PHF results in equal or (sometimes dramatically) better alignment across all three tasks (Fig. 1) as well as improved adversarial robustness. These findings results suggest that it is more effective to train LMs to exhibit desirable behaviors from the outset, rather than having them learn undesirable behavior and then attempt to unlearn it. Our results challenge the standard practice of aligning LMs with human preferences during finetuning alone, suggesting that should we incorporate human preferences from the very beginning of training.The code and datasets accompanying the paper are available at github.com/tomekkorbak/pretraining-with-human-feedback
Methods
Here we present five PHF objectives that we will evaluate in §4, in terms of various capabilities and alignment metrics for different tasks. In LM pretraining, we start with an LM with randomly initialized weights and an unlabeled dataset of documents . Each document is a sequence of segments (sentences or lines): . Each segment is a sequence of tokens: , where . Tokens come from a fixed vocabulary . In PHF, we additionally assume access to a segment-level reward function that takes a document segment and outputs a scalar score indicating how preferable is. For instance, could be the negative likelihood that a sentence would be harmful to civil conversation. At a high-level, pretraining can be posed as maximizing some pretraining objective across documents: . In the rest of the section we will describe MLE, the standard objective, followed by five PHF objectives.
Maximum likelihood estimation (MLE; Bengio et al., 2003; Mikolov & Zweig, 2012; Radford & Narasimhan, 2018; Brown et al., 2020) is the dominant approach to pretraining and finetuning LMs. This objective boils down to the log likelihood of training documents:
where can be decomposed autoregressively as
where denotes all segments in a document prior to and denotes all tokens in a document prior to .
MLE with Filtering
Dataset filtering (Solaiman & Dennison, 2021; Wang et al., 2022) corresponds to an objective identical to MLE except it is zero for documents such that their document-level reward is below a threshold :
is a hyperparameter we set to a certain percentile of document-level rewards in the training data (see Appendix A for values used in experiments and an ablation study). In practice, we train with this objective by discarding documents with rewards below and training for multiple epochs on the remaining ones at a fixed budget of training tokens.
Conditional Training
Conditional training (Ficler & Goldberg, 2017; Fan et al., 2018; Keskar et al., 2019) extends MLE by prepending documents with control tokens associated with properties of . It has been shown to be successful across tasks as diverse as as controllable language generation (Peng et al., 2018; Dai et al., 2019), mitigating toxicity (Gehman et al., 2020; Xu et al., 2020; Lu et al., 2022) and robotic control (Chen et al., 2021a; Janner et al., 2021). In contrast with prior work (e.g. Keskar et al., 2019), we found it to work substantially better when control tokens are prepended at a finer level of segments. Concretely, we prepend each segment with a control token based on that segment’s reward :
We use two control tokens: <|good|> if and <|bad|> otherwise. The threshold is a hyperparameter. At inference time, we sample from . See Appendix A for details.
Unlikelihood
Unlikelihood training (Welleck et al., 2020) follows MLE in maximizing the likelihoods of segments exceeding a certain reward threshold . However, for segments with rewards below the threshold, we use token-level unlikelihood instead. The unlikelihood of a token is the total log probability of all other tokens in the vocabulary on position of segment . This gives rise to the objective:
The threshold and , a coefficient scaling the second (unlikelihood) term, are hyperparameters.
RWR
Reward-weighted regression (RWR; Peters & Schaal, 2007) extends MLE by reweighting each segment by a term proportional to exponentiated reward:
, the coefficient controlling how much reward affects the loss, is a hyperparameter.
AWR
Advantage-weighted regression (AWR; Peng et al., 2019) extends RWR by subtracting a token-level value estimate from each segment-level reward . Value estimates are produced by a value function that shares parameters with the LM but is trained to minimize the mean-squared error between token-level value estimate and ground-truth returns . The LM and the value head are trained jointly to maximize:
where is the advantage. The two hyperparameters are (controlling the trade-off between value loss and policy loss) and (again, controlling the amount of reweighting). We implement the value function as a linear head on top of the LM ; they share the parameters of all other layers.
Experimental Setup
Here, we describe the setup of our pretraining (§4) and finetuning experiments (§5), which we use to compare MLE and various PHF objectives on both capabilities and alignment.
We evaluate PHF objectives on three tasks: (i) avoiding offensive content, (ii) avoiding leaking personally identifiable information (PII), and (iii) generating Python code following PEP8, the style guide for Python (van Rossum et al., 2001). Each task is associated with a reward function and a dataset as defined in §2. For evaluation, we use misalignment scores equal to the negative rewards.
LMs can generate highly harmful language, including insults, profanities and threats (Sap et al., 2019; Gehman et al., 2020; Abid et al., 2021). Following Welbl et al. (2021), we group these harms under the name of “toxicity,” understood as “a rude, disrespectful, or unreasonable comment that is somewhat likely to make you leave a discussion or give up on sharing your perspective” (Borkan et al., 2019). To obtain toxicity scores, we follow Askell et al. (2021) and use Detoxify (Hanu & Unitary team, 2020), a toxic comment classifier. We used the unbiased model, based on the 124M parameter RoBERTa (Liu et al., 2019) and trained on the Jigsaw Unintended Bias in Toxicity Classification dataset (Borkan et al., 2019). We define our reward as negative probability of toxicity according to Detoxify and misalignment score as the probability of toxicity. Since Detoxify was trained on short documents (predominantly comments), we first segment our training documents using a SpaCy (Honnibal et al., 2020) sentence segmenter and score them at sentence level. When scoring LM samples during evaluation, we skip segmentation.
PII
LMs sometimes generate text that occurs verbatim in their training data (Carlini et al., 2019; Perez et al., 2022). This poses privacy risks if the text contains confidential information identifying living people (PII) such as email addresses or social security numbers (Henderson et al., 2018). To detect such PII, we use Scrubadub,github.com/LeapBeyond/scrubadub a PII detector using both pattern matching rules and a pretrained SpaCy (Honnibal et al., 2020) named entity recognizer. We use pattern matching for detecting emails, addresses and postal codes, phone numbers, credit card numbers, US social security numbers, vehicle plates numbers, dates of birth, URLs and login credentials. The named entity recognizer detects mentions of people names, locations and organizations. We define our reward as the negative number of detected PII instances per character. Similarly to toxicity, we score training documents at sentence-level.
PEP8
While LMs are highly successful at generating code, the generated code is not always aligned with user intent (Chen et al., 2021b). For instance, prompted with low-quality code, LMs are likely to produce a low-quality completion even if user’s intent is to write high-quality code. We explore alignment failures in the context of code by requiring compliance with PEP8 (van Rossum et al., 2001), the style guide for Python. To detect PEP8 violations, we use pycodestyle, a popular static code analysis tool.github.com/PyCQA/pycodestyle Our reward function is the negative number of PEP8 violations per character. We assign rewards to individual lines of training documents, but note that the presence of PEP8 violations on a particular line does depend on previous lines.
2 Model Architecture and Hyperparamers
All of our LMs use the neural network architecture of gpt2-small (124M parameters; Radford et al., 2019). We keep the original hyperparameters of gpt2-small except for learning rate and batch size, which we tune for each task-objective pair based on train loss. If an objective has it own hyperparameters (e.g. , or ), we tune learning rate and batch size separately for each configuration considered and then chose the best configuration based on misalignment score of LM samples and the KL divergence from GPT-3 (§4.1). See Appendix A for hyperparameters used in experiments and ablations on them.
3 Training Data
We fixed training set size to 3.32B tokens which is compute-optimal for our model size according to the scaling laws from Hoffmann et al. (2022). For toxicity and PII, we prepared training data by subsampling 1.95M documents (totaling 3.32B tokens) from the Pile (Gao et al., 2020). For code generation, we subsampled 1.5M Python files (again totaling 3.32B tokens) from a cleaned and filtered version of the GitHub dataset from Google BigQuery released by Tunstall et al. (2022).GitHub on BigQuery
Pretraining Experiments
In this section, we investigate how PHF affects the alignment and capabilities of resulting models. In §4.1 we introduce two primary metrics: misalignment score (indicating how well unconditional samples from an LM satisfy human preferences) and the KL divergence from GPT3 (indicating general capabilities), and discuss the Pareto frontier of the capability-alignment trade-off. We additionally evaluate alignment by analyzing LM behavour when conditioned on adversarial prompts (“red-teaming”; §4.2) and evaluate capabilities by reporting performance on downstream tasks (§4.3). Finally, we measure diversity of LM samples (§4.4).
To estimate the frequency of undesirable content in text generated by an LM, we obtain a set of samples from it by nucleus sampling (Holtzman et al., 2020) with temperature and top-, constraining sequence length to be between 10 and 128 tokens. Unless specified otherwise, we generate unconditionally, i.e. only condition on a special <|endoftext|> token (or on <|endoftext|><|good|> when using conditional training). We then score those samples using the same scorers that had been used as reward functions during training. We report misalignment scores averaged across samples. In Appendix D, we also report metrics tracking the worst-case tail of misalignment score distribution.
KL from GPT-3
As a measure of an LM’s general capabilities, we estimate the Kullback-Leibler (KL) divergence of its output distribution from that of a highly capable model, GPT-3 (Brown et al., 2020). Lower divergence from GPT-3 likely translates into an increase in capabilities. We qualitatively found KL from GPT-3 to be sensitive to the most egregious failure modes of PHF, e.g., degeneration (Holtzman et al., 2020), repetition or reduced sample diversity. Note that KL from GPT-3 favors models trained like GPT-3, namely with MLE and without any alignment-relevant constraints; such constraints may cause the distribution to change in ways that do not impact a model’s performance on downstream tasks.
We estimate by computing , where are samples from GPT-3 obtained using its public APIopenai.com/api/ and is the LM being evaluated. We generate unbiased (temperature 1, top- 1) samples of at most 64 tokens, using <|endoftext|> as a stop token. To decrease variance due to the stochasticity of sampling we used the same set of samples for all evaluations. For toxicity and PII experiments, we use GPT-3 (175B; davinci) as . For PEP8, we use a 12B Codex model (code-cushman-001; Chen et al., 2021b). In prior experiments, we found that using InstructGPT (textdavinci-002; Ouyang et al., 2022) as a target distribution gives very similar results.
Results
We present our main results in Fig. 2. All PHF objectives are able to reduce the amount of undesirable content significantly, sometimes by an order of magnitude. For instance, on toxicity the average misalignment score of an MLE LM reaches 0.0141; conditional pretraining instead reaches 0.0011. These order-of-magnitude drops persist for metrics tracking the right tail of the misalignment score distribution (worst case), see Figs. 14-13(b) in Appendix D. Conditional training shifts the right tail furthest left (Fig. 14). Moreover, for conditional training and filtering, the misalignment score decreases consistently through training time, with no clear signs of a plateau. This scaling behavior suggests that increasing training set size further would lead to even lower scores.
Among PHF objectives, conditional training offers the best trade-off between misalignment score reduction and KL overhead. It is strictly Pareto-optimal in toxicity (leftmost and bottommost in Fig. 2, first column, first row) and on the Pareto frontier in PII and PEP8. It is also the only PHF method that is always on the Pareto frontier across all three tasks. In terms of score, it is only outperformed (by filtering) on PEP8. Filtering turns out to be a strong baseline; it is either second-best or best in terms of alignment. However, on two out of three tasks (PII and PEP8) it pays a significant capabilities penalty (the largest among all methods). RWR and AWR tend to obtain similar, rather poor, performance. They improve upon MLE’s misalignment score only slightly, while reducing capabilities significantly compared to MLE. Finally, the success of unlikelihood training is highly task-dependent; it reduces the misalignment score significantly for toxicity but only slightly for PII and PEP8.
2 Robustness to Red-Teaming
In addition to measuring how aligned our LMs are for unconditional generation, we also study their responses to prompts chosen by an adversary. The adversary tries to elicit misaligned behavior of the target LM , a procedure known as “red-teaming” (Perez et al., 2022). We use prompted InstructGPT (text-davinci-002; Ouyang et al., 2022) to simulate an adversary, extending the stochastic few-shot generation approach to red-teaming introduced by Perez et al. (2022). We start with an initial pool of human-written adversarial prompts and iteratively apply the following steps:
Assign each new adversarial prompt with for , where is the target LM.
Sample adversarial prompts from the pool, , with weights proportional to .
Instruct InstructGPT to generate text likely to elicit a particular alignment failure (offensive reply, leaking PII or violating PEP8). In addition to the instruction, InstructGPT is provided with as few shot examples. We sample independent completions and add them to the pool .
We repeat steps (1)-(3) for ten rounds. For each model and each task, we conduct ten separate trials of the procedure. We report average and standard deviation across ten trials. For more details, see Appendix B.
Results
We show the average misalignment score of all adversarial prompts in the pool, , throughout ten rounds of red-teaming in Fig. 3 (see also Figs. 11-11 in Appendix B for other metrics). The main trend is consistent with misalignment scores from §4.1: conditional training and filtering are the most robust objectives in terms of their their final misalignment scores. On toxicity and PII even after ten rounds of red-teaming conditional training outperforms MLE by up to an order of magnitude. Unlikelihood’s performance is heavily task-dependent; it is the most robust method (by a wide margin) for toxicity while being the least robust for PII. We verified that its unsually high robustness on toxicity persists when, instead of actively red-teaming, we compute misalignment scores for generation conditioned on a fixed set of challenging RealToxicityPrompts (Gehman et al., 2020), see Fig. 13(c) in Appendix D. Overall, all LMs pretrained with feedback (except for unlikelihood-trained LM in PII) are significantly more robust to adversaries than MLE-trained LMs.
On the other hand, all PHF objectives leave LMs with vulnerabilities that an adversary with black box access can exploit. For all PHF objectives, subsequent iterations of red-teaming increase the average score of target LM responses, with no clear plateau even after 10 iterations. This result highlight the limitations of PHF; while it results in LMs significantly more robust than after MLE pretraining, the resulting LMs are not completely aligned or safe in all deployment scenarios.
3 Downstream Benchmarks
We supplement KL from GPT-3 as a measure of LM capabilities, by measuring the performance of trained models on tasks without additional training or examples (zero-shot). We choose tasks for which a 124M parameter MLE-trained LMs should be able to achieve non-trivial performance. For toxicity and PII, we evaluate models on LAMBADA (Paperno et al., 2016), a passage understanding task that evaluates an LM’s accuracy and perplexity at predicting the final word in a passage. For PEP8, we report pass@10 and pass@100 on HumanEval (Chen et al., 2021b) which tasks models with generating code to solve a given problem, and evaluates the correctness of the generated code using test cases.
GLUE
We also study the performance of PHF-trained LMs on various natural language understanding tasks, after finetuning on those tasks. In this way, we evaluate the effectiveness of various pretraining objectives at representation learning. In contrast with metrics from previous subsections, this kind of evaluation does not involve any generation; it tests PHF affects representations acquired during pretraining rather than how it affects the distribution over LM outputs. Here, we use the GLUE benchmark (Wang et al., 2018), a suite of text classification tasks related to question answering, sentiment analysis and recognizing textual entailment, among others. We conduct single-model single-task evaluation, i.e. to evaluate a given pretrained LM, we finetune it on the training set of each GLUE task separately and report test set scores averaged across tasks. To control for the variance of results, we restart each finetuning three times and report standard deviation of scores as error bars. We omit GLUE evaluation for PEP8 models because they are trained on code rather than natural language (used in GLUE tasks). See Appendix C for details.
Results
We present the results of zero-shot evaluation in Fig. 4. Conditional training slightly exceeds MLE’s performance in terms of accuracy on both tasks. Other PHF objectives suffer from decreased accuracy, especially for toxicity. Unlikelihood also matches MLE accuracy, but only for PII; it obtains very low accuracy on toxicity (recall that we found similar task-sensitivity in §4.1 and §4.2). GLUE results paint a similar picture; conditional training most closely matches MLE scores. The second-best objective using feedback is Filtering (on toxicity) or unlikelihood (on PII). For results on individual GLUE tasks, see Appendix C. Finally, on HumanEval, the capabilities gap between MLE and PHF methods is wider. This gap is only closed – in terms of pass@100 – by filtering. Conditional training is no longer the best PHF method; it is outperformed or matched by filtering, AWR and RWR. Unlikelihood consistently obtains the lowest scores.
4 Diversity
Constraining an LM to be aligned with human preferences can result in decreased entropy or increased degeneration of LM samples (Korbak et al., 2022b), e.g. due to repeated tokens (Holtzman et al., 2020). To control for this, we supplement our capabilities evaluation with an examination of the diversity and rate of degeneration of LM samples. We measure diversity in terms of entropy over unigrams expected in a set of LM samples and degeneration in terms of the ratio of all unigrams and distinct unigrams within an average sample (Li et al., 2016). In Appendix E we also report Self-BLEU-5, a measure of text diversity across samples (Zhu et al., 2018), bigram entropy and fraction of distinct bigrams.
Results
The results for toxicity and PII, shown on Fig. 5, reveal two patterns of behavior. Unlikelihood, AWR and RWR tend to match MLE diversity but suffer from slightly increased degeneration. Conditional training and, to a degree, filtering, show the reverse trend; decreased diversity but more closely matching MLE’s fraction of distinct unigrams. In absolute terms, however, none of the PHF objectives cause significant degeneration or entropy collapse.
Finetuning with Human Feedback
As discussed in §1, the standard approach to aligning LMs with human preferences involves pretraining an LM using MLE and finetuning it using an objective involving human feedback, e.g., RL with KL penalties (Ziegler et al., 2019; Ouyang et al., 2022) or supervised finetuning (Solaiman & Dennison, 2021; Chung et al., 2022). In this section, we compare PHF to supervised finetuning with human feedback using PHF objectives, but only after MLE pretraining.We also experimented with finetuning using RL with KL penalties, but decided to exclude these experiments because we did not obtain results competitive with supervised finetuning. We are also interested in understanding whether pretraining with MLE and then finetuning with feedback is better than using PHF from scratch. To address this question, we compare finetuning runs against PHF with conditional training, the PHF objective we identified as the best in §4.
To ensure comparability, we use checkpoints of MLE runs from §4 trained either 50% of the training data (i.e. 1.66B tokens) or 90% of the training data (i.e. 2.97B tokens). We then continue finetuning them for another 1.66B or 300M tokens, respectively, using each of five objectives using feedback.It is worth noting that the fraction of the training budget we allocate to finetuning (50% or 10%) is already very high (e.g. compared to 1.6%-0.2% in (Chung et al., 2022) or 0.1% in (Tay et al., 2022)). This experiment design allows us to interpolate between pretraining and finetuning. We conduct separate hyperparameter sweeps over learning rate and batch size for each task and finetuning objective. Following standard practice for finetuning a pretrained model, we reset the learning rate schedule used during pretraining. Our setup is otherwise identical to that from §4, e.g., finetuning runs use the same order and batches of training data as pretraining runs from §4.
Results
We present the comparison of PHF and finetuning with human feedback in Fig. 6. PHF achieves scores that are always better, typically dramatically better, than finetuning with feedback. On toxicity and PII there is a significant gap between pretraining using conditional training and the best finetuning objective. For instance, in PII, aligning the LM during pretraining is two to three times more effective than finetuning on 300M tokens; conditional pretraining converges to misalignment score 0.0013 compared to 0.0018 (finetuning on 1.6B tokens) and 0.0023 (finetuning on 3.3B tokens). The gap between PHF and finetuning with feedback only widens as fewer tokens are available for finetuning (dashed vs dotted line in Fig. 6).
The size of this gap and its persistence across two tasks provides evidence that PHF is more effective than MLE pretraining followed by finetuning with feedback. We also present a head-to-head comparison of pretraining and finetuning performance of each objective on Fig. 17 in Appendix F; we find that the improvement from PHF over only finetuning with feedback tends to increase with how effective the PHF objective is at reducing scores in general. Cconditional training works well for both pretraining and finetuning (see Fig. 17 for a direct comparison with capabilities-alignment of trade-offs of all objectives during finetuning for 1.6B tokens).
Finally, we repeated the red-teaming procedure from §4.2 to compare adversarial robustness of LMs pretrained with conditional training and LMs only finetuned with conditional training (Fig. 7). Once again, low misalignment scores from unconditional sampling indicates increased robustness, and we found LMs pretrained with human feedback to be significantly more robust to red-teaming (on toxicity and PII). For instance, on PII, ten rounds of red-teaming of PHF-trained LMs are required to reach the misalignemnt score that a finetuned LM has just after one iteration. Overall, our findings demonstrate that alignment of an LM is closely tied to the quantity of human feedback it receives during training. Involving human feedback throughout the entire pretraining process (as in PHF) results in substantially better alignment than the standard practice of incorporating feedback for only a small portion of the training budget.
Related Work
In this paper, we tackled the problem of training an LM on (potentially undesirable) content annotated with feedback while constraining the LM not to imitate undesirable content at inference time. This setting is closely related to offline RL which addresses training an optimal policy on (possibly suboptimal) demonstrations annotated with rewards (Levine et al., 2020). Most work in offline RL has focused on pretraining policies for robotic control environments (Nair et al., 2020; Kumar et al., 2020; Emmons et al., 2022). However, offline RL techniques were recently used for finetuning pretrained LMs to be aligned with human preferences in dialog tasks (Jaques et al., 2020; Jang et al., 2022; Snell et al., 2022). Conditional training has recently emerged as an effective apporoach to offline RL (Schmidhuber, 2019; Kumar et al., 2019) and demonstrated strong results when paired with transformers (Chen et al., 2021a; Janner et al., 2021). For instance, decision transformer (Chen et al., 2021a) consists of training a sequence model on (reward, state, action) pairs and, at inference time, sampling an action conditioned on high reward. This approach mirrors our conditional training approach: training an LM on (control token, sentence) pairs and, at inference time, sampling tokens when conditioned on an <|good|> control token.
LM alignment during finetuning
While we focus on pretraining, aligning LMs is frequently approached through finetuning an MLE-pretrained LM. In addition to RLHF (Ziegler et al., 2019), alternative finetuning objectives included divergence from a target distribution (Khalifa et al., 2021; Korbak et al., 2022a; Go et al., 2023; Chen et al., 2023) or supervised finetuning on data generated by other LMs (Scheurer et al., 2022) or highly curated collections of tasks phrased as instructions (Sanh et al., 2022; Chung et al., 2022). For instance, instruction finetuning (Chung et al., 2022) improves usability and mitigates some potential harms (such as toxic responses or gender bias), suggesting that augmenting LM training distribution with demonstrations can have effects similar to finetuning for instruction-following using RLHF.
Conclusion
In the paper, we challenged the practice of aligning LMs during finetuning and advocated for utilizing human feedback during pretraining itself. Out of five PHF objectives we evaluated, conditional training consistently outperforms the alternatives in terms of both capabilities and alignment (with two notable exceptions: unlikelihood is more robust to red-teaming on toxicity and filtering achieves better HumanEval results). The fact that conditional training tends to match MLE’s capabilities while enjoying much better alignment corroborates previous findings (Bai et al., 2022) that alignment and capabilities might not be at odds with each other on many tasks of practical importance. While PHF requires additional overhead of annotating the training data with a reward model, the computational cost of reward model inference is low compared to the total pretraining cost. This is because the reward model (i) can be much significantly than the LM being pretrained (reducing its size doesn’t hurt performance much in RLHF experiments, see Bai et al., 2022) and (ii) optimized for efficient inference using techniques such as distillation (Tang et al., 2019) or very low-bit precision (e.g., 4-bit; Dettmers & Zettlemoyer, 2023). Moreover, recent follow-up work obtained good results for toxicity by including control tokens for only a fraction of the pretraining data (Anil et al., 2023). Overall, incorporating human preferences in pretraining leads to capable models that generate text more aligned with human preferences, even under adversarial attacks.
Acknowledgments
We are grateful to Adam Gleave, Ajeya Cotra, Alex Havrilla, Andy Jones, Asa Cooper Stickland, Beth Barnes, Charlie Snell, Claudia Shi, Daniel Ziegler, David Dohan, David Krueger, David Lindner, Euan McLean, Evan Hubinger, Ian McKenzie, Jérémy Scheurer, Kath Lupante, Kyle McDonell, Laria Reynolds, Leo Gao, Łukasz Kuciński, Michael Janner, Piotr Miłoś, Sean Welleck, Scott Emmons, and Xiang Pan for helpful conversations and feedback. Tomasz Korbak was supported by the Leverhulme Doctoral Scholarship and Open Philantropy. Angelica Chen was supported by the National Science Foundation Award no. 1922658. Sam Bowman was supported by Eric and Wendy Schmidt (by recommendation of the Schmidt Futures program), Open Philanthropy, Apple, and the National Science Foundation under Grant Nos. 1922658 and 2046556. Ethan Perez was supported by the National Science Foundation and Open Philanthropy. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author and do not necessarily reflect the views of the National Science Foundation. We also thank NYU HPC Center for providing access to computational resources and OpenAI for providing access and credits to their models via the API Academic Access Program.
References
Appendix A Hyperparameters and Implementation Details
We implement conditional training by prepending control tokens <|good|> (if ) and <|bad|> to segments (sentences or lines) in training documents. However, we do not prepend them at random to 1% of sentences. We found this intervention to slightly improve capabilities (measured in terms of KL from GPT-3) while incurring a negligible alignment penalty. We conjecture the capabilities penalty is due to the fact that text generated by GPT-3, not containing special tokens, is out-of-distribution for an LM trained with conditional training. Exposing the LM to sentences not prepended with special tokens likely alleviates this problem.
When generating unconditionally from the LM, we condition it only on <|endoftext|><|good|>. For toxicity and PII, we also block both special tokens (<|good|> and <|bad|>) by setting their probability to zero. For PEP8, we only block the <|bad|> token, allowing <|good|> tokens to be generated before each new line; instead, we remove them in a post-processing step. Similarly, during sampling as part of HumanEval evaluation, we use the <|good|> as a prefix and block <|bad|> and <|good|> for evaluation.
When evaluating KL from GPT-3, we measure it against a conditional distribution . We implement that by prepending samples from GPT-3 with a special token <|good|>. For PEP8, we additionally insert a infix <|good|> between each line generated by Codex.
In our finetuning experiments, conditional training requires extending the vocabulary of a pretrained LM. To minimize the effect of distribution shift, we follow Hewitt (2021) and initialize the embeddings of <|good|> and <|bad|> to the mean of the remaining embeddings plus a small amount () of Gaussian noise. Despite this intervention, a notable drop in alignment and capabilities can still be seen for the first 100m tokens after we start finetuning with new tokens, see Fig. 17 in Appendix F.
Hyperparameters
As discussed in §3, we keep the original hyperparameters of gpt2-small except for learning rate and batch size. We tune learning rate and batch size for each task-objective pair based on train loss. If an objective has it own hyperparameters (e.g. , or ), we first tune learning rate and batch size for each configuration considered and then chose the best configuration based on misalignment score of LM samples and KL from GPT-3 (§4.1). We swept over a fixed set of learning rates and batch sizes, the same for each task-objective pair. See Fig. 8 for an ablation study showing the effect of threshold on capabilities-alignment trade-off in conditional training and filtering. We report hyperparameters we used in our experiments in Tables 1-3.
Appendix B Details on the red-teaming procedure
We use InstructGPT text-davinci-002Model index for researchers, via the API, as the red LM that few-shot-generates adversarial prompts. After the red LM is given a task specific-instruction (see Tab. 4), we sample from it with temperature and top-. We set the number of few-shot examples and the number of adversarial prompts sampled from red LM . These hyperparameters were tuned empirically to maximize misalignment score of MLE-trained model’s responses.
Target LMs
We sample from target LMs () with temperature and top-, consistently with unconditional generation results. We additionally require the length of generated responses to be between 10 and 64 tokens. We set the number of completions per prompt to . When generating from a target LM trained with conditional training, we condition it first on a <|good|> control token, then on an adversarial prompt and generate a response blocking the bad (i.e. setting its probability to zero at each step of generation).
Scoring
We use the same setup for scoring LM samples as for scoring unconditional samples, described in Sec. 3. We only measure the misalignment score of target LM’s response except for PEP8 where we measure the score of the prompt concatenated with the response.
Prompt pool
For toxicity, we bootstrap the prompt pool with prompts from the “challenging” subset of RealToxicityPrompts (Gehman et al., 2020). For PII and PEP8, we bootstrap the pool using hand-written prompts. For toxicity, the temperature of sampling from the adversarial prompt pool is ; for PII and PEP8, .
Metrics
To measure target LM’s robustness to red-teaming, we track the following metrics over ten rounds:
average misalignment score of adversarial prompts found in the pool, at the end of the procedure (Fig. 3),
average misalignment score of adversarial prompt generated in a given round (Fig. 11),
misalignment score of the best adversarial found in the pool, at the end of the procedure (Fig. 11),
misalignment score of the best adversarial prompt generated in a given round (Fig. 11),
Reported results are averages from ten independent trials, each involving ten rounds of red-teaming. We also present examples of best adversarial prompts for each task in Tables 5-7.
Appendix C Details on GLUE evaluation
We select eight tasks from the GLUE benchmark (Wang et al., 2018): CoLA (Warstadt et al., 2018), SST-2 (Socher et al., 2013), MRPC (Dolan & Brockett, 2005), STS-B (Cer et al., 2017), QQP,quoradata.quora.com/First-Quora-Dataset-Release-Question-Pairs MNLI (Williams et al., 2018), QNLI (Rajpurkar et al., 2016), and RTE (Dagan et al., 2005; Bar-Haim et al., 2006; Giampiccolo et al., 2007; Bentivogli et al., 2009). Following prior work (Devlin et al., 2019), we drop one GLUE task from our evaluation: WNLI (Levesque, 2011). We directly finetune each our our pretrained LMs for toxicity and PII on each of the eight selected GLUE tasks and report test set performance. Due to domain mismatch, we leave out LMs we pretrained for PEP8. To use our LMs for classifcation and regression tasks, we add sequence classification heads on top of them, and we set the number of output labels correspondingly for each task.
Training
We sweep hyperparameters for each GLUE task based on toxicity MLE-pretrained LM’s dev set scores. We sweep across learning rates {5e-4,1e-4,5e-5,2e-5} and batch sizes {32,64,128}. We then transfer the optimal task configurations to all other runs. We train each LM for each GLUE task for a maximum of 6 epochs with early stopping based on dev scores. To account for variance, we conduct 3 random restarts for each experiment. Other hyper-parameters follow the default settings in a script provided by (Wolf et al., 2020).https://github.com/huggingface/transformers/blob/main/examples/pytorch/text-classification/run_glue.py
Results
For STS-B task, we clip the predicted scalars to range to satisfy GLUE leaderboard submission format. We obtain test set performance and aggregate the results. For tasks with two metrics (for example, F1 and accuracy), we take the average of two. We average the accuracy of MNLI-matched and MNLI-mismatched test set and report them as MNLI. We then average scores across three random seeds (restarts of the finetuning) and report average scores (and their standard deviations) in Table 8 and Table 9. As baselines, in Table 10 we also report the performance of OpenAI-pretrained GPT-2 (gpt2-small from HuggingFace Hub; Radford et al., 2019) and a randomly initialized GPT-2 model trained from scratch for GLUE tasks. Hyperparameters for these baselines we were tuned separately.
Appendix D Additional results on scores of LM samples
MLE Conditional Filtering Unlikelihood RWR AWR
Appendix E Additional results for diversity evaluation
Appendix F Additional results for finetuning experiments
Pretraining Finetuning from MLE for 1.6B tokens Finetuning from MLE for 300M tokens