Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, Yejin Choi

Introduction

The ultimate aim of language technology is to interact with humans. However, most language models are trained without direct signals of human preference, with supervised target strings serving as (a sometimes crude) proxy. One option to incorporate user feedback is via human-in-the-loop, i.e., a user would be expected to provide feedback for each sample online as the model trains, but this degree of dense supervision is often prohibitive and inefficient. Automated metrics offer a promising compromise: models of human preference like pairwise learned preference models (Ouyang et al., 2022), BERTScore (Zhang et al., 2019), BLEURT (Sellam et al., 2020) have significantly improved correlation with human judgment compared to earlier metrics (BLEU, METEOR, etc.), and are cheap to evaluate. But — these functions are usually not per-token differentiable: like humans, metrics can only offer quality estimates for full generations. Reinforcement Learning (RL) offers a natural path forward for optimizing non-differentiable, scalar objectives for LM-based generation when it is cast as a sequential decision-making problem. However, Goodhart’s LawStrathern (1997) paraphrases: When a measure becomes a target, it ceases to be a good measure. looms: particularly in the case of imperfect metrics that use neural networks, it is easy to find nonsense samples that achieve high-quality estimates. Recent works have shown promising results in aligning LMs to human preferences via RL by constraining preference-based rewards to incorporate notions of fluency (Wu et al., 2021a; Ouyang et al., 2022) but progress in this line of work is heavily hindered by a lack of open-source benchmarks and algorithmic implementations—resulting in perception that RL is a challenging paradigm for NLP (Choshen et al., 2020; Kreutzer et al., 2021).

To facilitate research in building RL algorithms to better align LMs, we release a library, a benchmark, and an algorithm. First, we release the RL4LMs library, which enables generative HuggingFace models (e.g., GPT-2 or T5) to be trained using a variety of existing RL methods like PPO/A2C/etc. Next, we apply models trained using RL4LMs to the new GRUE (General Reinforced-language Understanding Evaluation) benchmark: GRUE is a collection of 7 contemporary NLP tasks (see Table1 for details); in contrast to other benchmarks, instead of supervised training, we pair each task with reward function(s). GRUE challenges models to optimize these reward functions while remaining fluent language generators. We train language models via RL—both with and without task specific supervised pre-training—to optimize rewards. Finally, beyond existing RL methods, we introduce a novel on-policy RL algorithm called NLPO (Natural Language Policy Optimization), that dynamically learns task-specific constraints over the distribution of language at a token level.

Experiments on GRUE and human evaluations show that NLPO better balances learning preference rewards while maintaining language fluency compared to alternatives, including PPO (Figure 1). We find that using RL to learn from scalar reward feedback can be more: (1) data efficient than using additional expert demonstrations via supervised learning (though a combination of both is best)—a learned reward function enables greater performance when used as a signal for an RL method than a supervised method trained with 5 times more data, and (2) parameter efficient—enabling a 220 million parameter model trained with a combination of supervision and NLPO to outperform a 3 billion supervised model. We hope that the benchmarks, baselines, and building blocks we release serve to drive forward research in aligning LMs to human preferences.

Related Work

Imitation learning for NLP. Algorithms such as Schedule Sampling (SS) (Bengio et al., 2015), Parallel SS (Duckworth et al., 2019), SS for Transformers (Mihaylova & Martins, 2019), Diffential SS (Goyal et al., 2017), LOLS (Lampouras & Vlachos, 2016; Chang et al., 2015), TextGAIL (Wu et al., 2021b), and SEARNN (Leblond et al., 2017), have been inspired by DAGGER (Ross et al., 2011) and SEARN (Daumé et al., 2009). However, these algorithms are known to suffer from exposure bias in generation (Chiang & Chen, 2021; Arora et al., 2022) and the cliff MDP problem (Huszár, 2015; Agarwal et al., 2019; Swamy et al., 2021).

RL for Large Action Spaces. MIXER (Ranzato et al., 2016) combined ideas from schedule sampling and REINFORCE (Williams, 1992). Bahdanau et al. (2016) proposed an actor-critic algorithm to address the variance/large action space problems when using REINFORCE for language generation; follow-up works such as KG-A2C (Ammanabrolu & Hausknecht, 2020), TrufLL (Martin et al., 2022), AE-DQN (Zahavy et al., 2018), and GALAD (Ammanabrolu et al., 2022) addressed similar issues by attempting to eliminate and reduce the action space during exploration.

RL for NLP. RL, often in the form of bandit learning, has been used to improve models in machine translation (Wu et al., 2016; Nguyen et al., 2017; Kiegeland & Kreutzer, 2021), summarization (Stiennon et al., 2020; Paulus et al., 2017), dialogue (Li et al., 2016; Zhou et al., 2017; Jaques et al., 2020), image captioning (Rennie et al., 2017), question generation (Pang & He, 2021), text-games (Narasimhan et al., 2015; Hausknecht et al., 2020), and more (Ranzato et al., 2016; Snell et al., 2022). Lu et al. (2022) adapt reward-conditioned transformers (Chen et al., 2021) for several language generation tasks. RL has been the focus of efforts to align LMs with human preferences (Stiennon et al., 2020; Wu et al., 2021a; Nakano et al., 2021; Ziegler et al., 2019), e.g., Ouyang et al. (2022) fine-tuned large language model with PPO Schulman et al. (2017) to align with models of human preference, but their non-public dataset doesn’t enable comparison. Though RL has been successful in some of the use cases described above, it has simultaneously been critiqued for being significantly less stable than supervised LM training (Choshen et al., 2020). As a result, there is relatively little consensus if RL is a worthwhile consideration for training LMs compared to, say, collecting additional supervised data.

RL4LMs: A Library for Training LMs with RL

We introduce RL4LMs, an open-source library with building blocks for fine-tuning and evaluating RL algorithms on LM-based generation. The library is built on HuggingFace (Wolf et al., 2020) and stable-baselines-3 (Raffin et al., 2021), combining important components from their interfaces. RL4LMs can be used to train any decoder only or encoder-decoder transformer models from HuggingFace with any on-policy RL algorithm from stable-baselines-3. Furthermore, we provide reliable implementations of popular on-policy RL algorithms that are tailored for LM fine-tuning such as PPO (Schulman et al., 2017), TRPO (Schulman et al., 2015a), A2C (Mnih et al., 2016), and our own NLPO (§4). The library is modular, which enables users to plug-in customized environments, reward functions, metrics, and algorithms. In the initial release, we provide support for 6 different NLP tasks, 16 evaluation metrics and rewards, and 4 RL algorithms.

2 Reward Functions and Evaluation Metrics

Because RL4LMs provides a generic interface for per-token or per-sequence generation rewards, it is possible to quickly apply a wide array of RL algorithms to a similarly diverse range of textual metrics-as-rewards. Specifically, we provide interfaces to 1) n-gram overlap metrics metrics such as ROUGE (Lin, 2004), BLEU (Papineni et al., 2002), SacreBLEU (Post, 2018), METEOR (Banerjee & Lavie, 2005); (2) model-based semantic metrics such as BertScore (Zhang et al., 2019) and BLEURT (Sellam et al., 2020) which generally provide higher correlation with human judgment; 3) task-specific metrics such as CIDER (Vedantam et al., 2015), SPICE (Anderson et al., 2016) (for captioning/commonsense generation), PARENT (Dhingra et al., 2019) (for data-to-text) and SummaCZS (Laban et al., 2022) (for factuality of summarization); 4) diversity/fluency/naturalness metrics such as perplexity, Mean Segmented Type Token Ratio (MSSTR) (Johnson, 1944), Shannon entropy over unigrams and bigrams (Shannon, 1948), the ratio of distinct n-grams over the total number of n-grams (Distinct-1, Distinct-2) and count of n-grams that appear only once in the entire generated text (Li et al., 2015); 5) task-specific, model-based human preference metrics such as classifiers trained on human preference data collected in the methodology of Ouyang et al. (2022).

3 On-policy Actor-critic Algorithms

Given an input-output pair (x,y)({\bm{x}},{\bm{y}}) and generation predictions from our agent; because the environment rewards are sequence-level and sparse, following Wu et al. (2021a) we regularize the reward function using a token-level KL penalty for all on-policy algorithms, to prevent the model from deviating too far from the initialized LM π0\pi_{0}. Formally, the regularized reward function is:

where R^\hat{R} is the regularized KL reward, y{\bm{y}} is gold-truth predictions, KL(πθ(at∣st)∣∣π0(at∣st))=(log⁡π0(at∣st)−log⁡πθ(at∣st))\text{KL}\left(\pi_{\theta}(a_{t}|\bm{s}_{t})||\pi_{0}(a_{t}|\bm{s}_{t})\right)=(\log\pi_{0}(a_{t}|\bm{s}_{t})-\log\pi_{\theta}(a_{t}|\bm{s}_{t})) and the KL coefficient β\beta is dynamically adapted (Ziegler et al., 2019). Further details on actor-critic methods can be found in Appendix A.

NLPO: Natural Language Policy Optimization

Language generation action spaces are orders of magnitude larger than what most discrete action space RL algorithms are designed for (Ranzato et al., 2016; Ammanabrolu, 2021), e.g., GPT-2/3 and T5 have a vocabulary size of 50K and 32K respectively. We hypothesize that the size of the action space is a core cause of instability when training LMs with existing RL methods. To address this issue, we introduce NLPO (Natural Language Policy Optimization), which is inspired by work on action elimination/invalid-action masking (Zahavy et al., 2018; Huang & Ontañón, 2020; Ammanabrolu & Hausknecht, 2020). NLPO, a parameterized-masked extension of PPO, learns to mask out less relevant tokens in-context as it trains. NLPO accomplishes this via top-pp sampling, which restricts tokens to the smallest possible set whose cumulative probability is greater than the probability parameter pp (Holtzman et al., 2018).

Specifically, NLPO maintains a masking policy πψ\pi_{\psi}: the masking policy is a copy of the current policy (πθ\pi_{\theta}), but is updated only every μ\mu steps. A parameterized-invalid-mask is created from πψ\pi_{\psi} by first selecting the top-pp tokens from the vocabulary, πψ\pi_{\psi} could be trained with alternate sampling techniques like top-kk or beam search (or even hard-coded via rules by domain experts), though we find top-pp sampling to be most effective in practice. and then applying an invalid-mask to the remaining tokens—i.e. setting their probabilities to zero when sampling actions from πθ\pi_{\theta} during training; this periodic updating policy πψ\pi_{\psi} is inspired by off-policy Q-learning algorithms (Andrychowicz et al., 2017), providing the policy πθ\pi_{\theta} with an additional constraint that balances between the benefits of containing more task relevant information than the KL penalty derived from π0\pi_{0} and the risk of reward hacking. We provide pseudocode in Algorithm 1 (green portions highlight the differences with PPO).

GRUE (General Reinforced-language Understanding Eval)

GRUE is a collection of 7 generative NLP tasks. To combat reward hacking for any single metric, each task is evaluated at test time according to a task-specific mix of metrics, detailed in Table 1. The metrics span two categories. Task preference metrics capture how well the models produce generations that satisfy the desiderata of the specific generation task, e.g., for Commongen, if the generations contain all the required words, or for IMDB, how positive the generated completions are. Naturalness metrics capture fluency, readability, etc. and provide perspective on factors beyond semantics. At training time, there are no special restrictions: models are free to use the supervised data, compute metrics on intermediate generations, etc. Train/val/test splits follow the original works. All results are averaged over multiple seeds, with exact counts being found in Appendix B.

Experimental Setup. We use RL4LMs to test a large range of algorithms on the GRUE benchmark. Specifically: We compare 3 algorithms for direct fine-tuning — Supervised, PPO,We consider PPO representative of the present state-of-the-art — in particular, we do not consider the popular REINFORCE (Willianms, 1988; Williams, 1992), as recent works have shown PPO to be strictly superior to REINFORCE in multiple domains (Schulman et al., 2017) and NLPO. In addition, we consider a hybrid approach of supervised learning and our RL methods by applying PPO and NLPO on checkpoints that have been fine-tuned in a supervised fashion—we call these Supervised+PPO, Supervised+NLPO. As an additional baseline, we additionally run zero-shot evaluations where we design prompts which aim to elicit task-specific generations, but with no training data or parameter updates.

For each task, to isolate the effect of training method, we select a single pre-trained LM backbone. For IMDB text continuation we use GPT-2 (117m parameters), and for the rest of the tasks we use T5-base (220m parameters). For our RL models (PPO, NLPO, Supervised+PPO, Supervised+NLPO), for a thorough investigation of how reward-hacking might interplay with GRUE, we run a separate set of experiments optimizing multiple task rewards for each task independently, e.g., for Commongen which has 6 task rewards (CIDER, ROUGE-2, ROUGE-L, BLEU-3, BLEU-4, METEOR) we run 6 different experiments optimizing each metric independently and report all possible metrics seen in Table 1 regardless of which individual metric was being optimized for.

Human Participant Study. We gather human judgments for five of the tasks in GRUE. In doing so, our goals are 1) to validate that the automated metrics we selected for GRUE correlate with human judgments with respect to relative ranking between models; and 2) to provide additional empirical comparisons regarding NLPO vs. PPO, ablations to study the effects of the KL naturalness penalty, etc. We specifically consider IMDB, Commongen, ToTTo, DailyDialog, and CNN Daily Mail. For each individual sample in a task, we ask 3 unique human raters to provide Likert judgments of 1) quality, i.e., for the specific task, how correct/appropriate is the generation, given the context, and 2) fluency, i.e., how well-written is the generation. We used Amazon Mechanical Turk, and paid crowdworkers a minimum of $15/hr. More details, including qualification information, interface screenshots, instructions, etc. are given in the corresponding Appendicies.

Figures 2(a), 2(b) present the results on GRUE, split into task metrics and naturalness metrics, and Tables 5, 3 highlight key results via ablation studies. Full results are available in Appendix B. For text continuation and summarization, with non-trivial zero-shot performance, RL tends to perform better than supervised training, but for tasks like Commongen and ToTTo, which have very low zero-shot performance, supervised training performs best—with both approaches outperforming zero-shot.

However, using RL+Supervised learning in conjunction works best; NLPO+supervised and PPO+supervised usually always outperforms NLPO/PPO (or supervised in isolation) across both task metrics and naturalness metrics. Supervised warm-starting is particularly effective for Commongen and ToTTo, which our results suggest are more prone to reward hacking. The one exception to this trend is DailyDialog where the RL models outperform warm-started Supervised+RL models likely due to the low performance of the Supervised models. We note that Supervised+NLPO using a T5-base (220m parameter) LM currently outperforms all the models on the ToTTo leaderboard, many of which have ≥\geq 3b parameter supervised models—suggesting that RL is parameter efficient as well. In these cases, it is critical that the initial policy already contain (some) signal for the task due to it being used as a KL constraint and masking constraint in NLPO. If the mask contains no initial priors about task specific language, it will be eliminating the wrong actions—a better initial policy leads to better RL performance downstream.

Human agreement with automated metrics. As human judgments can be noisy, we run additional statistical analysis such as measuring inter-annotator agreement, via Krippendorf’s alpha score, and using a one-way ANOVA followed by a post-hoc Tukey HSD test to measure if differences in means of average scores between pairs of models are significant. We find that trends in our human evaluations generally match those seen in the automated metrics for both task and naturalness metrics (see Figures 2(c), 2(d) which summarize Appendix Tables 10,15,21,26, 35—Supervised+NLPO >> Supervised ≥\geq Supervised+PPO >> NLPO ≥\geq PPO >> Zero-shot—with the exception of Supervised outperforming Supervised+PPO on 2 out of 5 tasks when automated metrics would indicate that Supervised+PPO outperforms Supervised on all of the tasks. We draw two conclusions from this: (1) if the generated text is above a certain threshold of naturalness, the automated metrics usually correlate with human judgements; (2) usually but not always as seen in the relative performance of Supervised and Supervised+PPO, potentially indicating reward hacking behaviors undetected by automated metrics but caught by human preference feedback.

2 Preference Reward Learning, Selection, and Hacking

While the GRUE benchmark’s metric for each task is an average over several measures, the RL models we trained optimized only a single metric independently. Thus, we can empirically investigate which metric for which GRUE produces the best results. We observe that many possible single metric rewards provide task performance gains over supervised methods (results shown in Fig. 3(a), 2(c) are averaged across these reward functions) with the condition that the text is also coherent and natural.

Which constraints best prevent reward hacking? The reward function in Equation 1 balances a task-specific reward with a KL constraint — models are penalized from straying too far from a base LM in their pursuit of high reward (Table 3 and Appendix Table 5) clearly show that if KL constraints are removed entirely, models reward hack). But which model works best as a base regularizing LM? When the initial policy (i.e., the raw, pretrained model) has low performance on the task, the KL penalty pushes the policy towards nonsense, e.g. on Commongen and ToTTo the trained policy learns to simply repeat portions of the input (as seen in Tables B.4.5, B.6.4). This behavior is mitigated if the base regularizing LM is the supervised model—the reward encourages the policy to balance the task-specific reward and a more reasonable regularization term. Deriving KL penalties from warm-started initial policies is critical for performance on such tasks.

PPO vs. NLPO. Figure 2 shows that NLPO generally outperforms PPO and supervised, especially when applied after supervised training. We hypothesize that the primary reason for NLPO’s improved performance and stability is because the masking policy provides an additional constraint for the current policy. This constraint is not based on the initial untuned policy like the KL penalty but of the policy from μ\mu iterations ago and likely contains more task-relevant information learned during RL training. Table 3 (and Appendix Table 8) shows how performance increases up to a point and then decreases as pp in top-pp sampling is increased for the masking policy, relaxing the constraint by eliminating less tokens at each step, implying that there is a balance to be found in how much the model should be constrained during RL training.

Human Preference Reward Learning. To this point, our experiments have largely focused on optimizing evaluation metrics that correlate with human judgments, e.g., METEOR. Here: we additionally test how well preferences can be learned from direct human feedback. For this, we focus on Commongen — a GRUE dataset well-suited for displaying differences due to human preferences. First, we randomly select prompts from the Commongen train dataset and sample a single completion from both the Supervised and Supervised+NLPO models. We then present the prompt and the two completion candidates to 3 unique crowdworkers and ask them to select which one they prefer with respect to commonsense/fluency for 417 unique pairs (Krippendorf α=.28\alpha=.28). We use this data to train a reward model, T5-11B Raffel et al. (2020), on the balanced binary classification task of predicting which of the pair was preferred by a majority of 3 annotators, conditioned on the prompt and completion. The resulting model achieved 69.5 test ROC AUC suggesting it indeed captures average human preferences. Additional details on this process are found in Appendix B.4.4. We train Supervised+RL with a METEOR-only reward as a baseline, and compare it to a reward function that uses the fine-tuned T5-11B model. Finally, we rerun the same pairwise preference collection procedure—this time sampling from Commongen test—with human participants to compare the generations from a preference optimized RL policy to the previously best Supervised+NLPO policy. Comparing the METEOR-only to the preference model, the generations produced by the human feedback model are preferred in 682 cases, compared to the METEOR-only model which is preferred in 587 cases (p<0.01p<0.01 the models are equally preferred). This implies that this pipeline of collecting preferences, training a reward, and further tuning the policy improves alignment to human preferences.

3 Data Budget: Improve your Reward or Gather More Demonstration?

Given a fixed data collection budget, is it more efficient to gather feedback to improve a learned reward function or to gather more expert demonstrations? We use the IMDB text continuation task as a case study. In the IMDB task, a model is given a partial movie review as a prompt, and is asked to continue it as positively as possible (even if the prompt was negative). The original dataset consists of movie reviews and sentiment labels of positive, negative, or neutral. A DistilBERT (Sanh et al., 2019) classifier is trained on these labels and used to provide sentiment scores on how positive a given piece of text is, which serves as the task reward. The trade-off is between gathering more: 1) sentiment labels (improving the reward); or 2) positive sentiment reviews (improving supervised training).

We train a classifier on varying amounts of training data and evaluate on the held out test dataset—finding as expected that more training data improves test accuracy and so results in a higher quality reward. We then use each of these rewards of varying quality during RL training, and evaluate using the same metric as GRUE (i.e., a classifier trained with the entire training set). As seen in Table 3, we find that improving the reward quality improves LM performance as well. Further, we trained a supervised model with at least as many samples used to train each of these reward classifiers. We find that a learned reward function enables greater performance when used as a signal for an RL method than a supervised method trained with 5 times more data. This implies that improving reward models can be more data efficient than collection expert demonstrations for a task—and that’s not accounting for the fact that assigning sentiment labels is likely a simpler task than writing full demonstrations. Further details on this ablation are found in Appendix Table 7.

4 Practical Considerations: Which Implementation Details Matter Most?

Generation as a token-level MDP, not a bandit environment. Most recent works that tune LMs using RL do so by calculating a reward for all the tokens in the sentence (Wu et al., 2021a; Ouyang et al., 2022; Lu et al., 2022). This setting is equivalent to a bandit feedback environment where the action space is the space of all possible generations for the task (Sutton & Barto, 2018). This type of environment can be simulated within our RL formulation by setting the discount factor γ=1\gamma=1. Table 3 (and Appendix Table 6) shows that this causes instability in training with respect to naturalness in both PPO and NLPO for IMDB. Our standard setting is γ=0.95\gamma=0.95 when calculating discounted rewards-to-go in the token-level MDP formulation, which reduces the magnitude of the reward that is applied to tokens selected at the beginning. The sentiment scores are approximately the same between both settings but the naturalness of language in the bandit setting is significantly less—indicating that discounting rewards with γ<1\gamma<1 via a token-level MDP formulation is at least sometimes more effective for language generation.

Dropout and Sampling. We found two other implementation details to be critical for stability of RL training. The first is dropout, which in its standard form was found to cause instability in policy gradient methods in continuous control settings by Hausknecht & Wagener (2022). We find a similar effect when using dropout when RL training LMs as well, with training loss often diverging for dropout >0>0 in training. The second important detail, particularly affecting the machine translation task, is sampling methods. We find that using the same sampling methods during exploration and inference is critical to translating training performance to test performance–else the model exhibits high train rewards but low test metrics.

Conclusions

We’re hopeful that the GRUE benchmark and the RL4LMs library can push progress in aligning language models to human preferences via RL methods by providing the community with a standard means of comparing methods. Furthermore, we’re optimistic that, as the stability and consistency of training improves, our methods provide a path towards iterative improvement of language technologies, with deployment, user feedback collection, and re-optimization enabling better user experiences when interacting with generative models.

Acknowledgements

We’d like to acknowledge the support of DARPA MCS program through NIWC Pacific (N66001-19-2-4031), Google Cloud Compute, and the ReViz team at the Allen Institute for AI. KB is supported by NSF under grant No. 2127309 to the Computing Research Association for the CIFellows Project.

References

Appendix A On-policy Algorithm Implementation Details

Given discussion and equations in Section 3.3, we further note that we follow (Ziegler et al., 2019) and dynamically adapt the KL coefficient β\beta during training where,

where KLtarget\text{KL}_{\text{target}} is user-specified KL divergence between initial model hh and current policy π\pi and Kβ\text{K}_{\beta} is rate of update which we generally set to 0.20.2 in our experiments.

To increase stability during training, we further use Generalized Advantage Estimation (GAE) (Schulman et al., 2015b) and define the advantage estimator A^(sn,an)\hat{A}(s_{n},a_{n}) based on the Temporal Difference residual as:

where λ\lambda provides the trade-off between bias and variance.

A.2 NLPO Details

NLPO learns to mask irrelevant language by maintaining a masking policy πψ\pi_{\psi}: the masking policy is a copy of the current policy (πθ\pi_{\theta}), but is updated only every μ\mu steps. Given Z(πθ)=∑a∈Vπθ0(a∣s)Z(\pi_{\theta})=\sum_{a\in\mathcal{V}}\pi_{\theta_{0}}(a|s) the normalization value of the sum of probabilities of all action a∈Aa\in\mathcal{A} given a particular State s∈Ss\in\mathcal{S}, let the parameterized top-pp vocabulary Vπθp⊂V\mathcal{V}_{\pi_{\theta}}^{p}\subset\mathcal{V} be the subset of the vocab, consisting of the top-pp highest probability vocabulary tokens with respect to πθ\pi_{\theta}. Formally, let ZpZ^{p} be the normalization value for the parameterized top-pp vocabulary, can be defined as the subset of tokens that maximizes Zp(πθ)=∑a∈Vπθkπθ(a∣s)Z^{p}(\pi_{\theta})=\sum_{a\in\mathcal{V}^{k}_{\pi_{\theta}}}\pi_{\theta}(a|s). Then optimizing a policy according to the parameterized top-pp vocabulary can be defined as:

Appendix B Experimental Details

We ran a qualification round using the IMDB task. We opened the qualification around to users from {\{AU, CA, NZ, GB, US}\} with 5K prior approved HITs and a minimum acceptance rate of 97% on their previous HITs. We gathered judgments over 600 generations from 3 annotators per generation. One of the authors of this paper also completed 17 random HITs to serve as a proxy for “ground truth." After gathering these annotations, we selected workers who: 1) didn’t significantly disagree with other annotators on the same instance more than 20% of the time; 2) who completed at least 5 HITs; 3) who didn’t disagree with the author annotator on the 17 HITs by more than 1 point; and 4) (likely) spent a reasonable amount of time reading the instructions/examples provided. In the end, 56 annotators were qualified. Additional per-task details are provided in the per-task sections of the Appendix.

As per Amazon Mechanical Turk policy, annotators were compensated on a per-HIT basis. In addition, we used a timing script to estimate hourly wages to ensure our target of $15/hr was met. In cases where this minimum hourly rate was not met, we manually assigned bonuses.

B.2 GRUE Experiment Setup

We benchmark 5 training algorithms on 6 tasks (see Table 1) using either an encoder model (eg. GPT-2) or encoder-decoder model (eg. T5). We train policies using PPO, NLPO with variations of whether supervised pre-training is applied before RL fine-tuning and compare against supervised policy. The choice of LM is based on the type of task. For IMDB text continuation, we use GPT-2 and T5 for rest of the tasks. We use two separate LM models as actor and critics networks (i.e. no shared layers) in which the critic network has an additional linear layer mapping last token’s hidden representation to a scalar value. We use AdamW optimizer Loshchilov & Hutter (2017) with fixed learning rate and no scheduling.

B.3 IMDB

We consider IMDB dataset for the task of generating text with positive sentiment. The dataset consists of 25k training, 5k validation and 5k test examples of movie review text with sentiment labels of positive and negative. The input to the model is a partial movie review text (upto 64 tokens) that needs to be completed (generating 48 tokens) by the model with a positive sentiment while retaining fluency. For RL methods, we use a sentiment classifier Sanh et al. (2019) that is trained on pairs of text and labels as a reward model which provides sentiment scores indicating how positive a given piece of text is. For supervised Seq2Seq baselines, we consider only the examples with positive labels. We chose GPT-2 as LM for this task as it is more suited for text continuation than encoder-decoder LMs (eg. T5). We use top-k sampling with K=50K=50 as the decoding method and for fair comparison, we keep this setting for all methods. For PPO and NLPO models, we train for 64k64k steps in total and update policy and value networks every 12801280 steps with a mini-batch size of 6464 and epochs of 55 per update. We apply adaptive KL controllers with different target KLs of 0.02,0.05,0.1,inf⁡0.02,0.05,0.1,\inf with an initial KL co-efficient of β=0.1\beta=0.1. Table 4 provides an in-depth summary of all hyperparameters and other implementation details.

B.3.2 Results and Discussion

Fig 4 shows learning curves for PPO and NLPO in terms of episodic training reward, corpus level sentiment scores and perplexity scores on validation set averaged for 5 random seeds. It is seen that higher target KL of 0.10.1 is desired to achieve higher rewards but results in drifting away from pre-trained LM and loses fluency. Therefore, a lower target KL (0.02 or 0.05) is required to keep the LM closer to original LM. This is also seen in Table 5 where we presented a comparative analysis of final performance of all models.

We vary the amount of data used to train the reward classifier and the supervised baseline model to understand whether it is more efficient to gather data to improve reward model or to gather expert demonstrations for supervised learning. As observed in Table 7, improving the quality of reward function increases the performance on the overall task better than training with more data for supervised training, indicating that improving reward models is efficient than collect expert demonstrations for supervised training from a data efficiency perspective.

To understand the effect of discounted vs undiscounted (bandit) environments, we report sentiment and perplexity scores for different values of discount factor (0.50.5, 0.950.95 and 1.01.0) in Table 6 and observe that using a bandit environment (discount factor of 1.01.0) results in performance loss in the case of NLPO and reward hacking in the case of PPO, indicating that discounted setting (with 0.950.95) is desired.

Table. 8 shows ablation on different hyperparameters in NLPO algorithm.

B.3.3 Human Participant Study

Figure 5 shows the IMDB instructions, example, and interface used both for the qualification round, and then later, for the human evaluation experiments. Tables 9, 10 show averaged results, annotator agreement, and the results of statistical significance tests to determine which models output better generations when rated by humans.

B.3.4 Qualitative Results

We show sample generations from each of the algorithms for three randomly picked prompts below.

B.4 CommonGen

CommonGen (Lin et al., 2020) deals with task of generating coherent sentences describing an input set of concepts (eg. "a man is throwing a frisbee"). For training RL methods, we consider 3 traditional lexical rewards namely Rouge-1, Rouge-avg (which is an average of Rouge-1, 2 and L) and meteor. Additionally, we also train with task-specific rewards such as CIDEr (Vedantam et al., 2015), SPICE (Anderson et al., 2016) and SPiDer (Liu et al., 2017) which is a just a linear combination of both with equal weights. We chose T5-base as the base LM since it is well-suited for structure to text tasks. We additionally note that concept set inputs are prefixed with "generate a sentence with:" to encourage exploration.

During our initial experiments when fine-tuning directly on LM, we observed that policy learns to repeat the prompted concepts in order to maximize rewards resulting in a well-known problem of reward hacking. To mitigate this, we add a penalty score of −1-1 to final task reward if the n-grams of prompt text overlaps with generated text. In contrast, when initialized with a supervised policy, this problem is not seen and hence penalty score is not applied. We use beam search as the decoding method during evaluation whereas for rollouts, we use top k sampling to favor exploration over exploitation. Table 11 provides an in-depth summary of setting of hyperparameter values along with other implementation details.

B.4.2 Results and Discussion

Tables 13, 12 presents our benchmarking results with 6 reward functions along with supervised baseline performances on dev and test sets respectively. Our main finding is that warm-started initial policies are crucial for learning to generate coherent sentences with common sense. Without warm-start, policies suffer from reward hacking despite application of repetition penalty and task-specific metrics such as CIDer etc. Further, we find that RL fine-tuned models obtain very high concept coverage which is also seen in Table B.4.5. Supervised models often tend to miss few concepts in its generation compared to RL methods.

B.4.3 Human Participant Study

Figure 6 shows the commongen instructions, examples, and interface used for the human evaluation experiments. Different from the other human evaluations, we didn’t provide any prompt because knowing the set of words to be used isn’t required for rating either of the axes. Tables 14, 15 show averaged results, annotator agreement, and the results of statistical significance tests to determine which models output better generations when rated by humans.

B.4.4 Human Preference Learning Experiments

First, we randomly select prompts from the Commongen train dataset and sample a single completion from both the Supervised and Supervised+NLPO models. Next, we filter to prompts where both models at least attempted to use all input concepts. This filtration step was conducted because if a model fails to use all concepts, it may generate a more natural/fluent sentence, but, a priori, it shouldn’t be preferred by crowdworkers; instead of training crowdworkers to prefer sentences with all concepts, we perform this filter. Figure 7 shows the task presented to the crowdworkers. We then present the prompt and the two completion candidates to 3 unique crowdworkers and ask them to select which one they prefer with respect to commonsense/fluency; We gathered 3 annotations on 417 pairs (Krippendorf α=.28\alpha=.28), and split into 60/20/20 train/val/test split. We then trained a reward model, T5-11B Raffel et al. (2020), on the balanced binary classification task of predicting which of the pair was preferred by a majority of 3 annotators, conditioned on the prompt and completion. The resulting model achieved 69.5 test ROC AUC suggesting it indeed captures average human preferences. The model is then used as a reward function. We train Supervised+RL with a METEOR-only reward as a baseline, and compare it to a reward function that uses the fine-tuned T5-11B model. We design the reward function based on the preference model as r=meteor+pref/(1+∣miss∣)r=meteor+pref/(1+|miss|) where missmiss is a set of concepts not covered in the generated text, in an attempt to mimic the data collection process that humans are instructed to follow. This reward function accounts for both the task of using all concepts and also human’s preferences for how a sentence should look within the constraints stipulated by the task. Finally, we rerun the same pairwise preference collection procedure—this time sampling from Commongen test—with human participants to compare the generations from a preference optimized RL policy to the previously best Supervised+NLPO policy. Comparing the METEOR-only to the preference model head-to-head, the generations produced by the human feedback model are preferred in 682 cases, compared to the METEOR-only model which is preferred in 587 cases (p<0.01p<0.01 the models are equally preferred).

B.4.5 Qualitative Analysis

This section shows sample generations from different algorithms for three randomly picked prompts.

B.5 CNN Daily Mail

As a representative of the summarization task, we consider CNN/DM dataset consisting of long news articles and their highlights written by news authors. The dataset consists of 287k training, 13k validation and 11k test examples. We trained RL methods using 3 different automated metrics, namely Rouge-1, Rouge-avg and Meteor. We chose T5 as our base LM as it is pre-trained in a unified text-to-text framework and relishes Zero-Shot capabilities. For decoding, we use multinomial sampling with a temperature of 0.70.7 for all the models.

B.5.2 Results and Discussion

Table 17 presents benchmarking results on test set reporting a wide range of metrics: lexical, semantic, factual correctness and diversity metrics. As baselines, we report lead-3 which selects first three sentences as the summary, Zero-Shot and a supervised model. PPO and NLPO models are on par with supervised performance on several metrics including Rouge-2, Rouge-L, and Bleu. On fine-tuning on top of supervised model, performance improves consistently on all metrics indicating that RL fine-tuning is beneficial. Another interesting finding is that, RL fine-tuned models are factually consistent as measured by SummaCZS metric. For ablations on PPO params, NLPO params, we refer to Tables 18,19.

B.5.3 Human Participant Study

Figure 8 shows the summarization instructions and interface used for the human evaluation experiments. Participants weren’t required to read the entire article, but to encourage some reading, a minimum time on the window of 15s was enforced via hiding the sliders. Tables 20, 21 show averaged results, annotator agreement, and the results of statistical significance tests to determine which models output better generations when rated by humans.

B.5.4 Qualitative Analysis

We show sample generations from each of the algorithms for three randomly picked prompts below.

B.6 ToTTo

ToTTo (Parikh et al., 2020) is a controlled table-to-text generation task in which the goal is to produce one-sentence description of highlighted table cells. For training RL methods, we consider 5 different reward functions: BLEU, SacreBLEU, METEOR, PARENT and a combination of Meteor and PARENT. We chose T5 as our base LM here too, as they are more suitable for structure to text tasks. For decoding, we use beam search during inference and for generating rollouts, we use top k sampling. Other implementation details are captured in Table 22.

B.6.2 Results and Discussion

Tables 24, 23 presents our benchmarking results with 5 reward functions along with supervised baseline performances on dev and test sets respectively. Similar to other tasks, our main finding is that warm-started initial policies are crucial for learning to generate descriptions from highlighted cells. Without warm-start, policies suffer from reward hacking and resulting in sub-optimal solutions despite application of task-specific metrics such as PARENT etc. We find that Supervised+NLPO method outperforms all models on ToTTo leaderboard in terms of PARENT metric.

B.6.3 Human Participant Study

Figure 9 shows the ToTTo instructions, example, and interface used for the human evaluation experiments. We made small modifications to the original code release’s HTML renderer to make the tables display in our HITs. Tables 25, 26 show averaged results, annotator agreement, and the results of statistical significance tests to determine which models output better generations when rated by humans.

B.6.4 Qualitative Analysis

We show sample generations from each of the algorithms for three randomly picked prompts below.

B.7 Narrative QA

NarrativeQA (Kočiskỳ et al., 2018) deals with task of generating answers to questions about a given story. For training RL methods, we consider 2 traditional lexical rewards namely Rouge Combined and Rouge-L-Max. We chose T5-base as the base LM since it has been shown to do well at question answering in prior work (Khashabi et al., 2020). We note that the supervised models we use are trained on the UnifiedQA dataset, which contains other QA datasets, and is shown by Khashabi et al. (2020) to outperform supervised fine-tuning only on NarrativeQA. Hyperparams for our models can be found in Table 27.

B.7.2 Results and Discussion

Table 28 presents our benchmarking results with 2 reward functions along with supervised baseline performances on the NarrativeQA test set. Similar to other methods, our main finding is that warm-started initial policies are crucial for learning to generate answers that successfully use the input context.

B.7.3 Qualitative Results

We show sample generations from each of the algorithms for three randomly picked prompts below.

B.8 Machine Translation

WMT-16 We pick two languages, English and German, and frame this task similarly to other machine translation tasks—requiring the models to translate from English to German. We train models on 4 rewards: SacreBLEU, chRF, TER, and BertScore.

B.8.2 Results and Discussion

Tables 30, 31 presents our benchmarking results with 4 reward functions along with supervised baseline performances on test set. Our main finding is that NLPO + Supervised performs better than PPO and supervised models.

B.8.3 Qualitative Results

We show sample generations from each of the algorithms for three randomly picked prompts from IWSLT below.

B.9 Daily Dialog

We consider DailyDialog (Li et al., 2017) as the test bed for the dialogue generation task. The dataset includes conversations written by human on various topics. In addition, each utterance contains labels of intent and emotional information. For simplicity, we focus only on generating the next utterance, given the dialogue context. We chose a context window of size 55, which results in 3535k training, 33k and 33k utterances. The input to the model is dialogue history in which utterances are concatenated using a token. We picked GPT-2 as the LM as they are more suited for text continuation than encoder-decoder LMs. For a fair comparison, we use top-k sampling with k=20k=20 as the decoding method for all methods. For RL methods, we use a linear combination of meteor score and intent match score (whether the generated text’s intent matches with the reference’s intent) as the reward function. The coefficients for meteor and intent are chosen based on both lexical scores and intent accuracy on the validation set. For this purpose, we trained an intent classifier (fine-tuned RoBERTa (Liu et al., 2019)) that classifies given text into intent categories such as inform, question, directive and commisive, etc. Table 32 provides a summary of hyperparameters and implementation details.

B.9.2 Results and Discussion

Table 33 presents our benchmarking results of RL methods along with supervised baseline performances on test sets. Our main finding is that RL methods generally achieve better intent accuracy and automatic metric scores, in particular NLPO variants perform better than all other methods.

B.9.3 Human Participant Study

Figure 10 shows the Daily Dialogue instructions and interface used for the human evaluation experiments. Tables 34, 35 show averaged results, annotator agreement, and the results of statistical significance tests to determine which models output better generations when rated by humans.

B.9.4 Qualitative Analysis

We show sample generations from each of the algorithms for three randomly picked prompts below.