SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu
Introduction
While massively scaling model parameters and training compute of Transformer-based language models have led to impressive few-shot in-context learning , reinforcement learning from human feedback fine-tuning (RLHF) can significantly improve generation quality as judged by humans. This has been observed at all model scales for various downstream language generation tasks, such as abstractive summarization, dialogue, and creative writing .
For summarization in particular, multiple studies have shown that summaries generated by models tuned with RLHF are preferred over the reference summaries in commonly used datasets . Reference summaries are often mined from web documents and might not have the highest quality or preferred style. As a result, pure supervised learning, i.e. maximizing the likelihood of reference summaries given documents, is limited by the quality of reference summaries, thus additional feedback can improve models beyond the references. Commonly used reference-based metrics, such as ROUGE , only measure similarity between model generated and reference texts. These reference-based metrics cannot measure quality improvement beyond the reference summaries.
To implement RLHF, a reward model, , is trained on human preference data, , collected via side-by-side human evaluation, where raters are asked to judge which of the two summaries and is better for document , i.e. . If we denote the preferred summary as and the other , the human feedback becomes . One common option for the training loss of the reward model used by RLHF is:
Reinforcement learning algorithms such as PPO are then used to refine a supervised fine-tuned model (SFT) to maximize the expected reward assigned by the reward model . A KL-penalty term is typically added to the loss to prevent the RLHF model from diverging too far from the original supervised policy.
However, algorithms such as RLHF-PPO introduce significant complexity to the training process by adding separate value and reward networks that may be comparable in size to the policy network. They are typically kept in memory to maximize training speed, which for a given memory budget significantly reduces the maximum size of trainable model. Furthermore the optimization steps are significantly slower due to the use of roll-outs in the training loop, which involves sampling/decoding from the model. Hyper-parameter tuning and co-coordinating the PPO process is also more complex, requiring niche expertise.
Recently, another class of sequence-level contrastive methods seek to align model likelihood with an arbitrary, possibly non-differentiable reward, presenting an alternative to RL for optimizing the expected reward of samples. Zhao et al. proposed Sequence Likelihood Calibration (SLiC) to align a language model’s sequence likelihood, , over decoded sequences according to their similarity to reference sequences. The ranking calibration loss contrasts a positive sequence and a negative sequence , encouraging the model to assign more probability mass to positive compared to negative sequences:
While the original SLiC work used similarity to references as criteria for ranking, e.g. ROUGE and model embedding distances, it can be replaced by an arbitrary, reference-less ranking function, . Particularly in this work, we use human preference as the ranking function, either by using off-policy preference data directly, or by training a predictive ranking model from .
We call using SLiC with this human preference ranking function SLiC-HF and apply it using the human feedback data collected in Stiennon et al. . Our experiments show that SLiC-HF also leads to improved summarization quality on the Reddit TL;DR Summarization task as judged by humans, even though this feedback was collected for different models, similar to off-policy, offline RL. While our T5-Large (770M parameter) SFT model performs similarly to Stiennon et al. ’s 6B decoder-only SFT model, we are able to improve our model with SLiC-HF such that it performs at least as well as Stiennon et al. ’s 6B RLHF-PPO model as judged by humans. Furthermore, applying SLiC-HF to the T5-XXL 11B parameter SFT model significantly improves results.
The primary contributions of this paper are showing:
how to apply SLiC to learn from human preferences (SLiC-HF), a simpler, more efficient yet competitive alternative to RLHF
feedback/preference data from another model (off-policy) can be effectively leveraged by SLiC-HF, making it unnecessary to collect costly new feedback data for our model
providing a general SLiC-HF recipe based on open-sourced T5 model that outperforms RLHF on the Reddit TL;DR summarization task
Method
2 SLiC-HF with Sample and Rank
Zhao et al. samples candidates from ’s training split, from which (positive, negative) pairs are determined. We call this approach SLiC-HF-sample-rank. To determine the rank, we consider two text-to-text models trained from the human preference data :
Similar to Askell et al. , we binarize each ranked pair into a positive and a negative sequence, as shown in Figure 1. When training the reward model, input sequences are formatted as ‘[Context] … [Summary] … ’ and target sequences are either ‘Good’ or ‘Bad’. At inference time, we compute the probability of token ‘Good’ on the decoder side to score each of the candidates in a list, and sample positive/negative pairs from them.
As shown in Figure 1, we formulate the human feedback into a pairwise ranking problem with text-to-text format. When training the ranking model, input sequences are formatted as ‘[Context] … [Summary A] … [Summary B]’ and target sequences are among ‘A’ or ‘B’. At inference time, we use a tournament-style procedure to rank candidates in a list. For example, given a list of 4 candidates , we first rank and and then rank . Given candidates, the ranking model is called times and positive/negative pairs are yielded.
3 SLiC-HF Directly On Human Feedback
We also consider a straight-forward approach of directly calibrating on positive and negative sequences from the human feedback dataset, , without a ranking or reward model. We call this approach SLiC-HF-direct. The obvious advantage of this approach is increased simplicity and efficiency from not training or using a ranking/reward model. SLiC-HF-direct does not incur additional engineering costs in decoding from the SFT model and training a model to label the decodes. The drawback is that the off-policy human feedback data distribution might differ much from the SFT model’s decode distribution.
4 Regularization Term for Calibration
Experimental Results
We study SLiC-HF on Reddit TL;DR summarization datasets from Stiennon et al. . The dataset contains both fine-tune data , human feedback data , along with their SFT and RLHF model decodes which we use for comparison with our models. is a filtered version of Reddit TL;DR dataset . It contains 117k/6k/6k examples in train, validation and test splits. consists of 64k human preferences on decodes from multiple models.
2 Experimental Hyper-parameters
We conduct all experiments using T5 models in the T5x framework . In our ablation study, we choose a T5-large model (770M) as the generation model and T5-XXL (11B) as the ranking model and the reward modelWe find that smaller T5 ranking/reward models do not converge reliably in our setup.. We train all generation models with batch size of 32 and ranking/reward models with batch size of 128. Both are trained with default learning rate of .
We train the ranking model and the reward model on training split, and picked checkpoints that have the highest accuracy on validation split. We fine-tune T5 models on training split, and pick checkpoints that have the lowest perplexity on validation split.
In calibration, we use learning rate of and ranking margin of . When calibrating models on their own decodes with SLiC-HF-sample-rank, we sample 8 decodes with temperature of 0.7 and topk of 40 from fine-tuned only generation models.
When evaluating our models, we use beam-search with beam size 4. For automatic evaluation, we calculate the model decodes’ win rate against human references measured by the T5-XXL ranking model on validation dataset. Win rate is defined as the percentage of model decoded summaries preferred by the ranking model compared to human references.
3 Reward Model and Ranking Model Accuracy
Human feedback and human evaluation are done by raters comparing two summaries as it is more reliable than pointwise rating. We hypothesize that ranking model has an advantage over reward model because of its pairwise nature which aligns better with the task. We train and compare a T5-XXL ranking model and a T5-XXL reward model (subsection 2.2). Results shows that our ranking model has accuracy of 73.23% on validation, about 2% higher than our reward model which has accuracy of 71.34%Our ranking and reward models’ accuracy are similar to the 6B reward model in Stiennon et al. .
4 SLiC Ablation
We conduct a set of experiments to ablate SLiC-HF settings against baselines. We use the ranking model as the main metric because of its higher correlation with human preferences demonstrated in Stiennon et al. . Selected settings are later verified with our human evaluation experiments in subsection 3.5. We report ROUGE numbers just for reference purpose and do not use them to select models. It is expected to see a drop in ROUGE numbers when learning from human feedback because it has less incentive to be similar to the reference texts. Similar to RLHF in Stiennon et al. , we also observe an increase in average length of models and conduct a length controlled study in subsection 3.5.
A simple way to learn from human feedback data is to convert it into SFT dataset and continue fine-tuning on it. In general, we use the filtering approach which has similar performance to controlled generation approaches but is cleaner to implement. We consider three approaches to filter data for continued fine-tuning:
keep only positive human feedback sequences and discard negative ones.
decode 8 summaries from the SFT model, use the ranking model to select the best 1 out of 8 summaries by a tournament-style ranking approach.
decode 8 summaries from the SFT model, use the reward model to select the best 1 out of 8 summaries by scoring each and taking the one with the max score.
As shown in Table 1, on Reddit TL;DR dataset, continue fine-tune on positive human feedback data improves model win rate against reference slightly from 44.96% to 51.65%. In this experiment, we choose to use all human feedback without filtering for better models because this mimics a real world scenario where we have access to some human feedback data without the explicit knowledge of its quality. Continuing fine-tuning on best 1 out of 8 further improves win rate against reference to 60%+ and using pairwise ranking model is slightly better than pointwise reward model for filtering.
4.2 Apply SLiC-HF Directly On Human Feedback Data
With SLiC-HF-direct, we observed that even though calibration loss decreases as expected, sequence length keeps increasing and does not converge to a stable value. On the other hand, SLiC-HF-sample-rank robustly converges. We hypothesize that SLiC-HF-direct is prune to out-of-distribution decodes generated by other models in the human feedback data.
When using the ranking model to select for the best checkpoint for SLiC-HF-direct, it has moderate length increment and has 82.92% win rate against reference which is close to SLiC-HF-sample-rank. The engineering complexity of SLiC-HF-direct is almost the same as fine-tuning a model. Therefore, it is a good candidate for quick experimentation on human feedback.
4.3 Apply SLiC-HF on Ranked Model Decodes
As shown in Table 1, SLiC-HF-sample-rank using the ranking model have about 3% gain in win rate against reference compared to SLiC-HF-sample-rank using the reward model. This results aligns with the observation in subsection 3.3 that the ranking model has higher agreement to human preference than the reward model.
For SLiC-HF-sample-rank using the ranking or the reward model, using SFT targets or best ranked decodes as regularization doesn’t show much difference. This shows that SLiC-HF-sample-rank is applicable even when there is no ground truth reference available. The gain from continue fine-tuning on best ranked decodes in Table 1 is not additive to SLiC-HF.
5 Human Evaluation
We conduct side-by-side human evaluation between multiple systems using crowd-sourcing.We use Amazon Mechanical Turk to set up the task and hire the raters Given a document and 2-4 summaries, raters are tasked to assign a pointwise overall quality to each summary, select if the summary is factual or not, and choose the best summary.
Each task is replicated and judged by 3 different raters. To eliminate bias, we anonymize all the models and randomly shuffle order of summaries for each task. We aggregate pointwise metrics by averaging the ratings across all 3 crowd workers, and we aggregate the choice metric using majority vote.
The human evaluation template and the rating instructions can be found in Appendix A.
We conduct a 4-way side-by-side human evaluation to confirm the ablation results in Table 1. 100 examples from the validation set are sampled from reference, SFT model, continue fine-tuning model and SLiC-HF model (SLiC-HF-sample-rank, using ranking model, regularized on best decodes). As shown in Table 3, SLiC-HF is chosen as the best model 73% of the time, has significantly higher average quality, and is the most factual model. In general, the average quality aligns well with the ranker win-rate from Table 1.
Figure 2 shows the lengths controlled quality of SFT, continue fine-tuning and SLiC-HF models, which clearly shows SLiC-HF is preferred. Length controlled quality study is similar to studies conducted in Stiennon et al. , where mean scores are calculated among examples bucketed by their relative length to the reference.
5.2 SLiC-HF vs RLHF-PPO
Correctly implementing and tuning the right hyper-parameters for the RLHF-PPO algorithms in Stiennon et al. are non-trivial tasks. Instead of re-implementing the algorithms in our framework, we directly compare with the model decodes from Stiennon et al. .
We first benchmark our T5-large SFT model against their 6B decoder-only SFT model in a two-way side-by-side human evaluation. As shown in Figure 3, our SFT has slightly higher quality and win rate but it is not statistically significant.
Next we benchmark two variants of our T5-large SLiC-HF-sample-rank models against the decoder-only 6B RLHF-PPO model from Stiennon et al. . SLiC-HF-sample-rank with reward model has similar performance as RLHF-PPO and SLiC-HF-sample-rank with ranking model is better than the RLHF-PPO. The summaries from SLiC-HF models are slightly longer than the RLHF-PPO model, and their length controlled win rate is similar to RLHF-PPO as shown in Figure 3.
6 Scaling Up SLiC
We study 2 ways of scaling up the SLiC-HF-sample-rank: (1) scaling up generation model parameters, (2) scaling up number of decoded candidates . As shown in Table 4, scaling up generation model from 770M to 11B significantly improves both the SFT model and the SLiC-HF model. On the other hand, scaling up from 8 to 64 does not help much.
Further discussion on SLiC-HF vs. RLHF-PPO
We summarize the compute and memory efficiency differences between SLiC-HF and RLHF-PPO in Table 5.
In both RLHF-PPO and SLiC-HF-sample-rank we train an auxiliary ranking or reward model that is used to judge the quality of summaries. However, Stiennon et al. found that having separate policy and value networks worked significantly better, and thus contributes an extra auxiliary model, the same size as the reward model that is updated along-side policy updates.
Furthermore, the policy, value, reward, and SFT models (all the same size in Stiennon et al. ) are used within the PPO training loop. They are often held in hardware memory to ensure faster training steps. Whereas in SLiC-HF, the rewards can be computed completely in parallel and offline, thus using 1/4 the memory for model weights during training. Such memory savings could be re-purposed to train larger models.
Stiennon et al. report using 1M episodes to conduct RLHF training, which corresponds to roughly the same number of decoded samples used in SLiC-HF, ( per training example, 123,169 examples). However, in practice SLiC-HF decoding can be significantly faster because all the decoded samples use the same policy allowing for completely parallel decoding. In contrast, with PPO the policy is updated every batch, limiting decoding parallelism to each batch (512, in ) as subsequent decoding is blocked on policy updates. Furthermore, PPO decoding occurs within the training loop leading to much longer optimization step times. Whereas with SLiC-HF step times are similar to fine-tuning, which is significantly faster as there is no decoding in the training loop.
Beyond the significant decoding parallelism gains, SLiC-HF can make use of simple input encoding caching optimizations to reduce compute. Since the decodes are sampled from the same SFT policy, the input sequence encoded states can be cached rather than recomputed. In summarization and other tasks involving long contexts, this may be significant as input sequence length tends to be much longer than output.
SLiC-HF has similar parallelism advantages in computing rewards per episode compared to RLHF as the ranking can be computed outside the training loop instead of within.
2 Pairwise Ranking vs Reward model
RL algorithms seek to maximize the expected reward of trajectories, in this case the human judgement in quality of model summaries. This reward function typically is assumed to be pointwise, whereas human preference data is collected pairwise to improve reliability. Thus there is noise introduced in converting pairwise judgements into pointwise rewards, which can be estimated as the difference in ranking accuracy as in subsection 3.3. Since SLiC-HF only cares about the relative rank of two summaries, this pairwise-to-pointwise noise is avoided and we conjecture this helps SLiC-HF (Table 1, Figure 3).
3 The Value of States and Actions in Language
For many tasks tackled using RL, rewards may be collected at the end of a trajectory (as in many Atari games) and the attribution of final reward to specific actions may be very important in learning to solve a task. Typically when RL is applied to language as in the RLHF literature, the state is the prefix of the current text and the actions correspond to choosing the next token. The value function’s role is to estimate the goodness of a trajectory (e.g. summary) from a prefix/input, which is intuitively a very difficult task for human raters, and thus RL may also suffer from value function estimation noise. In contrast, SLiC-HF does not rely on such a sub-model and only uses the cleaner preference signal to drive parameter updates and leading to what we conjecture is more stable optimization.
Related work
RL has been used to optimize arbitrary reward in language generation such as BLEU for translation and ROUGE for summarization ; however, while those metrics improved, human judgement of quality suffered due to metrics misalignment.
In an effort to better align the reward function with human judgement, many works used RL to align language models with a reward model trained to predict carefully collected human judgements using summarization as an initial proof-of-concept. A KL penalty term, first used in Jaques et al. , is used as regularization to prevent the tuned model from departing from the initial supervised model, and is also used in SLiC .
Liu et al. propose BRIO, which has a similar intent as SLiC of rank-ordering model-generated decodes according to a reward function. BRIO trains models to align length normalized sequence probability of generated decodes to their similarity to reference as measured by ROUGE using a list-wise loss function. In contrast, and similar to RLHF, SLiC-HF adapts the technique to align with a model trained to predict human preference given two summaries instead of their similarity to the reference.
Bai et al. substitutes human preference data with judgements from a large language model, and calls it AI feedback (AIF). SLIC-HF can also be used with AIF exactly in the same way and is indifferent about the AI or human origin of the feedback.
Conclusion
In this work, we proposed SLiC-HF that calibrates sequence likelihood on human feedback data. Our experiments on the Reddit TL;DR summarization task show that SLiC-HF significantly improves supervised fine-tuning (SFT) baselines, and presents a competitive alternative to the RLHF-PPO implementation of past work while being simpler to implement, easier to tune and computationally efficient. Future work may include studying SLiC-HF on other language generation tasks using other reward functions and/or non-human feedback.
References
Appendix A Human Evaluation
See Figure 4 for an example of the human evaluation task with 4 summaries. Summaries are randomly shuffled for each example and models are anonymized.