RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, Tong Zhang
Introduction
Generative foundation models have exhibited a remarkable capacity to accomplish diverse tasks that were previously unattainable, showcasing their broad-ranging capabilities in natural language processing and computer vision tasks. Large language models (LLMs) (Brown et al., 2020; Scao et al., 2022; Chowdhery et al., 2022; Smith et al., 2022; Hoffmann et al., 2022; Touvron et al., 2023) and diffusion models (Ho et al., 2020; Song et al., 2020b; a; Dhariwal & Nichol, 2021; Ramesh et al., 2022; Rombach et al., 2022), the most popular models in natural language and computer vision, are capable of generating high-quality meaningful outputs that are often indistinguishable from outputs produced by humans. AI-generated content is a rapidly evolving field that is widely believed to have the potential to revolutionize the way we create and consume content, ultimately enhancing the productivity of humanity. However, there are also concerns about the ethical implications of these models (Bender et al., 2021; Bommasani et al., 2021; Ouyang et al., 2022), such as the potential for misuse and the implicit bias from the model. It is important for researchers and developers to continue exploring the limitations of these models and restrict the output generation.
One of the most direct limitations of current generative models is the high dependency on unsupervised large-scale datasets. Such datasets often contain inherent biases that can manifest in the models’ outputs, leading to inaccurate or unfair results. To address this challenge, pre-trained models are typically fine-tuned on the downstream tasks with custom data, either to improve performance in a specialized setting or to eliminate potential biases and toxicity in the original model. One approach is to fine-tune the pre-trained models in a supervised manner using labeled data, known as supervised fine-tuning (SFT). Instruction tuning (Wei et al., 2021) is the most widely used approach to make LLMs adapt downstream tasks. However, collecting new supervised samples can be expensive in practical applications, especially when expert participation is required to generate high-quality data. More recently, Reinforcement Learning from Human Feedback (RLHF) has emerged as a promising method for fine-tuning pre-trained generative models. In recent studies of LLMs, RLHF has been widely employed to fine-tune pre-trained models using policy-based deep reinforcement learning (DRL) algorithms, typically the Proximal Policy Optimization (PPO). The idea of RLHF is to align the language models with human preferences and social values by optimizing a reward function that reflects specific human preferences (e.g. moral, helpful, harmless). For instance, OpenAI (Ouyang et al., 2022) fine-tuned a version of GPT-3 using RLHF with a reward function that emphasized certain human values. It is noteworthy to indicate that the alignment process often exerts a deleterious effect on the performance of generation, commonly referred to as the “alignment tax” in the literature (Askell et al., 2021). Specifically, when the reward model assesses only certain specific aspects, it may neglect the quality of the generated output. There has also been another line of work attempting to execute RLHF on visual generative models (Hao et al., 2022; Lee et al., 2023; Wu et al., 2023). This alignment process can be achieved through prompt learning or fine-tuning the diffusion model. Unlike the LLMs, the image generation process is typically not sequential: the pixels are generated simultaneously. Consequently, PPO is not well adapted to the vision task, and numerous adaptations are required in these works to align the visual generative models.
Although PPO is a well-established DRL method with numerous studies showcasing its effectiveness (Schulman et al., 2017; Engstrom et al., 2020), it learns in a trial-and-error fashion by interacting with the environment and is generally significantly less stable and less efficient as compared to supervised learning (Choshen et al., 2019). Meanwhile, in the context of LLMs, the predominant framework outlined in Ouyang et al. (2022) requires loading multiple LLMs for the PPO training, including the model being trained, the reference model, the reward model, and the critic model, which imposes a heavy burden on the memory resource. Additionally, although the SFT is more stable and fast than the PPO algorithm, the performance from SFT on the pre-determined dataset is typically inferior compared to the PPO-aligned one (Ramamurthy et al., 2022). The fundamental motivation behind our algorithm falls in between these two scenarios. First of all, while it is usually infeasible to collect a large amount of new samples from expert participation, the LLM to align itself can generate a large number of samples that can be used for training. Besides, the reward function provides a useful criterion for selecting high-quality samples without the expansive human evaluations.
Contributions. We propose an alignment framework – RAFT, which iteratively alternates among three steps, 1) we sample a batch of samples from the generative models; 2) we use the reward function to score the samples get from step 1 and filter them to get a filtered subset of high rewards; and 3) we improve the generative models by fine-tuning on the filtered subset from step 2. The proposed framework RAFT provides the following advantages compared to the predominant PPO algorithm:
The proposed framework is based on SFT-like training and offers enhanced stability and robustness compared to conventional-RL-based PPO. Additionally, its limited hyper-parameters make it easier and more straightforward to tune and adjust;
The proposed framework reduces memory burden as the data generation and model fine-tuning are decoupled. Meanwhile, the decoupled nature brings us the flexibility in data resource and processing;
The approach is flexible to train arbitrary generative models if a reward model is available as the quality measure, including LLMs and diffusion models;
The framework prioritizes preferences over values and is resistant to reward scaling. Its preference-based objective is clear and interpretable given the filtered dataset, which helps to mitigate the problem of reward hackingThe reward model used in RLHF is far from perfect, and the imperfection can be exploited by the algorithms to chase for a high reward, leading to reward hacking. by monitoring the selected samples.
Related Work
Generative foundation model. Foundation models (Bommasani et al., 2021) are generally pre-trained on large data and adapted to a broad range of downstream tasks. The roadmap towards the foundation model reveals a transition pattern from discriminative models (e.g., BERT) to generative models (e.g., GPT-3) due to their great scalability. Generative foundation models have reshaped the landscape of natural language processing (NLP), some of which even demonstrate emergent capabilities (Wei et al., 2022a) in complex reasoning tasks. Similar trends are observed in image generation, where diffusion models (Bender et al., 2021; Bommasani et al., 2021; Ouyang et al., 2022) have shown great text-to-image generation abilities with the increase of high-quality data and training compute. In particular, diffusion models captures the path from standard Gaussian distribution to the data distribution, which is proven to be successful in a variety of vision tasks, such as image inpainting, super-resolution, text-to-image generation, image denoising (Ho et al., 2020; Dhariwal & Nichol, 2021). Although generative foundation models have pushed the state-of-the-art on various language and vision tasks, they are suffering from implicit biases, leading to inaccurate or unfair results.
Alignment of generative models. Alignment (Leike et al., 2018) was first proposed to build agents that behave in accordance with the human’s intention. By communicating with human, agents can get accurate supervised signals (Ziegler et al., 2019) by applying several scalable reward learning methods (Leike et al., 2018; Christiano et al., 2018; Irving et al., 2018). Alignment benefits many recent generative foundation models, like InstructGPT (Ouyang et al., 2022), Claude (Bai et al., 2022b) and Sparrow (Glaese et al., 2022), in achieving better performance. In language foundation model training (Ouyang et al., 2022; Stiennon et al., 2020; Nakano et al., 2021; Bai et al., 2022a; b; Glaese et al., 2022; Ziegler et al., 2019; Wu et al., 2021; Scheurer et al., 2023), alignment is often achieved by Reinforcement Learning from Human Feedback (RLHF). The main idea is learning a reward function to reflect human preferences with human annotations and optimize LLMs by RL methods like proximal policy optimization (PPO) (Schulman et al., 2017). By incorporating supervised finetuning (SFT), InstructGPT (Ouyang et al., 2022) successfully achieved alignment for GPT-3 (Brown et al., 2020). Besides, Claude (Askell et al., 2021; Bai et al., 2022b) and Sparrow (Glaese et al., 2022) stressed aligning language foundation models from helpful, honest, and harmless (HHH) human feedbacks. In visual generative models, several works (Hao et al., 2022; Lee et al., 2023; Wu et al., 2023) studied aligning them with human feedbacks. Models are expected to understand specific visual control signals like colors, counts, and backgrounds (Lee et al., 2023) more accurately after alignment. It is still challenging to achieve tradeoffs between aligning human preferences and generating high-fidelity images. RRHF (Yuan et al., 2023) is an independent work that is contemporaneous with ours, which shares similar spirits with us to filter samples to serve as training samples for alignment of the generative model. In comparison, RRHF involves a diverse range of sources to generate data, and then finetune the model on the high-reward subset of these collected samples, while our primary focus lies in the online generated samples of the trained model itself, consistent with the setup of RL, where the behavior policy used to collect data also improves along the line. Moreover, we also validate the possibility of RAFT on diffusion models beyond the LLMs. Our work is also closely related to (Wang et al., 2022), which also boosts the performance of LLMs by the samples from the model itself. We note that (Wang et al., 2022) focuses on instruction-tuning, while we mainly study the RLHF. Due to the difference in context, Wang et al. (2022) filters the samples still mainly in a heuristic manner (e.g. instruction is too long/short, instance output is a repetition of the input, instruction is similar to existing one). While in RLHF, a preference-based reward function is trained based on comparison data (Ouyang et al., 2022) and can be used to measure the quality of samples.
Algorithm
We consider an initial generative model with model parameter , which can take input and generate an output according to a distribution , where is a temperature parameter to control the diversity. We also assume that we have a reward function , which returns a reward for any input-output pair . Due to common usage conventions, we refer to the input as the “prompt”. We use the reward function to guide the model . Specifically, if we denote as the conditional distribution given associated with and consider a distribution of the training input , the objective is
2 RAFT: Reward rAnked FineTuning
In this subsection, we will introduce the RAFT based on the combination of ranking samples by rewards and SFT. For simplicity, we assume that the generative model is powerful enough to achieve the maximum at each prompt . Then, we can separately consider each Another reason why we consider each prompt separately is that for LLMs, the prominent reward modeling approach from Ouyang et al. (2022) is based on such a local comparison with the same prompt. See Appendix A.1 for a detailed illustration.. Thus, the solution of Eq. (1) is
In practice, it is generally infeasible to search the entire output space to find the optimal policy. However, we can enhance our policy by fine-tuning our models using a high-reward dataset. One natural choice is to do so with a pre-determined high-quality dataset. Unfortunately, previous studies have shown that SFT with a pre-determined dataset is usually of inferior performance Ramamurthy et al. (2022). The reason behind this observation lies in the offline RL theory (see, e.g., (Xie et al., 2021; Jin et al., 2021; Xiong et al., 2022)), which suggests that the model’s performance in offline learning heavily depends on the coverage of the offline dataset. Specifically, to compete with the optimal policy in Eq. (2), even for a finite-state and finite-action case, the dataset should well capture every state-action pair that the optimal policy may visitMathematically, the ratio between the visitation probability of the optimal policy and the empirical distribution of the dataset should be uniformly bounded for every state-action pair. See Assumption A of Xie et al. (2021) for details. Nonetheless, fulfilling this prerequisite is arduous in practice due to the exponentially vast number of potential outputs.
This motivates us to involve further explorations with the environment into algorithmic design. The idea is to utilize the trained generative model, to generate additional samples and reinforcing the dataset. To ensure the quality of these newly collected samples, for each prompt, we may sample responses from the model and take the response with the highest reward. Then, we can fine-tune our model with these best-of- samples to improve the model. This process can be iterated for multiple times as the improved generative model in turn provides a better approximation of Eq. (2), leading to further enhancements for the model.
Specifically, the learning process of RAFT can be divided into three steps. For each stage ,
Step 1: Data collection. We first sample a batch of prompts from and generate for each , where is the parameter to control the output diversity. Step 2: Data ranking. In this step, we first use the reward model to compute for each . Then, we simply take and go through all the prompts and collect a subset of size . Step 3: Model fine-tuning. Then, we simply fine-tune the current model on and the next stage begins.
We will iteratively alternate among these three steps until the reward converges. The proposed framework admits a minimal hyper-parameter configuration, as summarized in Table 1 and is also easy to implement. A clear and elegant interpretation of RAFT is that the model iteratively learns from the induced best-of- policy (Nakano et al., 2021; Cobbe et al., 2021), which samples K responses and selects the one with the highest reward as the final output. It has been observed that the best-of- policy is competitive with the RLHF baseline (Nakano et al., 2021) across diverse scenarios. The best-of- policy can be viewed as a way to guide the inference using the reward model, although it incurs high inference costs. Conversely, RAFT iteratively learns from the induced best-of- policy, thereby improving the model.
We also note a distinct feature of RAFT that the data filtering is based on reward ranking instead of the absolute reward value, making RAFT less sensitive to the reward scale. We further hypothesized that RAFT is more robust against reward noise (variance and bias), which are known to be critical for the performance of PPO (Engstrom et al., 2020). We provide some evidences for our hypothesis in Appendix A.3.
3 Extension
Fluency/diversity-related regularization. In practice, a typical compromise exists between reward learning and the response quality, as assessed by other criteria like fluency or diversity. It is possible to integrate these metrics into a loss function , which evaluates the quality of generator . Consequently, the overall objective function can be represented as
A commonly used regularizer (Ziegler et al., 2019) is the KL divergence to the initial model:
which is used to reduce the disagreement and prevent the model from overfitting reward. The reason why we choose KL divergence in this form instead of the symmetric Jensen-Shannon divergence or the inverse form is that to achieve a rather small KL divergence, Eq. (4) will not assign much probability to the responses where the initial model will output them with a small probability ( is small). In particular, if some response is impossible in the initial model, this form of KL will also inhibit the updated model from generating them. We can integrate such a regularizer into our framework by considering the following modified reward
where is the coefficient to balance the goal of reward learning and keeping a low KL divergence. To incorporate the KL divergence, we simply further query the logits of the samples in step with both the current model and the initial reference model, and then rank the samples using Eq. (5).
Computational consideration. A notable property of RAFT is that the data collection stage is completely decoupled from the model improvement stage. For instance, we do not keep track of all operations performed on the data collection stage for the subsequent backward propagation. This allows us to implement the three steps separately and load only one model at a time. Therefore, as long as the computation source and memory source permit SFT on some specific model, the alignment process can be done with RAFT. In contrast, the on-policy PPO algorithm typically requires loading 4 models at the same time, including the trained model, the reference model (for KL estimation), the critic model, and the reward model. Moreover, considering the implementation of RAFT, one can use batch inference and model parallelism to accelerate data collection.
LLM Experiments
Model, Dataset, and Setup. We perform the experiment with the LLaMA-7B model (Touvron et al., 2023) and the HH-RLHF (Helpful and Harmless) datasethttps://huggingface.co/datasets/Dahoas/full-hh-rlhf (Bai et al., 2022a), which is collected for model alignment according to human preferences. The dataset consists of 112K training samples and 12.5K test samples. Each sample of the HH-RLHF dataset consists of a prompt and two responses: “chosen” and “rejected” where is the preferred compared to . See Table 2 for an example of the dataset. All the experiments are conducted using 8A40 (48G) with 600G RAM, and half-precision training (bf16). The code will be publicly available on GitHub in the camera ready version.
We follow the training procedure outlined by Ouyang et al. (2022), including SFT, reward modeling, and RLHF. Specifically, we first fine-tune the LLaMA-7B model with the chosen responses in the 112K training samples for 1 epoch to obtain LLaMA-7B-SFT. Then, we train a reward model based on the Open-LLaMA-3B (Geng & Liu, 2023) following the method in Ouyang et al. (2022) (Appendix A.1). The obtained reward model achieves a validation accuracy of , outperforming the GPT-J-6B modelhttps://huggingface.co/Dahoas/gptj-rm-static with an accuracy of . Then, we conduct RLHF experiments using the LLaMA-7B-SFT as starting checkpoint.
Prompt dataset. We use a context window of 256 tokens and discard the prompts with more tokens to reduce the GPU memory cost. This results in a prompt set of 82147 samples (originally 112K).
Competitor. We use the prominent approach in RLHF, PPO (Schulman et al., 2017) as our baseline. We implement the PPO algorithm with the TRL packagehttps://github.com/lvwerra/trl, which requires loading multiple LLMs concurrently and thus requires a significant amount of memory. Even with half-precision training, the out-of-memory error happens when we compute intermediate values during the training (e.g. attention scores). Following TRL, we use Parameter-Efficient Fine-Tuning (PEFT) in our experiment with the peft library, and perform Low-Rank Adaptation (LoRA) (Hu et al., 2021) for PPO with all the experiments. Note that it is possible to train the reward model using a larger base model and achieve better accuracy. However, we encountered an out-of-memory error when attempting to train PPO using 8A40 (48G) with a 7B reward model. Notably, since the data generation, data ranking, and SFT in RAFT can be performed separately, as long as we can fine-tune the model, we can also align the model with RAFT.
Generation and test configuration. For the generation configuration, we allow the model to generate up to 128 new tokens given the prompt. For RAFT algorithm, we will try out different temperatures, which would be specified in the individual experiment. For PPO algorithm, we follow the setting in TRL package and do not tune the generation configuration because it seems that the KL estimation can fail when we use a more complicated generation configuration. For a fair comparison, we keep the test configuration for all methods and report the metrics on a hand-out test set of size 4608. The perplexity is evaluated on 6K hand-out samples with the chosen responses. The detailed configuration can be found in Appendix D.
Hyper-parameters. For RAFT, we fix the batch size as and the learning rate for SFT as . For each SFT stage, we train for 2 epochs and use a linear decay scheduler. Other hyper-parameters will be specified for each experiment. For PPO, we adopt most of the parameter settings in TRL package. It is known that for the PPO, an explicit KL penalty is crucial for the training stability and to mitigate reward hacking (Ramamurthy et al., 2022). Without the KL penalty, the fluency of the language model (perplexity) and the diversity of the output degrade significantly as reward increases. Therefore, we primarily tune the weight of the KL penalty due to the different output lengths between the TRL example and our setup where we search in the space of . We also tune the learning rate in . For the KL regularization, we follow (Ziegler et al., 2019) to set the KL coefficient to be dynamically adapted (default setting of TRL package). The full list of hyper-parameters can be found in Appendix D.
Evaluation Metrics. The mean reward evaluated on the hand-out dataset and the perplexity are the main criteria for us to evaluate models and we also take the diversity metrics (Table 3) (Ramamurthy et al., 2022) into consideration, including Mean Segmented Type Token Ratio (MSSTR) (Johnson, 1944), the Distinct-1, Distinct-2 (the ratio of distinct n-grams over all n-grams) and the Unique-1, Unique-2 (Li et al., 2015) (count each n-gram in the texts only once). The fluency of the LLM (perplexity) and the diversity of the output typically degrade as reward increases, which is referred to as the alignment tax in the literature (Askell et al., 2021). All the diversity metrics are evaluated using the public projecthttps://github.com/GEM-benchmark/GEM-metrics as in Ramamurthy et al. (2022).
Interpretation. We list the evaluation results in Table 3, which consists of the results of RAFT and the best PPO models as the baseline. As we can see, the LLaMA-7B-SFT achieves a reward of , outperforming the original LLaMA-7B model. Both the RAFT and PPO can further improve the rewards compared to their starting checkpoint LLaMA-7B-SFT and also the preferred responses in the original dataset (). Among them, the RAFT-aligned model achieve the highest mean reward , while preserving a moderate perplexity . This proves that RAFT can stably optimize the LLMs with respect to a given reward model. In comparison with PPO, the RAFT-aligned model achieves a better perplexity and tends to respond with more details as its average response lengths are longer than the PPO-aligned one (we provide examples in Appendix B.1). We also find that the RAFT-aligned model with temperature consistently outperforms the SFT model in terms of the diversity metrics, which suggests the potential to employ the proposed framework to performance improvement beyond the scope of alignment.
GPT-4 and Human Evaluation. In addition to the reward, we also use GPT-4 (OpenAI, 2023) and human evaluation to measure the performance of the aligned models on randomly sampled 100 test prompts, where the results are provided in Table 4. To mitigate the issue that the GPT-4 evaluation may be influenced by the order in which the responses are provided, we conduct two experiments by switching the input order. The detailed problem setup and prompts for GPT-4 are provided in the Appendix A.4. As we can see, both the GPT-4 and human evaluation results are consistent with the automatic metrics. We also found that human tends to give more feedback of “Tie”, while the feedback of GPT-4 is more decisive.
Learning curve. We use the RAFT with and temperature as an example and report the training curve in the left part of Figure 1. In this typical RAFT experiment, the agent (blue line) learns from the best-of- policy (orange line), and the reward gradually increases. Meanwhile, the induced best-of- policy also improves along the line of the RAFT agent, which in turn further boost the performance of the RAFT agent. We also find that the perplexity is rather stable across the RAFT training, while the perplexity of the PPO agent usually gets worse quickly as the reward increases. To demonstrate this, we report the test reward with respect to the perplexity in the right part of Figure 1 for RAFT-K32- and also two PPO baselines. As we can see, RAFT agent achieves a better balance between reward and perplexity after the reward exceeds the threshold of . While we do observe that SFT changes the model rather significantly at the initial stage, it may not outperform PPO if we expect slight model modification.
Computation overhead. We conducted experiments for both the RAFT and PPO algorithms without early stopping, and the model is considered to be convergent if it oscillates around a fixed reward level for three consecutive iterations. We report the wall-clock time of RAFT with sampling temperature and different , averaged over three independent runs. For , the wall-clock times are 5 hours, 6.05 hours, and 7.05 hours, respectively. As increases, the inference time grows, which is main reason why a larger leads to a longer overall training time. On the other hand, we note that and typically converge faster with 10-12 iterations, while takes about 15-18 iterations to converge. The faster convergence rate partially compensates for the extra inference cost and helps to mitigate the overhead associated with loading models when RAFT switches between different stages of RAFT training. In comparison, the fastest-performing PPO configuration, with a KL penalty of 0.01 and LoRA training, converged in approximately 8.7 hours, which is slower than all the RAFT experiments with full training. Moreover, we note that RAFT trains in an off-policy manner and the inference and policy improvement are decoupled (the policy to improve can be different from the policy to collected samples). Therefore, any techniques that can speed up the inference can be readily integrated into the proposed framework. One straightforward option is to leverage the speculative decoding (Leviathan et al., 2023) for potential 2X-3X acceleration in inference with certain Large Language Models (LLMs). In contrast, PPO may not benefit from these inference techniques as its backward propagations require the gradient record in the forward pass.
2 Impacts of Hyper-parameters and Data Ranking Criteria for RAFT
Impact of . Since RAFT approximates the response with highest reward across the whole space by independent samples from the current model, it is clear that a larger leads to better performance. Meanwhile, the mean reward of the best-of- policy is increasing in . Specifically, suppose that the reward function is bounded by , a direct application of standard concentration inequality (e.g., Exercise 12, Chapter 2 of Wainwright (2019)) implies that the mean reward of the best-of- policy satisfies
Therefore, a larger typically leads to a better objective for the RAFT agent to iteratively learn from. On the other hand, the upper confidence bound is proportional to , so the marginal benefit diminishes quickly, which motivates us to adopt the iterative framework. We compare the performances of RAFT under with temperature and report the learning curve and model summarization in Figure 2 and Table 5. As we expect, as increases, the obtained model tends to achieve a higher test reward on the hand-out set. Meanwhile, the diversity metrics of the RAFT-K32 are never worse compared to and . However, a larger means a longer inference process (including data generation and reward computation). Therefore, in practice, we may balance the training cost and model performance, and use the largest within the range that the computational resource permits.
Impact of Sampling Temperature. In addition to the choice of , we can also modify the sampling temperature to control the diversity of the output. In particular, a higher temperature means that sampled responses are more diverse. To test the effect of temperature, we conduct experiments with and with . We report the results in Table 6. We find that for all three choices of temperature, RAFT consistently improves the reward to a rather stable level. The final test reward slightly gets worse as the temperature increases because the learning objectives, i.e., the reward of best-of- policy decreases as increases, as shown by the forth column of Table 6. The impact on reward, however, is less than . This may be because the best-of- policies also improve as iteration increases and may also because the higher temperature leads to better generalization as we found that for , the test reward is much lower than the training one. We can always compensate this by a larger as demonstrated in the last line of Table 6. On the other hand, a larger temperature consistently leads to a more diverse output for the final models, as we can see the model aligned with achieves the best diversity metrics compared to other choices of temperature and also the SFT model. One may try out even higher temperature but due to the limitation of model capacity, the LLaMA-7B-SFT may generate some responses with random and weird symbols, leading to an unstable learning process. Therefore, in practice, we can tune the temperature parameter by inspecting the filtered dataset from the initial SFT model to ensure a stable generation quality. To achieve the best performance, we may use the largest one within the range of a reasonable generation process and use a larger to compensate the reward decreasing in the objective policy from the higher temperature.
KL-penalty. While we observe that even though we do not impose any explicit restrictions in model update, the RAFT-aligned model is stable in perplexity and diversity metrics, it is helpful to understand the impact of KL regularization in RAFT. We conduct experiments with and temperature , and with different KL coefficients . We report the trend of KL divergence between the current model and initial model in Figure 3, where for each experiment we stop when the best model is obtained, and report the model metrics on Table 7. Across all the KL penalties, the RAFT-aligned models consistently outperform LLaMA-7B-SFT except for the perplexity. We find that a larger KL penalty can prevent the aligned model from moving award too far from the initial model as it attains a smaller KL divergence in terms of the initial model. On the other hand, the reward learning would be also affected as the final test reward decreases as the KL coefficient increases. We also find that the perplexity and diversity metrics are rather stable across the different KL penalties in contrast to the PPO training where the KL penalty leads to a better perplexity. Therefore, the KL penalty mainly serves to balance the reward learning and model update. However, computing the KL requires additional forward operation to get the logits from both the trained model and initial model. In practice, one can decide whether to incorporate such a regularization according to their customized needs (whether there is an explicit KL constraint).
3 Distillation
We have explained that since the data generation and model fine-tuning are separated in RAFT, we only need to load one model at a time, in contrast to the four models loading requirement of PPO. Another advantage of this property is that the RAFT can be implemented in an off-policy manner, which means that the data sources can be quite diverse beyond the model itself. In particular, in practice, we may want to align a series of models with different sizes (e.g. LLaMA-7B, LLaMA-13B, and LLaMA-70B) and use different models according to the customized needs of the scenarios (typically, a trade-off between the inference speed and the response quality). In this case, we may only use the most powerful LLaMA-70B to generate data responses and use the same samples to train the three models.
We investigate such an idea using the GPT-Neo-2.7B as our base model and use the LLaMA-7B model as the teacher. Specifically, the the teacher starts with LLaMA-7B-SFT and uses and temperature . We report the test reward curves in Figure 4 and the model evaluation metrics in Table 8. The model following the RAFT-LLaMA-7B-K32 consistently outperforms the model trained with only its output in both reward learning and diversity metrics. Moreover, we find that the perplexities of the aligned model also improve compared to the starting checkpoint GPT-Neo-2.7B, where we speculate that because we do not perform SFT first and the starting checkpoint does not well capture the knowledge of HH-RLHF dataset. This may also suggest that we can use RAFT in a more general sense beyond the alignment scenario (e.g. boost the model performance in mathematics).
Diffusion Model Experiments
We consider to use Stable-diffusion v1.5 (SD-1.5) as our visual generative model (https://huggingface.co/runwayml/stable-diffusion-v1-5). For all experiments, we use AdamW optimizer with fixed learning rate. It should be noted that for image-related tasks, CLIP (Radford et al., 2021; Ilharco et al., 2021), as a text-image matching score function, can be effectively utilized as a reward function to evaluate the degree of a certain concept. When the prompt is not available, it is still feasible to improve the model with general score function, such as aesthetic score. For efficient fine-tuning, we use LoRA (Hu et al., 2021) in our experiments. All our experiments are performed on NVIDIA A100 (40G) and A40 (48G).
Resolution adaptation. Although Stable diffusion was initially trained on a resolution of , due to catastrophic forgetting, SD-1.5 struggles to generate images at this resolution. However, we emphasize that by using a small number of generated samples and the RAFT algorithm, we can restore SD’s ability to generate images at resolution. The reward function is chosen as the CLIP-based aesthetic predictor (https://github.com/LAION-AI/aesthetic-predictor). We use the CIFAR-10 labels as our prompts (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck). Figure 5 has clearly demonstrated that with proper reward function, RAFT algorithm can improve the image quality significantly. We also show that the out-of-domain prompts (such as CIFAR-100 labels) can also be improved significantly. Table 9 suggests that both in-domain and out-of-domain scores are significantly improved. We evaluated the task using the state-of-the-art DDPO alignment algorithm for diffusion models (Black et al., 2023). Although DDPO achieves performance metrics similar to our approach, its computational overhead is approximately higher. This discrepancy arises because, while we frame the challenge as a contextual bandit problem suitable for general generative modeling, DDPO defines the iterative diffusion process as a Markov decision process. Such specialization enhances DDPO’s adaptability to diffusion models but compromises its extensibility to non-iterative generative models. Despite DDPO’s design insights into specific generative models, the computational cost could be a limiting factor. Regarding the RAFT algorithm, there is potential to compute rewards for intermediate states and select the best from samples. Refining RAFT’s sampling and reward modeling remains an open area of exploration.
Text-Image alignment. For resolution, SD-1.5 generally produces satisfactory outcomes. The main determinant affecting the generated outputs of SD-1.5 lies in the presentation method of prompts, which is because of the inductive bias in training data. Thus, the observed bias in the generated samples is more directly associated with the prompt delivery process. For example, the generator usually puts too much importance on the “style” information and ignore the objects. In such cases, we employ CLIP to evaluate the generated results and utilize the RAFT algorithm to achieve better alignment between the image and text prompts. Specifically, we use the OpenCLIP score with prompt input as the reward function (https://github.com/mlfoundations/open_clip). Figure 6 provide an illustrative case to demonstrate the lack of proper alignment between SD-1.5 and textual data. It is fortunate that our proposed RAFT algorithm can facilitate the attainment of well-aligned outputs through fine-tuning.
Discussion and Conclusion
In this paper, we proposed a simple but effective alignment framework, Reward rAnked FineTuning (RAFT), for aligning generative models to human preference using a reward function. Compared to the popular PPO algorithm, RAFT is easy to implement and tune with a simple parameter configuration, and typically converges more robustly and faster than the DRL approach PPO because of the SFT-like training feature. Another notable distinction between RAFT and the on-policy PPO is the decoupling of data generation and fine-tuning processes. This decoupling enables RAFT to be implemented 1) with less GPU memory source and 2) flexibly in terms of data sources and collection strategies.
Another potential advantage of RAFT is its interpretability. We can interpret RAFT as iteratively learning from the induced best-of- policies. In our study, we have demonstrated that the performance of RAFT heavily depends on the quality of the data set derived from the best-of- policy, which depends on the hyper-parameter choices. In a broader context, any strategies for improving inference, such as prompt engineering and advanced generation strategies, can also be integrated into the RAFT framework to further boost the performance of aligned models. Furthermore, the clear learning objective of RAFT enables us to mitigate the fundamental issue of reward hacking (Michaud et al., 2020; Tien et al., 2022), which is a common concern in RLHF. By monitoring the filtered dataset, we can mitigate the imperfections of the reward model used in RLHF and prevent algorithms from exploiting these imperfections to chase high rewards. We hope that the RAFT framework will enrich the toolbox of RLHF, thereby catalyzing additional investigation and enhancement in the alignment of foundational generative models.
References
Appendix A Details of LLM Experiments
We follow the training procedure outlined by Ouyang et al. (2022). First, we perform SFT on the 112K positive training samples of the HH-RLHF dataset. Then, we use 112K pairwise samples and the first 6275 pairwise samples in the test set of the HH-RLHF dataset for reward modeling and use the rest of the HH-RLHF test set as a handout evaluation setWe involve the hand-out set into reward modeling for a more reliable test procedure when evaluating the aligned models from RAFT and PPO with the hand-out set..
We consider the Bradley-Terry (BT) model (Bradley & Terry, 1952), which gives
where is the sigmoid function. Given a dataset of consisting of where is preferred by the human, we can maximize the likelihood (MLE) by minimizing the following loss:
where is the predicted reward of the model for prompt and response , and is the empirical distribution of the training set.
We report the hyper-parameters in Table 10 where we adopt the same parameters for two reward models and report training curves in Figure 7.
We note that the Open-LLaMA-13B outperforms the Open-LLaMA-3B in terms of both evaluation loss and evaluation accuracy. However, the PPO model requires loading the language model and reward model at the same time. During our current implementation with TRL, we encountered an out-of-memory error when attempting to train the model using 8A40 (48G). Therefore, we choose the Open-LLaMA-3B as our reward model in this experiment. Notably, since the data generation, data ranking, and SFT in RAFT can be performed separately, we can run RAFT with the 13B reward model in our experiment setup.
In practice, we will subtract a scalar baseline so that the starting policy of PPO is approximately of reward (Gao et al., 2023). In our setup, we use for the Open-LLaMA-3B and for the open-LLaMA-13B, respectively. Note that recentering the reward function with a fixed baseline will not influence the RAFT as RAFT is based on ranking and is less sensitive to the scale of reward function. We adopt this recentering operation as this typically leads to a more stable training for PPO.
A.2 RAFT Extension and Variant
From the experiment results presented in Section 4, the performance of the RAFT-aligned models heavily relies on the quality of the generated data. In what follows, we discuss several potential approaches to further improve the quality of the generated samples for future study.
Advanced generation strategy. In Section 4, we mainly adjust the hyper-parameters of RAFT in our experimental setup. In a more general sense, any methods that can improve the data generation quality will also contribute to the performance of the aligned model. As an extension, we may consider more advanced search methods, including the beam search (Reddy, 1977), top- sampling (Fan et al., 2018), top- sampling (Holtzman et al., 2019), contrastive search (Su et al., 2022).
Postprocessing to avoid reward hacking. One distinct feature of RLHF compared to the standard RL setting is that the reward function modeled from human preference is far from perfect. In practice, this imperfection can be easily to be exploited by the reward optimization algorithm to chase for a high reward. In an earlier version of the LLM experiment, the reward model mistakenly favors the responses containing emoji and notation #. The model’s output probability then quickly collapses and tends to output emoji and # in random positions of the responses. This was detected by the quickly decreased diversity metrics of the filtered dataset. To address this issue, we simply further filtered the collected dataset of RAFT to either clean the samples or just delete these samples. This is applicable because of the decoupled nature between the data generation and fine-tuning in RAFT.
Global ranking. While we present the RAFT algorithm in a local ranking manner, meaning that we rank the samples under the same prompt, we may also implement RAFT in a global ranking manner. In this case, we sample a batch of prompts and generate response for each prompt. Then, we compute the rewards for each sample and take the percent of samples with the highest reward as the training samples . As the reward modeling in LLMs (Appendix A.1) is based on the ranking under the same prompt, the prompt has a large impact on the reward and the comparison across different prompts is meaningless. Therefore, we mainly adopt the local ranking in this version. However, global ranking is more sample-efficient than the local ranking and is applicable when the rewards comparison are meaningful with different prompts.
A.3 Reward Imperfection and Reward Over-optimization
While the main focus of this paper is to present an alternative alignment framework for generative foundation models given a pre-determined reward function, it is worth noting that in the literature, there has been a growing interest in the reward hacking issue due to the reward imperfection (Michaud et al., 2020; Tien et al., 2022; Gao et al., 2023; Casper et al., 2023). For completeness, in this subsection, we present an initial study of the the reward hacking issue (with RAFT), which should serve as a motivator for future works in this important topic as in Gao et al. (2023).
Reward calibration issue. Given that the reward model serves as an MLE estimator for the BT model, it allows us to determine the predicted probability of the response pair’s order. This facilitates an examination of the reward model’s calibration based on the constructed reward model, which is common for LLM (OpenAI, 2023). Intuitively, the predicted probability should align with the accuracy computed by the labels in the HH-RLHF dataset. This is particularly pertinent for PPO algorithms which are more sensitive to the scale of the reward signal. We plot the calibration curve of the Open-LLaMA-3B in Figure 8. The reward model demonstrates a tendency toward pessimism when the predicted probability is low, since the predicted probability being smaller than the actual accuracy achieved. Conversely, it exhibits overconfidence when the predicted probability is high. We hypothesized that RAFT would be also less sensitive to the scale mismatch from calibration issue compared to PPO. The reason is that RAFT only takes the ranking information, while PPO further depends on the scale of the reward signals. Indeed, it is known that the code-level optimizations (such as reward recentering, clipping, and normalization) are critical for the success of PPO (Engstrom et al., 2020). A thorough examination of the impact of reward calibration, as well as developing improved training methodologies for the reward model, is crucial for RLHF. However, these aspects surpass the scope of the current paper and are earmarked for subsequent work.
Noisy rewards. In practice, the reward signals can be noisy, stemming from either the reward modeling process, or the usage of the reward models. For instance, in the RLHF process of LLaMA-2 (Touvron et al., 2023) and GPT-4 (OpenAI, 2023), the alignment objective is split into different alignment goals (e.g. helpful and safety) and independent reward models are trained for each goal. Then these reward models are combined by (LLM-based) classifiers and human-written rules, which may lead to noisy rewards due to the classification errors. While the impact of reward calibration requires more involved and comprehensive experiments, we can test our hypothesis with random noise on the reward signals to provide some initial evidences. We run RAFT-K32- and PPO without recentering the reward function. Meanwhile, we add random noises to the reward function in the following three ways:
where denotes the random noise with mean and variance . Note that in the second case, we either use the original reward with probability or add noises for all the responses given the prompt with probability , while in the third case, the noise is sampled for each single response independently. The second case may resemble the classification error in terms of the prompt, and the third case may resemble the classification error taking both the prompt and response into consideration. We report the training curves in Figure 9. As we can see, the noisy environment and also the bias in the noises make the PPO converge much slower and also a deduction on the true reward. This is because any modification in the collected reward contributes to a noisy training in the critic and also the actor of PPO. Moreover, when a bias exists in the noise, the PPO eventually converges to another model. In comparison, RAFT-K32- is more stable compared to the run without noise. Intuitively, the noise only makes us select the sub-optimal samples into the training set in some cases but the collected training set is still of a much higher reward compared to the current model. Moreover, since RAFT is invariant to linear transformation, it is also less sensitive to the bias existing in the noise. In particular, when the bias (the mean of the noise) only depends on the prompt (the second case), RAFT performs well regardless of the presence of random noises.
Reward Over-optimization. Another general issue is that the reward model is always imperfect and the LLMs can exploit these imperfections to chase for a high reward, leading to reward hacking (Gao et al., 2023). To recover such an observation, we additionally train two reward models based on GPT2 (124M) (Radford et al., 2019), and GPT-Neo-1.3B (Gao et al., 2020), respectively, and report their accuracies in Table 11. We use the reward model based on Open-LLaMA-3B to approximate the gold reward model and call the other two sub-optimal reward models the proxy reward models. Then, we run RAFT and PPO with respect to the proxy reward models but also record the gold reward along the way. Since the reward models have different scales, we normalize the reward model by subtract the minimal reward across all experiments and then divide it by the maximal reward across all experiments so that it starts approximately from zero and with a largest value of zero. This is sufficient for us to observe the trends. We plot the results in Figure 10. As we expected, we observe the reward over-optimization issue (Gao et al., 2023) for both RAFT and PPO, where the gold reward first increases as the proxy reward increases in the first stage, and then the gold reward decreases even though the proxy reward still increases or oscillates at a fixed level. In comparison, with the GPT2-RM, RAFT achieves much better peak rewards for both the proxy RM and the gold RM, but the gold reward decreases significantly as the LLM further overfits the proxy RM. With a GPT-Neo-1.3B, RAFT also achieves slightly better peak rewards for both the proxy RM and the gold. Meanwhile, since the proxy RM and gold RM are more consistent, the degree of overfitting significantly reduces. Another observation is that the gold reward drops typically when the proxy reward oscillates at a fixed level or increases in a much slower rate for both RAFT and PPO, suggesting that we should early stop in practice to mitigate the overfittig issue.
In summary, training a more accurate and well-calibrated reward model plays a central role in RLHF and also applies to the RAFT algorithm. A more comprehensive study is necessary in the future work.
A.4 GPT-4 and Human Evaluation
For human evaluations, we have 7 human experts to evaluate the output pairs without seeing the label and the order was shuffled. For GPT, we use GPT-4-0613 API to compare the outputs. The prompt is as below.
System Message: Please act as an impartial judge and evaluate the quality of the responses provided by two AI assistants to the user question displayed below. You should choose the assistant that follows the user’s instructions and answers the user’s question better. Your evaluation should consider factors such as the helpfulness, relevance, accuracy, depth, creativity, and level of detail of their responses. Begin your evaluation by comparing the two responses and provide a short explanation. Avoid any position biases and ensure that the order in which the responses were presented does not influence your decision. Do not allow the length of the responses to influence your evaluation. Do not favor certain names of the assistants. Be as objective as possible. After providing your explanation, output your final verdict by strictly following this format: [[A]] if assistant A is better, [[B]] if assistant B is better, and [[C]] for a tie. Prompt Template: [User Question]\n\n{Question}\n\n[The Start of Assistant A’s Answer]\n{Answer A}\n[The End of Assistant A’s Answer]\n\n[The Start of Assistant B’s Answer]\n{Answer B}\n[The End of Assistant B’s Answer]
Appendix B Examples
B.2 Diffusion Model Samples
All the experiments of diffusion models are performed with Nvidia-3090 with 256G RAM. We will release the code and demos for our paper.
Specifically, Figure 11 depicts the samples generated during resolution adaptation without any cherry-picking involved. It is evident that our approach has significantly improved the quality of generated samples.
It is worth noting that in our experiments conducted at a resolution of 256256, significant improvements were observed not only for the prompts used during training but also for other prompts. For instance, when using CIFAR-10 labels as samples, notable improvements in the generated quality were observed when utilizing CIFAR-100 labels (Figure 12). This observation highlights the generalization capability of our RAFT algorithm in enhancing sample quality during the alignment process.
Furthermore, we have included additional examples of Text-Image Alignment in Figure 13, further demonstrating the crucial role of RAFT alignment in diffusion models.
Appendix C Usage of RAFT in LMFlow
LMFlow (Diao et al., 2023) (https://github.com/OptimalScale/LMFlow) is a public package, which aims to provide a general and easy-to-use framework for researchers and engineers to finetune/align models. To run example code of RAFT alignment in LMFlow, one may simply execute:
By default this aligns GPT-2 base model (Radford et al., 2019) with the proposed RAFT algorithm on IMDB dataset (Maas et al., 2011). To specify LLaMA as the model, one can change the following option in the script:
--model_name_or_path {path-to-downloaded-llama-model}
with an extra option “--use_lora 1” if LoRA is applied during the alignment process.
We also added the diffusion demos in the LMFlow package.