Large-scale Reinforcement Learning for Diffusion Models
Yinan Zhang, Eric Tzeng, Yilun Du, Dmitry Kislyuk
Introduction
Diffusion probabilistic models have revolutionized generative modeling, particularly for producing creative and photorealistic imagery when combined with pre-trained text encoders . However, the resulting image quality is highly dependent on the distribution of the pre-training dataset, which typically consists of web-scale text-image pairs. Although pre-training on massive weakly supervised tasks of this form is effective in exposing the text-to-image model to a wide range of prompts, downstream applications often observe weaknesses around the following properties:
Fidelity and controllability : failing to accurately depict the semantics of the text prompts (e.g. incorrect composition and relationships between objects)
Human aesthetic mismatch : producing outputs that humans do not perceive to be aesthetically pleasing
Bias and stereotypes : presenting or exaggerating societal bias and stereotypes
To address these challenges, several works have explored classic fine-tuning techniques for pre-trained diffusion models with curated data, either to improve the aesthetic quality of the model outputs with human-selected high-quality images , or to eliminate existing biases in the model with synthetic dataset augmentation . Another approach, which bypasses the labor-intensive dataset curation, involves intervention in the sampling process to achieve controllability, by utilizing auxiliary input or refining the intermediate representations . However, this form of inference-time guidance results in an increase in the sampling time without improving the inherent capability of the model. A recent direction, motivated by the success of reinforcement learning from human feedback (RLHF) in the language domain , proposes fine-tuning diffusion models through full-sample gradient backpropogation on human preference reward models, though these approaches are memory intensive and only work for differentiable reward functions. Finally, RL-based optimization has enabled fine-tuning with arbitrary objective functions, but these methods have so far been limited in scope by focusing on a small set of prompts in a narrow domain, and lack the scale to improve model performance generally.
In this paper, we propose a generic RL-based framework for fine-tuning diffusion models, which works at scale across millions of prompts and with an arbitrary combination of objective functions. Our contributions are as follows:
We present an effective large-scale RL training algorithm for diffusion models which allows training over millions of prompts across a diverse set of tasks.
We propose a distribution-based reward function for RL fine-tuning to improve the output diversity.
We demonstrate how to perform effective multi-objective RL-training and illustrate how we can improve a base model across all objectives, which can include human aesthetic preference, fairness, and object composition.
We conduct extensive experiments and analysis studies comparing our approach with existing reward optimization methods across a suite of tasks.
Related Work
Existing reward fine-tuning methods for diffusion models can be classified into three categories: either supervised with reward-weighted data , optimized through gradient-backpropogation on the reward function or through reinforcement learning . Our work builds on work training diffusion models with reinforcement learning, but while past work has focused on simple settings (DPOK uses a training set of 1 prompt per model, and DDPO using simple set of 45 common animals and 3 activities), we illustrate how we can use reinforcement learning training across the scale of millions of prompts and different objectives.
Despite their remarkable capacity, current state-of-the-art text-to-image models still struggle to generate images that faithfully align with the semantics of the text prompts due to their limited compositional capabilities . Existing work addresses this by either modifying the inference procedure or by using auxiliary conditioning inputs such as bounding boxes or spatial layouts . Our method instead focus on improving the fidelity of existing SD models without using additional layout guidance.
Text-to-image generative models perpetuate and even amplify the societal biases present in the massive pretraining datasets of uncurated image-text pairs . Existing work addresses this by either using balanced synthetic data , with textual guidance during inference or with reference images of a particular attribute . Different from prior work, our method does not require synthetic data collection or inference-time intervention.
Method
In this section, we describe our approach for applying large-scale RL training to diffusion models. Our goal is to fine-tune the parameters of an existing diffusion model to maximize the reward signal of the generated images from the sampling process:
where is the context distribution, is the sample distribution, and is the reward function that is applied to the final sample image.
Following Black et al. , we reframe the iterative denoising procedure of diffusion models as a multi-step Markov decision process (MDP), where the policy, action, state and reward at each timestep are defined as follows:
We treat the reverse sampling process of the diffusion model as the policy. Starting from a sampled initial state , the policy’s action at any timestep is the update that produces the sample for the next timestep . The reward is defined as at the final timestep, and otherwise.
The policy gradient estimates can be made using the likelihood ratio method (also known as REINFORCE) :
We also apply importance sampling to enable collecting samples from the old policy for improved training efficiency, and incorporate a clipped trust region to ensure that the new policy does not deviate too much from the old policy . The final clipped surrogate objective function can be written as:
Here is the hyper-parameter that determines the clip interval, and is the estimated advantage for the samples. To further prevent over-optimization of the reward function, we also incorporate the original diffusion model objective as part of the loss function. Our full training objective is thus
One additional detail is that reward values are typically normalized to zero mean and unit variance during gradient updates to increase training stability. In policy-based RL, a general approach is to subtract a baseline state value function from the reward to obtain the advantage function
In the original implementation of DDPO, Black et al. normalize the rewards on a per-context basis by keeping track of a running mean and standard deviation for each prompt independently . However, this approach remains impractical if the training set size is unbounded or unfixed.
In contrast to the limited size of their training prompts (up to only), our large-scale fine-tuning experiments involve millions of training prompts. We instead normalize the rewards on a per-batch basis using the mean and variance of each training batch.
2 Distribution-based Reward Functions
In the previously outlined formulation of the diffusion MDP, each generation is considered independently, and thus rewards incurred by generated samples are independent of each other. This formulation is a natural fit for reward functions that only care about the contents of a single image, such as image quality or text-image alignment. However, sometimes what we care about is not the contents of any particular image, but instead the output distribution of the diffusion model as a whole. For example, if our goal is to ensure our model generates diverse outputs, considering a single generation in isolation is insufficient—we must consider the set of all outputs in order to understand these distributional properties of our model.
To this end, we also investigate the use of distribution-level reward functions for reinforcement learning with diffusion models. However, it is intractable to construct the true generative distribution. Thus, we instead approximate the reward by computing it using empirical samples across minibatches during the reinforcement learning process. During training, the attained reward is computed on each minibatch, and the minibatch reward is then backpropagated across the samples to perform model updates. In Section 4.2 we validate this approach by learning via a distribution-level reward function that optimizes for fairness and diversity in generated samples.
3 Multi-task Joint Training
We also perform multi-task joint training to optimize a single model for a diverse set of objectives simultaneously. As detailed in the next section, we incorporate the reward functions from human preference, skintone diversity, object composition and perform joint-optimization all at once. Since each task involves a different distribution of training prompts, in every training iteration, we sample multiple prompts from all the tasks and run the sampling process independently. Each reward model is applied to the corresponding sample image with the prompt. Then the gradient step from equation 7 is executed for each task sequentially. We outline the training framework in Algorithm 1 with hyper-parameters available in Appendix A.
Reward Functions and Experiments
To validate our method across a wide variety of settings, we perform experiments on three separate reward functions: human preference, image composition, and diversity and fairness. We begin with an introduction of the different reward functions we applied our method to.
To optimize diffusion models to adhere to human preferences, we use an open-source reward model, ImageReward (IR), trained on a large number of human preference pairs . ImageReward takes a pair consisting of a text caption and a generated sample, then outputs a human preference score, which is then used as the reward during training:
Our results with this human preference reward function are detailed in Section 4.1.
In order to encourage fairness and diversity across the samples generated by our model, following previous work , we leverage statistical parity, a metric commonly adopted for measuring biases in models, as a distribution-level reward function for our fine-tuning experiments. Given the generated distribution and a classifier that identifies a spurious attribute, we measure the L2 norm between the empirical and uniform distributions:
The reward attained by the model is then simply the negation of the statistical parity, so as to encourage the model to produce diverse samples. As explained in Section 3.2, it is intractable to compute the reward over the full output distribution of the model, so we compute the reward over individual minibatches. We present the results for this experiment in Section 4.2.
To improve the compositional skills of diffusion models, we devise a new reward function that uses an auxiliary object detector. We construct a set of training prompts, each containing multiple different objects, and use an object detection model on the image to predict the confidence score for each object class. The reward score is then defined as the average confidence score of all the objects:
where is the detection confidence score for the object class given input image . Our results on compositionality are detailed in Section 4.3.
Finally, we also experiment with jointly optimizing over all three previously described reward functions, to train a model that satisfies all three criteria simultaneously. We present the results of our joint optimization in Section 4.4. For all our fine-tuning experiments, we use SDv2 as our base model. The output resolution is x, which we consider as a good tradeoff between compute efficiency and image quality.
To fine-tune a diffusion model with human preferences, we use ImageReward , which was trained on large-scale human assessments of text-image pairs. In total, the authors collected 137k pairs of expert judgments on images generated from real-world user prompts from the DiffusionDB dataset . Compared to other existing metrics such as CLIP , BLIP , or Aesthetic score , ImageReward is better aligned with human judgments, making it better suited as a reward function.
We use a training set of 1.5 million unique real user prompts from DiffusionDB, among which 2,000 prompts were split for testing. We use 128 A100 GPUs (80GB) for all experiments, including the baselines. Experimental details, hyperparameters, and additional results are provided in Appendix A.
Prior reward fine-tuning methods for diffusion models mainly fall under three categories: reward-based loss reweighting , dataset augmentation , and backpropagation through the reward model . We compare against a variety of baseline methods, including ReFL , RAFT , DRaFT and Reward-weighted , covering the three different methodologies. We reimplement all methods and fine-tune them on SDv2 using the same training set of 1.5M prompts until convergence.
We show the qualitative and quantitative results of all baseline methods in Figure 3 and Table 1. We also provide training curves in Appendix E and note that, except for RAFT which diverged, all online-learning methods exhibit steadily increasing sample rewards during training, eventually saturating at some maximum level, at which point we consider the models converged. In contrast to the common belief that RL training is inefficient and slow to converge, our approach converges in as few as 1,000 steps, compared to DRaFT, the gradient-based reward optimization approach which takes 4,000 steps to converge while only being able to optimize for differentiable rewards. We provide a comprehensive comparison of all the reward optimization methods in Table 2.
For RAFT, we found the model diverges as the number of training iterations increases, similar to the finding from Xu et al. . Since RAFT uses the model-generated images with the highest rewards for fine-tuning the model, it is constrained by the diversity of the latest model’s generation and thus prone to overfitting. The reward-weighted method uses a similar idea of augmenting the training data using model-generated images and weighting the training loss by the reward values, but all the images are generated from the original model (in contrast to RAFT’s online generation using the latest model) and thus is less prone to overfitting.
Next, we evaluate our trained model’s ability to generalize to an out-of-domain test set, PartiPrompts . PartiPrompts is a comprehensive benchmark for text-to-image models, with over 1,600 challenging prompts across a variety of categories. We report the ImageReward and Aesthetic scores in Table 1, along with human evaluation results in Figure 2. When compared against each baseline model, our approach achieves the highest Aesthetic score and human preference rate on both sets.
We also achieve the second highest result on ImageReward, but note that this metric alone is not a robust indicator of performance, since the model was directly trained against it. Reward hacking is a commonly observed phenomenon in which models optimizing for a single reward function often overoptimize for this single metric at the cost of overall performance. We believe the high ImageReward scores achieved by ReFL are a result of this, and show example generations in Figure 4. The reward hacking problem of ReFL was observed by Clark et al. in their DRaFT experiments as well, where their fine-tuned model optimizing for Aesthetic score collapses to generate very similar, high-reward images. We hypothesize that gradient-based optimization methods (i.e. ReFL and DRaFT) are more prone to reward hacking due to their direct access to the gradients of the reward model. In contrast, our wins on human preference rate indicate that our method is more robust to these effects.
2 Optimizing Fairness and Diversity
The training of diffusion models is highly data-driven, relying on billion-sized datasets that are randomly scraped from internet. As a result, the trained models may contain significant social bias and stereotypes. For example, it has been observed that text-to-image diffusion models commonly exhibit a tendency to generate humans with lighter skintones . We aim to mitigate this bias by explicitly guiding the model using a skintone diversity reward.
For fine-tuning, we collect a dataset of 240M human images from Pinterest and run BLIP to generate captions for each image. Only the text prompts are used during training, and the reward calculation is based on the generated samples. We further filter out the captions containing terms relating to ethnicity and race (e.g. African, Asian, Indian) to ensure that the training prompts are race agnostic. In each training iteration, we load 128 prompts and generate a minibatch of 16 images for each prompt, then run a pre-trained skintone classifier on the generated samples and calculate the statistical parity for each minibatch according to Equation 12. Since the classifier has 4 skintone categories ranging from dark to light, the optimal reward is achieved when the output distribution is entirely uniform (i.e. 4 samples in each skintone bucket).
We show our qualitative results in Figures 1 and 5 and quantitative results in Table 3. We construct two test sets: a set of 100 randomly sampled occupations, for which we add the prefix “a portrait of” to produce the final prompts (e.g.“a portrait of a police officer”), and another set of 200 prompts from HRSBench , which are descriptions of people with random objects. We note that both are out-of-domain test sets, as their distribution is different from that of the BLIP-generated training prompts.
Our fine-tuned model greatly reduces the skintone bias embedded in the pretrained SDv2 model, especially for occupations with more social stereotypes or biases inherent in the pretraining dataset. For example, in Figure 5, we show that the pretrained SDv2 model is biased towards light skintone for portraits of dentists and judges, whereas our finetuned model generates a much more balanced distribution.
3 Optimizing Compositionality
While diffusion models are able to generate diverse images, they often fail to accurately generate different compositions of objects in a scene . We further explore using our RL framework in ensuring compositionality with diffusion models. We collect a list of 532 common object classes (e.g. apple, backpack, book, balloon, avocado; the full list is available in Appendix H) and use 450 of them for training. The remaining classes are withheld for testing. We then construct training prompts by combining two different objects using one of five relationship terms: “and,” “next to,” “near,” “on side of” and “beside,” producing captions that designate a spatial relationship between two objects, e.g. “an apple next to an avocado.” In total we create a training set of over 1M prompts. In order to compute our object composition reward function (Eq. 13), we use UniDet , an object detector trained on multiple large-scale datasets that supports a wide range of object classes.
We present qualitative and quantitative results in Figure 6 and Table 4. To evaluate generalizability, we also generate samples with our fine-tuned model on 300 randomly sampled prompts from both unseen and seen objects. Our trained model adheres better to compositional constraints in text captions when compared to SDv2, and the learned compositional abilities also generalize to unseen objects.
4 Multi-reward Joint Optimization
As detailed in Algorithm 1, we also perform multi-reward RL with all three reward functions jointly, aiming to improve the model performance on all three tasks simultaneously. We compare the jointly-trained model with the base model and the models fine-tuned for each individual task. The quantitative results are shown in Table 5, with more qualitative results available in Appendix B. Following the same evaluation setting, we test the models on multiple datasets for each metric and report the average scores.
While the best score for each metric is achieved by the model fine-tuned specifically for that task, our jointly-trained model is able to satisfy over 80% (relative) performance of the individually fine-tuned models across all three metrics simultaneously. In addition, it significantly outperforms the original base model on all tasks.
We observed degraded performance for individually fine-tuned models on some of the tasks that the models were not fine-tuned for. For example, the model optimized for human preference exhibits a significant regression on statistical parity, indicating a drastic drop in skintone diversity. Similarly, the model optimized for skintone diversity degrades in terms of human preference as compared to the base model. This is akin to the “alignment tax” issue that has been observed during RLHF fine-tuning procedure of LLMs . Specifically, when models are trained with a reward function that is only concerned with one aspect of images, it may learn to neglect sample quality or overall diversity of outputs. Our jointly fine-tuned model, in contrast, is able to mitigate the alignment tax issue by incorporating multiple diverse reward functions during fine-tuning, thereby maintaining performance on all tasks in question.
Conclusion
We present a scalable RL training framework for directly optimizing diffusion models with arbitrary downstream objectives, including distribution-based reward functions. We conducted large-scale multi-task fine-tuning to improve the general performance of an SDv2 model in terms of human preferences, fairness, and object composition simultaneously, and found that joint training also mitigated the alignment tax issue common in RLHF. By evaluating our trained model against several baseline models on diverse out-of-domain test sets, we demonstrated our method’s generality and robustness. We hope our work inspires future research on targeted tuning of diffusion models, with potential future topics including addressing more complex compositional relationships and mitigating bias along other social dimensions.
References
Appendix A Experiment Details and Hyperparameters
All our experiments including baseline methods training were conducted on 128 80GB A100 GPUs. If a pretraining dataset is required, all fine-tuning methods use the same 12M subset of LAION-5B filtered by the aesthetic score predictor with a threshold of 6. For optimization, we use the AdamW optimizer with and a weight decay of for all the experiments. For inference, we run the diffusion process with 50 steps for each image with DDIM noise scheduler. We use the default guidance scale of 7.0 for classifier-free guidance .
For our RL fine-tuning experiments, we collect 16x128 samples per training iteration, with 50 samplings steps using DDIM scheduler . We randomly sample 5 training timesteps and perform a gradient update across all the samples in the batch for each of the timesteps, resulting in 5 gradient updates per iteration. We use a small clip range of for all the experiments.
For the baseline methods including ReFL , Reward-weighted , RAFT and DRaFT , we refer to the original implementation for the suggested hyperparameters and report our experiment details in Table 6. We use the same training set for all the baseline models training and fine-tune them until convergence. Since the experiments involve million-sized training prompts, for reward-weighted approach, instead of pre-generating the samples and storing the dataset offline, we generate the samples on the fly during training using the original SDv2 model and re-weigh them according to the reward values for fine-tuning. Following Xu et al. , we also map the reward values to the range of $$ using min-max normalization.
We note that DRaFT imposes a high memory burden by directly back-propagating the gradient from the reward model through the sampling process of diffusion model, allowing for a much smaller batch size compared to other optimization methods. We implement DRaFT-LV, which claimed to be the most efficient DRaFT variant.
Appendix B Additional Qualitative Results
We provide additional qualitative results in this section, including results from the models that were trained with single rewards (i.e., ImageReward , compositionality reward and skintone diversity reward), as well as the results from our model that was jointly trained with all three rewards simultaneously.
We show the visual samples from our model fine-tuned with ImageReward on real-user prompts in Figure 7. We also provide more qualitative comparison of our model with other reward optimization methods in Figure 8. More results on the out-of-domain test set PartiPrompts are availble in Figure 9. Our trained model generates more visually appealing images compared to the base SDv2 model, and it generalizes well to out-of-domain test sets with unseen text prompts that have a different distribution from that of the training prompts.
B.2 Results on Optimizing Diversity
We provide more qualitative results of our model fine-tuned with skintone diversity reward in Figure 10. Our trained model effectively mitigates the inherent bias and stereotypes in the base SDv2 model with increased skintone diversity in the generated human samples.
B.3 Results on Optimizing Compositionality
We provide more qualitative results of our model fine-tuned with object composition reward in Figure 11. Our fine-tuned model demonstrates improved compositional skills compared to the base SDv2 model and other baseline models.
B.4 Results on Multi-reward Joint Optimization
Next, we show more qualitative results from our jointly-fine-tuned model (with all three rewards simultaneously) on multiple test sets: DiffusionDB (Figure 12), object composition (Figure 13) and occupation prompts (Figure 14). We demonstrate that our jointly-trained model has quite significant improvement over the base SDv2 model in terms of all three objectives: human preferences, skintone diversity and object composition. We further note that since joint training utilizes multiple reward signals (including ImageReward which reflects human preferences) during training, for portraits of occupations, we also observe additional increase in the aesthetic quality of the samples compared to single-reward training which optimizes for the skintone diversity only; see Figure 14.
Appendix C Human Evaluation Templates
We provide the detailed human evaluation guidelines document that were used to train our hired human labelers in section C.1, including the judging criteria and concrete examples for making trade-offs in order to help the evaluators better understand the task and make fair judgments. We use the annotation documents from ImageReward as a reference. We also show our evaluation UI interface in section C.2.
You will be given a number of prompts/queries and there are several AI-generated images according to the prompt/query. Your annotation requirement is to evaluate these images in terms of Image Fidelity, Relevance to the Query, and Aesthetic Quality. Below are more details on each of the three mentioned factors.
Definition: The generated image should be true to the shape and characteristics of the object, and not generated haphazardly. Some examples of low-fidelity images are:
Dogs should have four legs and two eyes, generating an image with extra / fewer legs or eyes is considered low-fidelity.
“Spider-Man”” (or human) should only have two arms and five fingers each. Generating extra arms / fingers is considered low-fidelity.
“Unicorn” should only have one horn, generating an image with multiple horns is considered low-fidelity.
People eat noodles with utensils instead of grabbing them with their hands, generating an image of someone eating noodles with their hands is considered low-fidelity.
See Figure 15 for examples of low-fidelity generation. Images of low fidelity should be ranked as low preference.
C.1.2 Relevance to the Query
Definition: the generated image should match the text in the query. Another term used for“Relevance” is “Text-alignment”. Some examples of inconsistent image generation are:
The subject described in the text does not appear in the image generated, for example, “A cat dressed as Napoleon Bonaparte” generates an image without the word “cat”.
The object properties generated in the image are different from the text description, for example, generating an image of “a little girl sitting in front of a sewing machine” with a boy (or many little girls) is incorrect.
See Figure 16 for examples of low-relevance generation. Images of low relevance to the query should be ranked as low preference.
C.1.3 Aesthetic Quality
Definition: the generated images should look visually appealing and beautiful. Examples are provided in Figure 17, where two images are generated given the same text prompt and the one with higher aesthetic quality is highlighted.
C.1.4 Overall Preference Ranking
Guidelines for deciding boundary cases: which generated images would you prefer to receive from AI painters? Evaluating the output of the model may involve making trade-offs between the criteria we discussed. These trade-offs will depend on the task. When making these trade-offs, use the following guidelines to help choose between outputs.
For most tasks, fidelity & aesthetic quality are more important than image-text alignment. So, in most cases, the image having higher fidelity and aesthetic quality is rated higher than an output that is more image-text aligned.
clearly matches the text better than the other;
is only slightly lacking in the requirements of fidelity;
the content does not have significant artifacts that would cause psychological discomfort
then the more consistent result is rated higher.
We provide more examples below to illustrate how to make trade-off between the different criteria when making judgements.
In the example above (Figure 18), image A and B are the ones that match the text description best, and they are also the most aesthetically appealing (A is better than B in both regards). The animals in image C look unnatural and have artifacts, also C does not align with the text very well. Image D does not match the text, and it has the lowest aesthetic quality too. Thus the overall ranking should be .
In the example above (Figure 19), image A and B both match the text (they are wallpapers of some anime style), and image B looks slightly more appealing, so we rank . Note that Image C has a lots of noticeable artifacts in the body parts of the anime character and it might cause psychological discomfort , so it should be ranked as the lowest. The overall ranking should be (D is better than C because of the significant artifacts in C).
In the example above (Figure 20), all four images are depiction of astronaut on mars during sunset, so they all match with the text well. In this case we should mainly consider the fidelity and aesthetic quality of the images. Among the four images, Image A and C look the most beautiful (with C slightly better than A). Image D has the lowest aesthetic quality compared to others. So the overall ranking should be .
In the example above (Figure 21), image C is somehow a nonsense generation and does not match the text, so it is apparent that C should be ranked the lowest. Image A, B and D all match with the text, and in terms of fidelity and aesthetic quality, they should be ranked as (B looks the most appealing, followed by A, while D only shows part of the palace and is not as beautiful as B). The overall ranking should be .
In the example above (Figure 22), image A and C both have fire in it, and image A looks more visually appealing. Note that although C is more colorful, we think image A matches with the text well enough; since A is much more visually appealing than C, we rank . B and D both have lower image-text alignment and lower aesthetic quality, so we rank them as the lowest two. The overall ranking should be .
C.2 Evaluation Interface
To compare our fine-tuned model with the base SDv2 model and models tuned with other baseline approaches, we perform head-to-head comparison of two images generated from different sources using the same text prompt. The two images are generated using the same random seed for fair comparison. The human evaluators were trained using the guidelines provided in section C.1. During evaluation, we show two generated images and the associated text query, and ask the evaluators to choose the preferred one based on image fidelity and aesthetic quality, as well as image-text relevance. We show the evaluation interface in Figure 23.
Appendix D Additional Human Evaluation
For a more thorough evaluation on the effectiveness of our method on improving compositionality and diversity, we also perform human evaluation on our models trained with compositionality reward and skintone diversity reward and provide the results in Figure 24. For the compositionality evaluation, the annotators were asked to rate the samples based on image-text relevance (how well the generated images match the text).
D.2 Skintone Diversity Evaluation
We provide the guidelines we used to train the hired human evaluators on rating skintone diversity in section D.2.1 and the evaluation interface in section D.2.2.
You will be given a number of prompts and there are several AI-generated images according to the prompt. Your annotation requirement is to evaluate these outputs in terms of skintone diversity. What we mean by “skintone diversity” is that AI model should have minimum bias and stereotypes, so the generated images should have a diverse set of people with different skintone. For example, “a portrait of a police officer” is not supposed to generate only light skintone or dark skintone; what we want instead is a balanced distribution of light and dark skintone. For each prompt, there are two sets (set A and set B) of images, each with 6 images. Your job is to choose the one set with better skintone diversity. See the examples below:
Note that in Figure 25 Set B is more diverse for both examples of portraits of a chemist and a builder, because it has a balanced distribution of light and dark skintone, while set A has mostly light skintone in it.
D.2.2 Interface
For evaluating skintone diversity, we perform head-to-head comparison of two groups of images generated from different sources using the same text prompt. The two groups are generated using the same random seed for fair comparison and the evaluators were asked to choose the one that has better diversity. The human evaluators were trained using the guidelines provided in section D.2.1. We show the evaluation interface in Figure 26.
Appendix E Training Curve
We plot the training curves of our method and other online learning baseline methods in Figure 27 and note that, except for RAFT which diverged, all online-learning methods exhibit steadily increasing sample rewards during training, eventually saturating at some maximum level, at which point we consider the models converged. Our method converged pretty quickly in as few as 1,000 steps.
Appendix F Reward Hacking
We found that ReFL is prone to reward hacking, a well known issue in RLHF . Specifically, since the reward model trained from human annotation data is far from perfect, the imperfection can be exploited by the algorithms to chase for a high reward, leading to reward hacking . We provide more visual examples of reward hacking from ReFL in Figure 28.
Appendix G Effect of Pretraining Dataset
As discussed in the paper, we incorporate the pretraining denoising loss to stabilize the training and to prevent reward over-optimization. In practice, we observe that the model is more prone to reward-hacking (i.e. producing unnatural artifacts and decreased photo-realism) without the pretraining loss. We experiment with removing and show the comparison in Figure 29 .
Appendix H Full List of Occupations and Objects
We provide the full list of 100 occupations used to evaluate the skintone diversity of the generated samples in Table LABEL:table:100_occupation.
We also provide the full list of 532 common objects used to construct the training set for the compositionality experiments in Table LABEL:table:objects_532. During training, two objects were randomly sampled and combined using one of the five relationship terms: “and”, “next to”, “near”, “on side of”, and “beside”.