Human Preference Score: Better Aligning Text-to-Image Models with Human Preference
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, Hongsheng Li
Introduction
The recent progress in diffusion models has enabled impressive advancements in text-to-image generation, with many models now being deployed in real-world applications such as DALL·E and Stable Diffusion . However, public attention has also highlighted new issues, such as the awkward combinations of limbs and facial expressions of generated persons as shown in Fig. 2. The users usually need to cherry-pick results to avoid these artifacts. In other words, the generated images are misaligned with human preferences.
To further improve the quality of generated images, it is essential to track the ability of a model to generate human preferable images. However, it is uncertain whether the existing evaluation metrics, such as Inception Score (IS) and Fréchet inception distance (FID) , are correlated with human choices. These metrics perceive an image through a classification-based CNN trained on ImageNet , which has been shown to be biased towards image texture rather than general image contents , and thus may not align well with human perception. Also, both IS and FID are single-modal evaluation metrics, which do not take user intention into account. Some recent studies use the CLIP model as a proxy for human judgment to evaluate the alignment between generated images and text prompts. The CLIP model is trained on a rich dataset and is believed to capture subtle aspects of human intention better. However, it is uncertain whether CLIP can measure the quality of generated synthetic images, which may not adhere to the same constraints as real images, such as the example shown in Fig. 2.
In this study, we investigate the problem of human preference using a novel, large-scale dataset of human choices on images generated by Stable Diffusion using the same prompt. The dataset comprises 98,807 diverse images generated from user-provided prompts, along with 25,205 human choices. By evaluating on this dataset, we find that the Inception Score (IS) , the Fréchet Inception distance (FID) and the CLIP score does not fully match the human choice, which means that the human preference is a missing dimension of image quality that is not well tracked by existing mainstream metrics.
We further train a human preference classifier on this dataset by fine-tuning the CLIP model and define human preference score (HPS) based on it. We validate HPS’s alignment with human choices and its generalization capability towards other generative models through user studies. HPS can be utilized to guide generative models toward producing human-preferred images. To this end, we devise a simple yet effective method to adapt Stable Diffusion by LoRA with awareness of human preference. We conduct user studies to validate the effectiveness of our approach. The results show that the adapted model can better capture human intentions, and generate more preferable images, which significantly mitigates the kind of artifact shown in Fig. 2.
Our contributions are as follows: (1) We create a large-scale dataset for studying human preferences. To our best knowledge, this dataset is the first of its kind that contains massive human choices on images generated with the same prompt. (2) We find that human choices cannot be accurately predicted by the existing mainstream evaluation metrics, while it can be better predicted via fine-tuning CLIP on the proposed dataset. (3) We propose a simple yet effective method to guide the Stable Diffusion model toward generating images with better aesthetic quality and better alignment with human intention.
Related Works
Text-to-image generative models. Text-to-image generative models have long been an active research area. Mansimov et al. show that Deep Recurrent Attention Writer (DRAW) can be conditioned on captions to generate novel scene compositions. Generative Adversarial Networks (GANs) improve image fidelity by training a discriminator to provide supervision for the generative model. DALL·E firstly achieves open-domain text-to-image synthesis with the help of massive image-text pairs.
Diffusion models formulate the generative process as the inverse of the diffusion process , which was improved by Song and Ermon and Ho et al. . Dhariwal et al. firstly show the superiority of diffusion models over GANs on image generation. Several following works, including DALL·E 2 , GLIDE , Imagen , ERNIE-ViLG , Stable Diffusion , bring the magic of text-to-image generation to the public attention. Among these models, Stable Diffusion is an open-source model with an active user community.
Several recent works improve Stable Diffusion on different aspects. DreamBooth and ELITE explore customizing Stable Diffusion to a certain object. Feng et al. propose a training-free method to guide diffusion models for better compositional capabilities. It has been discovered that prompt engineering plays an important role in generating high-quality images. Hao et al. devise an automatic prompt engineering scheme via reinforcement learning. Our method focuses on the misalignment between the generated image and human preference, which is orthogonal to the above-mentioned topics.
Datasets of generated images. Datasets of generated images play a vital role in computer vision tasks that has difficulty in ground-truth acquisition, such as optical flow estimation . Thanks to active user communities of text-to-image models, several databases of images generated by diffusion models have been introduced. Lexica (lexica.art) is a large database of images generated by Stable Diffusion and Lexica Aperture. It also provides related information about the image, such as the prompt and guidance scale. However, the database is closed-source and only allows online browsing. DiffusionDB is a large-scale open-source database collected from the Stable Foundation Discord channel, containing the text prompt and parameters for each image. SAC is a dataset of images generated from Stable Diffusion and GLIDE , along with user ratings from an aesthetic survey. However, SAC only contains limited user choices compared to our dataset.
Learning from human feedback. Human feedback has long been used in a wide range of deep learning tasks. Christiano et al. and Arakawa et al. incorporate human feedback into RL training, which is proven to accelerate the model convergence. Krishna et al. propose “socially situated AI”, which significantly improves image recognition performance via interacting with human users on Instagram. InstructGPT fine-tunes GPT via a reward function trained on human feedback, establishing the foundation for the success of ChatGPT. and use similar methodology to improve text-to-image models, which are highly related to our work. In , this is achieved by augmenting the text prompt. is a concurrent work that focuses more on the exact alignment between text and image, while our work shows that the potential of human feedback is far beyond the exact alignment when the feedback takes into account the aesthetic preference of humans.
Human Preference Dataset
In order to get a better understanding of the human preferences on the images generated from prompts, and to improve text-to-image generation quality, we start by collecting a dataset of human choices.
Data collection. We utilize the “dreambot” channel on the Stable Foundation Discord server to gather human choice data. The chat history of these channels is obtained using the DiscordChatExporter tool, which downloads the full chat history of a Discord channel and stores it in JSON format. Among the chat messages, a discernible pattern of interaction is observed, as depicted in Fig. 4, which reveals human preferences. In this pattern, a user initiates a session by sending a text prompt to the bot, which generates several images in response. Then, the user selects a preferred image and sends it back to the bot, along with the original text prompt. The bot will return several refined images. This interaction follows a pre-defined grammar, which allows us to extract human choice and related images using simple pattern-matching techniques.
Data format and statistics. Finally, we obtain a total of 98,807 images generated from 25,205 prompts. Each prompt corresponds to several images, among which one image is chosen by the user as the preferred one, while others are non-preferred negatives. Each prompt corresponds with a varying number of images. 23,722 prompts have four images, 953 prompts have three images and 530 prompts have two images. The number of images for each prompt depends on the user’s specifications in the generation request. Notably, the dataset exhibits a high level of diversity, with images generated across a broad range of themes. The dataset consists of choices made by 2,659 different users, and each user contributes at most 267 choices. Examples of the collected dataset can be found in Fig. 3. For further details on the dataset, we refer the readers to Fig. 9 in the appendix.
Privacy and NSFW contents. We observe that a small portion of images is generated with image condition (the condition image may be either generated or uploaded by the user). Since user-uploaded images may contain sensitive information or privacy, we do not include them in our dataset. For the images with potential NSFW content, we use the channel bot’s NSFW detector to filter them out.
In this work, we utilize this dataset to study the existing metrics’ correlation with human preferences, which will be introduced in Sec. 4. The dataset also serves as the training data for our human preference classifier, which is to be introduced in Sec. 5.
Existing Metrics
In this section, we show that the current mainstream evaluation metrics are not well correlated with human preferences on our dataset.
Inception Score (IS) and Fréchet inception distance (FID) are two popular metrics used to evaluate the quality of generated images. Both of them perceive an image through an Inception Net trained on ImageNet . In this section, we investigate their correlation with human choices.
Inception Score (IS) measures the quality of generated images by computing the expected KL-divergence between the marginal class distribution over all generated images and the conditional distribution for a particular generated image, using the class probability predicted by the Inception Net. This metric is expected to capture both the fidelity and diversity of generated images. To determine the correlation between IS and human preferences, we compute IS for both the set of preferred and non-preferred images in our dataset. For each setting, we divide 20,000 images into 10 splits and reported the mean and standard deviation of IS computed on them. Our results, as shown in Tab. 1, indicate no significant difference between the preferred and non-preferred images.
Fréchet Inception Distance (FID) measures the similarity between the embedding feature of generated and real images. This is achieved by fitting the embedding features into a multivariate Gaussian distribution and computing their Fréchet distance. To define the target distribution, FID requires a set of real images. However, in the case of images generated from user-provided prompts, such as in our dataset, the target distribution is defined by users’ intention, which can only be inferred from text prompts. To address this, we randomly sample 10,000 text prompts from our dataset, and for each prompt, we query the LAION dataset via the official api to find the closest image, which is taken as a “pseudo ground truth” for that prompt. This provides a set of real images aligned with the users’ intentions. We randomly sample 10,000 images from both the preferred and non-preferred split of the collected dataset to compute FID with the real images. Our results, as shown in Tab. 1, reveal no significant difference between the preferred and non-preferred images in terms of FID. This suggests that FID may not be a reliable metric for evaluating human preference.
Discussion. IS and FID may suffer from the following three issues when evaluating human preference. Firstly, generated images often contain shape artifacts, as shown in Fig. 2. However, classification-based CNNs tend to be biased towards image texture rather than shape , making them be likely to ignore shape artifacts in generated images. Additionally, the domain gap can pose a problem. While the evaluation model is trained on real images from ImageNet , the generated images in our dataset exhibit a wide range of styles and themes, from oil painting portraits to digital art of cyborgs. As a result, the ImageNet-trained model may not have meaningful representations for these diverse images . Furthermore, these metrics are limited by their single-modal nature, which means that they cannot infer user intentions by accessing prompts, unless the target images are known or provided as we do.
2 Metrics by CLIP
Thanks to the large and diverse set of training data, CLIP is better at encoding images from various domains compared to ImageNet-trained models. Moreover, it can capture users’ intentions by encoding text prompts, making it a plausible choice for evaluating the alignment between a prompt and a generated image . Aesthetic Score Predictor is another CLIP-based tool for image quality evaluation, which has been utilized to filter the training data for Stable Diffusion . In this section, we evaluate the capability of these tools in predicting human choices, which is done by counting the accuracy of the human choice prediction task conducted on a split of 5,000 samples from our dataset.
CLIP score is derived as the cosine similarity between the prompt embedding and the image embedding computed by CLIP. We evaluated the performance of ViT-L/14 and RN50x64 models, which are the largest open-source CLIP models for transformer and CNN architecture. Our results, presented in Tab. 2, demonstrate that both CLIP models exhibit superior performance over random guessing. However, we will show in Sec. 7.1 that the CLIP score does not correlate well with human choices. Nevertheless, we will also show that it can be further fine-tuned on our dataset to better align with human preferences.
Aesthetic score is based on a pre-trained ViT-L/14 CLIP image encoder, which is adapted to the task of aesthetic score prediction by adding a MLP layer on top of the CLIP image encoder. The MLP is trained on several aesthetic datasets, including both real images and generated images (e.g., AVA , SAC ) to predict aesthetic scores ranging from 1 to 10. Unlike CLIP, the aesthetic classifier does not condition on the prompt, so the image with the highest predicted score is taken as the model choice. As shown in Tab. 2, the aesthetic classifier also exhibits better-than-chance accuracy in predicting user choice, indicating the importance of the aesthetic aspect of an image in human decision-making.
Human Preference Score
We first train a human preference classifier to predict the human choice based on the prompt, and then derive HPS based on the trained classifier.
Human preference classifier We fine-tune the ViT-L/14 version of CLIP on our dataset to better align with human preferences. Each sample in the training set contains one prompt along with images, among which only one image is preferred by the user. The model is trained to maximize the similarity between the embedding of the text prompt computed by the CLIP text encoder and the embedding of the preferred image computed by the CLIP visual encoder, while minimizing the similarity for non-preferred images. By fine-tuning on human choices of generated images, the model is encouraged to better align with human preferences.
Human preference score (HPS) is derived from the human preference classifier. We define HPS as:
where and are the visual encoder and the text encoder of the human preference classifier. We multiply the cosine similarity by a factor of 100 for better visualization.
Better Aligning Stable Diffusion with Human Preferences
HPS can be used to guide diffusion-based generative models to better align with human users. We argue that the misalignment between generated images and human preferences is a problem of missing “awareness” rather than model capacity. To address this issue, we propose to adapt the generative model by explicitly distinguishing preferred images from non-preferred ones. Our solution is straightforward and intuitive. We construct another dataset consisting of prompts and their newly generated images, which we categorize as either preferred or non-preferred using our previously trained human preference classifier. For the non-preferred images, we modify their corresponding prompts by prepending a special prefix. By adapting Stable Diffusion on this dataset via LoRA , we enhance the model’s ability to learn the concept of non-preferred images, which can subsequently be avoided during inference.
Constructing training data. We construct the training data from the “large_first_1m” split of DiffusionDB , and a subset of the pre-train dataset of Stable Diffusion (LAION-5B) for regularization. DiffusionDB is a large-scale dataset of generated images along with their text prompts. For images from DiffusionDB, we first compute HPS for each image-prompt pair. After that, we group the images by their prompts, and for each prompt , we add the image with the highest HPS into the training data if it passes the following criteria:
where is the number of images with the same prompt, and is a hyper-parameter that controls the selectivity. is given by:
where is the set of images with the same prompt. Similarly, we construct the non-preferred subset by the same criteria, but using negative HPS. Finally, we get a mixed dataset of generated images and real images, where the non-preferred generated images are identified by their prompt prefix.
Adapting Stable Diffusion. We adopt LoRA to adapt Stable Diffusion to the training data, in which the parameters of the original model are kept frozen, and the {key, query, value, out} projection matrices are augmented with a low-rank residual. LoRA does not add new parameters to the model, since the learned projection matrices can be merged into the base model once trained. During training, we use the prompt as the caption for generated images. For non-preferred images, we prepend a special identifier before each of their captions (we choose “Weird image.” as the special identifier in our case). During inference, the special identifier is used as the negative prompt for classifier-free guidance to avoid generating non-preferred images.
Experiments
In this section, we firstly validate the reliability of HPS in Sec. 7.1, and then in Sec. 7.2, we introduce our experiments of adapting Stable Diffusion.
Implementation details of human preference classifier. We use 20,205 samples from our dataset during training, which contains 20,205 prompts and 79,167 images. We use the ViT-L/14 version of CLIP in our experiments. We fine-tune the last 10 layers of the CLIP image encoder and the last 6 layers of the text encoder. The model is trained by the AdamW optimizer with a learning rate of for 1 epoch. The batch size is 5. The learning rate decays with a cosine learning rate schedule. Weight decay is set as . Instead of using the original data augmentation of random resized crop, we directly resize the longest edge of the image to 224, and then pad zeros to make the shorter edge increase to 224. We empirically find that fixing the aspect ratio of the image is beneficial. The hyper-parameters are tuned via Bayesian optimization.
Alignment with human. As shown in Tab. 2, the trained model significantly outperforms CLIP in the human choice prediction task. Due to the strong diversity of human preferences, the accuracy is even higher than our human participants.
Generalization. We evaluate HPS’ generalization capability towards other generative models by user studies. In this experiment, we let the human preference classifier and several human participants evaluate 398 pairs of images. In each pair, the images are generated by DALL·E and Stable Diffusion with the same text prompt. The prompts are randomly sampled from DiffusionDB , which is a large database of images and prompts sourced from the Stable Foundation Discord channel. We filter out the NSFW prompts by the indicator provided in DiffusionDB .
In Tab. 3, we evaluate the agreement between the predictions from humans, CLIP, and HPS. The agreement is computed by averaging the similarity of the prediction of each participant. HPS is better aligned with human preference compared to CLIP score, and its agreement with humans is close to the agreement between humans. It shows that HPS can generalize toward images generated by other models. We refer the readers to the supplementary material for a full list of images and choices made in this user study.
Correlation with CLIP score. In Fig. 6, we visualize the correlation between HPS and CLIP score. The text prompts are randomly sampled from the COCO Captions dataset, and the images are generated by Stable Diffusion . We can see that HPS has a positive correlation with CLIP score, but emphasizes more on the aesthetic quality of an image. However, HPS put less importance on the direct matching between image contents and text prompts, which can be interpreted as a visual analogy of “alignment tax” introduced in .
2 Better Aligning Stable Diffusion with Human Preferences
Implementation details. We use the Stable Diffusion v1.4 for all our experiments. is set to 2.0 for both preferred images and non-preferred images when constructing the training set. The constructed training set contains 37,572 preferred generated images and 21,108 non-preferred generated images. The regularization images are from a 625k subset of LAION-5B filtered by the aesthetic score predictor with a threshold of 6.5. 200,231 regularization images participate in training. We only fine-tune the UNet of Stable Diffusion, while keeping the VAE and the text encoder frozen during training. The rank is set to 32 in LoRA . The LoRA weights are trained for 10k iterations with the AdamW optimizer with a learning rate of and a weight decay of , which is kept constant during training. We use a batch size of 40 in our experiments. For inference, we run the diffusion process by 50 steps for each image with PNDM noise scheduler. We use the default guidance scale of 7.5 for classifier-free guidance .
Human evaluation. We compare our trained model with the original Stable Diffusion by conducting user studies. In this study, we randomly sample 100 user-provided prompts from DiffusionDB . For each prompt, we generate an image from both models with the same random seed for fair comparison, resulting in 100 pairs of generated images for the user study. We ask 20 participants to read the prompt, and then choose between the image generated by our trained model and the original Stable Diffusion based on their preference. In Fig. 3, we visualize our result by showing the percentages of images with different numbers of positive votes. The adapted model significantly outperforms the original model. 74% of the images generated by the adapted model has more than 10 votes, while the number is 22% for the original model. A screenshot of the user-study interface is presented in Fig. 12 in the appendix.
Qualitative Evaluation. In Fig 7, we show some typical cases of improvement. We compare the original model, the regularization-only model, and the adapted model. The adapted model is trained with both real regularization images and generated images with HPS preference labels. The regularization-only model is a head-to-head comparison with the adapted model, which is trained by removing the generated images from the training set and is trained exclusively on regularization images for the same number of steps. The results show that the adapted model can better capture the user intention from the prompt, as shown in the first row. The last three rows show that training with generated images mitigates the problem of unnatural limbs. We refer the readers to Fig. 7 and Fig. 11 in the appendix for more examples.
Quantitative Evaluation. In Tab. 4, we compare the adapted model with the baseline on FID, Aesthetic Score, CLIP Score and HPS. The FID is computed on 10k images from the LAION dataset. CLIP Score and HPS are computed on prompts from DiffusionDB .
Limitations
There are several limitations about the dataset. The collected dataset contains generated prompts and images of public figures. We choose to mark them out instead of removing them to keep the diversity of the dataset. Despite the diversity of the dataset, we are also aware that it only represents the preference of a small portion of people in the world, and it may be biased towards a certain group of people that are active in the Stable Foundation Discord channel. Another potential bias about this dataset is that a large portion of text prompts are written by experienced Stable Diffusion users. These prompts are very likely to be tweaked to activate the potential of Stable Diffusion and deviate from normal language habits.
Conclusion
In this work, we study human preferences on a large-scale dataset of generated images. We find that the previous evaluation metrics for generative models are not well aligned with human preferences, but the CLIP model can be fine-tuned into a human preference classifier to better align with human choices. Then, we show a simple yet effective method to adapt the generative model to generate more preferable images with the guidance of human preference score. We hope our work can inspire the community to explore new possibilities of human-aligned AI research.
Acknowledgement
This project is funded in part by National Key R&D Program of China Project 2022ZD0161100, by the Centre for Perceptual and Interactive Intelligence (CPII) Ltd under the Innovation and Technology Commission (ITC)’s InnoHK, by General Research Fund of Hong Kong RGC Project 14204021. Hongsheng Li is a PI of CPII under the InnoHK. This project is also supported by SenseTime Collaborative Research Grant.
References
Appendix A Datasheet
Why was the dataset created?
The dataset was created to facilitate future academic Computer Vision research about human aesthetic preference.
The dataset was created by researchers at MMLab, The Chinese University of Hong Kong.
A.2 Composition
The instances are prompts and generated images, along with human preference choices among the images generated by the same prompt.
Are relationships between instances made explicit in the data (e.g. social network links, user/movie ratings, etc.)?
Yes, instances generated by the same user are identified by the same user id, which is anonymized for privacy.
How many instances are there? (of each type, if appropriate)?
There are 25,205 instances in the dataset.
Each instance consists of image, one prompt and one human choice.
Yes, we omit the specific parameters for generating the images, such as diffusion steps and guidance scale. They are omitted because we are more interested in the users’ preference about the generated images, rather than how they are created. Also, since the same batch of images (among which users make comparisons) are always generated with the same set of parameters except the random seed, they are irrelevant variables when studying human preferences.
Is everything included or does the data rely on external resources?
Are there recommended data splits and evaluation measures? (e.g. training, development, testing; accuracy or AUC)
In our experiments, we use a training set of 20,205 instances and validation set of 5,000 images, which will be made public. We recommend using accuracy (%) with one decimal place.
Are there any errors, sources of noise, or redundancies in the dataset?
Yes. The users are not prompted to selected images fitting their preference, so there should be noise in the collected data.
No, the dataset is collected from the Stable Foundation Discord server, which is publicly available for any user with an account.
We collect images and their prompts from the Stable Foundation discord server. Even though the discord server has rules against users sharing any NSFW (not suitable for work, such as sexual and violent content) and illegal images, our dataset still contains some NSFW images and prompts that were not removed by the server moderators.
Does the dataset relate to people?
Yes, the prompts are written by users and the choices are made by users.
Does the dataset identify any subpopulations (e.g. by age, gender)?
The dataset may contain sensitive data, because the prompts written by users may contain sensitive information, such as public figures and religious beliefs.
What experiments were initially run on this dataset? Have a summary of those results.
It has been used to validate the correlation between human preference and several popular image quality evaluation metrics, and serve as the training data for a human preference classifier. The results show that the tested metrics do not correlate well with human preference, and the correlation of the ViT-L/14 version of CLIP can be improved via fine-tuning on the dataset.
A.3 Data Collection Process
How was the data associated with each instance acquired?
The prompts, images and human choices are directly observable from the Stable Foundation Discord server.
Automatic scraping procedures were used to collect the data.
The dataset is not a sample of a larger set.
The authors of this paper were solely involved in the data collection process.
Over what time-frame was the data collected?
The dataset covers the chat history of dreambot channels between Dec. 2022 and Jan. 2023.
Were any ethical review processes conducted (e.g. by an institutional review board)?
No official processes were conducted, due to the public nature of the data on Discord channel.
Does the dataset relate to people?
The data was obtained from public messages in the Discord server.
A.4 Data Preprocessing
No preprocessing is done on the images and prompts.
A.5 Uses
Has the dataset been used for any tasks already? If so, please provide a description.
As described in the paper, this dataset has been used for analysis about several image quality evaluation metrics and training the proposed human preference classifier.
Is there a repository that links to any or all papers or systems that use the dataset?
What (other) tasks could the dataset be used for?
It can be used for tasks related to human preference on generated images.
Yes. As discussed in Sec. 8, the dataset is biased towards the preference of the certain group of people that are active in the Stable Foundation Discord server.
Are there tasks for which the dataset should not be used?
A.6 Data Distribution
Yes. Researchers at academic institutions will be able to request access to the dataset.
We will provide download links for researchers on a GitHub repository.
When will the dataset be distributed?
We will provide a terms of use agreement with the dataset. The dataset as a whole will be distributed under a non-commercial license.
A.7 Dataset Maintenance
Who is supporting/hosting/maintaining the dataset?
The authors of this paper are maintainers of this dataset.
How can the owner/curator/manager of the dataset be contacted (e.g. email address)?
Is there an erratum?
At this time, we are not aware of errors in our dataset. However, we will create an erratum as errors are identified.
The dataset will be updated by the authors on an at-will basis (but no more than once a month).
Will older versions of the dataset continue to be supported/hosted/maintained?
There will not be a mechanism to build on top of the dataset.
Appendix B More Dataset Examples
Appendix C More Visualization
See Fig. 11 and Fig. 11 for more visualizations. We show that the adapted model generates images with less artifacts and are better aware of users’ intentions.