Aligning Text-to-Image Models using Human Feedback

Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Shixiang Shane Gu

Introduction

Deep generative models have recently shown remarkable success in generating high-quality images from text prompts (Ramesh et al. 2021; Ramesh et al. 2022; Saharia et al. 2022; Yu et al. 2022b; Rombach et al. 2022). This success has been driven in part by the scaling of deep generative models to large-scale datasets from the web such as LAION (Schuhmann et al. 2021; Schuhmann et al. 2022). However, major challenges remain in domains where large-scale text-to-image models fail to generate images that are well-aligned with text prompts (Feng et al. 2022; Liu et al. 2022a; Liu et al. 2022b). For instance, current text-to-image models often fail to produce reliable visual text (Liu et al. 2022b) and struggle with compositional image generation (Feng et al. 2022).

In language modeling, learning from human feedback has emerged as a powerful solution for aligning model behavior with human intent (Ziegler et al. 2019; Stiennon et al. 2020; Wu et al. 2021; Nakano et al. 2021; Ouyang et al. 2022; Bai et al. 2022a). Such methods first learn a reward function intended to reflect what humans care about in the task, using human feedback on model outputs. The language model is then optimized using the learned reward function by a reinforcement learning (RL) algorithm, such as proximal policy optimization (PPO; Schulman et al. 2017). This RL with human feedback (RLHF) framework has successfully aligned large-scale language models (e.g., GPT-3; Brown et al. 2020) with complex human quality assessments.

Motivated by the success of RLHF in language domains, we propose a fine-tuning method for aligning text-to-image models using human feedback. Our method consists of the following steps illustrated in Figure 1: (1) We first generate diverse images from a set of text prompts designed to test output alignment of a text-to-image model. Specifically, we examine prompts where pre-trained models are more prone to errors – generating objects with specific colors, counts, and backgrounds. We then collect binary human feedback assessing model outputs. (2) Using this human-labeled dataset, we train a reward function to predict human feedback given the image and text prompt. We propose an auxiliary task—identifying the original text prompt within a set of perturbed text prompts—to more effectively exploit human feedback for reward learning. This technique improves the generalization of reward function to unseen images and text prompts. (3) We update the text-to-image model via reward-weighted likelihood maximization to better align it with human feedback. Unlike the prior work (Stiennon et al. 2020; Ouyang et al. 2022) that uses RL for optimization, we update the model using semi-supervised learning to measure model-output quality w.r.t. the learned reward function.

We fine-tune the stable diffusion model (Rombach et al. 2022) using 27K image-text pairs with human feedback. Our fine-tuned model shows improvement in generating objects with specified colors, counts, and backgrounds. Moreover, it improves compositional generation (i.e., can better generate unseen objects We use the term “unseen objects” w.r.t. our human dataset, i.e., they are not included in the human dataset but pre-training dataset would contain them. given unseen combinations of color, count, and background prompts). We also observe that the learned reward function is better aligned with human assessments of alignment than CLIP score (Radford et al. 2021) on tested text prompts. We analyze several design choices, such as using an auxiliary loss for reward learning and the effect of using “diverse” datasets for fine-tuning.

We can summarize our main contributions as follows:

We propose a simple yet efficient fine-tuning method for aligning a text-to-image model using human feedback.

We show that fine-tuning with human feedback significantly improves the image-text alignment of a text-to-image model. On human evaluation, our model achieves up to 47% improvement in image-text alignment at the expense of mildly degraded image fidelity.

We show that the learned reward function predicts human assessments of the quality more accurately than the CLIP score (Radford et al. 2021). In addition, we show that rejection sampling based on our learned reward function can also significantly improve the image-text alignment.

Naive fine-tuning with human feedback can significantly reduce the image fidelity, despite better alignment. We find that careful investigations on several design choices are important in balancing alignment-fidelity tradeoffs.

Even though our results do not address all the failure modes of the existing text-to-image models, we hope that this work highlights the potential of learning from human feedback for aligning these models.

Related Work

Various deep generative models, such as variational auto-encoders (Kingma & Welling 2013), generative adversarial networks (Goodfellow et al. 2020), auto-regressive models (Van Den Oord et al. 2016), and diffusion models (Sohl-Dickstein et al. 2015; Ho et al. 2020) have been proposed for image distributions. Combined with the large-scale language encoders (Radford et al. 2021; Raffel et al. 2020), these models have shown impressive results in text-to-image generation (Ramesh et al. 2022; Saharia et al. 2022; Yu et al. 2022b; Rombach et al. 2022).

However, text-to-image models frequently struggle to generate images that are well-aligned with text prompts (Feng et al. 2022; Liu et al. 2022a; Liu et al. 2022b). Liu et al. 2022b show that current models fail to produce reliable visual text and often perform poorly w.r.t. compositional generation (Feng et al. 2022; Liu et al. 2022a). Several techniques, such as character-aware text encoders (Xue et al. 2022; Liu et al. 2022b) and structured representations of language inputs (Feng et al. 2022) have been investigated to address these issues. We study learning from human feedback, which aligns text-to-image models directly using human feedback on model outputs.

Fine-tuning with few images (Ruiz et al. 2022; Kumari et al. 2022; Gal et al. 2022) for personalization of text-to-image diffusion models is also related with our work. DreamBooth (Ruiz et al. 2022) showed that text-to-image models can generate diverse images in a personalized way by fine-tuning with few images, and Kumari et al. 2022 proposed a more memory and computationally efficient method. In this work, we demonstrate that it is possible to fine-tune text-to-image models using simple binary (good/bad) human feedback.

Learning with human feedback.

Human feedback has been used to improve various AI systems, from translation (Bahdanau et al. 2016; Kreutzer et al. 2018), to web question-answering (Nakano et al. 2021), to story generation (Zhou & Xu 2020), to training RL agents without the need for hand-designed rewards (MacGlashan et al. 2017; Christiano et al. 2017; Warnell et al. 2018; Ibarz et al. 2018; Lee et al. 2021, inter alia), and to more truthful and harmless instruction following and dialogue (Ouyang et al. 2022; Bai et al. 2022a; Ziegler et al. 2019; Wu et al. 2021; Stiennon et al. 2020; Liu et al. 2023; Scheurer et al. 2022; Bai et al. 2022b, inter alia). In relation to prior works that focus on improving language models and game agents with human feedback, our work explores using human feedback to align multi-modal text-to-image models with human preference. Many prior works on learning with human feedback consist of learning a reward function and maximizing reward weighted likelihood (often dubbed as supervised fine-tuning) (Ouyang et al. 2022; Ziegler et al. 2019; Stiennon et al. 2020, see e.g.) Inspired by their successes, we propose a fine-tuning method with human feedback for improving text-to-image models.

Evaluating image-text alignment.

To measure the image-text alignment, various evaluation protocols have been proposed (Madhyastha et al. 2019; Hessel et al. 2021; Saharia et al. 2022; Yu et al. 2022b). Most prior works (Ramesh et al. 2022; Saharia et al. 2022; Yu et al. 2022b) use the alignment score of image and text embeddings determined by pre-trained multi-modal models, such as CLIP (Radford et al. 2021) and CoCa (Yu et al. 2022a). However, since scores from pre-trained models are not calibrated with human intent, human evaluation has also been introduced (Saharia et al. 2022; Yu et al. 2022b). In this work, we train a reward function that is better aligned with human evaluations by exploiting pre-trained representations and (a small amount) of human feedback data.

Main Method

To improve the alignment of generated images with their text prompts, we fine-tune a pre-trained text-to-image model (Ramesh et al. 2022; Saharia et al. 2022; Rombach et al. 2022) by repeating the following steps shown in Figure 1. We first generate a set of diverse images from a collection of text prompts designed to test various capabilities of the text-to-image model. Human raters provide binary feedback on these images (Section 3.1). Next we train a reward model to predict human feedback given a text prompt and an image as inputs (Section 3.2). Finally, we fine-tune the text-to-image model using reward-weighted log likelihood to improve text-image alignment (Section 3.3).

To test specific capabilities of a given text-to-image model, we consider three categories of text prompts that generate objects with a specified count, color, or background. For simplicity, we consider a limited class of text categories in this work, deferring the study of broader and more complex categories to future work. For each category, we generate prompts by combining a word or phrase from that category with some object; e.g., combining green (or in a city) with dog. We also consider combinations of the three categories (e.g., two green dogs in a city). From each prompt, we generate up to 60 images using a pre-trained text-to-image model—in this work, we use Stable Diffusion v1.5 (Rombach et al. 2022).

Human feedback.

We collect simple binary feedback from multiple human labelers on the image-text dataset. Labelers are presented with three images generated from the same prompt and are asked to assess whether each image is well-aligned with the prompt (‘‘good’’) or not (‘‘bad’’). Labelers are instructed to skip a query if it is hard to answer. Skipped queries are not used in training. We use binary feedback given the simplicity of our prompts—the evaluation criterion are fairly clear. More informative human feedback, such as ranking (Stiennon et al. 2020; Ouyang et al. 2022), should prove useful when more complex or subjective text prompts are used (e.g., artistic or open-ended generation).

2 Reward Learning

To measure image-text alignment, we learn a reward function rϕ(x,z)r_{\phi}(\mathbf{x},\mathbf{z}) (parameterized by ϕ\phi) that maps the CLIP embeddings To improve generalization ability, we use CLIP embeddings pre-trained on various image-text samples. (Radford et al. 2021) of an image x\mathbf{x} and a text prompt z\mathbf{z} to a scalar value. It is trained to predict human feedback y∈{0,1}y\in\{0,1\} (1 = good, 0 = bad).

Formally, given the human feedback dataset Dhuman={(x,z,y)}\mathcal{D}^{\tt human}=\{\left(\mathbf{x},\mathbf{z},y\right)\}, the reward function rϕr_{\phi} is trained by minimizing the mean-squared-error (MSE):

Data augmentation can significantly improve the data-efficiency and performance of learning (Krizhevsky et al. 2017; Cubuk et al. 2019). To effectively exploit the feedback dataset, we design a simple data augmentation scheme and auxiliary loss for reward learning. For each image-text pair that has been labeled good, we generate N−1N-1 text prompts with different semantics than the original text prompt. For example, we might generate {Blue dog , …\ldots , Green dog} given the original prompt Red dog. We use a rule-based strategy to generate different text prompts (see Appendix E for more details). This process generates a dataset Dtxt={(x,{zj}j=1N,i′)}\mathcal{D}^{\tt txt}=\{(\mathbf{x},\{\mathbf{z}_{j}\}_{j=1}^{N},i^{\prime})\} with NN text prompts {zj}j=1N\{\mathbf{z}_{j}\}_{j=1}^{N}, including the original, for each image x\mathbf{x}, and the index i′i^{\prime} of the original prompt.

We use the augmented prompts in an auxiliary task, namely, classifying the original prompt for reward learning. Our prompt classifier uses the reward function rϕr_{\phi} as follows:

where T>0T>0 is the temperature. Our auxiliary loss is

where LCE\mathcal{L}^{\tt CE} is the standard cross-entropy loss. This encourages rϕr_{\phi} to produce low values for prompts with different semantics than the original. Our experiments show this auxiliary loss improves the generalization to unseen images and text prompts. Finally, we define the combined loss as

where λ\lambda is the penalty parameter.

The pseudo code of reward learning is in Appendix E.

3 Updating the Text-to-Image Model

We use our learned rϕr_{\phi} to update the text-to-image model pp with parameters θ\theta by minimizing the loss

where Dmodel\mathcal{D}^{\tt model} is the model-generated dataset (i.e., images generated by the text-to-image model on the tested text prompts), Dpre\mathcal{D}^{\tt pre} is the pre-training dataset, and β\beta is a penalty parameter. The first term in (2) minimizes the reward-weighted negative log-likelihood (NLL) on Dmodel\mathcal{D}^{\tt model}. To increase diversity, we collect an unlabeled dataset Dunlabel\mathcal{D}^{\tt unlabel} by generating more images from the text-to-image model, and use both the human-labeled dataset Dhuman\mathcal{D}^{\tt human} and the unlabeled dataset Dunlabel\mathcal{D}^{\tt unlabel} for training, i.e., Dmodel=Dhuman∪Dunlabel\mathcal{D}^{\tt model}=\mathcal{D}^{\tt human}\cup\mathcal{D}^{\tt unlabel}. By evaluating the quality of the outputs using a reward function aligned with the text prompts, this term improves the image-text alignment of the model.

Typically, the diversity of the model-generated dataset is limited, which can result in overfitting. To mitigate this, similar to Ouyang et al. 2022, we also minimize the pre-training loss, the second term in (2). This reduces NLL on the pre-training dataset Dpre\mathcal{D}^{\tt pre}. In our experiments, we observed regularization in the loss function L(θ)\mathcal{L}(\theta) in (2) enables the model to generate more natural images.

Different objective functions and algorithms (e.g., PPO; Schulman et al. 2017) could be considered for updating the text-to-image model similar to RLHF fine-tuning (Ouyang et al. 2022). We believe RLHF fine-tuning may lead to better models because it uses online sample generation during updates and KL-regularization over the prior model. However, RL usually requires extensive hyperparameter tuning and engineering, thus, we defer the extension to RLHF fine-tuning to future work.

Experiments

We describe a set of experiments designed to test the efficacy of our fine-tuning approach with human feedback.

For our baseline generative model, we use stable diffusion v1.5 (Rombach et al. 2022), which has been pre-trained on large image-text datasets (Schuhmann et al. 2021; Schuhmann et al. 2022). Our fine-tuning method can be used readily with other text-to-image models, such as Imagen (Saharia et al. 2022), Parti (Yu et al. 2022b) and Dalle-2 (Ramesh et al. 2022). For fine-tuning, we freeze the CLIP language encoder (Radford et al. 2021) and fine-tune only the diffusion module. For the reward model, we use ViT-L/14 CLIP model (Radford et al. 2021) to extract image and text embeddings and train a MLP using these embeddings as input. More experimental details (e.g., model architectures and the final hyperparameters) are reported in Appendix D.

Datasets.

From a set of 2700 English prompts (see Table 1 for examples), we generate 27K images using the stable diffusion model (see Appendix B for further details). Table 2 shows the feedback distribution provided by multiple human labelers, which has been class-balanced. We note that the stable diffusion model struggles to generate the number of objects specified by the prompt, but reliably generates specified colors and backgrounds.

We use 23K samples for training, with the remaining samples used for validation. We also use 16K unlabeled samples for the reward-weighted loss and a 625K subset https://huggingface.co/datasets/ChristophSchuhmann/improved_aesthetics_6.5plus of LAION-5B (Schuhmann et al. 2022) filtered by an aesthetic score predictor https://github.com/christophschuhmann/improved-aesthetic-predictor for the pre-training loss.

2 Text-Image Alignment Results

We measure human ratings of image alignment with 120 text prompts (60 seen text prompts and 60 unseen Here, “unseen” text prompts consist of “unseen” objects, which are not in our human dataset. text prompts), testing the ability of the models to render different colors, number of objects, and backgrounds (see Appendix B for the full set of prompts). Given two (anonymized) sets of images, one from our fine-tuned model and one from the stable diffusion model, we ask human raters to assess which is better w.r.t. image-text alignment and fidelity (i.e., image quality). We ask raters to declare a tie if they have similar quality. Each query is evaluated by 9 independent human raters. We show the percentage of queries based on the number of positive votes.

As shown in Figure 4, our method significantly improves image-text alignment against the original model. Specifically, 50% of samples from our model receive at least two-thirds vote (7 or more positive votes) for image-text alignment. However, fine-tuning somewhat degrades image fidelity (15% compared to 10%). We expect that this is because (i) we asked the labelers to provide feedback mainly on alignment, (ii) the diversity of our human data is limited, and (iii) we used a small subset of pre-training dataset for fine-tuning. Similar issue, which is akin to the alignment tax, has been observed in language domains (Askell et al. 2021; Ouyang et al. 2022). This issue can presumably be mitigated with larger rater and pre-training datasets.

Qualitative comparison.

Figure 2 shows image samples from the original model and our fine-tuned counterpart (see Appendix A for more image examples). While the original often generates images with missing details (e.g., color, background or count) (Figure 2(a)), our model generates objects that adhere to the prompt-specified colors, counts and backgrounds. Of special note, our model generates high-quality images on unseen text prompts that specify unseen objects (Figure 2(b)). Our model also generates reasonable images given unseen text categories, such as artistic generation (Figure 2(c)).

However, we also observe several issues of our fine-tuned models. First, for some specific text prompts, our fine-tuned model generates oversaturated and non-photorealistic images. Our model occasionally duplicates entities within the generated images or produces lower-diversity images for the same prompt. We expect that it would be possible to address these issues with larger (and diverse) human datasets and better optimization (e.g., RL).

3 Results on Reward Learning

We investigate the quality of our learned reward function by evaluating its prediction of with human ratings. Given two images from the same text prompt (x1,x2,z)(\mathbf{x}_{1},\mathbf{x}_{2},\mathbf{z}), we check whether our reward rϕr_{\phi} generates a higher score for the human-preferred image, i.e., rϕ(x1,z)>rϕ(x2,z)r_{\phi}(\mathbf{x}_{1},\mathbf{z})>r_{\phi}(\mathbf{x}_{2},\mathbf{z}) when rater prefers x1\mathbf{x}_{1}. As a baseline, we compare it with the CLIP score (Hessel et al. 2021), which measures image-text similarity in the CLIP embedding space (Radford et al. 2021).

Figure 3(a) compares the accuracy of rϕr_{\phi} and the CLIP score on unseen images from both seen and unseen text prompts. Our reward (green) more accurately predicts human evaluation than the CLIP score (red), hence is better aligned with typical human intent. To show the benefit of our auxiliary loss (prompt classification) in (1), we also assess a variant of our reward function which ignores the auxiliary loss (blue). The auxiliary classification task improves reward performance on both seen and unseen text prompts. The gain from the auxiliary loss clearly shows the importance of text diversity and our auxiliary loss in improving data efficiency. Although our reward function is more accurate than the CLIP score, its performance on unseen text prompts (∼80%\sim 80\%) suggests that it may be necessary to use more diverse and large human datasets.

Rejection sampling.

Similar to Parti (Yu et al. 2022b) and DALL-E (Ramesh et al. 2021), we evaluate a rejection sampling technique, which selects the best output w.r.t. the learned reward function. Parti and DALL-E use similarity scores of image and text embeddings from CoCa (Yu et al. 2022a) and CLIP (Radford et al. 2021), respectively. Specifically, we generate 16 images per text prompt from the original stable diffusion model and select the 4 with the highest reward scores. We compare these to 4 randomly sampled images in Figure 6(a). Rejection sampling significantly improves image-text alignment (46% with two-thirds preference vote by raters) without sacrificing image fidelity. This result illustrates the significant value of the reward function in improving text-to-image models without any fine-tuning.

We also compare our fine-tuned model to the original with rejection sampling in Figure 6(b). We remark that rejection sampling can also be applied on top of our fine-tuned model. Our fine-tuned model achieves a 10% gain in image-text alignment (20%-10% two-thirds vote) but sacrifices 17% in image fidelity (3%-20% two-thirds vote). However, as discussed in Section 4.2, we expect degradation in fidelity to be mitigated with larger human datasets and better hyper-parameters. Note also that rejection sampling has several drawbacks, including increased inference-time computation and the inability to improve the model (since it is output post-processing).

4 Ablation Studies

To investigate how human data quality affects reward learning, we conduct an ablation study, reducing the number of images per text prompt by half before training the reward function. Figure 3(b) shows that model accuracy decreases on both seen and unseen prompts as data size decreases, clearly demonstrating the importance of diversity and the amount of rater data.

Effects of using diverse datasets.

To verify the importance of data diversity, we incrementally include unlabeled and pre-training datasets during fine-tuning. We measure the reward score (image-text alignment) on 120 tested text prompts and FID score (Heusel et al. 2017)—the similarity between generated images and real images—on MS-CoCo validation data (Lin et al. 2014). Table 3 shows that FID score is significantly reduced when the model is fine-tuned using only human data, despite better image-text alignment. However, by adding the unlabeled and pre-training datasets, FID score is improved without impacting image-text alignment. We provide image samples from unseen text prompts in Figure 5. We see that fine-tuned models indeed generate more natural images when exploiting more diverse datasets.

Discussion

In this work, we have demonstrated that fine-tuning with human feedback can effectively improve the image-text alignment in three domains: generating objects with a specified count, color, or backgrounds. We analyze several design choices (such as using an auxiliary loss and collecting diverse training data) and find that it is challenging to balance the alignment-fidelity tradeoffs without careful investigations on such design choices. Even though our results do not address all the failure modes of the existing text-to-image models, we hope that our method can serve as a starting point to study learning from human feedback for improving text-to-image models.

Limitations and future directions. There are several limitations and interesting future directions in our work:

More nuanced human feedback. Some of the poor generations we observed, such as highly saturated image colors, are likely due to similar images being highly ranked in our training set. We believe that instructing raters to look for a more diverse set of failure modes (oversaturated colors, unrealistic animal anatomy, physics violations, etc.) will improve performance along these axes.

Diverse and large human dataset. For simplicity, we consider a limited class of text categories (count, color, background) and thus consider a simple form of human feedback (good or bad). Due to this, the diversity of our human data is bit limited. Extension to more subjective text categories (like artistic generation) and informative human feedback such as ranking would be an important direction for future research.

Different objectives and algorithms. For updating the text-to-image model, we use a reward-weighted likelihood maximization. However, similar to prior work in language domains (Ouyang et al. 2022), it would be an interesting direction to use RL algorithms (Schulman et al. 2017). We believe RLHF fine-tuning may lead to better models because (a) it uses online sample generation during updates and (b) KL-regularization over the prior model can mitigate overfitting to the reward function.

Acknowledgements

We thank Peter Anderson, Douglas Eck, Junsu Kim, Changyeon Kim, Jongjin Park, Sjoerd van Steenkiste, Younggyo Seo, and Guy Tennenholtz for providing helpful comments and suggestions. Finally, we would like to thank Sehee Yang for providing valuable feedback on user interface and constructing the initial version of human data, without which this project would not have been possible.

References

Appendix A Qualitative Comparison

Appendix B Image-text Dataset

In this section, we describe our image-text dataset. We generate 2774 text prompts by combining a word or phrase from that category with some object. Specifically, we consider 9 colors (red, yellow, green, blue, black, pink, purple, white, brown), 6 numbers (1-6), 8 backgrounds (forest, city, moon, field, sea, table, desert, San Franciso) and 25 objects (dog, cat, lion, orange, vase, cup, apple, chair, bird, cake, bicycle, tree, donut, box, plate, clock, backpack, car, airplane, bear, horse, tiger, rabbit, rose, wolf). We use the following 5 objects only for evaluation: bear, tiger, rabbit, rose, wolf. For each text prompt, we generate 60 or 6 images according to the text category. In total, our image-text dataset consists of 27528 image-text pairs. Labeling for training is done by two human labelers.

For evaluation, we use 120 text prompts listed in Table 4. Given two (anonymized) sets of 4 images, we ask human raters to assess which is better w.r.t. image-text alignment and fidelity (i.e., image quality). Each query is rated by 9 independent human raters in Figure 4 and Figure 6.

Appendix C Additional Results

Appendix D Experimental Details

Model architecture. For our baseline generative model, we use stable diffusion v1.5 (Rombach et al. 2022), which has been pre-trained on large image-text datasets (Schuhmann et al. 2021; Schuhmann et al. 2022). For the reward model, we use ViT-L/14 CLIP model (Radford et al. 2021) to extract image and text embeddings and train a MLP using these embeddings as input. Specifically, we use two-layer MLPs with 1024 hidden dimensions each. We use ReLUs for the activation function between layers, and we use the Sigmoid activation function for the output. For auxiliary task, we use temperature T=2T=2 and penalty parameter λ=0.5\lambda=0.5.

Training. Our fine-tuning pipeline is based on publicly released repository (https://github.com/huggingface/diffusers/tree/main/examples/text_to_image). We update the model using AdamW (Loshchilov & Hutter 2017) with β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, ϵ=1e−8\epsilon=1e-8 and weight decay 1e−21e-2. The model is trained in half-precision on 4 40GB NVIDIA A100 GPUs, with a per-GPU batch size of 8, resulting in a toal batch size of 512 (256 for pre-training data and 256 for model-generated data). In prior work (Ouyang et al. 2022), model is optimized with bigger batch for pre-training data. However, in our experiments, using a bigger batch does not make a big difference. We expect this is because small pre-training dataset is used in our work. It is trained for a total of 10,000 updates.

FID measurement using MS-CoCo dataset. We measure FID scores to evaluate the fidelity of different models using MS-CoCo validation dataset (i.e., val2014). There are a few caption annotations for each MS-CoCo image. We randomly choose one caption for each image, which results in 40,504 caption and image pairs. MS-CoCo images have different resolutions and they are resized to 256×256256\times 256 before computing FID scores. We use pytorch-fid Python implementation for the FID measurement (https://github.com/mseitzer/pytorch-fid).

Appendix E Pseudocode