DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback
Jiao Sun, Deqing Fu, Yushi Hu, Su Wang, Royi Rassin, Da-Cheng Juan, Dana Alon, Charles Herrmann, Sjoerd van Steenkiste, Ranjay Krishna, Cyrus Rashtchian
Introduction
Although we invite creative liberty when we commission art, we expect an artist to follow our instructions. Despite the advances in text-to-image (T2I) generation models , it remains challenging to obtain images that meticulously conform to users’ intentions . Current models often fail to compose multiple objects , bind attributes to the wrong objects , and struggle to generate visual text . In fact, the difficulty of finding effective textual prompts has led to a myriad of websites and forums dedicated to collecting and sharing useful prompts (e.g. PromptHero, Arthub.ai, Reddit/StableDiffusion). There are also online marketplaces for purchasing and selling useful such commands (e.g. PromptBase). The onus to generate aesthetic images that are faithful to a user’s desires should lie with the model and not with the user.
Today, there are efforts to address these challenges. For example, it is possible to manipulate attention maps based on linguistic structure to improve attribute-object binding ; or train reward models using human feedback to better align generations with user intent . Unfortunately, these methods either operate on a specific model architecture or require expensive labeled human data . Worse, most of these methods sacrifice aesthetic appeal when optimizing for faithfulness, which we confirm in our experiments.
We introduce DreamSync, a model-agnostic framework that improves T2I generation faithfulness while maintaining aesthetic appeal. Our approach extends work on fine-tuning T2I models for alignment, but does not require any human feedback. The key insight behind DreamSync is in leveraging the advances in vision-language models (VLMs), which can identify fine-grained discrepencies between the generated image and the user’s input text . Intuitively at a high level, our method can be thought of as a scalable version of reinforcement learning with human feedback (RLHF); just as LLaMA2 was iteratively refined using human feedback, DreamSync improves T2I models using feedback from VLMs, except without the need for reinforcement learning.
Given a set of textual prompts, T2I models first generates multiple candidate images per prompt. DreamSync automatically evaluates these generated images using two VLMs. The first one measures the generation’s faithfulness to the text , while the second one measures aesthetic quality . The best generations are collected and used to finetune the T2I model using parameter-efficient LoRA finetuning . With the new finetuned T2I model, we repeat the entire process for multiple iterations: generate images, curate a new finetuning set, and finetune again.
We conduct extensive experiments with latest benchmarks and human evaluation. We experiment DreamSync with two T2I models, SDXL and SD v1.4 . Results on both models show that DreamSync enhance the alignment of images to user inputs and retains their aesthetic quality. Specifically, quantitative results on TIFA and DSG demonstrate that DreamSync is more effective than all baseline alignment methods on SD v1.4, and can yield even bigger improvements on SDXL. Human evaluation on SDXL shows that DreamSync give consistent improvement on all categories of alignment in DSG. While our study primarily focuses on boosting faithfulness and aesthetic quality, DreamSync has broader applications: it can be used to improve other characteristics of an image as long as there is an underlying model that can measure that characteristic.
Related Work
Several prior works have proposed to use VQA models to evaluate text-to-image generation. The TIFA benchmark, which pioneered this approach for evaluation, consists of 4K prompts and 25K questions across 12 categories (e.g., object, count, material), enabling T2I model evaluation by using VQA models to answer questions about the generated images . TIFA prompts come from various resources, including DrawBench used in Imagen , PartiPrompt used in Parti , PaintSkill used in Dall-Eval, etc. DSG further improves TIFA’s realiability by examining their evaluation questions carefully. Another related benchmark is SeeTrue, which also uses VQA models to measure alignment . Before the VQA evaluation era, several other evaluation benchmarks were proposed focusing primarily on compositional text prompts for attribute binding (e.g., color, texture, shape) and object relationships (e.g., spatial). Examples include T2I-CompBench , C-Flowers , CC-500 and ABC-6K benchmarks . Aside from automated benchmarks, human evaluation for text-to-image generation is widely used in the community, although such annotations are notoriously costly to collect. In response, Xu et al. propose ImageReward, the first general purpose text-to-image human preference reward model to encode human preferences automatically. In our work, we use a collection of three evaluation methods to evaluate DreamSync: VQA evaluation for generated images on both TIFA and DSG benchmarks, human evaluation, and ImageReward for automatic human preference prediction.
Improving General T2I Alignment.
We roughly categorize the alignment methods for improving T2I alignment into two classes depending on if they involve training. For training-involved methods, several works use Reinforcement Learning from Human Feedback (RLHF) based on human rankings to maximize a reward and improve faithful generation . In a similar vein, Pick-a-Pic is a dataset of prompts and preferences that is used to train a CLIP-based scoring function . StyleDrop trains adapters to synthesize of images that follow a specific style , and T2I-Adapter trains adapters to improve the control for the color and structure of the generation results . DreamBooth and HyperDreamBooth improve personalized generation , and they have inspired more efficient methods such as SVDiff . Being orthogonal to training-involved methods, there is a body of work on training-free methods that make inference time adjustments to the model to improve alignment, such such as SynGen and StructuralDiffusion. . DreamSync leverages training but does not involve reinforcement learning. We compare DreamSync with two RL-based methods and two learning-free methods in our experiments. We find that DreamSync outperform all the baselines in terms of text-image alignment on both DSG and TIFA.
Iterative Bootstrapping.
Iterative Bootstrapping, also known as model self-training, is a semi-supervised learning approach that utilizes a teacher model to assign labels to unlabelled data, which is then used to train a student model . In our work, we adopt a self-training scheme where the teacher model are the VLMs and the student model is the T2I model we aim to improve. During training, the VLMs (teacher) are used to annotate and select aligned examples for the next batch finetuning (student).
DreamSync
Our method improves alignment and aesthetics in four steps (see Figure 1): Sample, Evaluate, Filter, and Finetune. The high level idea is that T2I models are capable of generating interesting and varied samples. These examples are further judged by VLMs to pass qualification as faithful and aesthetic candidates for further finetuning T2I models. We next dive into each component more formally.
Given a text prompt , the text-to-image generation model generates an image . Generation models are randomized, and running multiple times on the same prompt can produce different images, which we index as . To improve the model’s faithfulness to text guidance, our method collects faithful examples generated by . We use to generate samples of the same prompt , so that with some probability , a generated image is faithful. Note that we need samples for each prompt , and DreamSync is not expected to improve totally unaligned models (with ). Prior work estimates that 5–10 samples can yield a good image, and hence, can be thought of as roughly 0.1 to 0.2.
Evaluate.
and the Absolute score is the absolute success rate
Filter.
We combine text faithfulness and visual appeal (given by ) as rewards for filtering. For a text prompt and its corresponding synthetic image set , we select samples that pass both VQA and aesthetic filters:
To avoid an imbalanced distribution where easy prompts have more samples, which could cause adversely affected image quality, we select one representative image (denoted as ) having the highest visual appeal for each :
We apply this procedure to all text prompts in our finetuning prompt set with , where is a prompt distribution. After filtering, we collect a subset of examples, , that meet our aesthetic and faithfulness criteria. Note that it is possible for to be empty, and we empirically show what fraction of the training data is selected in Figure 4. We ablate other aspects of the selection procedure in § 2.
Finetune.
After obtaining a new subset of faithful and aesthetic text-image pairs, we fine-tune our generative model on this set. We denote the generative model after iterations of DreamSync as , such that denotes the baseline model. To obtain we fine-tune on data generated by after applying our filtering procedure as outlined above. We follow the same loss objective and fine-tuning dynamics as LoRA . Let denote all parameters of a model, then the hypothesis class at iteration is:
where denotes the rank of weight updates and in practice we choose to balance efficiency and image quality. Overall, the iterative training procedure is as follows:
The self-training process Eq. 1 can in principle be executed indefinitely. In practice, it repeats for three iterations at which point we observe diminishing returns.
Datasets and Evaluation
In this section, we will introduce our training data in § 4.1 and evaluation benchmark in § 4.2.
To obtain prompts, and corresponding question-answer pairs without human-in-the-loop, we utilize the in-context learning capability of Large Language Models (LLM). We choose PaLM 2https://ai.google/discover/palm2/ as our LLM and proceed as follows:
Prompt Generation. We provide five hand-crafted seed prompts as examples and then ask PaLM 2 to generate similar textual prompts. We include additional instructions that specify the prompt length, a category (randomly drawn from twelve desired categories as in , e.g., spatial, counting, food, animal/human, activity), no repetition, etc.In Section A.1, we show the complete instruction used to probe LLM for the first two steps: prompt generation and QA generation. We change the seed prompts and repeat the prompt generation three times.
QA Generation. Given prompts, we then use PaLM 2 again to generate question and answer pairs that we will use as input for VQA models as in TIFA .
Filtering. We finally use PaLM 2 once more to filter out unanswerable QA pairs. Here our instruction aims to identify three scenarios: the question has multiple answers (e.g., “black and white panda” where the object has multiple colors, each color could be the answer), the answer is ambiguous (e.g., “a lot of people”) or the answer is not valid to the question.
We showcase the diversity of PaLM 2 generated prompts in Figure 3 using qualitative examples and quantitive statistics of our generated prompts in Section A.2.
2 Evaluation Benchmarks
Using the previously generated prompts, we evaluate whether DreamSync can improve the T2I model performance on benchmarks that include general prompts. We consider the follow benchmarks.
To evaluate the faithfulness of the generated images to the textual input, TIFA uses VQA models to check whether, given a generated image, questions about its content are answered correctly. There are 4k diverse prompts and 25k questions spread across 12 categories in the TIFA benchmark. Although there is no overlap between our training data and TIFA, we use the TIFA attributes to constrain our LLM-based prompt generation. Therefore, we use TIFA to test DreamSync on in-distribution prompts. We follow TIFA and use BLIP-2 as the VQA model for evaluation.
Davidsonian Scene Graph (DSG).
DSG exhibits the same VQA-as-evaluator insight as TIFA’s and further improves its reliability. Specifically, DSG ensures that all questions are atomic, distinct, unambiguous, and valid. To comprehensively evaluate T2I images, DSG provides 1,060 prompts covering many concepts and writing styles from different datasets that are completely independent from DreamSync’s training data acquisition stage. Not only is DSG a strong T2I benchmark, it also enables further analysis of DreamSync with out-of-distribution prompts. Furthermore, DSG uses PaLI as the VQA model for evaluation, which is different from the VQA model that we use in training (i.e., BLIP-2) and lifts the concern of VQA model bias in evaluation. We use DSG QA both automatically (with PaLI) and with human raters (details in Appendix C).
Experiments
We explain our experimental setup in § 5.1, and showcase the efficacy of training with DreamSync and compare against other methods in § 5.2. § 2 analyzes our choice of rewards; § 5.4 reports results for a human study.
We evaluate DreamSync on Stable Diffusion v1.4 , which is also used in related work. Additionally, we consider SDXL , which is the current state-of-the-art open-sourced T2I model. For each prompt, we generate eight images per prompt, i.e., .
Fine-grained VLM Feedback.
Baselines.
We compare DreamSync with two types of methods that improve the faithfulness of T2I models: two training-free methods (StructureDiffusion and SynGen ) and two RL-based methods (DPOK and DDPO ). As the baselines use SD v1.4 as their backbone, we also use it with DreamSync for a fair comparison.
2 Benchmark Results
In Table 1 we compare DreamSync to various state-of-the-art approaches with four random seeds. In Appendices E and D we show more qualitative comparisons.
For SDXL , we show how three iterations of DreamSync improves the generation faithfulness by 1.7 point of mean score and 3.7 point of absolute score on TIFA. The visual aesthetic scores after performing DreamSync improved by 3.4 points. Due to the model-agnostic nature, it is straightforward to apply DreamSync to different T2I models. We also apply DreamSync to SD V1.4 . DreamSync improves faithfulness by 1.0 points of mean score and 1.7 points of absolute score on TIFA, together with a 0.3 points of VILA score improvement for aesthetics. Most prominently on DSG1K, DreamSync improve text faithfulness of SDXL by 2.9 points. We report fine-grained results for DSG in Appendix C.
DreamSync yields the best performance in terms of textual faithfulness on TIFA and DSG.
This is true without sacrificing the visual appearance as shown in Table 1. In Figure 4 we report TIFA and aesthetics scores for each iteration, where we observe how DreamSync gradually improves the alignment and aesthetics of the generated images. We highlight several qualitative examples in Figure 2.
3 Analysis & Ablations
We analyze whether using BLIP-2 as a VQA model for finetuning and for evaluation in TIFA might be the reason for the improvement by DreamSync that we have observed. To test this we use PaLI to replace the BLIP-2 as the VQA in TIFA. Using SDXL as the backbone, DreamSync improves the mean score from 90.09 to 92.02 on TIFA compared to the vanilla SDXL model. This results confirms that DreamSync is in fact able to improve the textual faithfulness of T2I models.
Ablating the Reward Models
In Table 2, we present the results for an ablation study where we remove one of the VLMs during filtering and evaluate SDXL after applying one iteration of DreamSync. It can be seen how training with a single pillar mainly leads to an improvement in the corresponding metric, while the combination of the two VLM models leads to strong performance for both text faithfulness and visual easthatics, justifying our approach. One interesting finding is that training with both rewards, rather than VILA only, gives the highest visual appeal score. Our possible explanation is that images that align with user inputs may have higher visual appeal.
ImageReward.
We next test whether DreamSync yields an improvement on human preference reward models, even though DreamSync is not trained to optimize them. We use ImageReward as an off-the-shelf human preference model for generated images. Table 3 shows that DreamSync plus either SD v1.4 or SDXL increases ImageReward scores on images based on both TIFA and DSG1K. Tuning with VLM-based feedback helps align the generated images with human preferences, at least according to ImageReward.
4 Human Evaluation
To corroborate the VQA-based results, we first conduct a preliminary human study to evaluate the faithfulness of generated images. It shows simply asking one question ‘‘Which image better aligns with the prompt?’’ yields poor inter-annotator agreement. We speculate that asking a single question encompassing the whole prompt makes the alignment difficult to evaluate.
To address this issue, we conduct a larger follow-up study based on DSG , where we ask approximately 8 fine-grained questions for each of 1060 images to external raters. These questions are divided into categories (entity, attribute, relation, global). Here in Figure 5, we observe consistent and statistically significant improvements comparing DreamSync to SDXL. In each category, images from DreamSync contain more components of the prompts, while excluding extraneous features. Overall, DreamSync’s images led to 3.4% more correct answers than SDXL images, from 70.9% to 74.3%. Full details and findings for both studies are in Appendix C.
Discussion
A key design choice behind DreamSync is to maintain simplicity and automation throughout each step of the pipeline. Despite this feature, our experimental results show that DreamSync can improve both SD v1.4 and SDXL on TIFA, DSG, and visual appeal. In the case of SD v1.4, this improvement holds true compared four different baseline models (two training-free and two RL-based). For SDXL, even though the base model achieves SoTA results among open-source models, DreamSync can still substantially improve both alignment and aesthetics.
The effectiveness of DreamSync’s self-training methodology opens the door for a new paradigm of parameter-efficient finetuning. Indeed, the DreamSync pipeline is easily generalizable. For the training prompts, we can construct a set with complex and non-conventional examples compared to standard web-scraped data. On the filtering and fine-tuning side, our framework shows that VLMs can provide effective feedback for T2I models. Together, these steps do not require human annotations, yet they can tailor a generative model toward desirable criteria.
Like prior methods, the performance of DreamSync is limited by the pre-trained model it starts with. As exemplified in “the eye of the planet Jupiter” in Figure 2, SDXL generates a human’s eye rather than Jupiter’s. DreamSync adds more features of the Jupiter in each iteration. Nevertheless, it did not manage to produce an image that is perfectly faithful to the prompt. This is also exemplified by the quantitative results in §5.2. Despite outperforming the baselines using SD v1.4 on TIFA and DSG, SD v1.4 + DreamSync still falls behind SDXL. Similarly, our human studies on DSG in §5.4 indicate that DreamSync improves SDXL from 70.9% accuracy to 74.3%. Nonetheless, there is still a 25.7% headroom to improve. We identify several common failure modes (e.g., attribute-binding) and conduct a detailed analysis in Appendix B. Future works may investigate if these challenges can be addressed by further scaling up DreamSync, or mixing it with large-scale pre-training.
Conclusion
We introduce DreamSync, a versatile framework to improve text-to-image (T2I) synthesis with feedback from image understanding models. Our dual VLM feedback mechanism helps in both the alignment of images with textual input and the aesthetic quality of the generated images. Through evaluations on two challenging T2I benchmarks (with over five thousand prompts), we demonstrate that DreamSync can improve both SD v1.4 and SDXL for both alignment and visual appeal. The benchmarks also show that DreamSync performs well in both in-distribution and out-of-distributions settings. Furthermore, human ratings and a human preference prediction model largely agree with DreamSync’s improvement on benchmark datasets.
For future work, one direction is to ground the feedback mechanism to give fine-grained annotations (e.g., bounding boxes to point out where in the image the misalignment lies). Another direction is to tailor the prompts used at each iteration of DreamSync to target different improvements: backpropagating VLM feedbacks to the prompt acquisition pipelines for continual learning.
List of Contributions
Jiao Sun: Jiao leads the project. She implemented DreamSync internally at Google, showcasing its success on various reward models. She also wrote the first paper draft.
Deqing Fu: Deqing initiated the idea of LLM-based prompt generation. He implemented DreamSync with open-source text-to-image models, conducting all the experiments. He also contributed to paper writing and drew all the figures.
Yushi Hu: Yushi conceived the idea of DreamSync and designed the experiments. He also implemented the VQA feedback, the baselines, and contributed to paper writing.
Jiao, Deqing, and Yushi completed the full development cycle of DreamSync. They contributed equally on designing technical directions, implementing, polishing and improving DreamSync from scratch.
S. Wang ran both automatic and human evaluations on DSG1K and contributed to corresponding paper sections.
R. Rassin contributed to the implementation of the baseline methods SynGen and DPOK.
J. Sun, C. Rashtchian, S. van Steenkiste, C. Herrmann, D-C. Juan, and D. Alon conceived of the initial project directions, e.g., T2I models struggle with compositionality and image quality.
R. Krishna, S. van Steenkiste, C. Herrmann, and D. Alon provided constructive feedback and suggested experiments.
R. Krishna and S. van Steenkiste helped frame the story via writing and polishing several sections of the paper.
C. Rashtchian served as senior project lead and manager, scoping technical directions and facilitating collaborations.
We thank Yi-Ting Chen, Otilia Stretcu, Yonatan Bitton, Fei Sha, Kihyuk Sohn, Chun-Sung Ferng, and Jason Baldridge for helpful project discussions and technical support. JS and DF would like to thank USC NLP group and YH would like to thank UW NLP group, for providing both additional GPU computational resources and fruitful discussions.
References
Appendix A Training Data Acquisition
Training Data Acquisition is the first step and the foundation of DreamSync as discussed in Section 4.1. We use PaLM 2 for each step of the training data acquisition, including prompt generation, QA generation and filtering. Here are the complete instructions that we use.
You are a large language model, trained on a massive dataset of text. You can generate texts from given examples. You are asked to generate similar examples to the provided ones and follow these rules:
Your generation will be served as prompts for Text-to-Image models. So your prompt should be as visual as possible.
Your generated examples should be as creative as possible.
Your generated examples should not have repetition.
Your generated examples should be as diverse as possible.
Do NOT include extra texts such as greetings.
Instruction for QA Generation.
Given a image descriptions, generate one or two multiple-choice questions that verifies if the image description is correct. Classify each concept into a type (object, human, animal, food, activity, attribute, counting, color, material, spatial, location, shape, other), and then generate a question for each type. We then provide fifteen prompts together with about ten question answer pairs as demonstration for PaLM 2. Table 4 shows an example of PaLM2-generated prompt and QA. Answer source and Answer Type are also automatically generated altogether, making it possible for us to get statistics of our training set below.
A.2 Statistics
Table 5 shows the statistics of the prompts and questions we obtained, and we list a few prompts from our training set and DreamSync’s generation in Figure 3. Prior work (e.g., TIFA, DSG) identifies that T2I models do not perform equally well for depicting different attribute categories; we verify the variety of attributes in our prompts by counting unique words (i.e., Answer Source in Table 4) in these categories (i.e., Answer Type in Table 4): counting (4179), object (3638), shape (973), human (945), location (1047), activity (2984), attribute (2925), color (3259), food (1086), spatial (1009), animal (645), material (1610), existence (3072), and other (878).
A.3 Images Generated by DreamSync for Finetuning Exhibit High Quality
Appendix B Failure Modes Analysis
Figure 7 presents several side-by-side examples showcasing common failure modes of DreamSync. For each example, we show the image generated by SDXL on the left, and the image of SDXL + DreamSync on the right. We also indicate some key directions for improvements.
Composing multiple objects and attributes is still challenging. As shown in (a), (b), (c), and (d), SDXL + DreamSync struggles to produce an image that is faithful to the prompt. In (a), DreamSync adds a bench in the image. However, the attributes of chairs and benches are mixed. In (b), DreamSync removes the extra glass in the background, but neither model is able to place the lemon wedge in the rim of the bottom. In (c), DreamSync adds purple fish in the image, but the counting is not correct. In (d), DreamSync produces four objects but they are cloud-keychain combinations.
We observe decline of texture details and shadows on some images. In (e), the alignment between the text and the bus significantly improves. However, the quality of the bus shadow declines. In (f), both images align well with the text. The main difference is in the details of the temple facade. Notice that for most images we observe DreamSync yields images with high quality and visual appeal, as illustrated in Section A.3.
Future work may explore if these challenges can be addressed by following extensions to DreamSync: (1) DreamSync could be used in tandem with RL-based method and training-free method to further improve text-to-image faithfulness; (2) prompt engineering methods in DALL-E 3 may help rewriting challenging prompts into simpler ones for models to synthesize; (3) scaling up DreamSync with a more diverse set of prompts and reward models; (4) mixing DreamSync with large-scale pre-training on real images. In summary, as discussed in §6.1, there is still plenty of headroom to improve.
Appendix C DSG and Human Rating Evaluation
Tab. 6 presents the data sources, quantity, and examples for DSG-1k. Fig. 8 summarizes the 4 broad and 14 detailed semantic categories covered in the benchmark.
Like TIFA, DSG falls into the Question Generation / Answering (QG/A) alignment evaluation framework. Unlike TIFA, DSG introduces a linguistically motivated question generation module to ensure the questions generated to hold 4 reliability traits: a) atomic: only queries about 1 semantic detail, for unambiguous interpretation; b) unique: no duplicated questions; c) dependency-aware: prevent invalid queries to VQA/human answerers, e.g. if the answer to a parent question “is there a bike?” is negative, then the child question “is the bike blue?” will not be queried; d) full semantic coverage: dovetailing the semantic content of a prompt, no more no less. DSG is powered by a large variant of PaLM 2 for QG and the SoTA VQA module PaLI for QA. For our evaluation task, we adopt DSG-1k (DSG’s 1,060 benchmark prompt set) which covers a balanced set of diverse semantic categories and writing styles – including 4 broad categories (e.g. entity/attribute/etc.) and 14 detailed categories (e.g. color/counting/texture/etc.).
Human QA protocol.
For human evaluation, we elicit 3 rating responses per prompt/question set (with 8 questions per set on average, and a total of 8183 questions). Fig. 9 exemplifies the UI the human raters see. Fig. 10 presents the annotation instructions used to guide the raters. The inner-annotator agreement for this study is 0.684. While the raters respond with YES/NO/UNSURE, we find it to be practically useful to numerically convert the answers – 1.0 point for YES, 0 for NO, and 0.5 for UNSURE as partial credit, with the justification that if a semantic detail can potentially be grounded in an image yet not necessarily so (e.g. “does this man dress like an engineer?” image: a male in a plain shirt; “is this a cat” image: a blob that may be interpreted as a cat), partial credit is fair for not completely failing.
Detailed Human Evaluation Results on DSG-1K.
We present detailed evaluations on DSG-1K by semantic categories listed in Figure 8. The results are shown in Figure 11. By applying DreamSync upon SDXL, the human evaluation on alignments improved on all categories.
Single-Question Human Evaluation.
Besides the large-scale human annotation, we also did a light-weight single-question human evaluation for text prompt alignment. This study was completed by three of the paper’s authors. Although this study yields a quite low inter-annotator agreement, we hope it would provide valuable insights on how to set up human evaluation for measuring textual faithfulness of generated images. For this study, we generated one image with SDXL and DreamSync. See Figure 12 for an example rating screen. We randomized the order of the images and prompts. Three authors were asked ‘‘Which image better aligns with the prompt?’’ They could choose Image 1 is better, Image 2 is better, or that they cannot tell (indicating a tie). We use 200 prompts in total with 100 prompts from TIFA and another 100 from DSG.
As mentioned, the inner-annotator agreement was quite low for this study. Only for 42.5% of the 200 prompts did the human raters all agree in their answers. This is likely due to the fact that it is hard to judge overall prompt alignment directly when given two side-by-side images. Indeed, the majority of prompts led to the raters choosing that they cannot tell which image is better. Using the scoring rules from the DSG study described above (with 1 point going to the model with a direct vote, and with 0.5 going to each model for a tie vote), then we have that DreamSync scores 50.08 while SDXL scores 49.92.
Key Takeaway from Human Evaluation.
Comparing the fine-grained large-scale human evaluation and the single-question human evaluation, we encourage researchers who are interested in evaluating the text-image alignment to ask annotators detailed and fine-grained questions. It yields significantly better inter-annotator agreement than asking a general single question about alignment. Our large-scale human evaluation with a better agreement suggests that DreamSync improves the textual faithfulness of SDXL on DSG-1k, resonating with our automatic evaluation.
Appendix D Randomly-Sampled SDXL+DreamSync Images
Appendix E Qualitative Comparison with SD v1.4-based Methods in Table 1
Among the 6 examples shown in Figure 15, DreamSync has 3 absolute successes, wheres SynGen, DDPO and StructureDiffusion each has 2, DPOK has 1 and the base model SD v1.4 has 0 absolute success. These results match well with Table 1.