DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

Yuang Peng, Yuxin Cui, Haomiao Tang, Zekun Qi, Runpei Dong, Jing Bai, Chunrui Han, Zheng Ge, Xiangyu Zhang, Shu-Tao Xia

Introduction

Driven by the significant advances in large-scale text-to-image (T2I) generative models , it is now possible to generate images conditioned on not only arbitrary text prompts but also by given reference images—personalized image generation . In general, to be useful as an artistic creation tool for inspiration or products , the following two basic criteria must be fulfilled: i) Prompt following (image & prompt consistency). Generated images must follow the prompt description, which is a requirement shared with vanilla T2I generation . ii) Concept preservation (image & image consistency). For personalized image generation, the concept of the reference image, i.e., the main subject’s semantic details (e.g., facial characters) or high-level abstractions (e.g., overall style), must be preserved in the generated image. For example, a user may want to “imagine his own dog traveling around the world” , and the generated dog must be the same as his but traveling.

To meet the aforementioned requirements, numerous efforts have been devoted. One line of fine-tuning-based works focuses on fine-tuning general T2I models to specialist personalization models by reproducing specific concepts present in training sets . Meanwhile, another line of encoder-based works, instead, achieves concept-preservation by training features adaptation to inject reference image features into a general T2I model . Although remarkable results have been achieved, one question arises: can we comprehensively evaluate these models to figure out which technical route is superior and where to head?

In this work, we aim to answer this question by developing a new benchmark that properly evaluates personalized T2I models driven by the above two requirements. We present DreamBench++, a comprehensive benchmark designed based on the following de-facto principled advantages:

Human-Aligned As shown in Fig. 2, traditional metrics like DINO and CLIP often result in significant discrepancies from humans. This is caused by the image similarity measurement nature of DINO and CLIP models, and thus crowd-sourced human evaluation is typically necessary for obtaining a correct quantitative understanding of generated images . Therefore, different from existing works that utilize CLIP and DINO as metrics that may be humanly misaligned, our DreamBench++ demonstrates surprisingly consistent evaluation results aligned with humans. For instance, by evaluating 7 modern models, DreamBench++ achieves 79.64% and 93.18% agreement with human’s evaluation in concept preservation and prompt following capabilities, respectively. Notably, it is +54.1% and +50.7% higher than traditional DINO and CLIP metrics.

Automated However, it is non-standardized and expensive to perform high-quality human evaluations, making human evaluations non-trivial to use. To address this challenge, DreamBench++ achieves automated but human-aligned evaluation by using advanced multimodal GPT models as metrics such as GPT-4o . The challenges lie in two aspects: i) prompt design and ii) reasoning procedure for scoring. We systematically standardize the automated GPT evaluation by first designing the evaluation instructions that provide overall task requirements, where language is a general interface for instructing human preference. Inspired by Self-Align , we instruct the GPT model to conduct internal thinking that aligns itself for better task and preference understanding. Then, the GPT model provides the summary & planning for the task and scoring criteria, and the final scores are provided with optional chain-of-thought (CoT) .

Diverse To avoid biased understanding due to the low diversity of evaluation data, DreamBench++ collects a large number of diverse images covering various levels of difficulty categories, from animals and styles that are relatively simple to more challenging human subjects, objects, and non-natural image styles (See Fig. 1). Compared to commonly used DreamBench that only contains 30 subjects and 25 prompts, DreamBench++ scales up the benchmarking data to 150 images and 1,350 prompts, i.e., 5×\times and 54×\times more than DreamBench. As a result, when evaluating models on this larger dataset, different but more robust conclusions have been obtained that demonstrate an unbiased and more comprehensive evaluation of DreamBench++.

Takeaways Here, we offer some insightful findings that we obtain by evaluating 7 modern personalized T2I models: i) DINO-based rating puts tremendous weight on resemblance in overall shape or color and neglects detailed visual features, which makes it a sub-optimal metric for personalized image generation evaluation; ii) The main goal of personalized image generation is to find the Pareto optimal solution that balances concept preservation and prompt following. Among these 7 models, DreamBooth is the best in overall performance, preserving highly detailed visual features while keeping close adherence with text prompt; iii) Currently, personalized T2I models perform well on animal and style category. However, they have not performed well in the human category due to sensitivity to facial details and in the object category due to its diversity. Existing works focuses on preserving facial features to address the former problem, but the latter has not been extensively explored.

We are presenting DreamBench++ with open-sourced codes and evaluation standardization to promote innovation within the research community. In addition, we believe our design of the human-aligned & automated evaluation using advanced foundation models is robust and transferrable to other domains and foundation models (e.g., GPT-5 in the future).

DreamBench++

We introduce DreamBench++, a human-aligned, automated, and diverse benchmark that evaluates the two capabilities of personalized image generation models. In the following, we describe how we construct DreamBench++ from two aspects: prompts and data.

It is challenging to obtain a solid quantitative understanding of generated models, especially when evaluating visual contents like images that rely on human evaluations . Thus, it is critical to achieve automated evaluation by utilizing multimodal GPT models, which are trained particularly in the principle of aligning with human preference . This is evidenced by the recent progress achieved by Wu et al., which demonstrates that GPT-4V can serve as a human-aligned text-to-3D generation evaluator. However, as pointed out by Zhang et al. and Ku et al., multimodal GPT models often fall short in evaluating personalized image generation—often more challenging when distinguishing subtle difference for concept preservation assessment using GPT—still underexplored. To tackle this issue, we detail how we systematically design the prompt of multimodal GPT (GPT-4o , by default) for human alignment reinforcement but also improve the reasoning progress that helps the GPT models to be more self-aligned, introduced as follows.

Compare or rate? There are typically two schemes for quantitatively evaluating generative models in human evaluations: rating and comparison . The rating scheme requires human reviewers to assign an absolute score to each instance, while the comparison scheme asks human reviewers to express a relative preference among different instances. Though effective as the comparison scheme is when humans are involved, we find that there are two critical issues. i) Positional Bias: the scoring results of GPT-4V/GPT-4o is sensitive to the order in which images are presented , making it unsuitable for comparison scheme. ii) Quadratic Complexity: As the number of methods increases, the number of essential evaluation runs for numerical rating increases linearly, while the number of comparative assessments increases quadratically. Therefore, direct numerical rating is more efficient and scalable when evaluating multiple methods. Hence, in this work, we adhere to the rating scheme, and we establish a 5-level rating scheme where scores are integers ranging from 0 (very poor) to 4 (excellent).

Evaluation Instructions The evaluation instructions serve as the meta-prompting that describes overall tasks, which is shown in Fig. 3. As stated in Section 1, there are two fundamental quality criteria to be evaluated: i) concept preservation and ii) prompt following. For each aspect, we use a similar prompt template that contains ❶ task description, ❷ scoring criteria explanation, ❸ scoring range definition, and ❹ format specification. Only the scoring criteria are tailored for different tasks: for concept preservation evaluation, we prompt GPT to focus on shape, color, texture, and facial features (if applicable), while for prompt following evaluation we requested for focus on relevance, accuracy, completeness and context.

Reasoning Instructions Given the evaluation instructions, it is crucial to reinforce the alignment with both the human instruction and itself to largely leverage the pretrained knowledge. To this end, we adopt a 2-step evaluation policy as follows: i) Internal Thinking: Inspired by Self-Align , we introduce internal thinking to strengthen task understanding and instruction following capabilities. Specifically, we prompt the GPT model by asking if it understands the task or not and let it summarize the task. ii) Summary & Planning: According to the given internal thinking instruction, the GPT will summarize and plan for the evaluation task itself. It can also be viewed as a generalized form of chain-of-thought reasoning . The complete procedure is illustrated in Fig. 3.

2 Scaling Up Personalized Image Generation Benchmarking

Pioneering works like DreamBooth and SuTI have successfully set up baseline datasets for the evaluation of personalized image generation, and DreamBench++ follows them to categorize images into three types: ❶ objects, ❷ living subjects, and ❸ styles. However, due to the small-scale nature of DreamBench, it is limited as some methods may converge well on its samples while performing unsatisfactorily on other data. To avoid this possible biased evaluation, we scale up the benchmarking data by increasing both image numbers and diversity.

Data Construction from Internet There are broad images on the Internet, and many datasets are constructed from it . DreamBench++ mainly collects images from Unsplash , Rawpixel , and Google Image Search , along with contributions from individuals with authorized permissions. Each image’s copyright status has been verified for academic suitability. As shown in Fig. 4, we collect and construct high-quality data in DreamBench++ by following 3 steps:

Keywords Generation First, we generate 200 relevant keywords using GPT-4o and join them with the 200 most frequent keywords from Unsplash. After filtering out duplicated keywords, seven human annotators will extend the list to around 300 based on their interests.

Internet Images Collection Given selected keywords, we retrieved corresponding images from Unsplash, Rawpixel, and Google Image Search. To filter out images unsuitable for personalized image generation, SAM is applied to identify subject regions in images and discard those with too small subject areas. Human annotators will then filter out images with noisy backgrounds. Curated images were cropped to centralize the subject, resulting in two images per keyword. Keywords that fail to yield suitable images will be discarded in this process.

Prompt Generation After image collection, 9 text prompts per image were generated using GPT-4o, designed to cover a range of difficulties: 4 prompts for ❶ photorealistic styles, 3 for ❷ non-photorealistic styles, and 2 for ❸ complicated & imaginative contents. To align with established evaluation methods, we use few-shot prompts selected from PartiPrompts . Human calibration ensures that all generated prompts are ethical and without flaws. As a result, the construction process finally yields 150 high-quality images and 1,350 prompts.

Diversity Visualization When collecting images, we notice that Internet images possess a bias towards photorealistic styles. To diversify, various non-photorealistic styles are enlisted, and human annotators are tasked to gather images for each style, including anime, sketches, traditional Chinese paintings, artworks, and cartoon characters from games. Then, a manual selection process ensures a balanced distribution of images across subject classes and between photorealistic and non-photorealistic styles. In Fig. 5(a), we visualize the t-SNE of images from DreamBench and DreamBench++, which demonstrates the superiority of DreamBench++ in diversity. Besides, Fig. 5(b) presents the detailed statistics of the image distribution in DreamBench++.

Experiments

In this section, we provide evidence that DreamBench++ can more accurately reflect the capabilities of existing personalized image generation methods with its diverse and comprehensive data set. Additionally, we demonstrate that our systematic prompts-based GPT is a superior evaluator compared to existing metrics by analyzing their degree of alignment with human ratings. In order to further facilitate benchmarking research with GPT models, we conduct a comprehensive ablation study to show the designing key points for more human-aligned automated GPT evaluation.

Reimplementation Details We mainly conduct experiments on the two mainstream methods: i) ∙\bullet Fine-tuning-based methods, including ❶ Textual Inversion (TI) , ❷ DreamBooth , and ❸ DreamBooth LoRA (DreamBooth-L) ; ii) ∙\bullet Encoder-based methods that trains feature adaptation, including ❹ BLIP-Diffusion (BLIP-D) , ❺ Emu2 , ❻ IP-Adapter-Plus ViT-H (IP-Adapt.-P) , and ❼ IP-Adapter ViT-G (IP-Adapt.) . All these methods are based on the base T2I models, including SD v1.5 and SDXL v1.0 . We stay true to the official implementations wherever possible and dedicate significant effort to tuning hyper-parameters to ensure that the performance of each method on DreamBench is consistent with results reported in original papers. More implementation details can be found in Appendix A.

Human Annotators We employ 7 human annotators to score for each instance in DreamBench++ to obtain ground truth human preference data. We provide human annotators with sufficient training to ensure they fully understand the personalized image generation task and can provide unbiased and discriminating scores. The scoring task and scheme given to humans are identical to those used for GPT, as described in Section 2. The GPT results and human results are isolated to avoid hindsight bias. Additionally, we ensure that each instance is rated by at least two humans to reduce noise.

2 Main Results

Quantitative Analysis Table 1 shows the overall evaluation results, including human and GPT-4o rating scores. From the results, we observe that: i) DreamBench++ aligns better with humans than DINO or CLIP models. Driven by our dedicatedly-designed prompts, GPT-4o used by DreamBench++ yields impressive alignment with humans. This is because humans and DreamBench++ are all advanced in evaluating facial and textural characters and producing scores with a balanced consideration of all aspects. ii) DINO-I and CLIP-I yield significant divergence from humans in evaluating concept preservation. This could be because DINO/CLIP scores show a preference for images that preserve shapes or overall styles, as shown in Fig. 6. For example, methods proficient in posture and detailed texture transformation, such as DreamBooth-L, rank lower than humans and DreamBench++. iii) Traditional CLIP-T scores are as effective as DreamBench++ in evaluating prompt following, showing strong alignment with humans.

Qualitative Analysis Fig. 7 shows selected qualitative results on DreamBench++, which helps provide an intuitive understanding of evaluated models. With a more comprehensive and diverse collection of images, we have discovered numerous intriguing characteristics of these generation methods that were not apparent on existing datasets such as DreamBench. Specifically, we observe that: i) Fine-tuning-based methods outperform encoder-based methods on images containing more subject-oriented information, such as an animal or object, as they preserve more intricate details in the generated images. However, for images containing a person, fine-tuning-based methods often fail to preserve facial and clothing features. This suggests that the personalized generation of human images is more demanding for visual concept preservation than for textual following capabilities, which is an advantage for encoder-based methods. ii) However, for style-oriented cases when subject details are less critical, encoder-based methods perform better than fine-tuning-based methods. This further highlights the strengths of encoder-based methods in that they are more adept at recognizing and extracting high-level visual semantics, including overall shape, style, and thematic features.

Leaderboard Table 2 shows the detailed results of the leaderboard with respect to the concept and prompt categories defined in Section 2. Note that: i) the human category shows the lowest average score of 0.482, which is -0.204 lower than the highest average score of animal. This category is very challenging in terms of concept preservation because humans are very sensitive to facial details, and many works are conducted specifically on it . ii) The object is also a relatively difficult category because the characteristics of each same or different object are often very diverse. In contrast, animals within the same category often share a large degree of visual similarity. iii) There exists a negative correlation between concept preservation and prompt following. The primary aim of personalized T2I evolution is to identify the Pareto optimum that balances both factors.

3 Ablation Study

Table 3 shows the ablation study of the prompt design influences on alignment with humans. From the results, we observe that: i) The proposed prompt designs are all necessarily effective, demonstrating the superiority of the prompting method in DreamBench++. For example, removing the proposed internal thinking leads to a significant drop, indicating the effectiveness of self-alignment for alignment with humans. ii) The capability of the multimodal GPT used is scalable. This shows that DreamBench++ has the potential to be improved with stronger GPT models in the future. iii) Some human prior knowledge, such as reminding the GPT not to consider background when assessing visual concept preservation, leads to performance degradation.

Discussions

Table 4 shows a more rigorous study of human alignment level using the mean Krippendorff’s alpha value . The results show that DreamBench++ is a highly human-aligned benchmark. Notably, DreamBench++ achieves 79.64% and 93.18% evaluation consistency with human’s evaluation in concept preservation and prompt following capabilities, respectively. This result is +54.1% and +50.7% higher than traditional DINO and CLIP metrics.

2 Is data diversity necessary?

To investigate the necessity of diverse and rich evaluation data, we compare the results on DreamBench and DreamBench++ using DINO and CLIP metrics. Table 5 demonstrates that diverse data in DreamBench++ is necessary to avoid biased evaluation. Although the results are overall consistent, especially for encoder-based methods, fine-tuning-based methods such as TI and DreamBooth both show a notable score drop. We argue that this is due to the reduced training steps and learning rate limited by the single-image fine-tuning, which may make the model not fully converged. However, unlike other encoder-based methods, Emu2 shows a significant score drop. Emu2 performs well only on natural images and simple text, while it cannot generate valid results when texts are complex or stylized or with anime reference images, as shown in Fig. 8.

3 Can we use free lunch to improve DreamBench++ evaluation?

Table 6 shows the result of utilizing free lunch techniques, including chain-of-thought (CoT) and In-Context Learning (ICL) . CoT indicates that GPT-4 articulates its reasoning process before scoring, and ICL indicates GPT-4o is provided with human-written few-shot examples.

Chain-of-Thought: i) CoT is effective in evaluating prompt following capability. Through CoT, the model more accurately discerns the significance of phrases such as “morphs into a mythical dragon”, allowing it to assign a more appropriate evaluation score. ii) CoT does not bring improvement in concept preservation evaluation. We argue that CoT may shift attention to unnecessarily important background or texture information, as shown in Fig. 9.

In-Context Learning: ICL counterintuitively leads to a drop in alignment. This could be attributed to the patching scheme, sample selection, or inherent bias within GPT-4o, making it non-trivial to prompt effectively. Thus, we provide our detailed prompt and hope to inspire future works.

Related Works

Personalized Image Generation aims to preserve concept consistency while accommodating the diverse contexts suggested by the instructions. In general, it can be traced back to early efforts on pixel-to-pixel (Pix2Pix) translation where the personalization orientation is free-form texts or predefined translation across styles, seasons, species, or plants, etc . Modern efforts go beyond Pix2Pix translation toward a free-form image generation conditioned on both reference images and prompts. Some works focus on fine-tuning techniques that turn a general T2I model into a specialist personalization model using LoRA or contrastive learning , learning the subject or style information by reconstructive autoencoding . However, the necessity to fine-tune for new subjects limits their scalability. In contrast, encoder-based methods can generate subject-guided or style-guided images or edit images following prompts with one shot. Encoder- or adapter-based methods train an encoder to encode the conditional image into embeddings, which are integrated into cross-attention mechanism in the diffusion process . Adapter-free methods extract the information, such as attention maps [36, RBModulation24] from reference images, and fuse them into the image generation process. Furthermore, multimodal large language models (MLLMs) that are trained on extensive multimodal sequences can also serve as general foundation models .

Benchmarking Image Generation involves a variety of metrics that focus on different aspects. Inception Score and FID judge image quality, while LPIPS , DreamSim , CLIP-I , and DINO Score measure perceptual similarity. In text-guided generation, prompt-image alignment can be assessed by CLIP-T , CLIPScore , and BLIP Score . However, these metrics often fall short of reflecting human perception. To address this, human-aligned metrics have been introduced, offering a more perceptive evaluation. Yet, they face limitations in scaling with the pace of new model developments. Thus, the necessity for automated and sustainable evaluation methods has emerged, with some leveraging reward-model-based methods to encode human preferences, while others use multimodal to automate the process and better mirror human tastes. While MLLM-based methods show promise in aligning with human preferences , automated personalized evaluation remains an unresolved issue. VIEScore assesses image generation quality by prompting GPT-4V and LLaVA , but is limited to four models in subject-driven tasks and obtains suboptimal results. Meanwhile, Dreambench , a common benchmark for personalized generative evaluation, only consists of 30 simple objects and lacks diversity comprehensiveness.

Conclusions

This paper introduces DreamBench++, a human-aligned personalized image generation benchmark. Extensive and comprehensive experiments have demonstrated our advantages in dataset diversity and complexity, as well as the alignment of our evaluation metrics with human preferences. In addition, we offer insights into prompt design for advanced multimodal GPTs, emphasizing the potential and challenges of enhancing GPT evaluation through chain-of-thought prompting and in-context learning. Our work aims to support future research on personalized image generation by providing a human-aligned benchmark and heuristics in utilizing advanced multimodal GPTs in visual evaluation.

References

Appendix A Implementation Details

The configurations for the training hyperparameters used in training-based methods on DreamBench and DreamBench++, are detailed in Table 7. During the inference stage, all methods employ a guidance_scale of 7.5 and execute 100 inference steps, with the exception of Emu2, which uses a guidance_scale of 3 and performs 50 inference steps. Furthermore, BLIP-Diffusion and IP-Adapter incorporate negative prompts, as demonstrated in Table 8. Specifically, IP-Adapter includes an additional parameter, ip_adapter_scale, set at 0.6.

We dedicate significant effort to tuning hyperparameters to ensure that the performance of each method on DreamBench is consistent with results reported in original papers. As shown in Table 9, the results of our reproduction are comparable to or even better than the official results reported by the original papers.

Appendix B Limitation & Future Work

Human-aligned evaluation & benchmarking is an emerging but challenging direction, and we have only made preliminary attempts at personalized image generation. Moreover, our evaluation results heavily rely on the advancements of multimodal large language models and require carefully designed system prompts. We believe that as visual world models continue to develop, the evaluation performance will be further optimized. Our future work will focus on more applications with human-aligned evaluation, such as 3D generation , video generation , autonomous driving , and even embodied visual intelligence .

Broader Impact

Powerful as the T2I generative models pretrained on large-scale web-scraped data, the models may be misused as illegal or unethical tools for generating NSFW content. This potential impact can also be brought by personalized T2I models as they are typically built on the pretrained T2I foundation models. As a result, it is critical to use tools such as NSFW detectors to avoid such content during both usage and evaluation. For example, the data used for evaluation must avoid the NSFW content by data filtering. In this paper, such contents are filtered out by human annotators.