Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Daiqing Li, Aleks Kamko, Ehsan Akhgari, Ali Sabet, Linmiao Xu, Suhail Doshi
Introduction
Great progress has been made in diffusion-based generative models since the success of better image modeling performance with ImageNet, as compared to performance with the previously dominate framework of generative adversarial networks (GAN) . Open-source models like SDXL have built on top of latent diffusion models (LDM) by scaling up text-to-image pre-training datasets and the latent UNet architecture. PixArt-alpha , on the other hand, explores Diffusion Transformer (DiT) as the latent backbone, showing better training efficiency and image quality. Playground v2 , an open-source model recently developed by us, focuses on the training recipe and aesthetic quality, achieving higher user preference compared to SDXL .
Playground v2 was open-sourced in December 2023 and we were pleased to see the open-source and research community take up our work and reference it. Notably, Playground v2 has amassed over 135,000 downloads in just the last month from HuggingFace, and our work has been cited in recent papers for state-of-the-art image models such as Stable Cascade . Following Playground v2 , we chose not to change the underlying model architecture for this work; instead, we focused on analyzing and improving our training recipe and pushing the model’s aesthetic quality to a new level.
We focus on three critical issues for image models: enhancing color and contrast (sec. 2.1), improving generation across multiple aspect ratios (sec. 2.2), and improving human-centric fine details (sec. 2.3). More generally, we aim to refine the model’s capabilities to produce more realistic and visually compelling outputs. To evaluate the efficacy of our enhancements, we conducted extensive user studies and benchmarked our model against previous state-of-the-art models (sec. 3). We also propose a new automatic-evaluation benchmark MJHQ-30K (sec. 3.5) to evaluate the model’s performance against 10 unique categories. When evaluating our model’s outputs on human preference, we are thrilled to report that Playground v2.5 surpasses state-of-the-art models, including Midjourney 5.2, DALLE 3 , Playground v2 , PIXART- , and SDXL (Fig 10). See sec. 3.2 for details. Playground v2.5 endeavors to surpass the performance of its predecessor and establish itself as a leading contender in the space of text-to-image generative models.
We open-source the weights for Playground v2.5, on HuggingFacehttps://huggingface.co/playgroundai/playground-v2.5-1024px-aesthetic with a licensehttps://huggingface.co/playgroundai/playground-v2.5-1024px-aesthetic/blob/main/LICENSE.md that makes it easy for research teams to use. We will also provide extensions for using our model in A1111 and ComfyUI, two popular community tools for using diffusion models. Given how much we have benefited from the research and open-source communities, it is important that we make multiple aspects of our work on Playground v2.5 available to the community.
Methods
Latent diffusion models have struggled to generate images with high color contrast and vibrant color range since the release of SD1.5. This is a known limitation . For example, SDXL is incapable of generating a pure black image or a pure white image, and fails to place subjects onto solid backgrounds (see Fig 2 (a)).
This issue stems from the noise scheduling of the diffusion process, as pointed out by . The signal-to-noise ratio of Stable Diffusion is too high, even when the discrete noise level reaches its maximum. Several works have attempted to fix this flaw. Guttenberg and CrossLabs propose offset noise. Lin et al. propose Zero Terminal SNR to ensure the last denoising step is pure Gaussian noise. SDXL adopts the strategy of adding offset noise in the last stage of the training, as does Playground v2. However, as can be seen in Fig 2 (b), we still notice SDXL exhibits muted color and contrast.
For Playground v2.5, we aimed to dramatically improve upon this issue. We wanted to achieve vivid color and contrast in imagery and be able to produce pure-colored backgrounds. To this end, we took a more principled approach and trained our models from scratch using the EDM framework, proposed by Karras et al .
EDM brings two distinct advantages: (1) Like Zero Terminal SNR, the EDM noise schedule exhibits a near-zero signal-to-noise ratio for the final “timestep”. This removes the need for Offset Noise and fixes muted colors. (2) EDM takes a first-principles approach to designing the training and sampling processes, as well as preconditioning of the UNet. This enables the EDM authors to make clear design choices that lead to better image quality and faster model convergence.
We were also inspired by Hoogeboom et al to skew the noise schedule towards an overall noisier one when training on high-resolution images.
In Fig 3, we show a qualitative comparison between Playground v2.5 and Playground v2, the latter of which uses offset noise and a DDPM noise schedule. We see in the first column that Playground v2.5 can generate a vivid portrait with enhanced color range, and it exhibits better prompt-image alignment, which enables v2.5 to generate a pure black background.
2 Generation Across Multiple Aspect Ratios
The ability to generate images of various different aspect ratios is an important feature in real-world applications of text-to-image models. However, common pre-training procedures for these models start by training only on square images in the early stages, with random or center cropping. This technique is standard practice from conditional generative models trained on ImageNet .
Theoretically, this should not pose a problem. A diffusion model like SDXL consisting of mostly convolution layers – mimicking a Convolutional Neural Network (CNN) – should work with any input resolution at inference time, even if it was not trained on that particular resolution. This is due to the transition-invariant property of CNNs. Unfortunately, in practice, diffusion models do not generalize well to other aspect ratios when only trained on square images, as pointed out by NovelAI .
To address this challenge, NovelAI proposes doing bucketed sampling, where images with similar aspect ratios are bucketed together in the same forward pass. SDXL adopted this bucketing strategy and also added additional conditioning to specify the source and target image sizes.
SDXL’s conditioning strategy forced the model to learn to place the image’s subject at the center under different aspect ratios. However, due to an unbalanced distribution of the aspect ratio buckets in SDXL’s dataset, i.e. the majority of the dataset’s images are square, SDXL also learned the bias of certain aspect ratios in its conditioning. Furthermore, images generated at non-square aspect ratios typically exhibit much lower quality than square images.
In Playground v2.5, one of our explicit goals was to make the model reliably produce high-quality images at multiple aspect ratios, as we had learned from the users that this is a critical element for a high-quality production-grade model. While we followed a bucketing strategy similar to SDXL’s, we carefully crafted the data pipeline to ensure a more balanced bucket sampling strategy across various aspect ratios. Our strategy avoids catastrophic forgetting and helps the model not be biased towards one ratio or another.
Fig 4 and Fig 5 show a qualitative comparison between SDXL and Playground v2.5 across portrait and landscape aspect ratios, respectively. Our model can generate high-quality images under various aspect ratios without errors like generating multiple objects or wrong composition.
3 Human Preference Alignment
Humans are particularly sensitive to visual errors on human features like hands, faces, and torsos. An image with perfect lighting, composition, and style will likely be voted as low-quality if the face of the human is malformed or the structure of the body is contorted.
Generative models, both language and image, are prone to hallucination. In image models, this can manifest as malformed human features. There are multiple reasons for hallucination, but one evident explanation is a misaligned training objective: generative models are trained to maximize the log-likelihood of the data rather than maximizing human preference. In LLMs, a common strategy to align pre-trained generative models with human preference is called supervised fine-tuning, or SFT. In short , SFT fine-tunes a pre-trained base model with a small but very-high-quality dataset. This simple technique often outperforms a more complicated approach like RLHF . However, the question of how to best curate an SFT alignment dataset from different sources to maximize performance on a downstream task remains an ongoing research problem .
One of our goals with Playground v2.5 was to reduce the likelihood of visual errors in human features, which is a common critique of open-source diffusion models more broadly, as compared to closed-source models like Midjourney. Emu introduces an alignment strategy similar to SFT for text-to-image generative models. Inspired by Emu, we developed a system that enables us to automatically curate a high-quality dataset from multiple sources via user ratings. Then, we took an iterative, human-in-the-loop training approach to select the best dataset candidates. We monitored the training progress by empirical evaluation, comparing our aligned models by looking at image grids generated from a fixed set of prompts, similar to .
Our novel alignment strategy enables us to excel over SDXL in at least four important human-centric categories:
Overall lighting, color, saturation, and depth-of-field
We chose to focus on these categories based on usage patterns in our product and from user feedback.
In Fig 6, we showcase some examples of the difference in fine details between images generated using Playground v2.5 and those using SDXL. In Fig 7, we compare our model with other SoTA methods in generating human-centric images.
Evaluations
Since we ultimately build our models to be used by our hundreds of thousands of users, it is critical for us to understand their preferences on model outputs. To this end, we conduct our user studies directly within our product (see Fig 8). We believe this is the best context to gather preference metrics and provides the harshest test of whether a model actually delivers on making something valuable for an end user.
For a given user study, we choose a fixed set of prompts and sample images from two models. We then show a pair of images with the prompt to the user (without showing them which model corresponds to which image) and ask them to pick the best one according to some attribute, e.g. aesthetic preference. Because a single user’s rating is prone to bias, we show each image pair to at least 7 unique users. To further reduce bias, an image pair only "wins" for a given model if its output is preferred by at least a 2-vote margin. A 1-vote margin is considered a tie. Lastly, we involve thousands of unique users on each user study. All user studies mentioned in this report are conducted through this interface.
We conducted studies to measure overall aesthetic preference, as well as for the specific areas we aimed to improve with Playground v2.5, namely generation across multiple aspect ratios and human preference alignment.
2 Overall Aesthetic Preference against other SoTA models
We used a prompt set called Internal-1K to compare Playground v2.5 against other state-of-the-art models. Internal-1K is a prompt set collected from real user prompts on Playground.com, making it representative of real users’ prompting styles. We showed image pairs to thousands of users, specifically focusing on aesthetic preference for this study. This is the same study setup as our previous release of Playground v2 . For reference, our previous studies demonstrated that images produced from Playground v2 were favored 2.5x more than those produced by SDXL. We aimed to surpass this for Playground v2.5 and succeeded: v2.5 is favored 4.8x over SDXL.
Fig. 10 shows our results against various publicly available text-to-image models. Across the board, Playground v2.5 dramatically outperforms the current state-of-the-art open source models of SDXL and PIXART- , as well as Playground v2 . Because the performance differential between Playground v2.5 and SDXL was so large, we also tested against state-of-the-art closed-source models like DALLE 3 and Midjourney 5.2, and found that Playground v2.5 still outperforms these models in aesthetic quality.
3 Evaluation of Generation Across Multiple Aspect Ratios
We report user preference study metrics on commonly-used aspect ratios using the Internal-1K prompt set. We conducted a separate user study for each aspect ratio, ranging from 9:16 to 16:9. For a given study, we used the same aspect ratio conditioning for both models on all images. Fig 11 shows our results. Our model outperforms SDXL in all aspect ratios by a large margin.
4 Evaluation on People-centric Prompts
As discussed in Section 2.3 about improving human preference alignment, people-related prompts are an important practical use-case for commercial text-to-image models. Indeed, they are quite popular in our product. To assess our model’s ability to generate people-related images, we curated 200 high-quality people-related prompts from real user prompts in our product. We call this the People-200 prompt set. We will release this prompt set to the community for benchmarking purposes.
We conducted our user study using portrait aspect ratio 3:2, since this is the most popular choice in the community for images showing people. We compared Playground v2.5 against two commonly-used baseline models: SDXL and RealStock v2, a community fine-tune of SDXL that was trained on a realistic people dataset.
Fig 12 shows that Playground v2.5 outperforms both baselines by a large margin.
5 Automatic Evaluation Benchmark
Lastly, we introduce a new benchmark, MJHQ-30K, for automatic evaluation of a model’s aesthetic quality. The benchmark computes Fréchet Inception Distance (FID) on a high-quality dataset to gauge aesthetic quality. By spot-checking the FID and ensuring it was trending lower, we were able to quickly gauge progress throughout the different stages of pre-training and alignment.
We curated a high-quality dataset from images made on Midjourney 5.2. The dataset covers 10 common categories, and each category has 3K samples. Following common practice, we used aesthetic score and CLIP score to ensure high image quality and high text-to-image alignment.
Furthermore, we took extra care to make the images and prompts well-varied within each category.
We report both the overall FID (Table 1) and per category FID (Fig 13). All FID metrics are computed at resolution 1024x1024. Our results show that Playground v2.5 outperforms both Playground v2 and SDXL in overall FID and all category FIDs, especially in the people and fashion categories. This is in line with the results of the user study, which indicates a correlation between human preferences and the FID score of the MJHQ30K benchmark.
We release this benchmark to the public on HuggingFace https://huggingface.co/datasets/playgroundai/MJHQ-30K and encourage the community to adopt it for benchmarking their models’ aesthetic quality during pre-training and alignment.
Conclusion
In this work, we share three insights for achieving state-of-the-art aesthetic quality in text-to-image generative models, and we analyze and empirically evaluate Playground v2.5 against SoTA models in various conditions and setups. Playground v2.5 demonstrates: (1) superior performance in enhancing image color and contrast, (2) ability to generate high-quality images under various aspect ratios, and (3) alignment to human preference for aesthetic quality in generated images, especially for fine details in images of humans.
We are excited to release Playground v2.5 to the public. The model is available today to use at our product websitehttps://playground.com for all users, and we have open-sourced the weights on HuggingFacehttps://huggingface.co/playgroundai/playground-v2.5-1024px-aesthetic. Furthermore, we will soon provide extensions for using Playground v2.5 in A1111 and ComfyUI, two popular community tools for using diffusion models.
For future works, we hope to tackle improving text-to-image alignment, enhancing the model’s variation capabilities, and exploring new architectures.
At Playground, our goal is to build a unified general-purpose vision system that deeply understands pixels and enables humans of all skill levels to masterfully generate and edit pixels. We see Playground v2.5 as a stepping stone towards this vision, and we encourage the community to build with us.