Improved Distribution Matching Distillation for Fast Image Synthesis

Tianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Fredo Durand, William T. Freeman

Introduction

Diffusion models have achieved unprecedented quality in visual generation tasks . But their sampling procedure typically requires dozens of iterative denoising steps, each of which is a forward pass through a neural network. This makes high resolution text-to-image synthesis slow and expensive. To address this issue, numerous distillation methods have been developed to convert a teacher diffusion model into an efficient, few-step student generator . However, they often result in degraded quality, as the student model is typically trained with a loss to learn the pairwise noise-to-image mapping of the teacher, but struggles to perfectly mimic its behavior.

Nevertheless, it should be noted that loss functions aimed at matching distributions, such as the GAN or the DMD loss, are not burdened with the complexity of precisely learning the specific paths from noise to image because their goal is to align with the teacher model in terms of distribution—by minimizing either a Jensen-Shannon (JS) or an approximate Kullback-Leibler (KL) divergence between the student and teacher output distributions.

In particular, DMD has demonstrated state-of-the-art results in distilling Stable Diffusion 1.5, yet it remains less investigated than GAN-based methods . A likely reason is that DMD still requires an additional regression loss to ensure stable training. In turn, this necessitates creating millions of noise-image pairs by running the full sampling steps of the teacher model, which is particularly costly for text-to-image synthesis. The regression loss also negates the key benefit of DMD’s unpaired distribution matching objective, because it causes the student’s quality to be upper-bounded by the teacher’s.

In this paper, we show how to do away with DMD’s regression loss, without compromising training stability. We then push the limits of distribution matching by integrating the GAN framework into DMD, and enable few-steps sampling with a novel training procedure, which we termed ‘backward simulation’. Taken together, our contributions lead to state-of-the-art fast generative models that outperform their teacher, using as few as 4 sampling steps. Our method, which we call DMD2, achieves state-of-the-art results in one-step image generation, setting a new benchmark with FID scores of 1.28 on ImageNet-64×64 and 8.35 on zero-shot COCO 2014. We demonstrate our approach’s scalability by distilling from SDXL to produce high-quality megapixel images, establishing new standards among few-step methods.

In short, our contributions are as follows:

We propose a new distribution matching distillation technique that does not require a regression loss for stable training, thereby eliminating the need for costly data collection, and allowing for more flexible and scalable training.

We show that training instability in DMD without regression loss stems from an insufficiently trained fake diffusion critic, and implement a two time-scale update rule to address this issue.

We integrate a GAN objective into the DMD framework, where the discriminator is trained to distinguish samples from the student generator vs. real images. This additional supervision operates at the distribution level, which better aligns with DMD’s distribution-matching philosophy than the original regression loss. It mitigates approximation errors in the teacher diffusion model and enhances image quality.

While the original DMD only supports one-step students, we introduce a technique to support multi-step generators. Unlike previous multi-step distillation methods, we avoid the domain mismatch between training and inference by simulating inference-time generator inputs during training, thus improving overall performance.

Related Work

Diffusion Distillation. Recent diffusion acceleration techniques have focused on speeding up the generation process through distillation . They typically train a generator to approximate the ordinary differential equation (ODE) sampling trajectory of a teacher model, in fewer sampling steps. Notably, Luhman et al. precompute a dataset of noise and images pairs, generated by the teacher using an ODE sampler, and use it to train the student to regress the mapping in a single network evaluation. Follow-up works like Progressive Distillation eliminate the need to precompute this paired dataset offline. They iteratively train a sequence of student models, each halving the number of sampling steps of its predecessor. A complementary technique, Instaflow straightens the ODE trajectories, so they are easier to approximate with a one-step student. Consistency Distillation , and TRACT , train student models so their outputs are self-consistent at any timesteps along the ODE trajectory, and thus consistent with the teacher.

GANs. Another line of research employs adversarial training to align the student with the teacher at a broader distribution level. In ADD , the generator, initialized with weights from a diffusion model, is trained using a projected GAN objective with an image-space classifier . Building on this, LADD utilizes a pre-trained diffusion model as the discriminator and operates in latent space, thus improving scalability and enabling higher-resolution synthesis. Inspired by DiffusionGAN , UFOGen introduces noise injection prior to the real vs. fake classification in the discriminator, to smooth out the distributions, which stabilizes the training dynamics. A few recent approaches combine adversarial objectives with a distillation loss that preserves the original sampling trajectory. For instance, SDXL-Lightning integrates a DiffusionGAN loss with a progressive distillation objective , while the Consistency Trajectory Model combines a GAN with an improved consistency distillation .

Score Distillation was initially introduced in the context of text-to-3D synthesis , utilizing a pre-trained text-to-image diffusion model as a distribution matching loss. These methods optimize a 3D object by aligning rendered views with a text-conditioned image distribution, using the scores predicted by a pretrained diffusion model. Recent works have extended score distillation to diffusion distillation . Notably, DMD minimizes an approximate KL divergence, with its gradient represented as the difference between two score functions: one, fixed and pretrained, for the target distribution and another, trained dynamically, for the output distribution of the generator.

DMD parameterizes both score functions using diffusion models. This training objective proved more stable than GAN-based methods and has demonstrated superior performance in one-step image synthesis. An important caveat, DMD requires a regression loss for stability, calculated using precomputed noise-image pairs, similar to Luhman et al. . Our work does away with this requirement. We introduce techniques to stabilize the DMD training procedure without the regression regularizer, thus significantly reducing the computational costs incurred by paired data precomputation. Furthermore, we extend DMD to support multi-step generation and integrate the strengths of both GANs and distribution matching approaches , leading to state-of-the-art results in text-to-image synthesis.

Background: Diffusion and Distribution Matching Distillation

This section gives a brief overview of diffusion models and distribution matching distillation (DMD).

Diffusion Models generate images through iterative denoising. In the forward diffusion process, noise is progressively added to corrupt a sample x∼prealx\sim p_{\text{real}} from the data distribution into pure Gaussian noise over a predetermined number of steps TT, so that, at each timestep tt, the diffused samples follow the distribution preal,t(xt)=∫preal(x)q(xt∣x)dxp_{\text{real},t}(x_{t})=\int p_{\text{real}}(x)q(x_{t}|x)dx, with qt(xt∣x)∼N(αtx,σt2I)q_{t}(x_{t}|x)\sim\mathcal{N}(\alpha_{t}x,\sigma_{t}^{2}\mathbf{I}), where αt,σt>0\alpha_{t},\sigma_{t}>0 are scalars determined by the noise schedule . The diffusion model learns to iteratively reverse the corruption process by predicting a denoised estimate μ(xt,t)\mu(x_{t},t), conditioned on the current noisy sample xtx_{t} and the timestep tt, ultimately leading to an image from the data distribution prealp_{\text{real}}. After training, the denoised estimate relates to the gradient of the data likelihood function, or score function of the diffused distribution:

Sampling an image typically requires dozens to hundreds of denoising steps .

Distribution Matching Distillation (DMD) distills a many-step diffusion models into a one-step generator GG by minimizing the expectation over tt of approximate Kullback-Liebler (KL) divergences between the diffused target distribution preal,tp_{\text{real},t} and the diffused generator output distribution pfake,tp_{\text{fake},t}. Since DMD trains GG by gradient descent, it only requires the gradient of this loss, which can be computed as the difference of 2 score functions:

where z∼N(0,I)z\sim\mathcal{N}(0,\mathbf{I}) is a random Gaussian noise input, θ\theta are the generator parameters, FF is the forward diffusion process (i.e., noise injection) with noise level corresponding to time step tt, and sreals_{\text{real}} and sfakes_{\text{fake}} are scores approximated using diffusion models μreal\mu_{\text{real}} and μfake\mu_{\text{fake}} trained on their respective distributions (Eq. (1)). DMD uses a frozen pre-trained diffusion model as μreal\mu_{\text{real}} (the teacher), and dynamically updates μfake\mu_{\text{fake}} while training GG, using a denoising score-matching loss on samples from the one-step generator, i.e., fake data .

Yin et al. found that an additional regression term was needed to regularize the distribution matching gradient (Eq. (2)) and achieve high-quality one-step models. For this, they collect a dataset of noise-image pairs (z,y)(z,y) where the image yy is generated using the teacher diffusion model, and a deterministic sampler , starting from the noise map zz. Given the same input noise zz, the regression loss compares the generator output with the teacher’s prediction:

where dd is a distance function, such as LPIPS in their implementation. While gathering this data incurs negligible cost for small datasets like CIFAR-10, it becomes a significant bottleneck with large-scale text-to-image synthesis tasks, or models with complex conditioning . For instance, generating one noise-image pair for SDXL takes around 5 seconds, amounting to about 700 A100 days to cover the 12 million prompts in the LAION 6.0 dataset , as utilized by Yin et al. . This dataset construction cost alone is already more than 4×4\times our total training compute (as detailed in Appendix F). This regularization objective is also at odds with DMD’s goal of matching the student and teacher in distribution, since it encourages adherence to the teacher’s sampling paths.

Improved Distribution Matching Distillation

We revisit multiple design choices in the DMD algorithm and identify significant improvements.

The regression loss used in DMD ensures mode coverage and training stability, but as we discussed in Section 3, it makes large-scale distillation cumbersome, and is at odds with the distribution matching idea, thus inherently limiting the performance of the distilled generator to that of the teacher model. Our first improvement is to remove this loss.

2 Stabilizing pure distribution matching with a Two Time-scale Update Rule

Naively omitting the regression objective, shown in Eq. (3), from DMD leads to training instabilities and significantly degrades quality (Tab. 4). For example, we observed that the average brightness, along with other statistics, of generated samples fluctuates significantly, without converging to a stable point (See Appendix C). We attribute this instability to approximation errors in the fake diffusion model μfake\mu_{\text{fake}}, which does not track the fake score accurately, since it is dynamically optimized on the non-stationary output distribution of the generator. This causes approximation errors and biased generator gradients (as also discussed in ).

We address this using the two time-scale update rule inspired by Heusel et al. . Specifically, we train μfake\mu_{\text{fake}} and the generator GG at different frequencies to ensure that μfake\mu_{\text{fake}} accurately tracks the generator’s output distribution. We find that using 5 fake score updates per generator update, without the regression loss, provides good stability and matches the quality of the original DMD on ImageNet (Tab. 4) while achieving much faster convergence. Further analysis are included in Appendix C.

3 Surpassing the teacher model using a GAN loss and real data

Our model so far achieves comparable training stability and performance to DMD without the need for costly dataset construction (Tab. 4). However, a performance gap remains between the distilled generator and the teacher diffusion model. We hypothesize this gap could be attributed to approximation errors in the real score function μreal\mu_{\text{real}} used in DMD, which would propagate to the generator and lead to suboptimal results. Since DMD’s distilled model is never trained with real data, it cannot recover from these errors.

We address this issue by incorporating an additional GAN objective into our pipeline, where the discriminator is trained to distinguish between real images and images produced by our generator. Trained using real data, the GAN classifier does not suffer from the teacher’s limitation, potentially allowing our student generator to surpass it in sample quality. Our integration of a GAN classifier into DMD follows a minimalist design: we add a classification branch on top of the bottleneck of the fake diffusion denoiser (see Fig. 3). The classification branch and upstream encoder features in the UNet are trained by maximizing the standard non-saturing GAN objective:

where DD is the discriminator, and FF is the forward diffusion process (i.e., noise injection) defined in Section 3, with noise level corresponding to time step tt. The generator GG minimizes this objective. Our design is inspired by prior works that use diffusion models as discriminators . We note that this GAN objective is more consistent with the distribution matching philosophy since it does not require paired data, and is independent of the teacher’s sampling trajectories.

4 Multi-step generator

With the proposed improvements, we are able to match the performance of teacher diffusion models on ImageNet and COCO (see Tab. 2 and Tab. 5). However, we found that larger scale models like SDXL remain challenging to distill into a one-step generator because of limited model capacity and a complex optimization landscape to learn the direct mapping from noise to highly diverse and detailed images. This motivated us to extend DMD to support multi-step sampling.

We fix a predetermined schedule with NN timestep {t1,t2,…tN}\{t_{1},t_{2},\ldots t_{N}\}, identical during training and inference. During inference, at each step, we alternate between denoising and noise injection steps, following the consistency model , to improve sample quality. Specifically, starting from Gaussian noise z0∼N(0,I)z_{0}\sim\mathcal{N}(0,\mathbf{I}), we alternate between denoising updates x^ti=Gθ(xti,ti)\hat{x}_{t_{i}}=G_{\theta}(x_{t_{i}},t_{i}), and forward diffusion steps xti+1=αti+1x^ti+σti+1ϵx_{t_{i+1}}=\alpha_{t_{i+1}}\hat{x}_{t_{i}}+\sigma_{t_{i+1}}\epsilon with ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,\mathbf{I}), until we obtain our final image x^tN\hat{x}_{t_{N}}. Our 4-step model uses the following schedule: 999, 749, 499, 249, for a teacher model trained with 1000 steps.

5 Multi-step generator simulation to avoid training/inference mismatch

Previous multi-step generators are typically trained to denoise noisy real images . However, during inference, except for the first step, which starts from pure noise, the generator’s input come from a previous generator sampling step x^ti\hat{x}_{t_{i}}. This creates a training-inference mismatch that adversely impacts quality (Fig. 4). We address this issue by replacing the noisy real images during training, with noisy synthetic images xtix_{t_{i}} produced by the current student generator running several steps, similar to our inference pipeline (§ 4.4). This is tractable because, unlike the teacher diffusion model, our generator only runs for a few steps. Our generator then denoises these simulated images and the outputs are supervised with the proposed loss functions. Using noisy synthetic images avoids the mismatch and improves overall performance (See Sec. 5.3).

A concurrent work, Imagine Flash , proposed a similar technique. Their backward distillation algorithm shares our motivation of reducing the training and testing gap by using the student-generated images as the input to the subsequent sampling steps at training time. However, they do not entirely resolve the mismatch issue, because the teacher model of the regression loss now suffers from the training–test gap: it is never trained with synthetic images. This error is accumulated along the sampling path. In contrast, our distribution matching loss is not dependent on the input to the student model, alleviating this issue.

6 Putting everything together

In summary, our distillation method lifts DMD stringent requirements for precomputed noise–image pairs. It further integrates the strength of GANs and supports multi-step generators. As shown in Fig. 3, starting from a pretrained diffusion model, we alternate between optimizing the generator GθG_{\theta} to minimize the original distribution matching objective as well as a GAN objective, and optimizing the fake score estimator μfake\mu_{\text{fake}} using both a denoising score matching objective on the fake data, and the GAN classification loss. To ensure the fake score estimate is accurate and stable, despite being optimized on-line, we update it with higher frequency than the generator (5 steps vs. 1).

Experiments

We evaluate our approach, DMD2, using several benchmarks, including class-conditional image generation on ImageNet-64×64 , and text-to-image synthesis on COCO 2014 with various teacher models . We use the Fréchet Inception Distance (FID) to measure image quality and diversity, and the CLIP Score to evaluate text-to-image alignment. For SDXL models, we additionally report patch FID , which measures FID on 299x center-cropped patches of each image, to assess high-resolution details. Finally, we conduct human evaluations to compare our approach with other state-of-the-art methods. Comprehensive evaluations confirm that distilled models trained using our approach outperform previous work, and even rival the performance of the teacher models. Detailed training and evaluation procedures are provided in the appendix.

Table 2 compares our model with recent baselines on ImageNet-64×64. With a single forward pass, our method significantly outperforms existing distillation techniques and even outperforms the teacher model using ODE sampler . We attribute this remarkable performance to the removal of DMD’s regression loss (Sec. 4.1 and 4.2), which eliminates the performance upper bound imposed by the ODE sampler, as well as our additional GAN term (Sec. 4.3), which mitigates the adverse impact of the teacher diffusion model’s score approximation error.

2 Text-to-Image Synthesis

We evaluate DMD2’s text-to-image generation performance on zero-shot COCO 2014 . Our generators are trained by distilling SDXL and SD v1.5 , respectively, using a subset of 3 million prompts from LAION-Aesthetics . Additionally, we collect 500k images from LAION-Aesthetic as training data for the GAN discriminator. Table 2 summarizes distillation results for the SDXL model. Our 4-step generator produces high quality and diverse samples, achieving a FID score of 19.3219.32 and a CLIP score of 0.3320.332, rivaling the teacher diffusion model for both image quality and prompt coherence. To further verify our method’s effectiveness, we conduct an extensive user study comparing our model’s output with those from the teacher model and existing distillation methods. We use a subset of 128 prompts from PartiPrompts following LADD . For each comparison, we ask a random set of five evaluators to choose the image that is more visually appealing, as well as the one that better represents the text prompt. Details about the human evaluation are included in Appendix H. As shown in Figure 5, our model achieves much higher user preferences than baseline approaches. Notably, our model outperforms its teacher in image quality for 24% of samples and achieves comparable prompt alignment, while requiring 25×25\times fewer forward passes (4 vs 100). Qualitative comparisons are shown in Figure 6. Results for SDv1.5 are provided in Table 5 in Appendix A. Similarly, one-step model trained using DMD2 outperforms all previous diffusion acceleration approaches, achieving a FID score of 8.35, representing a significant 3.143.14-point improvement over the original DMD method . Our results also surpass the teacher models that uses a 50-step PNDM sampler .

3 Ablation Studies

Table 4 ablates different components of our proposed method on ImageNet. Simply removing the ODE regression loss from the original DMD results in a degraded FID of 3.48 due to training instability (see further analysis in Appendix C). However, incorporating our Two Time-scale Update Rule (TTUR, Sec. 4.2) mitigates this performance drop, matching the DMD baseline performance without requiring additional dataset construction. Adding our GAN loss achieves a further 1.1-point improvement in FID. Our integrated approach surpasses the performance of using GAN alone (without distribution matching objective), and adding the two-timescale update rule to GAN alone does not improve it, highlighting the effectiveness of combining distribution matching with GANs in a unified framework.

In Table 4, we ablate the influence of the GAN term (Sec. 4.3), distribution matching objective (Eq. 2), and backward simulation (Sec. 4.4) for distilling the SDXL model into a four-step generator. Qualitative results are shown in Figure. 7. In the absence of the GAN loss, our baseline model produces oversaturated and oversmoothed images (Fig. 7 third column). Similarly, eliminating distribution matching objective (Eq. 2) reduces our approach to a pure GAN-based method, which struggles with training stability . Moreover, pure GAN-based methods also lack a natural way to incorporate classifier-free guidance , essential for high-quality text-to-image synthesis . Consequently, while GAN-based methods achieve the lowest FID by closely matching the real distribution, they significantly underperform in text alignment and aesthetic quality ( Fig. 7 second column). Likewise, omitting the backward simulation leads to worse image quality, as indicated by the degraded patch FID score.

Limitations

While achieving superior image quality and text alignment, our distilled generator experiences a slight degradation in image diversity compared to the teacher models (see Appendix B). Additionally, our generator still requires four steps to match the quality of the largest SDXL model. These limitations, while not unique to our model, highlight areas for further improvement. Like most previous distillation methods, we use a fixed guidance scale during training, limiting user flexibility. Introducing a variable guidance scale could be a promising direction for future research. Furthermore, our methods are optimized for distribution matching; incorporating human feedback or other reward functions could further enhance performance . Lastly, training large-scale generative models is computationally intensive, making it inaccessible for most researchers. We hope our efficient approach and optimized, user-friendly codebase will help democratize future research in this field.

Broader Impact

Our work on improving the efficiency and quality of diffusion model has several potential societal impacts, both positive and negative. On the positive side, the advancements in fast image synthesis can significantly benefit various creative industries. These models can enhance graphic design, animation, and digital art by providing artists with powerful tools to generate high-quality visuals efficiently. Additionally, improved text-to-image synthesis capabilities can be used in education and entertainment, enabling the creation of personalized learning materials and immersive experiences.

However, potential negative societal impacts must be considered. Misuse risks include generating misinformation and creating fake profiles, which could spread false information and manipulate public opinion. Deploying these technologies could result in biases that unfairly impact specific groups, especially if models are trained on biased datasets, potentially perpetuating or amplifying existing societal biases. To mitigate these risks, we are interested in developing monitoring mechanisms to detect and prevent misuse and methods to enhance output diversity and fairness .

Acknowledgements

We extend our gratitude to Minguk Kang and Seungwook Kim for their assistance in setting up the human evaluation. We also thank Zeqiang Lai for suggesting the timestep shift technique used in our one-step generator. Additionally, we are grateful to our friends and colleagues for their insightful discussions and valuable comments. This work was supported by the National Science Foundation under Cooperative Agreement PHY-2019786 (The NSF AI Institute for Artificial Intelligence and Fundamental Interactions, http://iaifi.org/), by NSF Grant 2105819, by NSF CISE award 1955864, and by funding from Google, GIST, Amazon, and Quanta Computer.

References

Appendix A SD v1.5 Results

Table 5 presents detailed comparisons between our one-step generator distilled from SD v1.5 and competing approaches.

Appendix B Text-to-Image Synthesis Further Analysis

Qualitative ablation results using SDXL backbone are shown in Figure 7. Additionally, we compare the image diversity of our 4-step generator with other competing approaches distilled from SDXL . We employ an LPIPS-based diversity score, similar to that used in multi-modal image-to-image translation . Specifically, we generate four images per prompt and calculate the average pairwise LPIPS distance . For this evaluation, we use the LADD subset of PartiPrompts . We also report the FID and CLIP score measured on 10K prompts from COCO 2014 on the side. Table 6 summarizes the results. Our model achieves the best image quality, indicated by the lowest FID and Patch FID scores. We also achieve text alignment comparable to SDXL-Turbo while attaining a better diversity score. While SDXL-Lightning exhibits a higher diversity score than our approach, it suffers from considerably worse text alignment, as reflected by the lower CLIP score and human evaluation (Fig. 5). This suggests that the improved diversity is partially due to random outputs lacking prompt coherence. We note that it is possible to increase the diversity of our model by raising the weights for the GAN objective, which aligns with the more diverse unguided distribution. Further investigation into finding the optimal balance between distribution matching and the GAN objective is left for future work.

Appendix C Two Time-scale Update Rule Further Analysis

In Section 4.2, we discuss that updating the fake score multiple times (5 updates) per generator update leads to better stability. Here, we provide further analysis. Figure 8 visualizes pixel brightness variations throughout training. The baseline approach, which omits the regression objective from DMD and uses just 1 fake score update, results in significant training instability, as evidenced by periodic fluctuations in pixel brightness. In contrast, our two time-scale update rule with 5 fake score updates per generator update stabilizes the training and leads to better sample quality, as shown in Tab. 4.

We further examine the influence of the update frequency for the fake diffusion model μfake\mu_{\text{fake}} in Figure 9. An update frequency of 1 fake diffusion update per generator update corresponds to the naive baseline (red line) and suffers from training instability. Although a frequency of 10 updates (magenta line) provides excellent stability, it significantly slows down the training process. We found that a moderate frequency of 5 updates (green line) achieves the best balance between stability and convergence speed on ImageNet. Our approach proves more effective than using asynchronous learning rates (cyan line) and converges significantly faster than the original DMD method that employs a regression loss (dark blue line). For new models and datasets, we recommend adjusting the iteration number to the smallest value that ensures the stability of general image statistics, such as pixel brightness.

Appendix D Additional Text-to-Image Synthesis Results

Additional visual comparisons for the 4-step distilled models are shown in Figure 10. Sample outputs from our one-step generator are presented in Figure 11.

Appendix E ImageNet Visual Results

In Figure 12, we present qualitative results obtained from our one-step distilled model trained on the ImageNet dataset.

Appendix F Implementation Details

This section provides a brief overview of the implementation details. All results presented can be easily reproduced using our open-source training and evaluation code.

Our GAN classifier design is inspired by SDXL-Lightning . Specifically, we attach a prediction head to the middle block output of the fake diffusion model. The prediction head consists of a stack of 4×44\times 4 convolutions with a stride of 2, group normalization, and SiLU activations. All feature maps are downsampled to 4×44\times 4 resolution, followed by a single convolutional layer with a kernel size and stride of 4. This layer pools the feature maps into a single vector, which is then passed to a linear projection layer to predict the classification result.

F.2 ImageNet

Our ImageNet implementation closely follows the DMD paper . Specifically, we distill a one-step generator from the EDM pretrained model , released under the CC BY-NC-SA 4.0 License. For the standard training setup, we use the AdamW optimizer with a learning rate of 2×10−62\times 10^{-6}, a weight decay of 0.01, and beta parameters (0.9, 0.999). We use a batch size of 280 and train the model on 7 A100 GPUs for 200K iterations, which takes approximately 2 days. The number of fake diffusion model update per generator update is set to 5. The weight for the GAN loss is set to 3×10−33\times 10^{-3}. For the extended training setup shown in Table 2, we first pretrain the model without GAN loss for 400K iterations. We then resume from the best checkpoint (as measured by FID), enable the GAN loss with a weight of 3×10−33\times 10^{-3}, reduce the learning rate to 5×10−75\times 10^{-7}, and continue training for an additional 150K iterations. The total training time for this run is approximately 5 days.

F.3 SD v1.5

We distill a one-step generator from the SD v1.5 model , released under the CreativeML Open RAIL-M license, using prompts from the LAION-Aesthetic 6.25+ dataset . Additionally, we collect 500K images from LAION-Aesthetic 5.5+ as training data for the GAN discriminator, filtering out images smaller than 1024×10241024\times 1024 and those containing unsafe content. Our training process involves two stages. In the first stage, we disable the GAN loss and use the AdamW optimizer with a learning rate of 1×10−51\times 10^{-5}, a weight decay of 0.01, and beta parameters of (0.9, 0.999). The fake diffusion model is updated 10 times per generator update. We set the guidance scale for the real diffusion model to be 1.75. We use a batch size of 2048 and train the model on 64 A100 GPUs for 40K iterations. In the second stage, we enable the GAN loss with a weight of 10−310^{-3}, reduce the learning rate to 5×10−75\times 10^{-7}, and continue training for an additional 5K iterations. The total training time is approximately 26 hours.

F.4 SDXL

We train both one-step and four-step generators by distilling from the SDXL model , released under the CreativeML Open RAIL++-M License. For the one-step generator, we observed similar block noise artifacts as reported in SDXL-Lightning and Pixart-Sigma . We addressed this by adopting the timestep shift technique from OpenDMD and Pixart-Sigma , setting the conditioning timestep to 399. Additionally, we initialized the one-step generator by pretraining it with a regression loss using a small set of 10K pairs for a short period. These adjustments are not necessary for the multi-step model or other backbones, suggesting this issue might be specific to SDXL. Similar to SD v1.5, we use prompts from the LAION-Aesthetic 6.25+ dataset and collect 500K images from LAION-Aesthetic 5.5+ as training data for the GAN discriminator, filtering out images smaller than 1024×10241024\times 1024 and those containing unsafe content. The generator is trained using the AdamW optimizer with a learning rate of 5×10−75\times 10^{-7}, a weight decay of 0.01, and beta parameters of (0.9, 0.999). The fake diffusion model is updated 5 times per generator update. We set the guidance scale for the real diffusion model to be 8. We use a batch size of 128 and train the model on 64 A100 GPUs for 20K iterations for the 4-step generator and 25K iterations for the 1-step generator, taking approximately 60 hours.

Appendix G Evaluation Details

For the COCO experiments, we follow the exact evaluation setup as GigaGAN and DMD . For the results presented in Table 5, we use 30K prompts from the COCO 2014 validation set and generate the corresponding images. The outputs are downsampled to 256×\times256 and compared with 40,504 real images from the same validation set using clean-FID . For the results presented in Table 2, we use a random set of 10K prompts from the COCO 2014 validation set and generate the corresponding images. The outputs are downsampled to 512×\times512 and compared with the corresponding 10K real images from the validation set with the same prompts. We compute the CLIP score using the OpenCLIP-G backbone. For the ImageNet results, we generate 50,000 images and calculate the FID statistics using EDM’s evaluation code .

Appendix H User Study Details

To conduct the human preference study, we use the Prolific platform (https://www.prolific.com). We use 128 prompts from the LADD subset of PartiPrompts . All approaches generate corresponding images, which are presented in pairs to human evaluators to measure aesthetic and prompt alignment preference. The specific questions and interface are shown in Figure 13. Consent is obtained from the voluntary participants, who are compensated at a flat rate of 12 dollars per hour. We manually verify that all generated images contain standard visual content that poses no risks to the study participants.

Appendix I Prompts for Figure 1, Figure 2, and Figure 11

We use the following prompts for Figure 1. From left to right, top to bottom:

A photo of an astronaut riding a horse in the forest.

a giant gorilla at the top of the Empire State Building

A close-up photo of a wombat wearing a red backpack and raising both arms in the air. Mount Rushmore is in the background.

An oil painting of two rabbits in the style of American Gothic, wearing the same clothes as in the original.

A sloth in a go kart on a race track. The sloth is holding a banana in one hand. There is a banana peel on the track in the background.

a teddy bear on a skateboard in times square

We use the following prompts for Figure 2. From left to right, top to bottom:

A television made of water that displays an image of a cityscape at night.

a portrait of a statue of the Egyptian god Anubis wearing aviator goggles, white t-shirt and leather jacket. The city of Los Angeles is in the background.

a capybara made of voxels sitting in a field

Cinematic photo of a beautiful girl riding a dinosaur in a jungle with mud, sunny day shiny clear sky. 35mm photograph, film, professional, 4k, highly detailed.

A still image of a humanoid cat posing with a hat and jacket in a bar.

A soft beam of light shines down on an armored granite wombat warrior statue holding a broad sword. The statue stands an ornate pedestal in the cella of a temple. wide-angle lens. anime oil painting.

A photograph of the inside of a subway train. There are red pandas sitting on the seats. One of them is reading a newspaper. The window shows the jungle in the background.

A close-up of a woman’s face, lit by the soft glow of a neon sign in a dimly lit, retro diner, hinting at a narrative of longing and nostalgia.

We use the following prompts for Figure 11. From left to right, top to bottom:

A close-up of a woman’s face, lit by the soft glow of a neon sign in a dimly lit, retro diner, hinting at a narrative of longing and nostalgia.

A television made of water that displays an image of a cityscape at night.

a portrait of a statue of the Egyptian god Anubis wearing aviator goggles, white t-shirt and leather jacket. The city of Los Angeles is in the background.

a capybara made of voxels sitting in a field

A soft beam of light shines down on an armored granite wombat warrior statue holding a broad sword. The statue stands an ornate pedestal in the cella of a temple. wide-angle lens. anime oil painting.

An oil painting of two rabbits in the style of American Gothic, wearing the same clothes as in the original.

A still image of a humanoid cat posing with a hat and jacket in a bar.