Adversarial Diffusion Distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, Robin Rombach
Introduction
Diffusion models (DMs) have taken a central role in the field of generative modeling and have recently enabled remarkable advances in high-quality image- and video- synthesis. One of the key strengths of DMs is their scalability and iterative nature, which allows them to handle complex tasks such as image synthesis from free-form text prompts. However, the iterative inference process in DMs requires a significant number of sampling steps, which currently hinders their real-time application. Generative Adversarial Networks (GANs) , on the other hand, are characterized by their single-step formulation and inherent speed. But despite attempts to scale to large datasets, GANs often fall short of DMs in terms of sample quality. The aim of this work is to combine the superior sample quality of DMs with the inherent speed of GANs.
Our approach is conceptually simple: We propose Adversarial Diffusion Distillation (ADD), a general approach that reduces the number of inference steps of a pre-trained diffusion model to 1–4 sampling steps while maintaining high sampling fidelity and potentially further improving the overall performance of the model. To this end, we introduce a combination of two training objectives: (i) an adversarial loss and (ii) a distillation loss that corresponds to score distillation sampling (SDS) . The adversarial loss forces the model to directly generate samples that lie on the manifold of real images at each forward pass, avoiding blurriness and other artifacts typically observed in other distillation methods . The distillation loss uses another pretrained (and fixed) DM as a teacher to effectively utilize the extensive knowledge of the pretrained DM and preserve the strong compositionality observed in large DMs. During inference, our approach does not use classifier-free guidance , further reducing memory requirements. We retain the model’s ability to improve results through iterative refinement, which is an advantage over previous one-step GAN-based approaches .
Our contributions can be summarized as follows:
We introduce ADD, a method for turning pretrained diffusion models into high-fidelity, real-time image generators using only 1–4 sampling steps.
Our method uses a novel combination of adversarial training and score distillation, for which we carefully ablate several design choices.
ADD significantly outperforms strong baselines such as LCM, LCM-XL and single-step GANs , and is able to handle complex image compositions while maintaining high image realism at only a single inference step.
Using four sampling steps, ADD-XL outperforms its teacher model SDXL-Base at a resolution of px.
Background
While diffusion models achieve remarkable performance in synthesizing and editing high-resolution images and videos , their iterative nature hinders real-time application. Latent diffusion models attempt to solve this problem by representing images in a more computationally feasible latent space , but they still rely on the iterative application of large models with billions of parameters.
In addition to utilizing faster samplers for diffusion models , there is a growing body of research on model distillation such as progressive distillation and guidance distillation . These approaches reduce the number of iterative sampling steps to 4-8, but may significantly lower the original performance. Furthermore, they require an iterative training process. Consistency models address the latter issue by enforcing a consistency regularization on the ODE trajectory and demonstrate strong performance for pixel-based models in the few-shot setting. LCMs focus on distilling latent diffusion models and achieve impressive performance at 4 sampling steps. Recently, LCM-LoRA introduced a low-rank adaptation training for efficiently learning LCM modules, which can be plugged into different checkpoints for SD and SDXL . InstaFlow propose to use Rectified Flows to facilitate a better distillation process.
All of these methods share common flaws: samples synthesized in four steps often look blurry and exhibit noticeable artifacts. At fewer sampling steps, this problem is further amplified. GANs can also be trained as standalone single-step models for text-to-image synthesis . Their sampling speed is impressive, yet the performance lags behind diffusion-based models. In part, this can be attributed to the finely balanced GAN-specific architectures necessary for stable training of the adversarial objective. Scaling these models and integrating advances in neural network architectures without disturbing the balance is notoriously challenging. Additionally, current state-of-the-art text-to-image GANs do not have a method like classifier-free guidance available which is crucial for DMs at scale.
Score Distillation Sampling also known as Score Jacobian Chaining is a recently proposed method that has been developed to distill the knowledge of foundational T2I Models into 3D synthesis models. While the majority of SDS-based works use SDS in the context of per-scene optimization for 3D objects, the approach has also been applied to text-to-3D-video-synthesis and in the context of image editing .
Recently, the authors of have shown a strong relationship between score-based models and GANs and propose Score GANs, which are trained using score-based diffusion flows from a DM instead of a discriminator. Similarly, Diff-Instruct , a method which generalizes SDS, enables to distill a pretrained diffusion model into a generator without discriminator.
Conversely, there are also approaches which aim to improve the diffusion process using adversarial training. For faster sampling, Denoising Diffusion GANs are introduced as a method to enable sampling with few steps. To improve quality, a discriminator loss is added to the score matching objective in Adversarial Score Matching and the consistency objective of CTM .
Our method combines adversarial training and score distillation in a hybrid objective to address the issues in current top performing few-step generative models.
Method
Our goal is to generate high-fidelity samples in as few sampling steps as possible, while matching the quality of state-of-the-art models . The adversarial objective naturally lends itself to fast generation as it trains a model that outputs samples on the image manifold in a single forward step. However, attempts at scaling GANs to large datasets observed that is critical to not solely rely on the discriminator, but also employ a pretrained classifier or CLIP network for improving text alignment. As remarked in , overly utilizing discriminative networks introduces artifacts and image quality suffers. Instead, we utilize the gradient of a pretrained diffusion model via a score distillation objective to improve text alignment and sample quality. Furthermore, instead of training from scratch, we initialize our model with pretrained diffusion model weights; pretraining the generator network is known to significantly improve training with an adversarial loss . Lastly, instead of utilizing a decoder-only architecture used for GAN training , we adapt a standard diffusion model framework. This setup naturally enables iterative refinement.
While we formulate our method in pixel space, it is straightforward to adapt it to LDMs operating in latent space. When using LDMs with a shared latent space for teacher and student, the distillation loss can be computed in pixel or latent space. We compute the distillation loss in pixel space as this yields more stable gradients when distilling latent diffusion model .
2 Adversarial Loss
For the discriminator, we follow the proposed design and training procedure in which we briefly summarize; for details, we refer the reader to the original work. We use a frozen pretrained feature network and a set of trainable lightweight discriminator heads . For the feature network , Sauer et al. find vision transformers (ViTs) to work well, and we ablate different choice for the ViTs objective and model size in Section 4. The trainable discriminator heads are applied on features at different layers of the feature network.
whereas the discriminator is trained to minimize
where denotes the R1 gradient penalty . Rather than computing the gradient penalty with respect to the pixel values, we compute it on the input of each discriminator head . We find that the penalty is particularly beneficial when training at output resolutions larger than px.
3 Score Distillation Loss
The distillation loss in Eq. (1) is formulated as
where sg denotes the stop-gradient operation. Intuitively, the loss uses a distance metric to measure the mismatch between generated samples by the ADD-student and the DM-teacher’s outputs averaged over timesteps and noise . Notably, the teacher is not directly applied on generations of the ADD-student but instead on diffused outputs , as non-diffused inputs would be out-of-distribution for the teacher model .
Experiments
For our experiments, we train two models of different capacities, ADD-M (860M parameters) and ADD-XL (3.1B parameters). For ablating ADD-M, we use a Stable Diffusion (SD) 2.1 backbone , and for fair comparisons with other baselines, we use SD1.5. ADD-XL utilizes a SDXL backbone. All experiments are conducted at a standardized resolution of 512x512 pixels; outputs from models generating higher resolutions are down-sampled to this size.
Our training setup opens up a number of design spaces regarding the adversarial loss, distillation loss, initialization, and loss interplay. We conduct an ablation study on several choices in Table 1; key insights are highlighted below each table. We will discuss each experiment in the following.
Discriminator feature networks. (Table LABEL:tab:abl0). Recent insights by Stein et al. suggest that ViTs trained with the CLIP or DINO objectives are particularly well-suited for evaluating the performance of generative models. Similarly, these models also seem effective as discriminator feature networks, with DINOv2 emerging as the best choice.
Student pretraining. (Table LABEL:tab:abl2). Our experiments demonstrate the importance of pretraining the ADD-student. Being able to use pretrained generators is a significant advantage over pure GAN approaches. A problem of GANs is the lack of scalability; both Sauer et al. and Kang et al. observe a saturation of performance after a certain network capacity is reached. This observation contrasts the generally smooth scaling laws of DMs . However, ADD can effectively leverage larger pretrained DMs (see Table LABEL:tab:abl2) and benefit from stable DM pretraining.
Loss terms. (Table LABEL:tab:abl3). We find that both losses are essential. The distillation loss on its own is not effective, but when combined with the adversarial loss, there is a noticeable improvement in results. Different weighting schedules lead to different behaviours, the exponential schedule tends to yield more diverse samples, as indicated by lower FID, SDS and NFSD schedules improve quality and text alignment. While we use the exponential schedule as the default setting in all other ablations, we opt for the NFSD weighting for training our final model. Choosing an optimal weighting function presents an opportunity for improvement. Alternatively, scheduling the distillation weights over training, as explored in the 3D generative modeling literature could be considered.
Teacher type. (Table LABEL:tab:abl4). Interestingly, a bigger student and teacher does not necessarily result in better FID and CS. Rather, the student adopts the teachers characteristics. SDXL obtains generally higher FID, possibly because of its less diverse output, yet it exhibits higher image quality and text alignment .
Teacher steps. (Table LABEL:tab:abl5). While our distillation loss formulation allows taking several consecutive steps with the teacher by construction, we find that several steps do not conclusively result in better performance.
2 Quantitative Comparison to State-of-the-Art
For our main comparison with other approaches, we refrain from using automated metrics, as user preference studies are more reliable . In the study, we aim to assess both prompt adherence and the overall image. As a performance measure, we compute win percentages for pairwise comparisons and ELO scores when comparing several approaches. For the reported ELO scores we calculate the mean scores between both prompt following and image quality. Details on the ELO score computation and the study parameters are listed in the supplementary material.
Fig. 5 and Fig. 6 present the study results. The most important results are: First, ADD-XL outperforms LCM-XL (4 steps) with a single step. Second, ADD-XL can beat SDXL (50 steps) with four steps in the majority of comparisons. This makes ADD-XL the state-of-the-art in both the single and the multiple steps setting. Fig. 7 visualizes ELO scores relative to inference speed. Lastly, Table 2 compares different few-step sampling and distillation methods using the same base model. ADD outperforms all other approaches including the standard DPM solver with eight steps.
3 Qualitative Results
To complement our quantitative studies above, we present qualitative results in this section. To paint a more complete picture, we provide additional samples and qualitative comparisons in the supplementary material. Fig. 3 compares ADD-XL (1 step) against the best current baselines in the few-steps regime. Fig. 4 illustrates the iterative sampling process of ADD-XL. These results showcase our model’s ability to improve upon an initial sample. Such iterative improvement represents another significant benefit over pure GAN approaches like StyleGAN-T++. Lastly, Fig. 8 compares ADD-XL directly with its teacher model SDXL-Base. As indicated by the user studies in Section 4.2, ADD-XL outperforms its teacher in both quality and prompt alignment. The enhanced realism comes at the cost of slightly decreased sample diversity.
Discussion
This work introduces Adversarial Diffusion Distillation, a general method for distilling a pretrained diffusion model into a fast, few-step image generation model. We combine an adversarial and a score distillation objective to distill the public Stable Diffusion and SDXL models, leveraging both real data through the discriminator and structural understanding through the diffusion teacher. Our approach performs particularly well in the ultra-fast sampling regime of one or two steps, and our analyses demonstrate that it outperforms all concurrent methods in this regime. Furthermore, we retain the ability to refine samples using multiple steps. In fact, using four sampling steps, our model outperforms widely used multi-step generators such as SDXL, IF, and OpenMUSE.
Our model enables the generation of high quality images in a single-step, opening up new possibilities for real-time generation with foundation models.
Acknowledgements
We would like to thank Jonas Müller for feedback on the draft, the proof, and typesetting; Patrick Esser for feedback on the proof and building an early model demo; Frederic Boesel for generating data and helpful discussions; Minguk Kang and Taesung Park for providing GigaGAN samples; Richard Vencu, Harry Saini, and Sami Kama for maintaining the compute infrastructure; Yara Wald for creative sampling support; and Vanessa Sauer for her general support.
References
Appendix
Appendix A SDS As a Special Case of the Distillation Loss
If we set the weighting function to where is the scaling factor from the weighted diffusion loss as in and choose , the distillation loss in Eq. (4) is equivalent to the score distillation objective:
Appendix B Details on Human Preference Assessment
For the evaluation results presented in Figures 5, 6 and 7, we employ human evaluation and do not rely on commonly used metrics for quality assessment of generative models such as FID and CLIP-score , since these have been shown to capture more fine grained aspects like aesthetics and scene composition only insufficiently . However these categories in particular have become more and more important when comparing current state-of-the-art text-to-image models. We evaluate all models based on 100 selected prompts from the PartiPrompts benchmark with the most relevant categories (excluding prompts from the category basic). More details on how the study was conducted Section B.1 and the rankings computed Section B.2 are listed below.
Given all models for one particular study (e.g. ADD-XL, OpenMUSEhttps://huggingface.co/openMUSE, IF-XLhttps://github.com/deep-floyd/IF, SDXL and LCM-XLhttps://huggingface.co/latent-consistency/lcm-lora-sdxl in Figure 7) we compare each prompt for each pair of models (1v1). For every comparison, we collect an average of four votes per task from different annotators, for both visual quality and prompt following. Human evaluators, recruited from the platform Prolifichttps://app.prolific.com with English as their first language, are shown two images from different models based on the same text prompt. To prevent biases, evaluators are restricted from participating in more than one of our studies. For the prompt following task, we display the text prompt above the two images and ask, “Which image looks more representative of the text shown above and faithfully follows it?” For the visual quality assessment, we do not show the prompt and instead ask, “Which image is of higher quality and aesthetically more pleasing?”. Performing a complete assessment between all pair-wise comparisons gives us robust and reliable signals on model performance trends and the effect of varying thresholds. The order of prompts and the order between models are fully randomized. Frequent attention checks are in place to ensure data quality.
B.2 ELO Score Calculation
To calculate rankings when comparing more than two models based on 1v1 comparisons we use ELO Scores (higher-is-better) which were originally proposed as a scoring method for chess players but have more recently also been applied to compare instruction-tuned generative LLMs . For a set of competing players with initial ratings participating in a series of zero-sum games the ELO rating system updates the ratings of the two players involved in a particular game based on the expected and and actual outcome of that game. Before the game with two players with ratings and , the expected outcome for the two players are calculated as
After observing the result of the game, the ratings are updated via the rule
where indicates the outcome of the match for player . In our case we have if player wins and if player looses. The constant can be see as weight putting emphasis on more recent games. We choose and bootstrap the final ELO ranking for a given series of comparisons based on 1000 individual ELO ranking calculations with randomly shuffled order. Before comparing the models we choose the start rating for every model as .
Appendix C GAN Baselines Comparison
For training our state-of-the-art GAN baseline StyleGAN-T++, we follow the training procedure outlined in . The main differences are extended training (2M iterations with a batch size of 2048, which is comparable to GigaGAN’s schedule ), the improved discriminator architecture proposed in Section 3.2, and R1 penalty applied at each discriminator head.
Fig. 11 shows that StyleGAN-T++ outperforms the previous best GANs by achieving a comparable zero-shot FID to GigaGAN at a significantly higher CLIP score. Here, we do not compare to DMs, as comparisons between model classes via automatic metrics tend to be less informative . As an example, GigaGAN achieves FID and CLIP scores comparable to SD1.5, but its sample quality is still inferior, as noted by the authors.
Appendix D Additional Samples
We show additional one-step samples as in Figure 1 in Figure 12. An additional qualitative comparison as in Figure 4 which demonstrates that our model can further refine quality by using more than one sampling step is provided in Figure 14, where we show that, while sampling quality with a single step is already high, more steps can give higher diversity and better spelling capabilities. Lastly, we provide an additional qualitative comparison of ADD-XL to other state-of-the-art one and few-step models in Figure 13.