SDXL-Lightning: Progressive Adversarial Diffusion Distillation

Shanchuan Lin, Anran Wang, Xiao Yang

Introduction

Diffusion models are a rising class of generative models that has achieved state-of-the-art results in a wide range of applications, such as text-to-image , text-to-video , and image-to-video , etc. However, the iterative generation process of diffusion models is slow and computationally expansive. How to generate high-quality samples faster is an actively researched area and is the main focus of our work.

Conceptually, the generation involves a probability flow that gradually transports samples between the data and the noise probability distribution. The diffusion model learns to predict the gradient at any location of this flow. The generation is simply transporting samples from the noise distribution to the data distribution by following the predicted gradient through the flow. Because the flow is complex and curved, the generation must take a small step at a time. Formally, the flow can be expressed as an ordinary differential equation (ODE) . In practice, generating a high-quality data sample requires more than 50 inference steps.

Different approaches to reduce the number of inference steps have been researched. Prior works have proposed better ODE solvers to account for the curving nature of the flow . Others have proposed formulations to make the flow straighter . Nonetheless, these approaches generally still require more than 20 inference steps.

Model distillation , on the other hand, can achieve high-quality samples under 10 inference steps. Instead of predicting the gradient at the current flow location, it changes the model to directly predict the next flow location much farther ahead. Existing methods can achieve good results using 4 or 8 inference steps, but the quality is still not production-acceptable using 1 or 2 inference steps. Our method falls under the model distillation umbrella and achieves much superior quality compared to existing methods.

Our method combines the best of both worlds from progressive and adversarial distillation . Progressive distillation ensures that the distilled model follows the same probability flow and has the same mode coverage as the original model. However, progressive distillation with mean squared error (MSE) loss produces blurry results under 8 inference steps and we provide theoretical analysis in our paper. To mitigate the issue, we use adversarial loss at every stage of the distillation to strike a balance between quality and mode coverage. Progressive distillation also brings an additional benefit, i.e., for multi-step sampling, our model predicts the next location on the ODE trajectory instead of jumping to the ODE trajectory endpoints every time by other distillation approaches . This better preserves the original model behavior and facilitates better compatibility with LoRA modules and control plugins .

Furthermore, our paper proposes innovative discriminator design, loss objectives, and stable training techniques. Specifically, we use the pre-trained diffusion UNet encoder as the discriminator backbone and fully operate in latent space. We propose two adversarial loss objectives to trade off sample quality and mode coverage. We investigate the implication of diffusion schedules and output formulation. We discuss techniques to stabilize the adversarial training.

Our distillation method produces new state-of-the-art SDXL models that support one-step/few-step generation at 1024px resolution. We open-source our distilled models as SDXL-Lightning.

Background

The forward diffusion process gradually transforms samples from the data distribution to the Gaussian noise distribution. Given a data sample x0x_{0}, noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,\mathbf{I}) and time t∼U(1,T)t\sim\mathcal{U}(1,T). The forward function is defined as the following, with αˉt\bar{\alpha}_{t} as the manually defined schedule :

Many prior works formulate the network to perform noise prediction , i.e. ϵ^=f(xt,t,c)\hat{\epsilon}=f(x_{t},t,c). We can use the conversion function x\mathbf{x} to convert the prediction to x^0\hat{x}_{0} space:

Alternatively, the network can be formulated to perform data sample prediction , i.e. x^0=f(xt,t,c)\hat{x}_{0}=f(x_{t},t,c). We can use the conversion function ϵ\boldsymbol{\epsilon} to convert the prediction to ϵ^\hat{\epsilon} space:

Regardless of the formulation, the network in essence predicts the gradient u^t\hat{u}_{t}. Given the gradient utu_{t} at any location xtx_{t}, we can move samples along the flow:

The generation process is simply moving sample xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\mathbf{I}) from t=Tt=T to t=0t=0 a small step at a time.

2 Latent Diffusion Model

Instead of directly generating samples at the data space, latent diffusion models (LDMs) propose to first train a Variational Autoencoder (VAE) that encodes the data to a more compact latent space. Diffusion models are trained to generate the latent codes, which are passed through the VAE decoder to generate the final data sample.

Latent diffusion models are widely adopted for high-resolution image and video generation due to their computational efficiency. SDXL is the state-of-the-art text-to-image generation model that can generate 1024px resolution images from 128px latent space.

3 Progressive Distillation

Progressive distillation trains the student to predict directions pointing to the next flow location as if the teacher has performed multiple steps.

Specifically, given data x0,cx_{0},c from the dataset, and noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,\mathbf{I}), we jump to arbitrary timestep tt:

Once the student model converges, it is used as the teacher model and the distillation process repeats. In theory, it can produce one-step generation models, but in practice, models produce blurry results. We analyze this issue in Section 3.1.

4 Adversarial Distillation

Adversarial training involves a minimax optimization between a discriminator network that aims to identify generated samples from real samples and a generator network that aims to fool the discriminator. It was originally proposed as Generative Adversarial Networks (GANs) , a standalone class of generative networks, but it suffers from issues such as mode collapse and instability. Recent studies have found that the adversarial objective can be incorporated in diffusion training and distillation .

SDXL-Turbo is the latest and the most popular open-source model using adversarial diffusion distillation. It follows prior works to use a pre-trained image encoder DINOv2 as the discriminator backbone to accelerate training. However, this brings several limitations. First, using an off-the-shelf vision encoder means it must operate in the pixel space instead of the latent space, which significantly increases computation, memory consumption, and training time, making high-resolution distillation impractical. This is likely the reason SDXL-Turbo only supports up to 512px resolution. Second, an off-the-shelf vision encoder only works at t=0t=0. The distilled model has to be trained to jump to ODE trajectory endpoints x0x_{0}, but since the quality for one-step inference is not good enough, random noises are added again for multi-step inference. This way of multi-step inference significantly alters the model behavior, making it less compatible with existing LoRA modules and control plugins . Third, off-the-shelf encoders may be hard to find for other datasets (anime, line arts, etc.) and modalities (video, audio, etc.). This reduces the generalizability of the distillation method. Lastly, the adversarial objective alone does not force the model to follow the same probability flow, so mode coverage is not enforced.

Our method uses the diffusion model’s U-Net encoder as the discriminator backbone. This allows us to efficiently distill in the latent space for high-resolution models, supports discrimination at all timesteps, and is generalizable to all datasets and modalities. Our method also allows control over the trade-off between quality and mode coverage, as later discussed in Sections 3.2 and 3.4.

5 Other Distillation Methods

We briefly discuss the advantages of our approach compared to other distillation methods.

Consistency Model (CM) also requires jumping to the ODE trajectory endpoints at every inference step. This causes large model behavior changes for multi-step sampling which reduces compatibility with LoRA modules and plugins. This method has been applied to SDXL but its generation quality is poor under 8 steps. Consistency Trajectory Model (CTM) adds adversarial loss and supports jumping to arbitrary flow locations, but the adversarial training is applied post-distillation, instead of during the distillation, and the method has not been applied to large-scale text-to-image models.

Rectified Flow (RF) straightens the flow by repeatedly training with deterministic data and noise pairs. However, its few-step generation quality is still poor. Additionally, since the model has only seen specific data and noise pairs during the distillation, it no longer supports data pairing with arbitrary noise. This impacts the ability for image editing such as SDEdit .

Score Distillation Sampling (SDS) has been used in SDXL-Turbo to stabilize adversarial training, yet its effect is minimal and it cannot be used as a distillation method alone. Variation Score Distillation (VSD) has recently been used in diffusion distillation . However, it requires training an additional score model of the negative distribution during the distillation process, and like the discriminator in adversarial training, it also involves a dynamic training target that can negatively affect training stability. There is no open-source model for comparison, and our preliminary experiments find our method achieves better quality.

6 LoRA

Low-Rank Adaptation (LoRA) is an efficient finetuning technique. It only trains a small number of additional parameters to the model and has become particularly popular for training stylization modules for existing text-to-image models.

LCM-LoRA is the first to show that model distillation can also be trained as a LoRA module. This ensures minimum parameter changes and can be conveniently plugged into the existing ecosystem.

Our work is inspired by this approach and we provide our distilled models both as LoRAs for convenient plug and play and as full models for even better quality.

Method

The learned probability flow is determined by the dataset, the forward function , the loss function , and the model capacity. Given finite training samples, the underlying data distribution is ambiguous. The maximum likelihood estimation (MLE) is a distribution that assigns even probability only to the observed samples and zero everywhere else. If the model has infinite capacity, it will learn a flow of this maximum likelihood estimation and overfit to always produce observed samples and generate no new data. In practice, diffusion models can generate new data because neural networks are not exact learners.

When the model is used in multi-step generations, it is stacked and has a higher Lipschitz constant and more non-linearities to approximate a more complex distribution. But when the model is used in few-step generations, it no longer has the same amount of capacity to approximate well the same distribution. This is evidenced by diffusion models can have very sharp changes in results despite small changes in the initial noises , but the distilled models have much smoother latent traversal. This explains why distillation with MSE loss produces blurry results. The student model simply does not have the capacity to match the teacher.

Additionally, neural network parameter optimization involves a complex landscape. Even models with the same capacity can hardly match output exactly since parameters can get stuck at different local minima.

We find that other distance metrics, e.g. L1 and perceptual loss , also produce undesirable results. On the other hand, we find adversarial objectives to be effective in mitigating this issue.

2 Adversarial Objective

We use non-saturated adversarial loss and train the discriminator and the student model in alternating steps. This encourages the student prediction x^t−ns\hat{x}_{t-ns} to be closer to the teacher prediction xt−nsx_{t-ns}:

The condition on xtx_{t} is important for preserving the probability flow. This is because the teacher’s generation of xt−nsx_{t-ns} is deterministic from xtx_{t}. By providing the discriminator both xt−nsx_{t-ns} and xtx_{t}, the discriminator learns the underlying probability flow and the student must also follow the same flow to fool the discriminator.

Our formulation is very similar to a prior work except we use it for distillation instead of training from scratch. Note that this approach only preserves the probability flow and ensures mode coverage when used in distillation.

3 Discriminator Design

A prior work has shown that a pre-trained diffusion model’s U-Net encoder can be used as a vision backbone. Such a pre-trained backbone is very suitable for our discriminator because it has been pre-trained on the target dataset, directly operates in the latent space, supports noised input at all timesteps, and supports text condition.

We follow the approach and copy the encoder and midblock of the pre-trained SDXL model as our discriminator backbone dd. We pass xt−nsx_{t-ns} and xtx_{t} independently through the shared backbone dd, concatenate the hidden features after the midblock in the channel dimension, and pass it to a prediction head. The prediction head consists of simple blocks of 4×44\times 4 convolution with a stride of 2, group normalization with 32 groups, and SiLU activation layers to further reduce the spatial dimension. The output is projected to a single value and clamped to $rangewithsigmoidrange with sigmoid\sigma(\cdot).Togethertheyformthecompletediscriminator. Together they form the complete discriminatorD$:

Note that the backbone is initialized with the pre-trained weights and we train the entire discriminator without freezing the backbone. We find our training stable without the need for expansive R1 regularization nor switching to L2 attention . Additional stabilization techniques are discussed in Section 3.7.

4 Relax the Mode Coverage

The adversarial objective above encourages the prediction to be both sharp and flow-preserving, but this does not change the fact that the student does not have enough capacity to perfectly match the teacher as discussed in Section 3.1. With the MSE objective, it manifests blurry results. With the adversarial objective, it manifests the “Janus” artifacts.

As shown in Figure 2, the teacher model can sometimes generate drastic layout changes for adjacent noise inputs, but the student model does not have the same capacity to make such sharp changes. As a result, the adversarial loss sacrifices semantic correctness in need to preserve the sharpness and the layout, manifesting artifacts that feature conjoined heads and bodies.

Semantic correctness is more important than mode coverage by human preference. Therefore, after training with the original adversarial objective, we relax the flow preservation requirement. Specifically, we further finetune the model without the condition on xtx_{t}:

We find that finetuning with this objective is effective in removing the “Janus” artifacts while still preserving the original flow to a great extent in practice. Therefore, at every stage of the progressive distillation, we first train with the conditional objective and then finetune with this unconditional objective. Since the unconditional objective only concerns per-sample quality, we use the skip-level teacher for distilling the one-step and two-step models to further retain quality and mitigate error accumulation.

5 Fix the Schedule

A prior work has shown that common diffusion schedules are flawed. Specifically, the schedule does not reach pure noise at t=Tt=T during training, yet pure noise is given during inference, causing a discrepancy. Unfortunately, SDXL uses this flawed schedule. The effect is less obvious under a large number of inference steps but is particularly detrimental for few-step generations.

A hacky way to circumvent the problem is to hard swap pure noise ϵ\epsilon as model input at t=Tt=T during training. This way the model is trained to expect pure noise as input at t=Tt=T and we still use Equation 3 with the old αˉ\bar{\alpha} schedule at inference to avoid singularity. It incurs minimum changes to the sampling procedure with existing software ecosystems . This approach is also used by SDXL-Turbo .

6 Distillation Procedure

First, we perform distillation from 128 steps directly to 32 steps with MSE loss. We find MSE is sufficient for the early stage. We also apply classifier-free guidance (CFG) only in this stage. We use a guidance scale of 6 without any negative prompts.

Then, we switch to using adversarial loss to distill the step count in this order: 32→8→4→2→132\rightarrow 8\rightarrow 4\rightarrow 2\rightarrow 1. At each stage, we first train with the conditional objective as in Section 3.2 to preserve the probability flow, and then train with the unconditional objective as in Section 3.4 to relax the mode coverage.

At each stage, we first train with LoRA using the two objectives, then we merge the LoRA and train the whole UNet further with the unconditional objective. We find finetuning the whole UNet can achieve even better performance, while the LoRA module can be used on other base models. Our LoRA settings are the same as LCM-LoRA , which uses rank 64 on all the convolution and linear weights except the input and output convolutions and the shared time embedding linear layers. We do not use LoRA on the discriminator. We re-initialize the discriminator at each stage.

We distill our models on a subset of LAION and COYO dataset. We select images to be greater than 1024px and LAION images with aesthetic scores above 5.5. We additionally filter images by sharpness using a Laplacian filter and clean up the text prompts. The distillation is conducted on a square aspect ratio, but we find it generalizes well to other aspect ratios at inference time.

We use batch size 512 across 64 A100 80G GPUs. For the first 128→32128\rightarrow 32 stage with MSE loss, we use learning rate 1e-5 with Adam β1=0.9,β2=0.999\beta_{1}=0.9,\beta_{2}=0.999. For the remaining stages with adversarial loss, we use learning rates 1e-6 with LoRA and 5e-7 without LoRA for both the student and the discriminator networks. The Adam optimizer uses β1=0,β2=0.99\beta_{1}=0,\beta_{2}=0.99 following prior works without weight decay . We use gradient accumulation, VAE slicing, BF16 mixed precision , flash attention , and zero redundancy optimizer to reduce the memory footprint.

7 Stable Training Techniques

For one-step and two-step distillations, we employ additional techniques to stabilize the training.

While we only need to train the one-step model at timestep {1000}, and the two-step model at timesteps {500, 1000} for complete image generation, we find training on more timesteps {250, 500, 750, 1000} improves stability. As an additional benefit, this allows our models to support SDEdit at different timesteps as illustrated below:

7.2 Train Discriminator at Multiple Timesteps

We find training the one-step model using the above discriminator formulation very unstable. We find the root reason is that our discriminator uses the pre-trained diffusion UNet encoder as the backbone, yet the diffusion encoder is trained to only focus on high-frequency details at lower timesteps and low-frequency structures at higher timesteps. For one-step generations, the student network directly predicts x^0\hat{x}_{0}. If we pass x^0\hat{x}_{0} and t=0t=0 to the discriminator backbone, it is not able to critique the image structure. This leads to images with bad shapes and even divergence.

Our solution is to add noise to both teacher-predicted x0x_{0} and student-predicted x^0\hat{x}_{0} to timesteps: {10, 250, 500, 750} randomly. This way the discriminator can critique the prediction on both high-frequency details and low-frequency structures.

Specifically, we first draw t∗←{10,250,500,750}t*\leftarrow\{10,250,500,750\} with uniform weighting 1:1:1:1 and sample a new noise ϵ∗∼N(0,I)\epsilon*\sim\mathcal{N}(0,\mathbf{I}). Then we apply the noise before passing it through the conditional and unconditional discriminators:

After the model is trained stable, we change the timesteps weighting to 5:1:1:1. This further improves details and removes noisy artifacts.

Note that this stabilization technique can also be viewed from the lens of bridging the distribution gap , discriminator augmentation , and multi-scale discriminator .

We find the one-step model with ϵ\epsilon-prediction formulation tends to generate noise artifacts likely due to numerical instability. We change the one-step model to x0x_{0}-prediction and it resolves the issue.

Specifically, we copy the network and convert the predicted ϵ^\hat{\epsilon} to x^0\hat{x}_{0} through conversion function x\mathbf{x} defined in Equation 3. We use MSE to gradually guide the online model to x0x_{0}-prediction.

As discussed in Section 3.1, MSE loss cannot convert our model perfectly. The converted model generates blurry results, but this will be fixed by the adversarial objectives.

After the conversion, the one-step model is trained with adversarial objectives in x0x_{0}-prediction formulation, while the teacher model still operates in ϵ\epsilon-prediction formulation. Due to the substantial formulation change, we do not provide LoRA for one-step generation.

Evaluation

Table 1 shows the specification of our distilled models compared to others.

2 Qualitative Comparison

Figure 3 compares our method against other open-source distillation models: SDXL-Turbo and LCM . Our method is substantially better in overall quality and details. Our method is also substantially better in the preservation of the style and layout of the original model. Furthermore, we find our 4-step and 8-step model can often outperform the original SDXL model for 32 steps. This is because our progressive distillation starts all the way from 128 steps.

Figure 4 compares our LoRA models against the fully trained models. We find that fully trained models have better structures and details. This is less noticeable on 8-step models, but more observable on 2-step models.

3 Quantitative Comparison

Table 2 shows Fréchet Inception Distance (FID) and CLIP score . Following the convention, we generate images using the first 10K prompts from the COCO validation dataset. FID metric is computed against the corresponding ground truth images from COCO.

FID is normally computed by resizing the whole image to 299px for the InceptionV3 network . This only assesses the high-level sample quality and diversity. The metric shows that our model achieves similar performance as other distillation techniques. All distillation methods have worse FID compared to the original SDXL likely due to the reduction in diversity.

We additionally propose to calculate FID on patches of images to assess high-resolution details. Specifically, we calculate FID on the 299px center-cropped patch of every image. For Turbo, we resize the 512px to 1024px before the crop for a fair comparison. The metric shows that our models have much better high-resolution details compared to other methods. Additionally, the metric shows that our model has better high-resolution details compared to original SDXL models for 32 steps because our distillation starts from 128 steps. It also shows that the quality degrades as the number of inference steps decreases.

CLIP score shows that our method achieves similar text-alignment performance compared to other methods.

Ablation

Figure 5 shows that our distillation LoRA model can be applied to different base models. Specifically, we test it on third-party cartoon , anime , and realistic base models. Our distillation LoRAs are able to keep the style and layout of the new base model to a great extent.

2 Inference with Different Aspect Ratios

Figure 6 shows that are models can mostly retain the ability to infer at different resolutions and aspect ratios despite the distillation is only performed on square images. However, we do notice an increasing amount of bad cases when performing 1-step and 2-step generations. This can be improved by distilling with multiple aspect ratios, which we leave for future improvements.

3 Compatibility with ControlNet

Figure 7 shows that our models are compatible with ControlNet . We test it on the canny edge and depth ControlNet. We observe that our models follow the condition correctly, with some quality degradation as the number of inference steps decreases.

Limitation

Unlike other methods having a single distilled checkpoint that supports multiple inference step settings, our method produces separate checkpoints for each corresponding inference step setting. This is usually not an issue in production when the number of inference steps is fixed. In case the number of inference steps must be flexible, our LoRA modules can mitigate the checkpoint switching issue.

Our method produces distilled student models with the same architecture as the teacher model. However, we believe that the UNet architecture is not optimal for one-step generation. We inspect the feature maps at each UNet layer and find that most of the generation is carried out by the decoder. We leave this problem to future improvements.

Conclusion

To sum up, we have presented SDXL-Lightning, our state-of-the-art one-step/few-step text-to-image generative models resulting from our novel progressive adversarial diffusion distillation method. In our evaluation, we have found that our models produce superior image quality compared to prior works. We are open-sourcing our models to advance the research in generative AI.

References