One-step Diffusion with Distribution Matching Distillation

Tianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman, Fredo Durand, William T. Freeman, Taesung Park

Introduction

Diffusion models have revolutionized image generation, achieving unprecedented levels of realism and diversity with a stable training procedure. In contrast to GANs and VAEs , however, their sampling is a slow, iterative process that transforms a Gaussian noise sample into an intricate image by progressive denoising . This typically requires tens to hundreds of costly neural network evaluations, limiting interactivity in using the generation pipeline as a creative tool.

To accelerate sampling speed, previous methods distill the noise→\rightarrowimage mapping, discovered by the original multi-step diffusion sampling, into a single-pass student network. However, fitting such a high-dimensional, complex mapping is certainly a demanding task. A challenge is the expensive cost of running the full denoising trajectory, just to realize one loss computation of the student model. Recent methods mitigate this by progressively increasing the sampling distance of the student, without running the full denoising sequence of the original diffusion . However, the performance of distilled models still lags behind the original multi-step diffusion model.

In contrast, rather than enforcing correspondences between noise and diffusion-generated images, we simply enforce that the student generations look indistinguishable from the original diffusion model. At high level, our goal shares motivation with other distribution-matching generative models, such as GMMN or GANs . Still, despite their impressive success in creating realistic images , scaling up the model on the general text-to-image data has been challenging . In this work, we bypass the issue by starting with a diffusion model that is already trained on large-scale text-to-image data. Concretely, we finetune the pretrained diffusion model to learn not only the data distribution, but also the fake distribution that is being produced by our distilled generator. Since diffusion models are known to approximate the score functions on diffused distributions , we can interpret the denoised diffusion outputs as gradient directions for making an image “more realistic”, or if the diffusion model is learned on the fake images, “more fake”. Finally, the gradient update rule for the generator is concocted as the difference of the two, nudging the synthetic images toward higher realism and lower fakeness. Previous work , in a method called Variational Score Distillation, shows that modeling the real and fake distributions with a pretrained diffusion model is also effective for test-time optimization of 3D objects. Our insight is that a similar approach can instead train an entire generative model.

Furthermore, we find that pre-computing a modest number of the multi-step diffusion sampling outcomes and enforcing a simple regression loss with respect to our one-step generation serves as an effective regularizer in the presence of the distribution matching loss. Our method draws upon inspiration and insights from VSD , GANs , and pix2pix , showing that by (1) modeling real and fake distributions with diffusion models and (2) using a simple regression loss to match the multi-step diffusion outputs, we can train a one-step generative model with high fidelity.

We evaluate models trained with our Distribution Matching Distillation procedure (DMD) across various tasks, including image generation on CIFAR-10 and ImageNet 64×\times64 , and zero-shot text-to-image generation on MS COCO 512×\times512 . On all benchmarks, our one-step generator significantly outperforms all published few-steps diffusion methods, such as Progressive Distillation , Rectified Flow , and Consistency Models . On ImageNet, DMD reaches FIDs of 2.62, an improvement of 2.4×2.4\times over Consistency Model . Employing the identical denoiser architecture as Stable Diffusion , DMD achieves a competitive FID of 11.49 on MS-COCO 2014-30k. Our quantitative and qualitative evaluations show that the images generated by our model closely resemble the quality of those generated by the costly Stable Diffusion model. Importantly, our approach maintains this level of image fidelity while achieving a 100×100\times reduction in neural network evaluations. This efficiency allows DMD to generate 512×512512\times 512 images at a rate of 20 FPS when utilizing FP16 inference, opening up a wide range of possibilities for interactive applications.

Related Work

Diffusion models have emerged as a powerful generative modeling framework, achieving unparalleled success in diverse domains such as image generation , audio synthesis , and video generation . These models operate by progressively transforming noise into coherent structures through a reverse diffusion process . Despite state-of-the-art results, the inherently iterative procedure of diffusion models entails a high and often prohibitive computational cost for real-time applications. Our work builds upon leading diffusion models and introduces a simple distillation pipeline that reduces the multi-step generative process to a single forward pass. Our method is universally applicable to any diffusion model with deterministic sampling .

Diffusion Acceleration

Accelerating the inference process of diffusion models has been a key focus in the field, leading to the development of two types of approaches. The first type advances fast diffusion samplers , which can dramatically reduce the number of sampling steps required by pre-trained diffusion models—from a thousand down to merely 20-50. However, a further reduction in steps often results in a catastrophic decrease in performance. Alternatively, diffusion distillation has emerged as a promising avenue for further boosting speed . They frame diffusion distillation as knowledge distillation , where a student model is trained to distill the multi-step outputs of the original diffusion model into a single step. Luhman et al. and DSNO proposed a simple approach of pre-computing the denoising trajectories and training the student model with a regression loss in pixel space. However, a significant challenge is the expensive cost of running the full denoising trajectory for each realization of the loss function. To address this issue, Progressive Distillation (PD) train a series of student models that halve the number of sampling steps of the previous model. InstaFlow progressively learn straighter flows on which the one step prediction maintains accuracy over a larger distance. Consistency Distillation (CD) , TRACT , and BOOT train a student model to match its own output at a different timestep on the ODE flow, which in turn is enforced to match its own output at yet another timestep. In contrast, our method shows that the simple approach of Luhman et al. and DSNO to pre-compute the diffusion outputs is sufficient, once we introduce distribution matching as the training objective.

Distribution Matching

Recently, a few classes of generative models have shown success in scaling up to complex datasets by recovering samples that are corrupted by a predefined mechanism, such as noise injection or token masking . On the other hand, there exist generative methods that do not rely on sample reconstruction as the training objective. Instead, they match the synthetic and target samples at a distribution level, such as GMMD or GANs . Among them, GANs have shown unprecedented quality in realism , particularly when the GAN loss can be combined with task-specific, auxiliary regression losses to mitigate training instability, ranging from paired image translation to unpaired image editing . Still, GANs are a less popular choice for text-guided synthesis, as careful architectural design is needed to ensure training stability at large scale .

Lately, several works drew connections between score-based models and distribution matching. In particular, ProlificDreamer introduced Variational Score Distillation (VSD), which leverages a pretrained text-to-image diffusion model as a distribution matching loss. Since VSD can utilize a large pretrained model for unpaired settings, it showed impressive results at particle-based optimization for text-conditioned 3D synthesis. Our method refines and extends VSD for training a deep generative neural network for distilling diffusion models. Furthermore, motivated by the success of GANs in image translation, we complement the stability of training with a regression loss. As a result, our method successfully attains high realism on a complex dataset like LAION . Our method is different from recent works that combine GANs with diffusion , as our formulation is not grounded in GANs. Our method shares motivation with concurrent works that leverage the VSD objective to train a generator, but differs in that we specialize the method for diffusion distillation by introducing regression loss and showing state-of-the-art results for text-to-image tasks.

Distribution Matching Distillation

Our goal is to distill a given pretrained diffusion denoiser, the base model, μbase\mu_{\text{base}}, into a fast “one-step” image generator, GθG_{\theta}, that produces high-quality images without the costly iterative sampling procedure (Sec. 3.1). While we wish to produce samples from the same distribution, we do not necessarily seek to reproduce the exact mapping.

By analogy with GANs, we denote the outputs of the distilled model as fake, as opposed to the real images from the training distribution. We illustrate our approach in Figure 2. We train the fast generator by minimizing the sum of two losses: a distribution matching objective (Sec. 3.2), whose gradient update can be expressed as the difference of two score functions, and a regression loss (Sec. 3.3) that encourages the generator to match the large scale structure of the base model’s output on a fixed dataset of noise-image pairs. Crucially, we use two diffusion denoisers to model the score functions of the real and fake distributions, respectively, perturbed with Gaussian noise of various magnitudes. Finally, in Section 3.4, we show how to adapt our training procedure with classifier-free guidance.

Our distillation procedure assumes a pretrained diffusion model μbase\mu_{\text{base}} is given. Diffusion models are trained to reverse a Gaussian diffusion process that progressively adds noise to a sample from a real data distribution x0∼prealx_{0}\sim p_{\text{real}}, turning it into white noise xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\mathbf{I}) over TT time steps ; we use T=1000T=1000. We denote the diffusion model as μbase(xt,t)\mu_{\text{base}}(x_{t},t). Starting from a Gaussian sample xTx_{T}, the model iteratively denoises a running noisy estimate xtx_{t}, conditioned on the timestep t∈{0,1,...,T−1}t\in\{0,1,...,T-1\} (or noise level), to produce a sample of the target data distribution. Diffusion models typically require 10 to 100s steps to produce realistic images. Our derivation uses the mean-prediction form of diffusion for simplicity but works identically with ϵ\epsilon-prediction with a change of variable (see Appendix H). Our implementation uses pretrained models from EDM and Stable Diffusion .

Our one-step generator GθG_{\theta} has the architecture of the base diffusion denoiser but without time-conditioning. We initialize its parameters θ\theta with the base model, i.e., Gθ(z)=μbase(z,T−1),∀zG_{\theta}(z)=\mu_{\text{base}}(z,T-1),\forall z, before training.

2 Distribution Matching Loss

Ideally, we would like our fast generator to produce samples that are indistinguishable from real images. Inspired by the ProlificDreamer , we minimize the Kullback–Leibler (KL) divergence between the real and fake image distributions, prealp_{\text{real}} and pfakep_{\text{fake}}, respectively:

Computing the probability densities to estimate this loss is generally intractable, but we only need the gradient with respect to θ\theta to train our generator by gradient descent.

Gradient update using approximate scores. Taking the gradient of Eq. (1) with respect to the generator parameters:

where sreal(x)=∇xlog preal(x)s_{\text{real}}(x)=\nabla_{x}\text{log}~{}p_{\text{real}}(x), sfake(x)=∇xlog pfake(x)s_{\text{fake}}(x)=\nabla_{x}\text{log}~{}p_{\text{fake}}(x) are the scores of the respective distributions. Intuitively, sreals_{\text{real}} moves xx toward the modes of prealp_{\text{real}}, and −sfake-s_{\text{fake}} spreads them apart, as shown in Figure 3(a, b). Computing this gradient is still challenging for two reasons: first, the scores diverge for samples with low probability — in particular prealp_{\text{real}} vanishes for fake samples, and second, our intended tool for estimating score, namely the diffusion models, only provide scores of the diffused distribution. Score-SDE provides an answer to these two issues.

By perturbing the data distribution with random Gaussian noise of varying standard deviations, we create a family of “blurred” distributions that are fully-supported over the ambient space, and therefore overlap, so that the gradient in Eq. (2) is well-defined (Figure 4). Score-SDE then shows that a trained diffusion model approximates the score function of the diffused distribution.

Accordingly, our strategy is to use a pair of diffusion denoisers to model the scores of the real and fake distributions after Gaussian diffusion. With slight abuse of notation, we define these as sreal(xt,t)s_{\text{real}}(x_{t},t) and sfake(xt,t)s_{\text{fake}}(x_{t},t), respectively. Diffused sample xt∼q(xt∣x)x_{t}\sim q(x_{t}|x) is obtained by adding noise to generator output x=Gθ(z)x=G_{\theta}(z) at diffusion time step tt:

where αt\alpha_{t} and σt\sigma_{t} are from the diffusion noise schedule.

Real score. The real distribution is fixed, corresponding to the training images of the base diffusion model, so we model its score using a fixed copy of the pretrained diffusion model μbase(x,t)\mu_{\text{base}}(x,t). The score given a diffusion model is given by Song et al. :

Dynamically-learned fake score. We derive the fake score function, in the same manner as the real score case:

However, as the distribution of our generated samples changes throughout training, we dynamically adjust the fake diffusion model μfakeϕ\mu_{\text{fake}}^{\phi} to track these changes. We initialize the fake diffusion model from the pretrained diffusion model μbase\mu_{\text{base}}, updating parameters ϕ\phi during training, by minimizing a standard denoising objective :

where Ldenoiseϕ\mathcal{L}_{\text{denoise}}^{\phi} is weighted according to the diffusion timestep tt, using the same weighting strategy employed during the training of the base diffusion model .

Distribution matching gradient update. Our final approximate distribution matching gradient is obtained by replacing the exact score in Eq. (2) with those defined by the two diffusion models on the perturbed samples xtx_{t} and taking the expectation over the diffusion time steps:

where z∼N(0;I)z\sim\mathcal{N}(0;\mathbf{I}), x=Gθ(z)x=G_{\theta}(z), t∼U(Tmin,Tmax)t\sim\mathcal{U}(T_{\text{min}},T_{\text{max}}), and xt∼qt(xt∣x)x_{t}\sim q_{t}(x_{t}|x). We include the derivations in Appendix F. Here, wtw_{t} is a time-dependent scalar weight we add to improve the training dynamics. We design the weighting factor to normalize the gradient’s magnitude across different noise levels. Specifically, we compute the mean absolute error across spatial and channel dimensions between the denoised image and the input, setting

where SS is the number of spatial locations and CC is the number of channels. In Sec. 4.2, we show that this weighting outperforms previous designs . We set Tmin=0.02TT_{\text{min}}=0.02\hskip 1.42262ptT and Tmax=0.98TT_{\text{max}}=0.98\hskip 1.42262ptT, following DreamFusion .

3 Regression loss and final objective

The distribution matching objective introduced in the previous section is well-defined for t≫0t\gg 0, i.e., when the generated samples are corrupted with a large amount of noise. However, for a small amount of noise, sreal(xt,t)s_{\text{real}}(x_{t},t) often becomes unreliable, as preal(xt,t)p_{\text{real}}(x_{t},t) goes to zero. Furthermore, as the score ∇xlog(p)\nabla_{x}\text{log}(p) is invariant to scaling of probability density function pp, the optimization is susceptible to mode collapse/dropping, where the fake distribution assigns higher overall density to a subset of the modes. To avoid this, we use an additional regression loss to ensure all modes are preserved; see Figure 3(b), (c).

This loss measures the pointwise distance between the generator and base diffusion model outputs, given the same input noise. Concretely, we build a paired dataset D={z,y}\mathcal{D}=\{z,y\} of random Gaussian noise images zz and the corresponding outputs yy, obtained by sampling the pretrained diffusion model μbase\mu_{\text{base}} using a deterministic ODE solver . In our CIFAR-10 and ImageNet experiments, we utilize the Heun solver from EDM , with 18 steps for CIFAR-10 and 256 steps for ImageNet. For the LAION experiments, we use the PNDM solver with 50 sampling steps. We find that even a small number of noise–image pairs, generated using less than 1% of the training compute, in the case of CIFAR10, for example, acts as an effective regularizer. Our regression loss is given by:

Final objective. Network μfakeϕ\mu_{\text{fake}}^{\phi} is trained with Ldenoiseϕ\mathcal{L}_{\text{denoise}}^{\phi}, which is used to help calculate ∇θDKL\nabla_{\theta}D_{KL}.

For training GθG_{\theta}, the final objective is DKL+λregLregD_{KL}+\lambda_{\text{reg}}\mathcal{L}_{\text{reg}}, using λreg=0.25\lambda_{\text{reg}}=0.25 unless otherwise specified. The gradient ∇θDKL\nabla_{\theta}D_{KL} is computed in Eq. (7), and gradient ∇θLreg\nabla_{\theta}\mathcal{L}_{\text{reg}} is computed from Eq. (9) with automatic differentiation. We apply the two losses to distinct data streams: unpaired fake samples for the distribution matching gradient and paired examples described in Section 3.3 for the regression loss. Additional training details are provided in Appendix B.

4 Distillation with classifier-free guidance

Classifier-Free Guidance is widely used to improve the image quality of text-to-image diffusion models. Our approach also applies to diffusion models that use classifier-free guidance. We first generate the corresponding noise-output pairs by sampling from the guided model to construct the paired dataset needed for regression loss Lreg\mathcal{L}_{\text{reg}}. When computing the distribution matching gradient ∇θDKL\nabla_{\theta}D_{KL}, we substitute the real score with that derived from the mean prediction of the guided model. Meanwhile, we do not modify the formulation for the fake score. We train our one-step generator with a fixed guidance scale, following InstaFlow and LCM-LoRA .

Experiments

We assess the capabilities of our approach using several benchmarks, including class-conditional generation on CIFAR-10 and ImageNet . We use the Fréchet Inception Distance (FID) to measure image quality and CLIP Score to evaluate text-to-image alignment. First, we perform a direct comparison on ImageNet (Sec. 4.1), where our distribution matching distillation substantially outperforms competing distillation methods with identical base diffusion models. Second, we perform detailed ablation studies verifying the effectiveness of our proposed modules (Sec. 4.2). Third, we train a text-to-image model on the LAION-Aesthetic-6.25+ dataset with a classifier-free guidance scale of 3 (Sec. 4.3). In this phase, we distill Stable Diffusion v1.5, and we show that our distilled model achieves FID comparable to the original model, while offering a 30×\times speed-up. Finally, we train another text-to-image model on LAION-Aesthetic-6+, utilizing a higher guidance value of 8 (Sec. 4.3). This model is tailored to enhance visual quality rather than optimize the FID metric. Quantitative and qualitative analysis confirm that models trained with our distribution matching distillation procedure can produce high-quality images rivaling Stable Diffusion. We describe additional training and evaluation details in the appendix.

We train our model on class-conditional ImageNet-64×64 and benchmark its performance with competing methods. Results are shown in Table 1. Our model surpasses established GANs like BigGAN-deep and recent diffusion distillation methods, including the Consistency Model and TRACT . Our method remarkably bridges the fidelity gap, achieving a near-identical FID score (within 0.3) compared to the original diffusion model, while also attaining a 512-fold increase in speed. On CIFAR-10, our class-conditional model reaches a competitive FID of 2.66. We include the CIFAR-10 results in the appendix.

2 Ablation Studies

We first compare our method with two baselines: one omitting the distribution matching objective and the other missing the regression loss in our framework. Table 2 (left) summarizes the results. In the absence of distribution matching loss, our baseline model produces images that lack realism and structural integrity, as illustrated in the top section of Figure 5. Likewise, omitting the regression loss leads to training instability and a propensity for mode collapse, resulting in a reduced diversity of the generated images. This issue is illustrated in the bottom section of Figure 5.

Table 2 (right) demonstrates the advantage of our proposed sample weighting strategy (Section 4). We compare with σt/αt\sigma_{t}/\alpha_{t} and σt3/αt\sigma_{t}^{3}/\alpha_{t}, two popular weighting schemes utilized by DreamFusion and ProlificDreamer . Our weighting strategy achieves a healthy 0.9 FID improvement as it normalizes the gradient magnitudes across noise levels and stabilizes the optimization.

3 Text-to-Image Generation

We use zero-shot MS COCO to evaluate our model’s performance for text-to-image generation. We train a text-to-image model by distilling Stable Diffusion v1.5 on the LAION-Aesthetics-6.25+ . We use a guidance scale of 3, which yields the best FID for the base Stable Diffusion model. The training takes around 36 hours on a cluster of 72 A100 GPUs. Table 3 compares our model to state-of-the-art approaches. Our method showcases superior performance over StyleGAN-T , surpasses all other diffusion acceleration methods, including advanced diffusion solvers , and diffusion distillation techniques such as Latent Consistency Models , UFOGen , and InstaFlow . We substantially close the gap between distilled and base models, reaching within 2.72.7 FID from Stable Diffusion v1.5, while running approximately 30×\times faster. With FP16 inference, our model generates images at 20 frames per second, enabling interactive applications.

For text-to-image generation, diffusion models typically operate with a high guidance scale to enhance image quality . To evaluate our distillation method in this high guidance-scale regime, we trained an additional text-to-image model. This model distills SD v1.5 using a guidance scale of 8 on the LAION-Aesthetics-6+ dataset . Table 4 benchmarks our approach against various diffusion acceleration methods . Similar to the low guidance model, our one-step generator significantly outperforms competing methods, even when they utilize a four-step sampling process. Qualitative comparisons with competing approaches and the base diffusion model are shown in Figure 6.

Limitations

While our results are promising, a slight quality discrepancy persists between our one-step model and finer discretizations of the diffusion sampling path, such as those with 100 or 1000 neural network evaluations. Although the impact on quality is minimal in light of the substantial speed gains achieved, this indicates there is room for further enhancements in efficient one-step generative models. Additionally, the performance of our models is inherently limited by the capabilities of the teacher model. In line with the limitations observed in Stable Diffusion v1.5, our one-step generator struggles with rendering legible text and detailed depictions of small faces and people. We anticipate that distilling more advanced models, like SDXL , could yield significant improvements. Our model employs a fixed guidance scale during training. Introducing a variable guidance scale option, akin to the conditional guidance approach used in Guided-Distillation , could provide users with greater flexibility in image generation during inference.

Acknowledgements

This work was started while TY was an intern at Adobe Research. We are grateful for insightful discussions with Yilun Xu, Guangxuan Xiao, and Minguk Kang. This work is supported by NSF grants 2105819, 1955864, and 2019786 (IAIFI), by the Singapore DSTA under DST00OECI20300823 (New Representations for Vision), as well as by funding from GIST and Amazon.

References

Appendix A Qualitative Speed Comparison

In the accompanying video material, we present a qualitative speed comparison between our one-step generator and the original stable diffusion model. Our one-step generator achieves comparable image quality with the Stable Diffusion model while being around 30×30\times faster.

Appendix B Implementation Details

Algorithm 1 details our distribution matching distillation training procedure. For a comprehensive understanding, we include the implementation specifics for constructing the KL loss for the generator GG in Algorithm 2 and training the fake score estimator parameterized by μfake\mu_{\text{fake}} in Algorithm 3.

We distill our one-step generator from EDM pretrained models, specifically utilizing “edm-cifar10-32x32-cond-vp” for class-conditional training and “edm-cifar10-32x32-uncond-vp” for unconditional training. We use σmin=0.002\sigma_{\text{min}}=0.002 and σmax=80\sigma_{\text{max}}=80 and discretize the noise schedules into 1000 binshttps://github.com/openai/consistency_models/blob/main/cm/karras_diffusion.py#L422. To create our distillation dataset, we generate 100,000 noise-image pairs for class-conditional training and 500,000 for unconditional training. This process utilizes the deterministic Heun sampler (with Schurn=0S_{\text{churn}}=0) over 18 steps . For the training phase, we use the AdamW optimizer , setting the learning rate at 5e-5, weight decay to 0.01, and beta parameters to (0.9, 0.999). The model training is conducted across 7 GPUs, achieving a total batch size of 392. Concurrently, we sample an equivalent number of noise-image pairs from the distillation dataset to calculate the regression loss. Following Song et al. , we incorporate the LPIPS loss using a VGG backbone from the PIQ library . Prior to input into the LPIPS network, images are upscaled to a resolution of 224×224 using bilinear upsampling. The regression loss is weighted at 0.25 (λreg=0.25\lambda_{\text{reg}}=0.25) for class-conditional training and at 0.5 (λreg=0.5\lambda_{\text{reg}}=0.5) for unconditional training. The weights for the distribution matching loss and fake score denoising loss are both set to 1. We train the model for 300,000 iterations and use a gradient clipping with a L2 norm of 10.

B.2 ImageNet-64×\times64

We distill our one-step generator from EDM pretrained models, specifically utilizing “edm-imagenet-64x64-cond-adm” for class-conditional training. We use a σmin=0.002\sigma_{\text{min}}=0.002 and σmax=80\sigma_{\text{max}}=80 and discretize the noise schedules into 1000 bins. Initially, we prepare a distillation dataset by generating 25,000 noise-image pairs using the deterministic Heun sampler (with Schurn=0S_{\text{churn}}=0) over 256 steps . For the training phase, we use the AdamW optimizer , setting the learning rate at 2e-6, weight decay to 0.01, and beta parameters to (0.9, 0.999). The model training is conducted across 7 GPUs, achieving a total batch size of 336. Concurrently, we sample an equivalent number of noise-image pairs from the distillation dataset to calculate the regression loss. Following Song et al. , we incorporate the LPIPS loss using a VGG backbone from the PIQ library . Prior to input into the LPIPS network, images are upscaled to a resolution of 224×\times224 using bilinear upsampling. The regression loss is weighted at 0.25 (λreg=0.25\lambda_{\text{reg}}=0.25), and the weights for the distribution matching loss and fake score denoising loss are both set to 1. We train the models for 350,000 iterations. We use mixed-precision training and a gradient clipping with a L2 norm of 10.

B.3 LAION-Aesthetic 6.25+

We distill our one-step generator from Stable Diffusion v1.5 . We use the LAION-Aesthetic 6.25+ dataset, which contains around 3 million images. Initially, we prepare a distillation dataset by generating 500,000 noise-image pairs using the deterministic PNMS sampler over 50 steps with a guidance scale of 3. Each pair corresponds to one of the first 500,000 prompts of LAION-Aesthetic 6.25+. For the training phase, we use the AdamW optimizer , setting the learning rate at 1e-5, weight decay to 0.01, and beta parameters to (0.9, 0.999). The model training is conducted across 72 GPUs, achieving a total batch size of 2304. Simultaneously, noise-image pairs from the distillation dataset are sampled to compute the regression loss, with a total batch size of 1152. Given the memory-intensive nature of decoding generated latents into images using the VAE for regression loss computation, we opt for a smaller VAE network for decoding. Following Song et al. , we incorporate the LPIPS loss using a VGG backbone from the PIQ library . The regression loss is weighted at 0.25 (λreg=0.25\lambda_{\text{reg}}=0.25), and the weights for the distribution matching loss and fake score denoising loss are both set to 1. We train the model for 20,000 iterations. To optimize GPU memory usage, we implement gradient checkpointing and mixed-precision training. We also apply a gradient clipping with a L2 norm of 10.

B.4 LAION-Aesthetic 6+

We distill our one-step generator from Stable Diffusion v1.5 . We use the LAION-Aesthetic 6+ dataset, comprising approximately 12 million images. To prepare the distillation dataset, we generate 12,000,000 noise-image pairs using the deterministic PNMS sampler over 50 steps with a guidance scale of 8. Each pair corresponds to a prompt from the LAION-Aesthetic 6+ dataset. For training, we utilize the AdamW optimizer , setting the learning rate at 1e-5, weight decay to 0.01, and beta parameters to (0.9, 0.999). To optimize GPU memory usage, we implement gradient checkpointing and mixed-precision training. We also apply a gradient clipping with a L2 norm of 10. The training takes two weeks on approximately 80 A100 GPUs. During this period, we made adjustments to the distillation dataset size, the regression loss weight, the type of VAE decoder, and the maximum timestep for the distribution matching loss computation. A comprehensive training log is provided in Table 5. We note that this training schedule, constrained by time and computational resources, may not be the most efficient or optimal.

Appendix C Baseline Details

This baseline adheres to the training settings outlined in Sections B.1 and B.2, with the distribution matching loss omitted.

C.2 w/o Regression Loss Baseline

Following the training protocols from Sections B.1 and B.2, this baseline excludes the regression loss. To prevent training divergence, the learning rate is adjusted to 1e-5.

C.3 Text-to-Image Baselines

We benchmark our approach against a variety of models, including the base diffusion model , fast diffusion solvers , and few-step diffusion distillation baselines .

Stable Diffusion We employ the StableDiffusion v1.5 model available on huggingfacehttps://huggingface.co/runwayml/stable-diffusion-v1-5, generating images with the PNMS sampler over 50 steps.

Fast Diffusion Solvers We use the UniPC and DPMSolver++ implementations from the diffusers library , with all hyperparameters set to default values.

LCM-LoRA We use the LCM-LoRA SDv1.5 checkpoints hosted on Hugging Facehttps://huggingface.co/latent-consistency/lcm-lora-sdv1-5. As the model is pre-trained with guidance, we do not apply classifier-free guidance during inference.

Appendix D Evaluation Details

For zero-shot evaluation on COCO, we employ the evaluation code from GigaGAN https://github.com/mingukkang/GigaGAN/tree/main/evaluation. Specifically, we generate 30,000 images using random prompts from the MS-COCO2014 validation set. We downsample the generated images from 512×\times512 to 256×\times256 using the PIL.Lanczos resizer. These images are then compared with 40,504 real images from the same validation set to calculate the FID metric using the clean-fid library. Additionally, we employ the OpenCLIP-G backbone to compute the CLIP score. For ImageNet and CIFAR-10, we generate 50,000 images for each and calculate their FID using the EDM’s evaluation code https://github.com/NVlabs/edm/blob/main/fid.py.

Appendix E CIFAR-10 Experiments

Following the setup outlined in Section B.1, we train our models on CIFAR-10 and conduct comparisons with other competing approaches. Table 6 summarizes the results.

Appendix F Derivation for Distribution Matching Gradient

We present the derivation for Equation 7 as follows:

Appendix G Prompts for Figure 2

We use the following prompts for Figure 2. From left to right:

A DSLR photo of a golden retriever in heavy snow.

A professional portrait of a stylishly dressed elderly woman wearing very large glasses in the style of Iris Apfel, with highly detailed features.

Medium shot side profile portrait photo of a warrior chief, sharp facial features, with tribal panther makeup in blue on red, looking away, serious but clear eyes, 50mm portrait, photography, hard rim lighting photography.

A hyperrealistic photo of a fox astronaut; perfect face, artstation.

Appendix H Equivalence of Noise and Data Prediction

The noise prediction model ϵ(xt,t)\epsilon(x_{t},t) and data prediction model μ(xt,t)\mu(x_{t},t) could be converted to each other according to the following rule

Appendix I More Qualitative Results

We provide additional qualitative results on ImageNet (Fig. 7), LAION (Fig. 8, 9, 10, 11), and CIFAR-10 (Fig. 12, 13).