One-Step Diffusion Distillation through Score Implicit Matching
Weijian Luo, Zemin Huang, Zhengyang Geng, J. Zico Kolter, Guo-jun Qi
Introduction
Over the past years, diffusion models (DMs) have shown significant advancements across a broad spectrum of applications, ranging from data synthesis , to density estimation , text-to-image generation, text-to-3D creation , image editing , and beyond . From a high level point of view, diffusion models, also framed as score-based diffusion models, use diffusion processes to corrupt the data distribution. They are then trained to approximate the score functions of the noisy data distributions across varying noise levels. Diffusion models have multiple advantages, such as training flexibility, scalability, and the ability to produce high-quality samples, making them a favored choice for modern AIGC models. After training, the learned score functions can be used to reverse the data corruption process, which can be implemented by numerically solving the associated stochastic differential equation. Such a data generation mechanism usually requires many neural network evaluations, which leads to a significant limitation of DMs: the generation performance of DMs degrades substantially when the number of sampling steps is reduced. This shortcoming restricts the practical deployment of DMs, particularly where quick inference is crucial, such as on devices with limited computational capacities like mobile phones and edge devices, or in applications requiring rapid response times.
This challenge has spurred a variety of approaches aimed at expediting the sampling process of diffusion models while preserving their robust generative capabilities. Distillation approaches, in particular, focus on applying distillation algorithms to transition the knowledge from pre-trained, teacher diffusion models to efficient student-generative models which are capable of producing high-quality samples within a few generation steps.
Some works have studied the diffusion distillation algorithm through the lens of probability divergence minimization. For instance, Luo et al. , Yin et al. have studied the algorithms that minimize the KL divergence between teacher and one-step student models. Zhou et al. have explored distilling with Fisher divergences, resulting in impressive empirical performances. Though these studies have contributed to the community in both theoretical and empirical aspects with applicable single-step generator models, their theories are built upon specific divergences, namely the Kullback-Leibler divergence and the Fisher divergence, which potentially restrict the distillation performances. A more general framework for understanding and improving diffusion distillation is still lacking.
In this work, we introduce Score Implicit Matching (SIM), a novel framework for distilling pre-trained diffusion models into one-step generator networks while maintaining high-quality generations. To do so, we propose a wide and flexible class of score-based divergences between the (intractable) score function of the generator model and that of the original diffusion model, for arbitrary distance functions between the two score functions. The key technical insight of this work is that although such divergences cannot be computed explicitly, the gradient of these divergences can be computed exactly using a result we call the score-gradient theorem, leading to an implicit minimization of the divergence. This lets us efficiently train models based on such divergences.
We evaluate the performance of SIM compared to previous approaches, using different choices of distance functions to define the divergence. Most relatedly, we compare SIM with the Diff-Instruct (DI) method, which uses a KL-based divergence term, and the Score Identity Distillation (SiD) method , which we show to be a special case of our approach when the distance function is simply chosen to be the squared distance (though derived in an entirely different fashion). We also show empirically that SIM with a specially-designed Pseudo-Huber distance function shows faster convergences and stronger robustness to hyper-parameters than distance, making the resulting method substantially strong than previous approaches.
Finally, we show that SIM obtains very strong empirical performance in absolute terms relative to past work in the field on CIFAR10 image generation and text-to-image generation. On the CIFAR10 dataset, SIM shows a one-step generative performance with a Frechet Inception Distance (FID) of 2.06 for unconditional generation and 1.96 for class-conditional generation. More qualitatively, distilling a leading diffusion-transformer-based text-to-image diffusion model results in an extremely capable one-step text-to-image generator which we show is almost lossless in terms of generative performances as teacher diffusion model. Particularly, by applying SIM to PixelArt- , a single-step generator is distilled that reaches an outstanding aesthetic score of with no performance decline over the original multi-step diffusion model. This remarkably outperforms the other one-step text-to-image generators including SDXL-TURBO of 5.33, SDXL-LIGHTNING of 5.34 and HYPER-SDXL of 5.85. Such a result not only marks a new direction for one-step text-to-image generation but also motivates further studies of distilling diffusion-transformer-based AIGC models in other domains such as video generation.
Diffusion Models
In this section, we introduce preliminary knowledge and notations about diffusion models and diffusion distillation. Assume we observe data from the underlying distribution . The goal of generative modeling is to train models to generate new samples . The forward diffusion process of DM transforms any initial distribution towards some simple noise distribution,
where is a pre-defined drift function, is a pre-defined scalar-value diffusion coefficient, and denotes an independent Wiener process. A continuous-indexed score network is employed to approximate marginal score functions of the forward diffusion process (2.1). The learning of score networks is achieved by minimizing a weighted denoising score matching objective ,
Here the weighting function controls the importance of the learning at different time levels and denotes the conditional transition of the forward diffusion (2.1). After training, the score network is a good approximation of the marginal score function of the diffused data distribution. High-quality samples from a DM can be drawn by simulating SDE which is implemented by the learned score network . However, the simulation of an SDE is significantly slower than that of other models such as one-step generator models.
Score Implicit Matching
In this section, we introduce Score Implicit Matching which is a general method tailored for the one-step distillation of score-based diffusion models. We first introduce the problem setup and notations, then introduce a general family of score-based probability divergences and show how SIM can be used to minimize the mentioned divergences. We finally discuss specific choices of the method, such as the choice of distance function, and explore the effect this has on the distillation.
Our starting point is a pre-trained diffusion model specified by the score function
where ’s are the underlying distribution diffused at time according to (2.1). We assume that the pre-trained diffusion model provides a sufficiently good approximation of data distribution, and thus will be the only item of consideration for our approach.
The student model of interest is a single-step generator network , which can transform an initial random noise to obtain a sample ; this network is parameterized by network parameters . Let denote the data distribution of the student model, and denote the marginal diffused data distribution of the student model with the same diffusion process (2.1). The student distribution implicitly induces a score function
and evaluating it is generally performed by training an alternative score network as elaborated later.
1 General Score-based Divergences
where and denote the marginal densities of the diffusion process (2.1) at time initialized with and respectively. is an integral weighting function. Clearly, we have if and only if all marginal score functions agree, which implies that .
2 Score Implicit Matching
Based upon this motivation, we would like to minimize the integral score-based divergence between and in order to train the student model, i.e.,
where we assume that the distribution has no parameter dependence of , such as . Taking the gradient with respect to , we have
where denotes the derivative of wrt. its inputs, i.e. . Unfortunately, because the score function is not tractable, it is impossible to compute directly, rendering such a direct approach impractical.
Fortunately, a key finding of our paper is if we choose the sampling distribution to the diffused implicit distribution, i.e. where the notation denotes the stop gradient operator that cuts off the parameter dependence of , the loss function (3.4) along with its intractable gradient (3.5) can be minimized efficiently via an gradient-equivalent loss. This relies on our Theorem 3.1.
If distribution satisfies some mild regularity conditions, we have for any score function , the following equation holds for all parameter :
The key observation here is that we replace the intractable gradient of the score function on the left-hand side of (3.6) with a much affordable evaluation of the score function on the right-hand side, the latter of which can be accomplished much more easily using a separate approximation network. This theorem can be proved by using score-projection identity which was first introduced to bridge denoising score matching with denoising auto-encoders. However, the key in proving Theorem 3.1 is a proper choice of -parameter (in)dependence by appropriately stopping the gradients shown in this theorem. We provide the detailed proof in Appendix A.1.
Now it is ready to reveal the objective we will use to train the implicit generator . A direct result of (3.6) is the gradient (3.5) can be realized via minimizing a tractable loss function
with . By Theorem 3.1, this alternative loss has an identical gradient to that of the original loss without the need to access the gradient of the score network.
In practice, we can use another online diffusion model to approximate the generator model’s score function pointwise, which was also done in previous works such as Luo et al. , Zhou et al. , and Yin et al. . We name the distillation method that minimizes the objective in (3.7) the Score Implicit Matching (SIM) because the learning process implicitly matches the intractable marginal score function of the implicit student model with the explicit score function of the pre-trained diffusion model .
The complete algorithm for SIM is shown in Algorithm 1, which trains the student model through two alternative phases between learning the marginal score function , and updating the generator model with gradient (3.7). The former phase follows the standard DM learning procedure, i.e., minimizing the denoising score matching loss function (2.2), with a slight change that the sample is generated from the generator. The resulting provides a good pointwise estimation of . The latter phase updates the generator’s parameter by minimizing the loss function (3.7), where two needed functions are provided by pretrained DM and learned DM .
3 Instances of Score Implicit Matching.
The previous section introduced the SIM algorithm without choosing a specific distance function . Here we discuss different choices and their influence on the distillation process. We also show that in the SIM framework, the SiD can be viewed as a special case.
Clearly, various choices of distance function result in different distillation algorithms. Perhaps the most natural choice of the distance function is a simple squared distance, i.e. . The corresponding derivative term writes . In fact, such a loss function recovers the delta loss studied in SiD , in which the authors empirically find that such a loss function works satisfactorily (though through a very different derivation). Thus, SiD is in fact a special case of SIM, though the derivation of SiD there does not suggest how alternative losses may be employed. A direct generalization of the quadratic form is the -power of the -norm where and is even. In this case, the distance function writes and the resulting loss function is summarized in Table 4 in Appendix A.3.
The Pseudo-Huber distance function.
Different from powered norms, we introduce SIM with the Pseudo-Huber distance function, which is defined with , where is a pre-defined positive constant. The corresponding distillation objective writes
In the rest of this paper, we will use the Pseudo-Huber distance as the default choice of the distance, unless specified otherwise. Due to the limited space, we summarize different choices of distance function and the corresponding loss functions in Table 4 as well as their derivations, along with more discussions in Appendix A.3.
Particularly, unlike SiD (the case in Table 4), with the Pseudo-Huber distance in the SIM, we observe that the vector is naturally normalized adaptively by dividing by a squared root of the vector. Such a normalization can stabilize the training loss, resulting in a robust and fast-converging distillation process. In section 4.1, we conduct empirical experiments to show three advantages: robustness to large-learning rate, fast convergence, and improved performances.
4 Related Works
Diffusion distillation is a research area that aims to reduce generation costs using teacher diffusion models. It involves three primary distillation methods: 1) Trajectory Distillation: This method trains a student model to mimic the generation process of diffusion models with fewer steps. Direct distillation () and progressive distillation () variants predict less noisy data from noisy inputs. Consistency-based methods () minimize the self-consistency metric. These require true data samples for training. 2) Distributional Matching: It focuses on aligning the student’s generation distribution with that of a teacher diffusion model. Among them are adversarial training methods () requiring real data for distilling diffusion models. Another important line of methods attempts to minimize divergences like KL () such as Diff-Instruct (DI) and Fisher divergence such as Score identity Distillation (SiD) (), often without needing real data. Though SIM has gotten inspiration from SiD and DI, the gap between SIM and SiD and DI is significant. SIM not only offers solid mathematical foundations which may lead to a deep understanding of diffusion distillation, but also provides substantial flexibility in using different distance functions, resulting in strong empirical performances when using specific Pseudo-Huber distance. 3) Other Methods: Methods like operator learning () and ReFlow () provide alternative insights into distillation. Moreover, many works made outstanding efforts to scale up diffusion distillation to one-step text-to-image generation and beyond
Experiments
In this experiment, we apply SIM to distill the pre-trained EDM diffusion models into one-step generator models on the CIFAR10 dataset. We follow the same setting as DI and SiD to distill the diffusion model into a one-step generator. Details can be found in Appendix B.2. We refer to the high-quality codebase of SiD https://github.com/mingyuanzhou/SiD to reproduce its results by closely referring to its configurations on our devices. We also re-implement the DI under the same experiment settings.
Performances.
We evaluate the performance of the trained generator via Frechet Inception Distance (FID) , which is the lower the better. We refer to the evaluation protocols in for comparison https://github.com/pkulwj1994/diff_instruct. Table 2 and 2 summarize the FID of generative models on CIFAR10 datasets. We reproduce the SiD and the DI with the same computing environments and evaluation protocol as SIM for a fair comparison. Models in the upper part of the table have different architectures or diffusion models from the EDM model, while the models in the lower part of the tables share exactly the same architecture and the teacher EDM diffusion models, which thus are directly comparable.
As shown in Table 2, for the CIFAR10 unconditional generation task, the proposed SIM achieves a decent FID of with only one-step generation, outperforming SiD and DI in the same evaluation setup. It is on par with the CTM and the SiD’s official implementation has yet to be released. For the CIFAR10 class-conditional generation in Table 2, the SIM achieves an FID of 1.96, acting among top-performing models.
The CIFAR-10 generation tasks are much toyish as merely performed with diffusion models of limited capacities on a simple dataset. We will perform experiments to distill from top-performing transformer-based diffusion models for text-to-image generation tasks. We will show that the one-step T2I generator distilled by SIM demonstrates state-of-the-art results over other industry-level models. Before that let us further look into some advantages of SIM – robustness to large learning rate and faster convergences – over SiD and DI on CIFAR-10, which will shed some light on how distillation methods scale up to more complex tasks with much larger neural networks.
Robustness to large learning rate.
We apply SIM, SiD, and DI under the same settings to distill from EDM (details in Appendix) on the CIFAR10 unconditional generation task, with a learning rate of , and plot the Fretchet Inception Distance (FID) and the Inception Score in Figure 2. Both the DI and the SiD are unstable even in the early training phase, while the SIM can steadily converge even with a large learning rate. The potential reason is that SIM naturally normalizes the loss objective to keep its scale from changing abruptly along the training process. This distinguishes SIM from SiD in practice for training large models, because training modern large models is so expensive that researchers often have few chances to adjust the hyperparameters within budget.
Fast convergence.
The second advantage of SIM is its faster convergence than SiD We find that the DI converges fast but suffers from mode-collapse issues. So we do not compare with it.. To show this, we follow the same setting as SiD on CIFAR10 unconditional generation. As shown in Figure 2, under all configurations, the SIM consistently shows better FID and Inception Scores under the same training iterations. Due to page limitations, we put more details in Appendix B.2.
Experiments on CIFAR10 generation show that SIM is a strong, robust, yet fast converging one-step diffusion distillation algorithm. However, the power of SIM is not restricted to a toy CIFAR-10 benchmark. In section 4.2, we apply the SIM to distill a 0.6B DiT ) based text-to-image diffusion model and obtain the state-of-the-art transformer-based one-step generator.
2 Transformer-based One-step Text-to-Image Generator
In recent years, transformer-based text-to-X generation models have gained great attention across image generations such as Stable Diffusion V3 and video generation such as Sora . In this section, we apply SIM to distill one of the leading open-sourced DiT-based diffusion models that have gained lots of attention recently: the 0.6B PixelArt- model , which is built upon with DiT model , resulting in the state-of-the-art one-step generator in terms of both quantitative evaluation metric and subjective user studies.
Experiment Settings and Evaluation Metrics.
The goal of one-step distillation is to accelerate the diffusion model into one-generation steps while maintaining or even outperforming the teacher diffusion model’s performances. To verify the performance gap between our one-step model and the diffusion model, we compare four quantitative values: the aesthetic score, the PickScore, the Image Reward, and our user-studied comparison score. On the SAM-LLaVA-Caption10M, which is one of the datasets the original PixelArt- model is trained on, we compare the SIM one-step model, which we called the SIM-DiT-600M, with the PixelArt- model with a 14-step DPM-Solver to evaluate the in-data performance gap. We also compare the SIM-DiT-600M and PixelArt- with other few-step models, such as LCM , TCD , PeReflow , and Hyper-SD series on the widely used COCO-2017 validation dataset. We refer to Hyper-SD’s evaluation protocols to compute evaluation metrics. Table 3 summarizes the evaluation performances of all models. For the human preference study against PixArt- and SIM-DiT-600M, we randomly select 17 prompts from the SAM Caption dataset and generate images with both PixArt- and SIM-DiT-600M, then ask the studied user to choose their preference according to image quality and alignments with the prompts. Figure 1 shows a visualization of our user study cases, in which it is difficult to distinguish the images from PixArt- and SIM-DiT-600M.
Almost lossless one-step distillation.
It is surprising that SIM-DiT-600M achieves almost no performance loss compared to teacher diffusion models. For instance, on the SAM Caption dataset in Table 3, SIM-DiT-600M recovers aesthetic score of PixArt- model and PickScore. However, the SIM-DiT-600M shows a slightly smaller Image Reward, which can be potentially optimized with more training computes. When compared with leading few-step text-to-image models such as SDXL-Turbo, SDXL-lightning, and Hyper-SDXL, the SIM-DiT-600M shows a dominant aesthetic score with a significant margin, together with a decent Image Reward and Pick Score.
Besides the top performance, the training cost of SIM-DiT-600M is surprisingly cheap. Our best model is trained (data-freely) with 4 A100-80G GPUs for 2 days, while other models in Table 3 require hundreds of A100 GPU days. We summarize the distillation costs in Table 3, marking that SIM is a super efficient distillation method with astonishing scaling ability. We believe such efficiency comes from two properties of SIM. First, the SIM is data-free, making the distillation process not need ground truth image data. Second, the use of the Pseudo-Huber distance function (3.3) adaptively normalizes the loss function, resulting in robustness to hyper-parameters and training stability.
Qualitative comparison.
Figure 3 qualitatively compares SIM-DiT-600M against other leading few-step text-to-image generative models. It is obvious that SIM-DiT-600M generates images with higher aesthetic performances than other models. This reflects the quantitative results in Table 3 where the SIM-DiT-600M reaches a high aesthetic score. Both the quantitative and qualitative results showcase the SIM-DiT-600M as the top-performing one-step text-to-image generator. Please check our supplementary materials for more qualitative evaluations.
Failure Cases of One-step SIM-DiT Model.
Though the SIM-DiT one-step model shows impressive performances, it inevitably has limitations. For instance, we find that the 0.6B SIM-DiT one-step model sometimes fails to generate high-quality tiny human faces and proper human arms and fingers. Besides, the model sometimes generates a wrong number of objects and contents that do not strictly follow the prompts. We believe that scaling up the model size and teacher diffusion models will help to address these issues. Please refer to Figure 4 for visualization of failure cases.
Conclusion and Future Works
This paper presents a novel diffusion distillation method, the score implicit matching (SIM), which enables to transform pre-trained multi-step diffusion models into one-step generators in a data-free fashion. The theoretical foundations and practical algorithms introduced in this paper can enable more affordable deployment of single-step generators across various domains and applications at scale without compromising the performance of underlying generative models.
Nonetheless, SIM has its limitations that call for further research. First, with the abundance of other powerful pre-trained generative models such as flow-matching models, it is worth exploring to reveal if it is possible to generalize the application of SIM to such a broader family of generative models. Second, even though data-free is an important feature of SIM, incorporating new data in the SIM can further boost the quality of generated images failed by the teacher model. This potential benefit has yet to be explored. We hope this could ease the training of large generative models.
Acknowledgement
Zhengyang Geng is supported by funding from the Bosch Center for AI. Zico Kolter gratefully acknowledges Bosch’s funding for the lab.
We would like to acknowledge constructive suggestions from reviewers and ACs/SACs/PCs of NeurIPS 2024. We also acknowledge the authors of Diff-Instruct and Score-identity Distillation for their great contributions to high-quality diffusion distillation Python code. We appreciate the authors of PixelArt- for making their DiT-based diffusion model public.
References
Appendix A Theory Parts
The proof of Theorem 3.1 is based on the so-called Score-projection identity which was first found in Vincent to bridge denoising score matching and denoising auto-encoders. Later the identity is applied by Zhou et al. for deriving distillation methods based on Fisher divergences. We appreciate the efforts of Zhou et al. and re-write the score-projection identity here without proof. Readers can check Zhou et al. for a complete proof of score-projection identity.
Let be a vector-valued function, using the notations of Theorem 3.1, under mild conditions, the identity holds:
We prove a more general result. Let be a vector-valued function, the so-called score-projection identity holds,
Taking gradient on both sides of identity (A.1), we have
Therefore we have the following identity:
which holds for arbitrary function and parameter . If we set
A.2 Pytorch style pseudo-code of Score Implicit Matching
In this section, we give a PyTorch style pseudo-code for algorithm 1, with the Pseudo-Huber distance function. For a detailed algorithm on CIFAR10 with EDM model, please check Algorithm 2.
A.3 Instances of SIM with different distance functions
In section 3.3, we have discussed the powered normed as distance functions. Other choices, such as the Huber distance, which is defined as
For other choices of distance functions, such as norm and exponential with powered norms, we put them in Table 4.
Appendix B Empirical Parts
The answer to the human preference study in Figure 1 is
the middle image of the first row is generated by one-step SIM-DiT-600M;
the leftmost image of the second row is generated by one step SIM-DiT-600M;
the leftmost image of the third row is generated by one-step SIM-DiT-600M.
B.2 Experiment details on CIFAR10 dataset
We follow the experiment setting of SiD and DI on CIFAR10. We start with a brief introduction to the EDM model .
The EDM model depends on the diffusion process
Samples from the forward process (B.1) can be generated by adding random noise to the output of the generator function, i.e., where is a Gaussian vector. The EDM model also reformulates the diffusion model’s score matching objective as a denoising regression objective, which writes,
Where is a denoiser network that tries to predict the clean sample by taking noisy samples as inputs. Minimizing the loss (B.2) leads to a trained denoiser, which has a simple relation to the marginal score functions as:
Under such a formulation, we actually have pre-trained denoiser models for experiments. Therefore, we use the EDM notations in later parts.
Let be pretrained EDM denoiser models. Owing to the denoiser formulation of the EDM model, we construct the generator to have the same architecture as the pre-trained EDM denoiser with a pre-selected index , which writes
We initialize the generator with the same parameter as the teacher EDM denoiser model.
Time index distribution.
When training both the EDM diffusion model and the generator, we need to randomly select a time in order to approximate the integral of the loss function (B.2). The EDM model has a default choice of distribution as log-normal when training the diffusion (denoiser) model, i.e.
In our algorithm, we follow the same setting as the EDM model when updating the online diffusion (denoiser) model.
In SiD, they propose to use a special discrete time distribution, which writes
They proposed to choose uniformly from
We name such a time distribution the distribution in Figure 2 because such a schedule was originally proposed in Karras’ EDM work for sampling.
However, in practice, we find that distribution (B.8) empirically does not work well. Instead, we find that a modified log-normal time distribution when updating the generation with SIM works better than distribution. Our SIM time distribution writes:
Weighting function.
As we have said, we use the same (B.7) weighting function as EDM when updating the denoiser model. When updating the generator, SiD uses a specially designed weighting function, which writes:
The notation means stop-gradient, and is the data dimensions. They claim such a weighting function helps to stabilize the training. However, in our experiments, since the SIM itself has normalized the loss (see section 4), we do not use such ad-hoc weighting functions. Instead, we just set the weighting function to be 1 for all time. We call the SiD’s weighting function the in Figure 2, and our weighting the in Figure 2.
In Figure 2, we compare the SiD and SIM with different time distribution and weighting functions. We find that SIM+nowgt+lognormal time distribution gives the best performances significantly, therefore our final experiment tasks such a configuration. Table 5 records the detailed configurations we use for SIM on CIFAR10 EDM distillation.
With the optimal setting and EDM formulation, we can rewrite our algorithm in an EDM style in Algorithm 2.
B.3 Experiment details on Text-to-Image Distillation
In the Text-to-Image distillation part, in order to align our experiment with that on CIFAR10, we rewrite the PixArt- model in EDM formulation:
Here, following the iDDPM+DDIM preconditioning in EDM, PixArt- is denoted by , is the image data plus noise with a standard deviation of , for the remaining parameters such as and , we kept them unchanged to match those defined in EDM. Unlike the original model, we only retained the image channels for the output of this model. Since we employed the preconditioning of iDDPM+DDIM in the EDM, each value is rounded to the nearest 1000 bins after being passed into the model. For the actual values used in PixArt-, beta_start is set to 0.0001, and beta_end is set to 0.02. Therefore, according to the formulation of EDM, the range of our noise distribution is [0.01, 156.6155], which will be used to truncate our sampled . For our one-step generator, it is formulated as:
Here following SiD and , we observed in practice that larger values of lead to faster convergence of the model, but the difference in convergence speed is negligible for the complete model training process and has minimal impact on the final results.
We utilized the SAM-LLaVA-Caption10M dataset, which comprises prompts generated by the LLaVA model on the SAM dataset. These prompts provide detailed descriptions for the images, thereby offering us a challenging set of samples for our distillation experiments.
All experiments in this section were conducted on 4 A100-40G GPUs with bfloat16 precision, using the PixArt-XL-2-512x512 model version, employing the same hyperparameters. For both optimizers, we utilized Adam with a learning rate of 5e-6 and betas=[0, 0.999]. Additionally, to enable a batch size of 1024, we employed gradient checkpointing and set the gradient accumulation to 8. Finally, regarding the training noise distribution, instead of adhering to the original iDDPM schedule, we sample the from a log-normal distribution with a mean of -2.0 and a standard deviation of 2.0, we use the same noise distribution for both optimization steps and set the two loss weighting to constant 1. Our best model was trained on the SAM Caption dataset for approximately 16k iterations, which is equivalent to less than 2 epochs. This training process took about 2 days on 4 A100-40G GPUs.
We also tested the impact of different noise distributions on the distillation process. When the noise distribution is highly concentrated around smaller values, we observed a phenomenon where the generated samples appear excessively dark. On the other hand, when we used slightly larger noise distributions, we found that the structure of the generated samples tended to be unstable.
B.4 Instruction for Human Preference Study
Our user study primarily focuses on comparing the outputs of the distilled model and the teacher model. Each image has undergone rigorous manual review to ensure the safety of survey participants. We conducted the study using questionnaires, where users were presented with two randomly ordered images generated by the distilled model and teacher model and asked to select the sample that best matched the text description and had higher image quality. Finally, we used the collected votes for the distilled model and the teacher model as indicators of user preference. The questionnaire website used for conducting these evaluations are shown in Figure 5.
To be more specific, we randomly selected 17 prompt words and generated images of resolution 512x512 using both the student model and the teacher model. To facilitate comparison, we presented the two images side by side in random order. In the questionnaire, we provided the complete prompt words for reference in addition to the generated images. In the end, we collected approximately 30 survey responses in total.
B.5 Generated Samples on CIFAR10
B.6 FID Convergence on CIFAR10 Unconditional Generation
B.7 Prompts for Figure 3
prompt for first row of Figure 3: A small cactus with a happy face in the Sahara desert.
prompt for second row of Figure 3: An image of a jade green and gold coloured Fabergé egg, 16k resolution, highly detailed, product photography, trending on artstation, sharp focus, studio photo, intricate details, fairly dark background, perfect lighting, perfect composition, sharp features, Miki Asai Macro photography, close-up, hyper detailed, trending on artstation, sharp focus, studio photo, intricate details, highly detailed, by greg rutkowski.
prompt for third row of Figure 3: Baby playing with toys in the snow.