BOOT: Data-free Distillation of Denoising Diffusion Models with Bootstrapping
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Lingjie Liu, Josh Susskind
Introduction
Diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Nichol & Dhariwal, 2021; Song et al., 2020b) have become the standard tools for generative applications, such as image (Dhariwal & Nichol, 2021; Rombach et al., 2021; Ramesh et al., 2022; Saharia et al., 2022), video (Ho et al., 2022b, a), 3D (Poole et al., 2022; Gu et al., 2023; Liu et al., 2023b; Chen et al., 2023), audio (Liu et al., 2023a), and text (Li et al., 2022; Zhang et al., 2023) generation. Diffusion models are considered more stable for training compared to alternative approaches like GANs (Goodfellow et al., 2014a) or VAEs (Kingma & Welling, 2013), as they don’t require balancing two modules, making them less susceptible to issues like mode collapse or posterior collapse. Despite their empirical success, standard diffusion models often have slow inference times (around slower than single-step models like GANs), which poses challenges for deployment on consumer devices. This is mainly because diffusion models use an iterative refinement process to generate samples.
To address this issue, previous studies have proposed using knowledge distillation to improve the inference speed (Hinton et al., 2015). The idea is to train a faster student model that can replicate the output of a pre-trained diffusion model. In this work, we focus on learning single-step models that only require one neural function evaluation (NFE). However, conventional methods, such as Luhman & Luhman (2021), require executing the full teacher sampling to generate synthetic targets for every student update, which is impractical for distilling large diffusion models like StableDiffusion (SD, Rombach et al., 2021). Recently, several techniques have been proposed to avoid sampling using the concept of "bootstrap". For example, Salimans & Ho (2022) gradually reduces the number of inference steps based on the previous stage’s student, while Song et al. (2023) and Berthelot et al. (2023) train single-step denoisers by enforcing self-consistency between adjacent student outputs along the same diffusion trajectory (see Fig. 2). However, these approaches rely on the availability of real data to simulate the intermediate diffusion states as input, which limits their applicability in scenarios where the desired real data is not accessible.
In this paper, we propose BOOT, a data-free knowledge distillation method for denoising diffusion models based on bootstrapping. BOOT is partially motivated by the observation made by consistency model (CM, Song et al., 2023) that all points on the same diffusion trajectory (also known as PF-ODE (Song et al., 2020b)) have a deterministic mapping between each other. Unlike CM, which seeks self-consistency from any to , BOOT predicts all possible given the same noise point and a time indicator . Since our model always reads pure Gaussian noise, there is no need to sample from real data. Moreover, learning all from the same enables bootstrapping: it is easier to predict if the model has already learned to generate where . However, formulating bootstrapping in this way presents additional challenges, such as noisy sample prediction, which is non-trivial for neural networks. To address this, we learn the student model from a novel Signal-ODE derived from the original PF-ODE. We also design objectives and boundary conditions to enhance the sampling quality and diversity. This enables efficient inference of large diffusion models in scenarios where the original training corpus is inaccessible due to privacy or other concerns. For example, we can obtain an efficient model for synthesizing images of "raccoon astronaut" by distilling the text-to-image model with the corresponding prompts (shown in Fig. 3), even though collecting such data in reality is difficult.
In the experiments, we first demonstrate the efficacy of BOOT on various challenging image generation benchmarks, including unconditional and class-conditional settings. Next, we show that the proposed method can be easily adopted to distill text-to-image diffusion models. An illustration of sampled images from our distilled text-to-image model is shown in Fig. 1.
Preliminaries
where and for . By default, the signal-to-noise ratio (SNR, ) decreases monotonically with . A diffusion model learns to reverse the diffusion process by denoising , which can be easily sampled given the real data with :
Here, is the weight used to balance perceptual quality and diversity. The parameterization of typically involves U-Net (Ronneberger et al., 2015; Dhariwal & Nichol, 2021) or Transformer (Peebles & Xie, 2022; Bao et al., 2022). In this paper, we use to represent signal predictions. However, due to the mathematical equivalence of signal, noise, and v-predictions (Salimans & Ho, 2022) in the denoising formulation, the loss function can also be defined based on noise or v-predictions. For simplicity, we use for all cases in the remainder of the paper.
One can use ancestral sampling (Ho et al., 2020) to synthesize new data from the learned model. While the conventional method is stochastic, DDIM (Song et al., 2020a) demonstrates that one can follow a deterministic sampler to generate the final sample , which follows the update rule:
with the boundary condition . As noted in Lu et al. (2022), Eq. 2 is equivalent to the first-order ODE solver for the underlying probability-flow (PF) ODE (Song et al., 2020b). Therefore, the step size needs to be small to mitigate error accumulation. Additionally, using higher-order solvers such as Runge-Kutta (Süli & Mayers, 2003), Heun (Ascher & Petzold, 1998), and other solvers (Lu et al., 2022; Jolicoeur-Martineau et al., 2021) can further reduce the number of function evaluations (NFEs). However, these approaches are not applicable in single-step.
2 Knowledge Distillation
Orthogonal to the development of ODE solvers, distillation-based techniques have been proposed to learn faster student models from a pre-trained diffusion teacher. The most straightforward approach is to perform direct distillation (Luhman & Luhman, 2021), where a student model is trained to learn from the output of the diffusion model, which is computationally expensive itself:
Here, ODE-solver refers to any solvers like DDIM as mentioned above. While this naive approach shows promising results, it typically requires over 50 steps of evaluations to obtain reasonable distillation targets, which becomes a bottleneck when learning large-scale models.
Alternatively, recent studies (Salimans & Ho, 2022; Song et al., 2023; Berthelot et al., 2023) have proposed methods to avoid running the full diffusion path during distillation. For instance, the consistency model (CM, Song et al., 2023) trains a time-conditioned student model to predict self-consistent outputs along the diffusion trajectory in a bootstrap fashion:
where , typically with a single-step evaluation using Eq. 2. In this case, represents an exponential moving average (EMA) of the student parameters , which is important to prevent the self-consistency objectives from collapsing into trivial solutions by always predicting similar outputs. After training, samples can be generated by executing with a single NFE. It is worth noting that Eq. 4 requires sampling from the real data sample , which is the essence of bootstrapping: the model learns to denoise increasingly noisy inputs until . However, in many tasks, the original training data for distillation is inaccessible. For example, text-to-image generation models require billions of paired data for training. One possible solution is to use a different dataset for distillation; however, the mismatch in the distributions of the two datasets would result in suboptimal distillation performance.
Method
In this section, we present BOOT, a novel distillation approach inspired by the concept of bootstrapping without requiring target domain data during training. We begin by introducing signal-ODE, a modeling technique focused exclusively on signals (§ 3.1), and its corresponding distillation process (§ 3.2). Subsequently, we explore the application of BOOT in text-to-image generation (§ 3.3). The training pipeline is depicted in Fig. 3, providing an overview of the process.
We utilize a time-conditioned student model in our approach. Similar to direct distillation (Luhman & Luhman, 2021), BOOT always takes random noise as input and approximates the intermediate diffusion model variable: . This approach eliminates the need to sample from real data during training. The final sample can be obtained as . However, it poses a challenge to train effectively, as neural networks struggle to predict partially noisy images (Berthelot et al., 2023), leading to out-of-distribution (OOD) problems and additional complexities in learning accurately.
To overcome the aforementioned challenge, we propose an alternative approach where we predict . In this case, represents the low-frequency "signal" component of , which is easier for neural networks to learn. The initial noise for diffusion is denoted by . This prediction target is reasonable since it aligns with the boundary condition of the teacher model, where . Furthermore, we can derive an iterative equation from Eq. 2 for consecutive timesteps:
where , and represents the "negative half log-SNR." Notably, the noise term automatically cancels out in Eq. 5, indicating that the model always learns from the signal space. Moreover, Eq. 5 demonstrates an interpolation between the current model prediction and the diffusion-denoised output. Similar to the connection between DDIM and PF-ODE (Song et al., 2020b), we can also obtain a continuous version of Eq. 5 by letting as follows:
2 Learning with Bootstrapping
Our objective is to learn as a single-step prediction model using neural networks, rather than solving the signal-ODE with Eq. 6. By matching both sides of Eq. 6, we can readily obtain the loss function:
In Eq. 7, we use to estimate , and represents the corresponding noisy image. Instead of using forward-mode auto-differentiation, which can be computationally expensive, we can approximate the above equation with finite differences due to the 1-dimensional nature of . The approximate form is similar to Eq. 5:
Unlike CM-based methods, such as those mentioned in Eq. 4, we do not require an exponential moving average (EMA) copy of the student parameters to avoid collapsing. This avoids potential slow convergence and sub-optimal solutions. As shown in Eq. 8, the proposed objective is unlikely to degenerate because there is an incremental improvement term in the training target, which is mostly non-zero. In other words, we can consider as an exponential moving average of , with a decaying rate of . This ensures that the student model always receives distinguishable signals for different values of .
A critical challenge in learning BOOT is the "error accumulation" issue, where imperfect predictions of on large can propagate to subsequent timesteps. While similar challenges exist in other bootstrapping-based approaches, it becomes more pronounced in our case due to the possibility of out-of-distribution inputs for the teacher model, resulting from error accumulation and leading to incorrect learning signals. To mitigate this, we employ two methods: (1) We uniformly sample throughout the training time, despite the potential slowdown in convergence. (2) We use a higher-order solver (e.g., Heun’s method (Ascher & Petzold, 1998)) to compute the bootstrapping target with better estimation.
Boundary Condition
In theory, the boundary can have arbitrary values since , and the value of does not affect the value . However, is unbounded at , leading to numerical issues in optimization. As a result, the student model must be learned within a truncated range . This necessitates additional constraints at the boundaries to ensure that follows the same distribution as the diffusion model. In this work, we address this through an auxiliary boundary loss:
Here, we enforce the student model to match the initial denoising output. In our early exploration, we found that the boundary condition is crucial for the single-step student to fully capture the modeling space of the teacher, especially in text-to-image scenarios. Failure to learn the boundaries tends to result in severe mode collapse and color-saturation problems.
The overall learning objective combines , where is a hyper-parameter. The algorithm for student model distillation is presented in Appendix Algorithm 1.
3 Distillation of Text-to-Image Models
Our approach can be readily applied for distilling conditional diffusion models, such as text-to-image generation (Ramesh et al., 2022; Rombach et al., 2021; Balaji et al., 2022), where a conditional denoiser is learned with the same objective given an aligned dataset. In practice, inference of these models requires necessary post-processing steps for augmenting the conditional generation. For instance, one can perform classifier-free guidance (CFG, Ho & Salimans, 2022) to amplify the conditioning:
Pixel or Latent
Our method can be easily adopted in either pixel (Saharia et al., 2022) or latent space (Rombach et al., 2021) models without specific code change. For pixel-space models, it is sometimes critical to apply clipping or dynamic thresholding (Saharia et al., 2022) over the denoised targets to avoid over-saturation. Similarly, we also clip the targets in our objectives Eqs. 8 and 9. Pixel-space models (Saharia et al., 2022) typically involve learning cascaded models (one base model + a few super-resolution (SR) models) to increase the output resolutions progressively. We can also distill the SR models with BOOT into one step by conditioning both the SR teacher and the student with the output of the distilled base model.
Experiments
We begin by evaluating the performance of BOOT on diffusion models trained on standard image generation benchmarks: FFHQ (Karras et al., 2017), class-conditional ImageNet (Deng et al., 2009) and LSUN Bedroom (Yu et al., 2015). To ensure a fair comparison, we train all teacher diffusion models separately on each dataset using the signal prediction objective. Additionally, for ImageNet, we test the performance of CFG where the student models are trained with random conditioning on (see the effects of in Fig. 7).
For text-to-image generation scenarios, we directly apply BOOT on open-sourced diffusion models in both pixel-space (DeepFloyd-IF (IF), Saharia et al., 2022) https://github.com/deep-floyd/IF and latents space (StableDiffusion (SD), Rombach et al., 2021) https://github.com/Stability-AI/stablediffusion. Thanks to the data-free nature of BOOT, we do not require access to the original training set, which may consist of billions of text-image pairs with unknown preprocessing steps. Instead, we only need the prompt conditions to distill both models. In this work, we consider general-purpose prompts generated by users. Specifically, we utilize diffusiondb (Wang et al., 2022), a large-scale prompt dataset that contains million images generated by StableDiffusion using prompts provided by real users. We only utilize the text prompts for distillation.
Implementation Details
Similar to previous research (Song et al., 2023), we use student models with architectures similar to those of the teachers, having nearly identical numbers of parameters. A more comprehensive architecture search is left for future work. We initialize the majority of the student parameters with the teacher model , except for the newly introduced conditioning modules (target timestep and potentially the CFG weight ), which are incorporated into the U-Net architecture in a similar manner as how class labels were incorporated. It is important to note that the target timestep is different from the original timestep used for conditioning the diffusion model, which is always set to for the student model. Based on the actual implementation of the teacher models, we initialize the student output accordingly to accommodate the pretrained weights: , where represents “or” and correspond to the pre-trained teacher networks using the signal, noise or velocity (Salimans & Ho, 2022) parameterization, respectively. We include additional details in the Appendix C.
Evaluation Metrics
For image generation, results are compared according to Fréchet Inception Distance ((FID, Heusel et al., 2017), lower is better), Precision ((Prec., Kynkäänniemi et al., 2019), higher is better), and Recall ((Rec., Kynkäänniemi et al., 2019), higher is better) over real samples from the corresponding datasets. For text-to-image tasks, we measure the zero-shot CLIP score (Radford et al., 2021) for measuring the faithfulness of generation given randomly sampled captions from COCO2017 (Lin et al., 2014) validation set. In addition, we also report the inference speed measured by fps with batch-size 1 on single A100 GPU.
2 Results
We first evaluate the proposed method on standard image generation benchmarks. The quantitative comparison with the standard diffusion inference methods like DDPM (Ho et al., 2020) and the deterministic DDIM (Song et al., 2020a) are shown in Table 1. Despite lagging behind the -step DDIM inference, BOOT significantly improves the performance -step inference, and achieves better performance against DDIM with around denoising steps, while maintaining speed-up. Note that, the speed advantage doubles if the teacher employs guidance.
We also conduct quantitative evaluation on text-to-image tasks. Using the SD teacher, we obtain a CLIP-score of on COCO2017, a slight degradation compared to the -step DDIM results (), while it generates orders of magnitude faster, rendering real-time applications.
Visual Results
We show the qualitative comparison in Figs. 5 and 6 for image generation and text-to-image, respectively. For both cases, navïe -step inference fails completely, and the diffusion generally outputs grey and ill-structured images with fewer than NFEs. In contrast, BOOT is able to synthesize high-quality images that are visually close (Fig. 5) or semantically similar (Fig. 6) to teacher’s results with much more steps. Unlike the standard benchmarks, distilling text-to-image models (e.g., SD) typically leads to noticeably different generation from the original diffusion model, even starting with the same initial noise. We hypothesize it is a combined effect of highly complex underlying distribution and CFG. We show more results including pixel-space models in the appendix.
3 Analysis
The significance of incorporating the boundary loss is demonstrated in Fig. 8 (a) and (b). When using the same noise inputs, we compare the student outputs based on different target timesteps. As tracks the signal-ODE output, it produces more averaged results as approaches 1. However, without proper boundary constraints, the student outputs exhibit consistent sharpness across timesteps, resulting in over-saturated and non-realistic images. This indicates a complete failure of the learned student model to capture the distribution of the teacher model, leading to severe mode collapse.
Progressive v.s. Uniform Time Training
We also compare different training strategies in Fig. 8 (c) and (d). In contrast to the proposed approach of uniformly sampling , one can potentially achieve additional efficiency with a fixed schedule that progressively decreases as training proceeds. This progressive training strategy seems reasonable considering that the student is always initialized from and gradually learns to predict the clean signals (small ) during training. However, progressive training tends to introduce more artifacts (as observed in the visual comparison in Fig. 8). We hypothesize that progressive training is more prone to accumulating irreversible errors.
Controllable Generation
In Fig. 9, we visualize the results of latent space interpolation, where the student model is distilled from the pretrained IF teacher. The smooth transition of the generated images demonstrates that the distilled student model has successfully learned a continuous and meaningful latent space. Additionally, in Fig. 10, we provide an example of text-controlled generation by fixing the noise input and only modifying the prompts. Similar to the original diffusion teacher model, the BOOT distilled student retains the ability of disentangled representation, enabling fine-grained control while maintaining consistent styles.
Related Work
Speeding up inference of diffusion models is a broad area. Recent works and also our work (Luhman & Luhman, 2021; Salimans & Ho, 2022; Meng et al., 2022; Song et al., 2023; Berthelot et al., 2023) aim at reducing the number of diffusion model inference steps via distillation. Aside from distillation methods, other representative approaches include advanced ODE solvers (Karras et al., 2022; Lu et al., 2022), low-dimension space diffusion (Rombach et al., 2021; Vahdat et al., 2021; Jing et al., 2022; Gu et al., 2022), and improved diffusion targets (Lipman et al., 2023; Liu et al., 2022). BOOT is orthogonal and complementary to these approaches, and can theoretically benefit from improvements made in all these aspects.
Knowledge Distillation for Generative Models
Knowledge distillation (Hinton et al., 2015) has seen successful applications in learning efficient generative models, including model compression (Kim & Rush, 2016; Aguinaldo et al., 2019; Fu et al., 2020; Hsieh et al., 2023) and non-autoregressive sequence generation (Gu et al., 2017; Oord et al., 2018; Zhou et al., 2019). We believe that BOOT could inspire a new paradigm of distilling powerful generative models without requiring access to the training data.
Discussion and Conclusion
BOOT is a knowledge distillation algorithm, which by nature requires a pre-trained teacher model. Also by design, the sampling quality of BOOT is upper bounded by that of the teacher. Besides, BOOT may produce lower quality samples compared to other distillation methods (Song et al., 2023; Berthelot et al., 2023) where ground-truth data are easy to use, which can potentially be remedied by combining methods.
Future Work
As future research, we aim to investigate the possibility of jointly training the teacher and the student models in a manner that incorporates the concept of diffusion into the distillation process. By making the diffusion process "distillation aware," we anticipate improved performance and more effective knowledge transfer. Furthermore, we find it intriguing to explore the training of a single-step diffusion model from scratch. This exploration could provide insights into the applicability and benefits of BOOT in scenarios where a pre-trained model is not available.
Conclusion
In summary, this paper introduced a novel technique BOOT to distill diffusion models into single step. The method did not require the presence of any real or synthetic data by learning a time-conditioned student model with bootstrapping objectives. The proposed approach achieved comparable generation quality while being significantly faster than the diffusion teacher, and was also applicable to large-scale text-to-image generation, showcasing its versatility.
Acknowledgement
We thank Tianrong Chen, Miguel Angel Bautista, Navdeep Jaitly, Laurent Dinh, Shiwei Li, Samira Abnar, Etai Littwin for their critical suggestions and valuable feedback to this project.
References
Appendix A Algorithm Details
In this paper, we use to represent the diffusion model that denoises the noisy sample into its clean version, and we derive the DDIM sampler (Eq. 2) following the definition of Song et al. (2020a): we deterministically synthesize based on the following update rule:
where . Here we use ODE-Solver to represent the DDIM sampling from a random noise , and iteratively obtain the sample at step . In practice, we can generalize to higher-order ODE-solvers for better efficiency.
For distillation, we define the student model with which approximates along the diffusion trajectory above. To avoid directly predicting the noisy samples with neural networks, we re-parameterize where the noise part is constant throughout except the scale factor . In this way, the learning goal is to predict a new variable : the “signal” part of the original variable .
A.2 Derivation of Signal-ODE
Based on the definition of , we can derive the following equations from Eq. 11:
where we use the auxiliary variable for simplifying the equations. As mentioned in § 3.1, we can further obtain the continuous form of Eq. 12 by assigning . That is, Eq. 12 is equivalent to that shown in the following:
A.3 Bootstrapping Objectives
The bootstrapping objectives in Eq. 8 can be easily derived by taking the finite difference of Eq. 4. Here we use to estimate , and use to represent the noisy image obtained from .
Using Heun’s method essentially doubles the evaluations of the teacher model during training, while the add-on overheads are manageable as we stop the gradients to the teacher model.
A.4 Training Algorithm
We summarize the training algorithm of BOOT in Algorithm 1, where by default we assume conditional diffusion model with classifier-free guidance and DDIM solver. Here, for simplicity, we write . For unconditional models, we can simply remove the context sampling part.
Appendix B Connections to Existing Literature
Physics-Informed Neural Networks (PINNs, Raissi et al., 2019) are powerful approaches that combine the strengths of neural networks and physical laws to solve ODEs. Unlike traditional numerical methods, which rely on discretization and iterative solvers, PINNs employ machine learning techniques to approximate the solution of ODEs. The key idea behind PINNs is to incorporate physics-based constraints directly into the training process of neural networks. By embedding the governing equations and available boundary or initial conditions as loss terms, PINNs can effectively learn the underlying physics while simultaneously discovering the solution. This ability makes PINNs highly versatile in solving a wide range of ODEs, including those arising in fluid dynamics, solid mechanics, and other scientific domains. Moreover, PINNs offer several advantages, such as automatic discovery of spatio-temporal patterns and the ability to handle noisy or incomplete data.
Although motivated from different perspectives, BOOT shares similarities with PINNs at a high level, as both aim to learn ODE/PDE solvers directly through neural networks. In the domain of PINNs, solving ODEs can also be simplified into two objectives: the differential equation (DE) loss (Eq. 7) and the boundary condition (BC) loss (Eq. 9). The major difference lies in the focus of the two approaches. PINNs primarily focus on learning complex ODEs/PDEs for single problems, where neural networks serve as universal approximators to address the discretization challenges faced by traditional solvers. Moreover, the data space in PINNs is relatively low-dimensional. In contrast, BOOT aims to learn single-step generative models capable of synthesizing data in high-dimensional spaces (e.g., millions of pixels) from random noise inputs and conditions (e.g., labels, prompts). To the best of our knowledge, no existing work has applied similar methods in generative modeling. Additionally, while standard PINNs typically compute derivatives (Eq. 7) directly using auto-differentiation, in this paper, we employ the finite difference method and propose a bootstrapping-based algorithm.
B.2 Consistency Models / TRACT
The most related previous works to our research are Consistency Models (Song et al., 2023) and concurrently TRACT (Berthelot et al., 2023), which propose bootstrapping-style algorithms for distilling diffusion models. These approaches map an intermediate noisy training example at time step to the teacher’s -step denoising outputs using the DDIM inference procedure. The training target for the student is constructed by running the teacher model with one step, followed by the self-teacher with steps. As illustrated in Fig. 2, BOOT takes a different approach to bootstrapping. It starts from the Gaussian noise prior and directly maps it to an intermediate step in one shot. This change has significant modeling implications, as it does not require any training data and can achieve data-free distillation, a capability that none of the prior works possess.
B.3 Single-step Generative Models
BOOT is also related to other single-step generative models, including VAEs (Kingma & Welling, 2013) and GANs (Goodfellow et al., 2014b), which aim to synthesize data in a single forward pass. However, BOOT does not require an encoder network like VAEs. Thanks to the power of the underlying diffusion model, BOOT can produce higher-contrast and more realistic samples. In comparison to GANs, BOOT does not require a discriminator or critic network. Furthermore, the distillation process of BOOT enables better-controlled exploration of the text-image joint space, which is explored by the pretrained diffusion models, resulting in more coherent and realistic samples in text-guided generation. Additionally, BOOT is more stable to learn compared to GANs, which are challenging to train due to the adversarial nature of maintaining a balance between the generator and discriminator networks.
Appendix C Additional Experimental Settings
While the proposed method is data-free, we list the additional dataset information that used to train our teacher diffusion models:
FFHQ (https://github.com/NVlabs/ffhq-dataset) contains 70k images of real human faces in resolution of . In most of our experiments, we resize the images to a low resolution at for early-stage benchmarking.
LSUN (https://www.yf.io/p/lsun) is a collection of large-scale image dataset containing 10 scenes and 20 object categories. Following previous works (Song et al., 2023), we choose the category Bedroom (M images), and train an unconditional diffusion teacher. All images are resized to with center-crop. We use LSUN to validate the ability of learning in relative high-resolution scenarios.
ImageNet-1K (https://image-net.org/download.php) contains M images across classes. We directly merge all the training images with class labels and train a class-conditioned diffusion teacher. All images are resized to with center-crop. To support test-time classifier-free guidance, the teacher model is trained with unconditional probability.
As we do not need to train our own teacher models for text-to-image experiments, no additional text-image pairs are required in this paper. However, our distillation still requires the text conditions for querying the teacher diffusion. To better capture and generalize the real user preference of such diffusion models, we choose to adopt the collected prompt datasets:
DiffusionDB (https://poloclub.github.io/diffusiondb/) contains M images generated by Stable Diffusion using prompts and hyperparameters specified by users. For the purpose of our experiments, we only keep the text prompts and discard all model-generated images as well as meta-data and hyperparameters so that they can be used for different teacher models. We use the same prompts for both latent and pixel space models.
C.2 Text-to-Image Teachers
We directly choose the recently open-sourced large-scale diffusion models as our teacher models. More specifically, we looked into the following models:
StableDiffusion (SD) (https://github.com/Stability-AI/stablediffusion) is an open-source text-to-image latent diffusion model (Rombach et al., 2021) conditioned on the penultimate text embeddings of a CLIP ViT-H/14 (Radford et al., 2021) text encoder. Different standard diffusion models, SD performs diffusion purely in the latent space. In this work, we use the checkpoint of SD v2.1-Base (https://huggingface.co/stabilityai/stable-diffusion-2-1-base) as our teacher which first generates in latent space, and then directly upscaled to resolution with the pre-trained VAE decoder. The teacher model was trained on subsets of LAION-5B (Schuhmann et al., 2022) with noise prediction objective.
DeepFloyd IF (IF) (https://github.com/deep-floyd/IF) is a recently open-source text-to-image model with a high degree of photorealism and language understanding. IF is a modular composed of a frozen text encoder and three cascaded pixel diffusion modules, similar to Imagen (Saharia et al., 2022): a base model that generates image based on text prompt and two super-resolution models (). All stages of the model utilize a frozen text encoder based on the T5 (Raffel et al., 2020) to extract text embeddings, which are then fed into a UNet architecture enhanced with cross-attention and attention pooling. Models were trained on 1.2B text-image pairs (based on LAION (Schuhmann et al., 2022) and few additional internal datasets) with noise prediction objective. In this paper, we conduct experiments on the first two resolutions () with the checkpoints of IF-I-L-v1.0 (https://huggingface.co/DeepFloyd/IF-I-L-v1.0) and IF-II-M-v1.0 (https://huggingface.co/DeepFloyd/IF-II-M-v1.0).
C.3 Model Architectures
We follow the standard U-Net architecture (Nichol & Dhariwal, 2021) for image generation benchmarks and adopt the hyperparameters similar in f-DM (Gu et al., 2022). For text-to-image applications, we keep the default architecture setups from the teacher models unchanged. As mentioned in the main paper, we initialize the weights of the student models directly from the pretrained checkpoints and use zero initialization for the newly added modules, such as target time and CFG weight embeddings. We include additional architecture details in the Table 2.
C.4 Training Details
All models for all the tasks are trained on the same resources of NVIDIA A100 GPUs for K updates. Training roughly takes days to converge depending on the model sizes. We train all our models with the AdamW (Loshchilov & Hutter, 2017) optimizer, with no learning rate decay or warm-up, and no weight decay. Standard EMA to the weights is also applied for student models. Since our methods are data-free, there is no additional overhead on data storage and loading except for the text prompts, which are much smaller and can be efficiently loaded into memory.
Learning the boundary loss requires additional NFEs during each training step. In practice, we apply the boundary loss less frequently (e.g., computing the boundary condition every iterations and setting the loss to be otherwise) to improve the overall training efficiency. Note that distilling from the class-conditioned / text-to-image teachers requires multiple forward passes due to CFG, which relatively slows down the training compared to unconditional models.
Distilling from the DeepFloyd IF teacher requires learning from two stages. In this paper, we can easily achieve that by first distilling the first-stage model into single-step with BOOT, and then distilling the upscaler model based on the output of the first-stage student. Following the original paper (Saharia et al., 2022), noise augmentation is also applied on the first-stage output where we set the noise-level as https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/deepfloyd_if/pipeline_if_superresolution.py#L715. For more training hyperparameters, please refer to Table 2.
Appendix D Additional Samples from BOOT
Finally, we provide additional qualitative comparisons for the unconditional models of FFHQ (Fig. 14), LSUN (Fig. 15), the class-conditional model of ImageNet (Fig. 16), and comparisons for text-to-image generation based on DeepFloyd-IF ( in Figs. 17 and 20, in Figs. 1, 13, 11 and 12) and StableDiffusion ( in Figs. 19 and 21).