InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, Qiang Liu
Introduction
Modern text-to-image (T2I) generative models, such as DALL-E , Imagen , Stable Diffusion , StyleGAN-T , and GigaGAN , have demonstrated the remarkable ability to synthesize realistic, artistic, and detailed images based on textual descriptions. These advancements are made possible through the assistance of large-scale datasets and models .
However, despite their impressive generation quality, these models often suffer from excessive inference time and computational consumption . This can be attributed to the fact that most of these models are either auto-regressive or diffusion models . For instance, Stable Diffusion, even when using a state-of-the-art sampler , typically requires more than 20 steps to generate acceptable images. As a result, prior works have proposed employing knowledge distillation on these models to reduce the required sampling steps and accelerate their inference. Unfortunately, these methods struggle in the small step regime. In particular, one-step large-scale diffusion models have not yet been developed. The existing one-step large-scale T2I generative models are StyleGAN-T and GigaGAN , which rely on generative adversarial training and require careful tuning of both the generator and discriminator.
In this paper, we present a novel one-step generative model derived from the open-source Stable Diffusion (SD). We observed that a straightforward distillation of SD leads to complete failure. The primary issue stems from the sub-optimal coupling of noises and images, which significantly hampers the distillation process. To address this challenge, we leverage Rectified Flow , a recent advancement in generative models that utilizes probabilistic flows . In Rectified Flow, a unique procedure known as reflow is employed. Reflow gradually straightens the trajectory of the probability flows, thereby reducing the transport cost between the noise distribution and the image distribution. This improvement in coupling significantly facilitates the distillation process.
Consequently, we succeeded in training the first one-step SD model capable of generating high-quality images with remarkable details. Quantitatively, our one-step model achieves a state-of-the-art FID score of on the MS COCO 2017 dataset (5,000 images) with an inference time of only second per image. It outperforms the previous fastest SD model, progressive distillation , which achieved an one-step FID of . For MS COCO 2014 (30,000 images), our one-step model yields an FID of in second, surpassing one of the recent large-scale text-to-image GANs, StyleGAN-T ( in ). Notably, this is the first time a distilled one-step SD model performs on par with GAN, with pure supervised learning.
Related Works
Diffusion models have achieved unprecedented results in various generative modeling tasks, including image/video generation , audio generation , point cloud generation , biological generation , etc.. Most of the works are based on stochastic differential equations (SDEs), and researchers have explored techniques to transform them into marginal-preserving probability flow ordinary differential equations (ODEs) . Recently, propose to directly learn probability flow ODEs by constructing linear or non-linear interpolations between two distributions. These ODEs obtain comparable performance as diffusion models, but require much fewer inference steps. Among these approaches, Rectified Flow introduces a special reflow procedure which enhances the coupling between distributions and squeezes the generative ODE to one-step generation. However, the effectiveness of reflow has only been examined on small datasets like CIFAR10, thus raising questions about its suitability on large-scale models and big data. In this paper, we demonstrate that the Rectified Flow pipeline can indeed enable high-quality one-step generation in large-scale text-to-image diffusion models, hence brings ultra-fast T2I foundation models with pure supervised learning.
Early research on text-to-image generation focused on small-scale datasets, such as flowers and birds . Later, the field shifted its attention to more complex scenarios, particularly in the MS COCO dataset , leading to advancements in training and generation . DALL-E was the pioneering transformer-based model that showcased the amazing zero-shot text-to-image generation capabilities by scaling up the network size and the dataset scale. Subsequently, a series of new methods emerged, including autoregressive models , GAN inversion , GAN-based approaches , and diffusion models . Among them, Stable Diffusion is an open-source text-to-image generator based on latent diffusion models . It is trained on the LAION 5B dataset and achieves the state-of-the-art generalization ability. Additionally, GAN-based models like StyleGAN-T and GigaGAN are trained with adversarial loss to generate high-quality images rapidly. Our work provides a novel approach to yield ultra-fast, one-step, large-scale generative models without the delicate adversarial training.
Despite the impressive generation quality, diffusion models are known to be slow during inference due to the requirement of multiple iterations to reach the final result. To accelerate inference, there are two categories of algorithms. The first kind focuses on fast post-hoc samplers . These fast samplers can reduce the number of inference steps for pre-trained diffusion models to 20-50 steps. However, relying solely on inference to boost performance has its limitations, necessitating improvements to the model itself. Distillation has been applied to pre-trained diffusion models , squeezing the number of inference steps to below 10. Progressive distillation is a specially tailored distillation procedure for diffusion models, and has successfully produced 2/4-step Stable Diffusion . Consistency models are a new family of generative models that naturally operate in a one-step manner, but their performance on large-scale text-to-image generation is still unclear. Instead of employing direct distillation like previous works, we adopt Rectified Flow , which utilizes the reflow procedure to refine the coupling between the noise distribution and the image distribution, thereby improving the performance of distillation.
Methods
Recently, various of diffusion-based text-to-image generators have emerged with unprecedented performance. Among them, Stable Diffusion (SD) , an open-sourced model trained on LAION-5B , gained widespread popularity from artists and researchers. It is based on latent diffusion model , which is a denoising diffusion probabilistic model (DDPM) running in a learned latent space. Because of the recurrent nature of diffusion models, it usually takes more than 100 steps for SD to generate satisfying images. To accelerate the inference, a series of post-hoc samplers have been proposed . By transforming the diffusion model into a marginal-preserving probability flow, these samplers can reduce the necessary inference steps to as few as 20 steps . However, their performance starts to degrade noticeably when the number of inference steps is smaller than 10. For the 10 step regime, progressive distillation is proposed to compress the needed number of inference steps to 2-4. Yet, it is still an open problem if it is possible to turn large diffusion models, like SD, into an one-step model with satisfying quality.
2 Rectified Flow and Reflow
Rectified Flow learns to transfer to via an ordinary differential equation (ODE), or flow model
Different specific choices of the interpolation process result in different algorithms. As shown in , the commonly used denoising diffusion implicit model (DDIM) and the probability flow ODEs of correspond to with specific choices of time-differentiable sequences (see for details). In rectified flow, however, the authors suggested a simpler choice of
which favors straight trajectories that play a crucial role in fast inference, as we discuss in sequel.
In practice, the ODE in (1) need to be approximated by numerical solvers. The most common approach is the forward Euler method, which yields
where we simulate with a step size of and completes the simulation with steps.
Obviously, the choice yields a cost-accuracy trade-off: large approximates the ODE better but causes high computational cost.
For fast simulation, it is desirable to learn the ODEs that can be simulated accurately and fast with a small . This leads to ODEs whose trajectory are straight lines. Specifically, we say that an ODE is straight (with uniform speed) if
In this case, Euler method with even a single step () yields perfect simulation; See Figure 4. Hence, straightening the ODE trajectories is an essential way for reducing the inference cost.
where is learned using the same rectified flow objective (2), but with the linear interpolation (3) of pairs constructed from the previous .
The key property of reflow is that it preserves the terminal distribution while straightening the particle trajectories and reducing the transport cost of the transport mapping:
1) The distribution of and coincides; hence transfers to if does so.
2) The trajectories of tend to be straighter than that of . This suggests that it requires smaller Euler steps to simulate than . If is a fixed point of reflow, that is, , then must be exactly straight.
In text-to-image generation, the velocity field should additionally depend on an input text prompt to generate corresponding images. The reflow objective with text prompts is
In this paper, we set to be the velocity field of a pre-trained probability flow ODE model (such as that of Stable Diffusion, ), and denote the following as -Rectified Flow.
Theoretically, it requires an infinite number of reflow steps (5) to obtain ODEs with exactly straight trajectories. However, it is not practical to reflow too many steps due to high computational cost and the accumulation of optimization and statistical error. Fortunately, it was observed in that the trajectories of becomes nearly (even though not exactly) straight with even one or two steps of reflows. With such approximately straight ODEs, one approach to boost the performance of one-step models is via distillation:
Classifier-Free Guidance has a substantial impact on the generation quality of SD. Similarly, we can define the following velocity field to apply Classifier-Free Guidance on the learned Rectified Flow,
where trades off the sample diversity and generation quality. When , reduces back to the original velocity field . We provide analysis on in Section 6.
Preliminary Observations on Stable Diffusion 1.4
In this section, we conduct experiments with Stable Diffusion 1.4 to examine the effectiveness of the Rectified Flow framework and the reflow procedure.
The goal of the experiments in this section is to:
1) examine whether straightforward distillation can be effective for learning a one-step model from pre-trained large-scale T2I prbobility flow ODEs;
2) examine whether text-conditioned reflow can enhance the performance of distillation.
Our experiment concludes that: Reflow significantly eases the learning process of distillation, and distillation after reflow successfully produces a one-step model.
In this section, we use the pre-trained Stable Diffusion 1.4 provided in the official open-sourced repositoryhttps://github.com/CompVis/stable-diffusion to initialize the weights, since otherwise the convergence is unbearably slow.
1.1 Direct Distillation Fails
Our investigation starts from directly distilling the velocity field of Stable Diffusion 1.4 with (7) without applying any reflow. To achieve the best empirical performance, we conduct grid search on learning rate and weight decay to the limit of our computational resources. Particularly, the learning rates are selected from and the weight decay coefficients are selected from . For all the 9 models, we train them for steps. We generate pairs of as the training set for distillation. We compute the Fréchet inception distance (FID) on captions from MS COCO 2017 following the evaluation protocol in , then we show the model with the lowest FID in Figure 5. For more experiment results, please refer to Appendix.
We observe that, after training steps, all the nine models converge. However, the learned one-step generative model is far from satisfying. As shown in Figure 5, there is a huge gap in FID between SD and SD+Distill. In fact, it is difficult for the student model (SD+Distill) to imitate the teacher model (25-step SD). On the right side of Figure 5, with the same random noise, SD+Distill generates image with substantial difference from the teacher SD. From the experiments, we conclude that: directly distillation from SD is a tough learning problem for the student one-step model, and this is hard to mitigate by simply tuning the hyperparameters.
1.2 Reflow Improves Couling and Eases Distillation
For fair comparison with distillation, we train for steps with the weights initialized from pre-trained SD, then perform distillation for another training steps continuing from the obtained . The learning rate for reflow is . To distill from 2-Rectified Flow, we generate pairs of with 25-step Euler solver. The results are also shown in Figure 5 for comparison with direct distillation. The guidance scale for 2-Rectified Flow is set to .
First of all, the obtained 2-Rectified Flow has similar FID with the original SD, which are and , respectively. It indicates that reflow can be used to learn generative ODEs with comparable performance. Moreover, 2-Rectified Flow refines the coupling between the noise distribution and the image distribution, and eases the learning process for the student model when distillation. This can be inferred from two aspects. (1) The gap between the 2-Rectified Flow+Distill and the 2-Rectified Flow is much smaller than SD+Distill and SD. (2) On the right side of Figure 5, the image generated from 2-Rectified Flow+Distill shares great resemblance with the original generation, showing that it is easier for the student to imitate. This illustrates that 2-Rectified Flow is a better teacher model to distill a student model than the original SD.
2 Quantitative Comparison and Additional Analysis
In this section, we provide additional quantitative and qualitative results with further analysis and discussion. 2-Rectified Flow and its distilled versions are trained following the same training configuration as in Section 4.1.2.
For distillation, we consider two network structures: (i) U-Net, which is the exact same network structure as the denoising U-Net of SD; (ii) Stacked U-Net, which is a simplified structure from direct concatenation of two U-Nets with shared parameters. Compared with two-step inference, Stacked U-Net reduces the inference time to 0.12s from 0.13s by removing a set of unnecessary modules, while keeping the number of parameters unchanged. Stacked U-Net is more powerful than U-Net, which allows it to achieve better one-step performance after distillation. More details can be found in the Appendix.
According to Eq. (6), the reflow procedure can be repeated for multiple times. We repeat reflow for one more time to get 3-Rectified Flow (), which is initialized from 2-Rectified Flow (). 3-Rectified Flow is trained to minimize Eq. (6) for steps. Then we get its distilled version by generating new pairs of and distill for another steps. We found that to stabilize the training process of 3-Rectified Flow and its distillation, we have to decrease the learning rate from to .
Because our Rectified Flows are fine-tuned from the publicly available pre-trained models, the training cost is negligible compared with other large-scale text-to-image models. On our platform, when training with batch size of 4 and U-Net, one A100 GPU day can process iterations using L2 loss, iterations using LPIPS loss; when generating pairs with batch size of 16, one A100 GPU day can generate data pairs. Therefore, to get 2-Rectified Flow + Distill (U-Net), the training cost is approximately (Data Generation) + (Reflow) + (Distillation) 24.65 A100 GPU days. For reference, the training cost for SD 1.4 from scratch is 6250 A100 GPU days ; StyleGAN-T is 1792 A100 GPU days ; GigaGAN is 4783 A100 GPU days . A lower-bound estimation of training the one-step SD in Progressive Distillation is 108.8 A100 GPU days (the details for the estimation can be found in the Appendix).
We compare the performance of our models with baselines on MS COCO . Our first experiment follows the evaluation protocol of , where we use 5,000 captions from the MS COCO 2017 validation set and generate corresponding images. Then we measure the FID score and CLIP score using the ViT-g/14 CLIP model to quantitatively evaluate the image quality and correspondence with texts. We also record the average running time of different models with NVIDIA A100 GPU to generate one image. For fair comparison, we use the inference time of standard SD on our computational platform for Progressive Distillation-SD as their model is not available publicly. The inference time contains the text encoder and the latent decoder, but does NOT contain NSFW detector. The results are shown in table 1. (Pre) 2-Rectified Flow and (Pre) 3-Rectified Flow can generate realistic images that yields similar FID with SD 1.4+DPMSolver using 25 steps (). Within 0.09s, (Pre) 3-Rectified Flow+Distill (U-Net) gets an FID of 29.3, and within 0.12s, (Pre) 2-Rectified Flow+Distill (Stacked U-Net) gets an FID of 24.6, surpassing the previous best distilled SD model (37.2 and 26.0, respectively) . In our second experiment, we use 30,000 captions from MS COCO2014, and perform the same evaluation on FID. The results are shown in Table 2. We observe that (Pre) 2-Rectified Flow+Distill (Stacked U-Net) obtains an FID of , which is much better than SD+Distill with Stacked U-Net (). We empirically find that SD+Distill has worse FID with the larger Stacked U-Net in both experiments, though it has better visual quality. This could be attributed to the instability of FID metric when the images deviate severely from the real images.
Towards Better One-Step Generation: Scaling Up on Stable Diffusion 1.5
Our preliminary results with Stable Diffusion 1.4 demonstrate the advantages of adopting the reflow procedure in distilling one-step generative models. However, since only 24.65 A100 GPU days are spent in training, it is highly possible that the performance can be further boosted with more resources. Therefore, we expand the training time with a larger batch size to examine if scaling up has a positive impact on the result. The answer is affirmative. With 199 A100 GPU days, we obtain the first one-step SD that generates high-quality images with intricate details in second, on par with one of the state-of-the-art GANs, StyleGAN-T .
We switch to Stable Diffusion 1.5, and keep the same as in Section 4. The ODE solver sticks to 25-step DPMSolver for . Guidance scale is slightly decreased to because larger guidance scale makes the images generated from 2-Rectified Flow over-saturated. Since distilling from 2-Rectified Flow yields satisfying results, 3-Rectiifed Flow is not trained. We still generate pairs of data for reflow and distillation, respectively. To expand the batch size to be larger than , gradient accumulation is applied. The overall training pipeline for 2-Rectified Flow+Distill (U-Net) is summarized as follows:
Reflow (Stage 1): We train the model using the reflow objective (6) with a batch size of 64 for 70,000 iterations. The model is initialized from the pre-trained SD 1.5 weights. (11.2 A100 GPU days)
Reflow (Stage 2): We continue to train the model using the reflow objective (6) with an increased batch size of 1024 for 25,000 iterations. The final model is 2-Rectified Flow. (64 A100 GPU days)
The total training cost for InstaFlow-0.9B is (Data Generation) + + + + = A100 GPU days.
Expanding the model size is a key step in building modern foundation models . To this end, we adopt the Stacked U-Net structure in Section 4, but abandon the parameter-sharing strategy. This gives us a Stacked U-Net with 1.7B parameters, almost twice as large as the original U-Net. Starting from 2-Rectified Flow, 2-Rectified Flow+Distill (Stacked U-Net) is trained by the following distillation steps:
During training, we made the following observations: (1) the 2-Rectified Flow model did not fully converge and its performance could potentially benefit from even longer training duration; (2) distillation showed faster convergence compared to reflow; (3) the LPIPS loss had an immediate impact on enhancing the visual quality of the distilled one-step model. Based on these observations, we believe that with more computational resources, further improvements can be achieved for the one-step models.
Although one-step Stacked U-Net and 2-step progressive distillation (PD) need similar inference time, they have two key differences: (1) 2-step PD additionally minimizes the distillation loss at , which may be unnecessary for one-step generation from ; (2) by considering the consecutive U-Nets as one model, we are able to examine and remove redundant components from this large neural network, further reducing the inference time by approximately (from to ).
Evaluation
In this section, we systematically evaluate 2-Rectified Flow and the distilled one-step models. We name our one-step model 2-Rectified Flow+Distill (U-Net) as InstaFlow-0.9B and 2-Rectified Flow+Distill (Stacked U-Net) as InstaFlow-1.7B.
We follow the experiment configuration in Seciton 4.2. The guidance scale for the teacher model, 2-Rectified Flow, is set to . In Table 3, our InstaFlow-0.9B gets an FID-5k of with an inference time of , which is significantly lower than the previous state-of-the-art, Progressive Distillation-SD (1 step). The training cost for Progressive Distillation-SD (1 step) is A100 GPU days, while the training cost of the distillation step for InstaFlow-0.9B is A100 GPU days. With similar distillation cost, InstaFlow-0.9B yields clear advantage. The empirical result indicates that reflow helps improve the coupling between noises and images, and 2-Rectified Flow is an easier teacher model to distill from. By increaseing the model size, InstaFlow-1.7B leads to a lower FID-5k of with an inference time of .
On MS COCO 2014, our InstaFlow-0.9B obtains an FID-30k of within , surpassing StyleGAN-T ( in ). This is for the first time one-step distilled SD performs on par with state-of-the-art GANs. Using Stacked U-Net with 1.7B parameters, FID-30k of our one-step model achieves . Although this is still higher than GigaGAN , we believe more computational resources can close the gap: GigaGAN spends over A100 GPU days in training, while our InstaFlow-1.7B only consumes A100 GPU days.
2 Analysis on 2-Rectified Flow
2-Rectified Flow has straighter trajectories, which gives it the capacity to generate with extremely few inference steps. We compare 2-Rectified Flow with SD 1.5-DPM Solver on MS COCO 2017. 2-Rectified Flow adopts standard Euler solver. The inference steps are set to . Figure 10 (A) clearly shows the advantage of 2-Rectified Flow when the number of inference steps . In Figure 11, 2-Rectified Flow can generate images much better than SD with 1,2,4 steps, implying that it has a straighter ODE trajectory than the original SD 1.5.
It is widely known that guidance scale is a important hyper-parameter when using Stable Diffusion . By changing the guidance scale, the user can change semantic alignment and generation quality. Here, we investigate the influence of the guidance scale for 2-Rectified Flow, which has straighter ODE trajectories. In Figure 10 (B), increasing from to increases FID-5k and CLIP score on MS COCO 2017 at the same time. The former metric indicates degradation in image quality and the latter metric indicates enhancement in semantic alignment. Generated examples are shown in Figure 12. While the trending is similar to the original SD 1.5, there are two key differences. (1) Even when (no guidance), the generated images already have decent quality, since we perform reflow on SD 1.5 with a guidance scale of and the low-quality images are thus dropped in training. Therefore, it is possible to avoid using classifier-free guidance during inference to save GPU memory. (2) Unlike original SD 1.5, changing the guidance scale does not bring drastic change in the generated contents. Rather, it only perturbs the details and the tone. We leave explanations for these new behaviors as future directions.
The learned latent spaces of generative models have intriguing properties. By properly exploiting their latent structure, prior works succeeded in image editing , semantic control , disentangled control direction discovery , etc.. In general, the latent spaces of one-step generators, like GANs, are usually easier to analyze and use than the multi-step diffusion models. One advantage of our pipeline is that it gives a multi-step continuous flow and the corresponding one-step models simultaneously. Figure 13 shows that the latent spaces of our distilled one-step models align with 2-Rectified Flow. Therefore, the one-step models can be good surrogates to understand and leverage the latent spaces of continuous flow, since the latter one has higher generation quality.
3 Fast Preview with One-Step Model
A potential use case of our one-step models is to serve as previewers. Typically, large-scale text-to-image models work in a cascaded manner : the user can choose one image from several low-resolution images, and then an expensive super-resolution model expands the chosen low-resolution image to higher resolution. In the first step, the composition/color tone/style/other components of the low-resolution images may be unsatisfying to the user and the details are not even considered. Hence, a fast previewer can accelerate the low-resolution filtering process and provide the user more generation possibilities under the same computational budget. Then, the powerful post-processing model can improve the quality and increase the resolution. We verify the idea with SDXL-Refiner , a recent model that can refine generated images. The one-step models, InstaFlow-0.9B and InstaFlow-1.7B, generate images, then these images are interpolated to and refined by SDXL-Refiner to get high-resolution images. Several examples are shown in Figure 14. The low-resolution images generated in one step determine the content, composition, etc. of the image; then SDXL-Refiner refines the twisted parts, adds extra details, and harmonizes the high-resolution images.
The Best is Yet to Come
The recurrent nature of diffusion models hinders their deployment on edge devices , harms user experience, and adds to the overall cost. In this paper, we demonstrate that a powerful one-step generative model can be obtained from pre-trained Stable Diffusion (SD) using the text-conditioned Rectified Flow framework. Based on our results, we propose several promising future directions for exploration:
Improving One-Step SD: The training of the 2-Rectified Flow model did not fully converge, despite investing 75.2 A100 GPU days. This is only a fraction of training cost of the original SD ( A100 GPU days). By scaling up the dataset, model size, and training duration, we believe the performance of one-step SD will improve significantly. Moreover, SOTA base models, e.g., SDXL , can be leveraged as teachers to enhance the one-step SD model.
One-Step ControlNet : By applying our pipeline to train ControlNet models, it is possible to get one-step ControlNets capable of generating controllable contents within milliseconds. The only required modification involves adjusting the model structure and incorporating the control modalities as additional conditions.
Personalization for One-Step Models: By fine-tuning SD with the training objective of diffusion models and LORA , users can customize the pre-trained SD to generate specific contents and styles . However, as the one-step models have substantial difference from traditional diffusion models, determining the objective for fine-tuning these one-step models remains a subject that requires further investigation.
Neural Network Structure for One-Step Generation: It is widely acknowledged in the research community that the U-Net structure plays a crucial role in the impressive performance of diffusion models in image generation. With the advancement of creating one-step SD models using text-conditioned reflow and distillation, several intriguing directions arise: (1) exploring alternative one-step structures, such as successful architectures used in GANs, that could potentially surpass the U-Net in terms of quality and efficiency; (2) leveraging techniques like pruning, quantization, and other approaches for building efficient neural networks to make one-step generation more computationally affordable while minimizing potential degradation in quality.
References
Appendix A Neural Network Structure
The whole pipeline of our text-to-image generative model consists of three parts: the text encoder, the generative model in the latent space, and the decoder. We use the same text encoder and decoder as Stable Diffusion: the text encoder is adopted from CLIP ViT-L/14 and the latent decoder is adopted from a pre-trained auto-encoder with a downsampling factor of 8. During training, the parameters in the text encoder and the latent decoder are frozen. On average, to generate 1 image on NVIDIA A100 GPU with a batch size of 1, text encoding takes 0.01s and latent decoding takes 0.04s.
By default, the generative model in the latent space is a U-Net structure. For reflow, we do not change any of the structure, but just fine-tune the model. For distillation, we tested three network structures, as shown in Figure 15. The first structure is the original U-Net structure in SD. The second structure is obtained by directly concatenating two U-Nets with shared parameters. We found that the second structure significantly decrease the distillation loss and improve the quality of the generated images after distillation, but it doubles the computational time.
To reduce the computational time, we tested a family of networks structures by deleting different blocks in the second structure. By this, we can examine the importance of different blocks in this concatenated network in distillation, remove the unnecessary ones and thus further decrease inference time. We conducted a series of ablation studies, including:
Remove ‘Downsample Blocks 1 (the green blocks on the left)’
Remove ‘Upsample Blocks 1 (the yellow blocks on the left)’
Remove ‘In+Out Block’ in the middle (the blue and purple blocks in the middle).
Remove ‘Downsample Blocks 2 (the green blocks on the right)’
Remove ‘Upsample blocks 2 (the yellow blocks on the right)’
The only one that would not hurt performance is Structure 3, and it gives us a 7.7% reduction in inference time ( ). This third structure, Stacked U-Net, is illustrated in Figure 15 (c).
Appendix B Additional Details on Experiments
Our training script is based on the official fine-tuning script provided by HuggingFacehttps://huggingface.co/docs/diffusers/training/text2image. We use exponential moving average with a factor of 0.9999, following the default configuration. We clip the gradient to reach a maximal gradient norm of 1. We warm-up the training process for 1,000 steps in both reflow and distillation. BF16 format is adopted during training to save GPU memory. To compute the LPIPS loss, we used its official 0.1.4 versionhttps://github.com/richzhang/PerceptualSimilarity and its model based on AlexNet.
Appendix C Estimation of the Training Cost of Progressive Distillation (PD)
Measured on our platform, when training with a batch size of 4, one A100 GPU day can process 100,000 iterations using L2 loss. We compute the computational cost according to this.
We refer to Appendix C.2.1 (LAION-5B 512 512) of and estimate the training cost. PD starts from 512 steps, and progressively applies distillation to 1 step with a batch size of 512. Quoting the statement ‘For stage-two, we train the model with 2000-5000 gradient updates except when the sampling step equals to 1,2, or 4, where we train for 10000-50000 gradient updates’, a lower-bound estimation of gradient updates would be 2000 (512 to 256) + 2000 (256 to 128) + 2000 (128 to 64) + 2000 (64 to 32) + 2000 (32 to 16) + 5000 (16 to 8) + 10000 (8 to 4) + 10000 (4 to 2) + 50000 (2 to 1) = 85,000 iterations. Therefore, one-step PD at least requires A100 GPU days. Note that we ignored the computational cost of stage 1 of PD and ‘2 steps of DDIM with teacher’ during PD, meaning that the real training cost is higher than A100 GPU days.
Appendix D Direct Distillation of Stable Diffusion
We provide additional results on direct distillation of Stable Diffusion 1.4, shown in Figure 16, 17, 18, 19 and Table 5. Although increasing the learning rate boosts the performance, we found that a learning rate of leads to unstable training and NaN errors. A small learning rate, like and , results in slow convergence and blurry generation after training steps.
Appendix E Additional Generated Images
We show uncurated images generated from 20 random LAION text prompts with the same random noises for visual comparison. The images from different models are shown in Figure 20, 21, 22, 23.