Post-training Quantization on Diffusion Models
Yuzhang Shang, Zhihang Yuan, Bin Xie, Bingzhe Wu, Yan Yan
Introduction
Recently, denoising diffusion (also dubbed score-based) generative models have achieved phenomenal success in various generative tasks, such as images , audio , video , and graphs . Besides these fundamental tasks, their flexibility of implementation on downstream tasks is also attractive, e.g., they are effectively introduced for super-resolution , inpainting , and image-to-image translation . Diffusion models (DMs) have achieved superior performances on most of these tasks and applications, both concerning quality and diversity, compared with historically SoTA Generative Adversarial Networks (GANs) .
A diffusion process transforms real data gradually into Gaussian noise, and then the process is reversed to generate real data from Gaussian noise (denoising process) . Particularly, the denoising process requires iterating the noise estimation (also known as a score function ) via a cumbersome neural network over thousands of time-steps. While it has a compelling quantity of images, its long iterative process and high inference cost for generating samples make it undesirable. Thus, increasing the speed of this generation process is now an active area of research . To accelerate diffusion models, researchers propose several approaches, which mainly focus on sample trajectory learning for faster sampling strategies. For example, Chen et al. and San-Roman et al. propose faster step size schedules for VP diffusions that still yield relatively good quality/diversity metrics; Song et al. adopt implicit phases in the denoising process; Bao et al. and Lu et al. derive analytical approximations to simplify the generation process.
Our study suggests that two orthogonal factors slow down the denoising process: i) lengthy iterations for sampling images from noise, and ii) a cumbersome network for estimating noise in each iteration. Previously DM acceleration methods only focus on the former , but overlook the latter. From the perspective of network compression, many popular network quantization and pruning methods follow a simple pipeline: training the original model and then fine-tuning the quantized/pruned compressed model . Particularly, this training-aware compression pipeline requires a full training dataset and many computation resources to perform end-to-end backpropagation. For DMs, however, 1) training data are not always ready-to-use due to privacy and commercial concerns; 2) the training process is extremely expensive. For example, there is no access to the training data for the industry-developed text-to-image models Dall·E2 and Imagen . Even if one can access their datasets, fine-tuning them also consumes hundreds of thousands of GPU hours. Those two obstacles make training-aware compression not suitable for DMs.
Training-free network compression techniques are what we need for DM acceleration. Therefore, we propose to introduce post-training quantization (PTQ) into DM acceleration. In a training-free manner, PTQ can not only speed up the computation of the denoising process but also reduce the resources to store the diffusion model weight, which is required in DM acceleration. Although PTQ has many attractive benefits, its implementation in DMs remains challenging. The main reason is that the structure of DMs is hugely different from previously PTQ-implemented structures (e.g., CNN and ViT for image recognition). Specifically, the output distributions of noise estimation networks change with time-step, making previous PTQ methods fail in DMs since they are designed for single-time-step scenarios.
This study attempts to answer the following fundamental question: How does the design of the core ingredients of the PTQ for the DMs process (e.g., quantized operation selection, calibration set collection, and calibration metric) affect the final performance of the quantized diffusion models? To this end, we analyze the PTQ and DMs individually and correlatedly. We find that simple generalizations of previous PTQ methods to DMs lead to huge performance drops due to output distribution discrepancies w.r.t.time-step in the denoising process. In other words, noise estimation networks rely on time-step, which makes their output distributions change with time-step. This means that a key module of the previous PTQ calibration, cannot be used in our case. Based on the above observations, we devise a DM-specific calibration method, termed Normally Distributed Time-step Calibration (NDTC), which first samples a set of time-steps from a skew normal distribution, and then generates calibration samples in terms of sampled time-steps by the denoising process. In this way, the time-step discrepancy in the calibration set is enhanced, which improves the performance of PTQ4DM. Finally, we propose a novel DM acceleration method, Post-Training for Diffusion Models (PTQ4DM) via incorporating all the explorations.
Overall, the contributions of this paper are three-fold: (i) To accelerate denoising diffusion models, we introduce PTQ into DM acceleration where noise estimation networks are directly quantized in a post-training manner. To the best of our knowledge, this is the first work to investigate diffusion model acceleration from the perspective of training-free network compression. (ii) After all-inclusively investigations of PTQ and DMs, we observe the performance drop induced by PTQ for DMs can be attributed to the discrepancy of output distributions in various time-steps. Targeting this observation, we explore PTQ from different aspects and propose PTQ4DM. (iii) Experimentally, PTQ4DM can quantize the pre-trained diffusion models to 8-bit without significant performance loss for the first time. Importantly, PTQ4DM can serve as a plug-and-play module for other SoTA DM acceleration methods, as shown in Fig. 1.
Related Work
Due to the long iterative process in conjunction with the high cost of denoising via networks, diffusion models cannot be widely implemented. To accelerate the diffusion probabilistic models (DMs), previous works pursue finding shorter sampling trajectories while maintaining the DM performance. Chen et al. introduce grid search and find an effective trajectory with only six time-steps. However, the grid search approach can not be generalized into very long trajectories subject to its exponentially growing time complexity. Watson et al. model the trajectory searching as a dynamic programming problem. Song et al. construct a class of non-Markovian diffusion processes that lead to the same training objective, but whose reverse process can be much faster to sample from. As for DMs with continuous timesteps (i.e., score-based perspective ), Song et al. formulate the DM in of form of an ordinary differential equation (ODE), and improve sampling efficiency via utilizing faster ODE solver. Jolicoeur-Martineau et al. introduce an advanced SDE solver to accelerate the reverse process via an adaptively larger sampling rate. Bao et al. estimate the variance and KL divergence using the Monte Carlo method and a pretrained score-based model with derived analytic forms, which are simplified from the score-function. In addition to those training-free methods, Luhman & Luhman compress the reverse denoising process into a single-step model; San-Roman dynamically adjust the trajectory during inference. Nevertheless, implementing those methods requires additional training after obtaining a pretrained DM, which makes them less desirable in most situations. In summary, all those DM acceleration methods can be categorized into finding effective sampling trajectories.
However, we show in this paper that, in addition to finding short sampling trajectories, diffusion models can be further accelerated through network compression for each noise estimation iteration. Note that our method PTQ4DM is an orthogonal path with those above-mentioned fast sampling methods, which means it can be deployed as a plug-and-play module for those methods. To the best of our knowledge, our work is the first study on quantizing diffusion models in a post-training manner.
2 Post-training Quantization
Quantization is one of the most effective ways to compress a neural network. There are two types of quantization methods: Quantization-aware training (QAT) and Post-training quantization (PTQ). QAT considers the quantization in the network training phase. While PTQ quantizes the network after training. As PTQ consumes much less time and computation resources, it is widely used in network deployment.
Most of the work of PTQ is to set the quantization parameters for weights and activcations in each layer. Take uniform quantization as an example, the quantization parameters include scaling factor and zero point . A floating-point value is quantized to integer value according to the parameters:
The clamp function clip the rounded value to the range of . In order to set quantization parameters for the weight tensor and the activation tensor in a layer, a simple but effective way is to select the quantization parameters that minimize the MSE of the tensors before and after quantization . Other metrics, such as L1 distance, cosine distance, and KL divergence, can also be used to evaluate the distance of the tensors before and after quantization .
In order to calculate the activations in the network, a small number of calibration samples should be used as input in PTQ. The selected quantization parameters are dependent with the selection of these calibration samples. demonstrate the effect of the number of the calibration samples. Zero-shot quantization (ZSQ) is a special case of PTQ. ZSQ generates the calibration dataset according to information recorded in the network, such as the mean and var in batch normalization layer. They generate the input sample by gradient descent method to make the distribution of the activations in network similar to the distribution of real samples. The image generation process from noise in diffusion model only uses the network inference, which is quite different from previous ZSQ methods.
PTQ on Diffusion Models
Diffusion Models. The diffusion probabilistic model (DPM) is initially introduced by Sohl-Dickstein et al. , where the DPM is trained by optimizing the variational bound . Here, we briefly review the diffusion model to illustrate the difference from traditional models. Here, we briefly review the diffusion model, especially its lengthy diffusion and denoising process. We highlight that those properties make it difficult to simply generalize common PTQ methods into diffusion models simply in Sec. 3.2.
Given a real data distribution , we define the diffusion process that gradually adds a small amount of isotropic Gaussian noise with a variance schedule to produce a sequence of latent , which is fixed to a Markov chain. When is sufficiently large and a well-behaved schedule of , is equivalent to an isotropic Gaussian distribution.
A notable property of the diffusion process admits us to sample at an arbitrary timestep via directly conditioned on the input . Let and :
Since depends on the data distribution , which is intractable. Therefore, we need to parameterize a neural network to approximate it:
We utilize the variational lower bound to optimize the negative log-likelihood.
The objective function of the variational lower bound can be further rewritten to be a combination of several KL-divergence and entropy terms (more details in ).
uses a separate discrete decoder derived from . does not depend on , it is close to zero if . The remain term is a KL-divergence to directly compare to diffusion process posterior that is tractable when is conditioned,
This is the training process of the diffusion model. After obtaining the well-trained noise estimation model in Eq. 5, given a random noise, we can generate samples through the denoising process by iterative sampling from until we receive . Detailed information can be found in the surveys . Since the iterative process for denoising from noise input to synthetic images is extremely long (e.g., pioneer work, DDPM requires 4000 steps for generating a sample from noise), as illustrated in Fig. 2 (left); and the networks for estimating the noise in each denoising iteration is very deep and complicated, as illustrated in Fig. 2 (right). The inference of the diffusion model is expensive.
Post-training Quantization takes a well-trained network and selects the quantization parameters for the weight tensor and activation tensor in each layer. We use the quantization parameters, scaling factor , and zero point to transform a tensor to the quantized tensorWe focus on uniform quantization since it is the most widely used.. One of the most widely used methods to select the parameters is to minimize the error caused by quantization. The quantization error is formulated as:
where is the de-quantized tensor, and Metric is the metric function to evaluate the distance of and the full-precision tensor . MSE, cosine distance, L1 distance, and KL divergence are commonly used metric functions. The quantization process can be formulated as:
We can directly quantize the weight to minimize the quantization error, but we cannot get the activation tensor and quantize it without input. In order to collect the full-precision activation tensor, a number of unlabeled input samples (calibration dataset) are used as input. The size of the calibration dataset (e.g., 128 randomly selected images) is much smaller than the training dataset.
In general, PTQ quantizes a network in three steps: (i) Select which operations in the network should be quantized and leave the other operations in full-precision. For example, some special functions such as softmax and GeLU often takes full-precision .Quantizing these operations will significantly increase the quantization error and they are not very computationally intensive; (ii) Collect the calibration samples. The distribution of the calibration samples should be as close as possible to the distribution of the real data to avoid over-fitting of quantization parameters on calibration samples; (iii) Use the proper method to select quantization parameters for weight tensors and activation tensors.
In the next sections, we will explore how to apply PTQ to the diffusion model step by step.
2 Exploration on Operation Selection
For the diffusion model, we will analyze the image generation process to determine which operations should be quantized. The diffusion model iteratively generate the from . At each timestep, the inputs of the network are and , and the outputs are the mean and variance . Then is sampled from the distribution defined as Eq 5. As shown in Figure 2, the network in the diffusion model often takes UNet-like CNN architecture. The same as most previous PTQ methods, the computation-intensive convolution layers and fully-connected layers in the network should be quantized. The batch normalization can be folded into the convolution layer. The special functions such as SiLU and softmax are kept in full-precision.
There are two more questions for the diffusion model: 1. whether the network’s outputs, and , can be quantized? 2. whether the sampled image can be quantized? To answer the two questions, we only quantize the operation generating , , or . As shown in Table 1, we observe that they are not sensitive to quantization and we indicate that they can be quantized.
3 Exploration on Calibration Dataset
The second step is to collect the calibration samples for quantizing diffusion models. The calibration samples can be collected from the training dataset for quantizing other networks. However, the training dataset in the diffusion model is , which is not the network’s input. The real input is the generated samples . Should we use the generated samples in diffusion process or the generated samples in denoising process? At what time-step , should the generated samples be collected? This section will explore how to make a good calibration dataset.
By all-inclusively investigating several intuitive PTQ baselines, we obtain four meaningful observations (Sec. 3.3.1), which accordingly guide the design of our method (Sec. 3.3.2). Experimental results demonstrate that our method is efficient and effective. Through devised PTQ4DM calibration, the 8-bit post-training quantized diffusion model can perform at the same performance level as its full-precision counterpart, e.g., 8-bit diffusion model reaches 23.9 FID and 15.8 IS, while 32-bit one has 21.6 FID and 14.9 IS.
As discussed in Sec. 3.2, we desire the distribution of the collected calibration samples should be as close as possible to the distribution of the real data. In this way, the calibration set can supervise the quantization by minimizing the quantization error. Since previous works are implemented on single-time-step scenarios (e.g., CNN and ViT for image recognition and object detection) , they can directly collect samples from the real training dataset for quantizing networks. Due to the small size of the calibration dataset, its collection is extremely sensitive. If the distribution of the collected dataset is not representative of the real dataset, it can easily lead to overfitting for the calibration task.
We encounter more challenges when calibrating PTQ for DM. Since the inputs of the to-be-quantized network are the generated samples , in which is a large number to maintain the diffusion process converging to isotropic Normal distribution. To quantize the diffusion model, we are required to design a novel and effective calibration dataset collection method in this particular multi-time-step scenario. We start by investigating both PTQ calibration and DMs, and then obtain the following instructive observations.
Observation 0: Distributions of activations changes along with time-step changing.
To understand the output distribution change of diffusion models, we investigate the activation distribution with respect to time-step. We would like to analyze the output distribution at different time-step, for example, given and , the output activation distributions of and . Theoretically, if the distribution changes w.r.t.time-step, it would be difficult to implement previous PTQ calibration methods, as they are proposed for temporally-invariant calibration . We first analyze the overall activation distributions of the noise estimation network via boxplot as did, and then we take a closer look at the layer-wise distributions via histogram. The results are shown in Fig. 3. We can observe that at different time-steps, the corresponding activation distributions have large discrepancies, which makes previous PTQ calibration methods inapplicable for multi-time-step models (i.e., diffusion models).
Observation 1: Generated samples in the denoising process are more constructive for calibration.
In general, there are two directions to generate samples for PTQ calibration in diffusion: raw images as input for diffusion process, and noise as input for denoising process. Previous PTQ methods use raw images, as raw images can serve as ground truth, representing the training set’s distribution. We conduct a pair of comparison experiments, in which we separately collect two calibration sets with raw images for diffusion process and Gaussian noise for denoising process, and use these two sets to calibrate quantized models. Another similar intuitive baseline is to use the training samples in the diffusion process as calibration data. Specifically, we randomly generate a timestep for each image , and use Eq. 4 according to to generate . In other word, collect calibration samples in a “Image + Gaussian Noise” manner. We name this scheme as training-mimic baseline. The results are listed in Tab. 2. We find that the input noises for diffusion process are more constructive for calibrating quantized DMs.
Observation 2: Sample close to real image is more beneficial for calibration.
Based on the aforementioned observations, we establish a baseline of PTQ calibration for DM based on , in which the quantized diffusion models are calibrated with samples at time-step , i.e., a set of . We refer to this straight-forward approach as a naive PTQ-for-DM baseline. Specifically, given a set of Gaussion noise , we use the diffusion model with the full-precision noise estimation network, in Eq. 5 to generate a set of as calibration set. Then as described in Sec. 3.1, we use this collected set to calibrate our quantized noise estimation network, , in which is the quantized parameters. We conduct a series of experiments with this calibration baseline in different time-steps, i.e., , where is the total denoising time-steps. The results are presented in Fig. 4. We can see that the 8-bit model calibrated by this naive baseline cannot synthesize satisfying images quantitatively and qualitatively.
Fortunately, there is a windfall from these experiments. The PTQ calibration helps more when the time-step approaches the real image . There is an intuitive explanation for this observation. In the denoising process, with decreasing, the distribution of outputs of network is similar to real images’ distribution, which is a more significant phase in the image generation process.
Observation 3: Instead of a set of samples generated at the same time-step, calibration samples should be generated with varying time-steps.
Since our calibration dataset is collected for a multi-time-step scenario, while the common methods are proposed for single-time-step scenarios. We hypothesize that the calibration dataset for diffusion models should contain the samples with various time-steps, i.e., the calibration set should reflect the discrepancy of sample w.r.t.time-step. A straightforward way to test this hypothesis is to generate a set of uniformly sampled over the range of time-steps, i.e.,
where is a uniform distribution between and , is the size of calibration set, and is the number of time-steps in denoising process. Then given a Gaussion noise and , we utilize the diffusion model with the full-precision noise estimation network, in Eq. 5 to generate a . Finally, we get the calibration set, . Calibration samples can thus cover a wide range of time steps. We testify the effectiveness of this collection method, and present the results in Tab. 3. The result validates our hypothesis that calibration samples should reflect the time-step discrepancy.
3.2 Normally Distributed Time-step Calibration
Based on the above-demonstrated calibration baselines and observations, we desire the calibration samples: (1) generated by the denoising process (from noise ) with the full-precision diffusion model; (2) relatively close to , far away from ; (3) covered by various time-steps. Note that (2) and (3) are a pair of trade-off conditions, which can not be satisfied simultaneously.
Considering all the conditions, we propose a DM-specific calibration set collection method, termed as Normally Distributed Time-step Calibration (NDTC). In this method, the calibration set are generated by the denoising process (for condition 1), where time-step are sampled from a skew Normal distribution (for balancing conditions 2 & 3). Specifically, we first generate a set of sampled following skew normal distribution over the time-step range (satisfying condition 3), i.e.,
where is a normal distribution with mean and standard deviation , is the size of calibration set, and is the number of time-steps in denoising process. As is less than or equal to the median of time-step, (satisfying condition 2). Then given a Gaussian noise and , we utilize the diffusion model with the full-precision noise estimation network, in Eq. 5 to generate a (satisfying condition 1). The above-mentioned process of sampling time-steps is presented in Fig. 5. Finally, we get the calibration set, . The detailed collection algorithm is presented in Alg. 1.
The effectiveness of NDTC is assessed by comparing it to the mentioned PTQ baselines and full-precision DMs. The results are presented in Tab. 3 and Fig. 6.
4 Exploration on Parameter Calibration
When the calibration samples are collected, the third step is selecting quantization parameters for tensors in the diffusion model. In this section, we explore the metric to calibrate the tensors. As shown in Table 4, the MSE is better than the L1 distance, cosine distance, and KL divergence. Therefore, we take MSE as the metric for quantizing the diffusion model.
More Experiments
We select the diffusion models that generating CIFAR10 images or ImageNet down-sampled images. We experiment on both DDPM (4000 steps) and DDIM (100 and 250 steps) to generate the images.
We use the proposed method in Section 3.3 to generate 1024 calibration samples. And we quantize the network to 8-bit. Then we sample 10,000 images for evaluation. The results are listed in Table 5. Note that the number of samples that we generate is only 10,000 (50,000 in several papers) in order to efficiently compare our methods with other baselines quantitatively. Thus some reported results in this paper are slightly different from the results in the original papers. There is an exciting result in these experiments. In the setting of using DDPM to generate images with the size of , the 8-bit DDPM quantized by our method outperforms the full-precision DDPM. As discussed in Sec. 1, there are two factors slowing down the denoising process: i) lengthy iterations for sampling images from noise, and ii) a cumbersome network for estimating noise in each iteration. The successes of previous DM acceleration methods validate the existence of model redundancy from the perspective of iteration length. With this exciting result, we uncover the redundancy from a previously unknown perspective, in which the noise estimation network is also redundant.
Conclusion
Two orthogonal factors slow down the denoising process: i) lengthy iterations for sampling images from noise, and ii) a cumbersome network for estimating noise in each iteration. Different from mainstream DM acceleration works focusing on the former, our work digs into the latter. In this paper, we propose Post-Training Quantization for Diffusion Models (PTQ4DM), in which a pre-trained diffusion model can be directly quantized into 8 bits without experiencing a significant degradation in performance. Importantly, our method can be added to other fast-sampling methods, such as DDIM .
References
Appendix
Notably, we have carefully chosen the nonuniform distribution for sampling timestep (Eq. 15). Specifically, except for the normal distribution in the paper, we also consider the Poisson and exponential distributions. Also, a series of hyperparameter selection experiments are conducted. More details are presented in Tab. 6.
2 Actual Acceleration
We test the latency(ms) of the original network (provided checkpoint) and the quantized network on Nvidia RTX A6000 GPU. The results in Table 7 show that the 8-bit quantization achieves about 2x speedup. The speedup can be more significant on NPU.