Post-training Quantization on Diffusion Models

Yuzhang Shang, Zhihang Yuan, Bin Xie, Bingzhe Wu, Yan Yan

Introduction

Recently, denoising diffusion (also dubbed score-based) generative models have achieved phenomenal success in various generative tasks, such as images , audio , video , and graphs . Besides these fundamental tasks, their flexibility of implementation on downstream tasks is also attractive, e.g., they are effectively introduced for super-resolution , inpainting , and image-to-image translation . Diffusion models (DMs) have achieved superior performances on most of these tasks and applications, both concerning quality and diversity, compared with historically SoTA Generative Adversarial Networks (GANs) .

A diffusion process transforms real data gradually into Gaussian noise, and then the process is reversed to generate real data from Gaussian noise (denoising process) . Particularly, the denoising process requires iterating the noise estimation (also known as a score function ) via a cumbersome neural network over thousands of time-steps. While it has a compelling quantity of images, its long iterative process and high inference cost for generating samples make it undesirable. Thus, increasing the speed of this generation process is now an active area of research . To accelerate diffusion models, researchers propose several approaches, which mainly focus on sample trajectory learning for faster sampling strategies. For example, Chen et al. and San-Roman et al. propose faster step size schedules for VP diffusions that still yield relatively good quality/diversity metrics; Song et al. adopt implicit phases in the denoising process; Bao et al. and Lu et al. derive analytical approximations to simplify the generation process.

Our study suggests that two orthogonal factors slow down the denoising process: i) lengthy iterations for sampling images from noise, and ii) a cumbersome network for estimating noise in each iteration. Previously DM acceleration methods only focus on the former , but overlook the latter. From the perspective of network compression, many popular network quantization and pruning methods follow a simple pipeline: training the original model and then fine-tuning the quantized/pruned compressed model . Particularly, this training-aware compression pipeline requires a full training dataset and many computation resources to perform end-to-end backpropagation. For DMs, however, 1) training data are not always ready-to-use due to privacy and commercial concerns; 2) the training process is extremely expensive. For example, there is no access to the training data for the industry-developed text-to-image models Dall·E2 and Imagen . Even if one can access their datasets, fine-tuning them also consumes hundreds of thousands of GPU hours. Those two obstacles make training-aware compression not suitable for DMs.

Training-free network compression techniques are what we need for DM acceleration. Therefore, we propose to introduce post-training quantization (PTQ) into DM acceleration. In a training-free manner, PTQ can not only speed up the computation of the denoising process but also reduce the resources to store the diffusion model weight, which is required in DM acceleration. Although PTQ has many attractive benefits, its implementation in DMs remains challenging. The main reason is that the structure of DMs is hugely different from previously PTQ-implemented structures (e.g., CNN and ViT for image recognition). Specifically, the output distributions of noise estimation networks change with time-step, making previous PTQ methods fail in DMs since they are designed for single-time-step scenarios.

This study attempts to answer the following fundamental question: How does the design of the core ingredients of the PTQ for the DMs process (e.g., quantized operation selection, calibration set collection, and calibration metric) affect the final performance of the quantized diffusion models? To this end, we analyze the PTQ and DMs individually and correlatedly. We find that simple generalizations of previous PTQ methods to DMs lead to huge performance drops due to output distribution discrepancies w.r.t.time-step in the denoising process. In other words, noise estimation networks rely on time-step, which makes their output distributions change with time-step. This means that a key module of the previous PTQ calibration, cannot be used in our case. Based on the above observations, we devise a DM-specific calibration method, termed Normally Distributed Time-step Calibration (NDTC), which first samples a set of time-steps from a skew normal distribution, and then generates calibration samples in terms of sampled time-steps by the denoising process. In this way, the time-step discrepancy in the calibration set is enhanced, which improves the performance of PTQ4DM. Finally, we propose a novel DM acceleration method, Post-Training for Diffusion Models (PTQ4DM) via incorporating all the explorations.

Overall, the contributions of this paper are three-fold: (i) To accelerate denoising diffusion models, we introduce PTQ into DM acceleration where noise estimation networks are directly quantized in a post-training manner. To the best of our knowledge, this is the first work to investigate diffusion model acceleration from the perspective of training-free network compression. (ii) After all-inclusively investigations of PTQ and DMs, we observe the performance drop induced by PTQ for DMs can be attributed to the discrepancy of output distributions in various time-steps. Targeting this observation, we explore PTQ from different aspects and propose PTQ4DM. (iii) Experimentally, PTQ4DM can quantize the pre-trained diffusion models to 8-bit without significant performance loss for the first time. Importantly, PTQ4DM can serve as a plug-and-play module for other SoTA DM acceleration methods, as shown in Fig. 1.

Related Work

Due to the long iterative process in conjunction with the high cost of denoising via networks, diffusion models cannot be widely implemented. To accelerate the diffusion probabilistic models (DMs), previous works pursue finding shorter sampling trajectories while maintaining the DM performance. Chen et al. introduce grid search and find an effective trajectory with only six time-steps. However, the grid search approach can not be generalized into very long trajectories subject to its exponentially growing time complexity. Watson et al. model the trajectory searching as a dynamic programming problem. Song et al. construct a class of non-Markovian diffusion processes that lead to the same training objective, but whose reverse process can be much faster to sample from. As for DMs with continuous timesteps (i.e., score-based perspective ), Song et al. formulate the DM in of form of an ordinary differential equation (ODE), and improve sampling efficiency via utilizing faster ODE solver. Jolicoeur-Martineau et al. introduce an advanced SDE solver to accelerate the reverse process via an adaptively larger sampling rate. Bao et al. estimate the variance and KL divergence using the Monte Carlo method and a pretrained score-based model with derived analytic forms, which are simplified from the score-function. In addition to those training-free methods, Luhman & Luhman compress the reverse denoising process into a single-step model; San-Roman dynamically adjust the trajectory during inference. Nevertheless, implementing those methods requires additional training after obtaining a pretrained DM, which makes them less desirable in most situations. In summary, all those DM acceleration methods can be categorized into finding effective sampling trajectories.

However, we show in this paper that, in addition to finding short sampling trajectories, diffusion models can be further accelerated through network compression for each noise estimation iteration. Note that our method PTQ4DM is an orthogonal path with those above-mentioned fast sampling methods, which means it can be deployed as a plug-and-play module for those methods. To the best of our knowledge, our work is the first study on quantizing diffusion models in a post-training manner.

2 Post-training Quantization

Quantization is one of the most effective ways to compress a neural network. There are two types of quantization methods: Quantization-aware training (QAT) and Post-training quantization (PTQ). QAT considers the quantization in the network training phase. While PTQ quantizes the network after training. As PTQ consumes much less time and computation resources, it is widely used in network deployment.

Most of the work of PTQ is to set the quantization parameters for weights and activcations in each layer. Take uniform quantization as an example, the quantization parameters include scaling factor ss and zero point zz. A floating-point value xx is quantized to integer value xintx_{i}nt according to the parameters:

The clamp function clip the rounded value ⌊xs⌉−z\lfloor\frac{x}{s}\rceil-z to the range of [pmin,pmax][p_{min},p_{max}]. In order to set quantization parameters for the weight tensor and the activation tensor in a layer, a simple but effective way is to select the quantization parameters that minimize the MSE of the tensors before and after quantization . Other metrics, such as L1 distance, cosine distance, and KL divergence, can also be used to evaluate the distance of the tensors before and after quantization .

In order to calculate the activations in the network, a small number of calibration samples should be used as input in PTQ. The selected quantization parameters are dependent with the selection of these calibration samples. demonstrate the effect of the number of the calibration samples. Zero-shot quantization (ZSQ) is a special case of PTQ. ZSQ generates the calibration dataset according to information recorded in the network, such as the mean and var in batch normalization layer. They generate the input sample by gradient descent method to make the distribution of the activations in network similar to the distribution of real samples. The image generation process from noise in diffusion model only uses the network inference, which is quite different from previous ZSQ methods.

PTQ on Diffusion Models

Diffusion Models. The diffusion probabilistic model (DPM) is initially introduced by Sohl-Dickstein et al. , where the DPM is trained by optimizing the variational bound LVLBL_{\text{VLB}}. Here, we briefly review the diffusion model to illustrate the difference from traditional models. Here, we briefly review the diffusion model, especially its lengthy diffusion and denoising process. We highlight that those properties make it difficult to simply generalize common PTQ methods into diffusion models simply in Sec. 3.2.

Given a real data distribution x0∼q(x0)x_{0}\sim q(x_{0}), we define the diffusion process that gradually adds a small amount of isotropic Gaussian noise with a variance schedule β1,...,βT∈(0,1)\beta_{1},...,\beta_{T}\in(0,1) to produce a sequence of latent x1,...,xTx_{1},...,x_{T}, which is fixed to a Markov chain. When TT is sufficiently large T∼∞T\sim\infty and a well-behaved schedule of βt\beta_{t}, xTx_{T} is equivalent to an isotropic Gaussian distribution.

A notable property of the diffusion process admits us to sample xtx_{t} at an arbitrary timestep tt via directly conditioned on the input x0x_{0}. Let αt=1−βt\alpha_{t}=1-\beta_{t} and αˉt=∏i=1Tαi\bar{\alpha}_{t}=\prod_{i=1}^{T}\alpha_{i}:

Since q(xt−1∣xt)q(x_{t-1}|x_{t}) depends on the data distribution q(x0)q(x_{0}), which is intractable. Therefore, we need to parameterize a neural network to approximate it:

We utilize the variational lower bound to optimize the negative log-likelihood. LVLB=L_{\text{VLB}}=

The objective function of the variational lower bound can be further rewritten to be a combination of several KL-divergence and entropy terms (more details in ).

L0L_{0} uses a separate discrete decoder derived from N(x0;μθ(x1,1),Σθ(x1,1))\mathcal{N}(\mathbf{x}_{0};\boldsymbol{\mu}_{\theta}(\mathbf{x}_{1},1),\boldsymbol{\Sigma}_{\theta}(\mathbf{x}_{1},1)). LTL_{T} does not depend on θ\theta, it is close to zero if q(xT∣x0)≈N(0,I)q(x_{T}|x_{0})\approx\mathcal{N}(0,I). The remain term Lt−1L_{t-1} is a KL-divergence to directly compare pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) to diffusion process posterior that is tractable when x0x_{0} is conditioned,

This is the training process of the diffusion model. After obtaining the well-trained noise estimation model pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) in Eq. 5, given a random noise, we can generate samples through the denoising process by iterative sampling xt−1\mathbf{x}_{t-1} from pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) until we receive x0\mathbf{x}_{0}. Detailed information can be found in the surveys . Since the iterative process for denoising from noise input to synthetic images is extremely long (e.g., pioneer work, DDPM requires 4000 steps for generating a sample from noise), as illustrated in Fig. 2 (left); and the networks for estimating the noise in each denoising iteration is very deep and complicated, as illustrated in Fig. 2 (right). The inference of the diffusion model is expensive.

Post-training Quantization takes a well-trained network and selects the quantization parameters for the weight tensor and activation tensor in each layer. We use the quantization parameters, scaling factor ss, and zero point zz to transform a tensor to the quantized tensorWe focus on uniform quantization since it is the most widely used.. One of the most widely used methods to select the parameters is to minimize the error caused by quantization. The quantization error LquantL_{quant} is formulated as:

where XsimX_{sim} is the de-quantized tensor, and Metric is the metric function to evaluate the distance of XsimX_{sim} and the full-precision tensor XfpX_{fp}. MSE, cosine distance, L1 distance, and KL divergence are commonly used metric functions. The quantization process can be formulated as:

We can directly quantize the weight to minimize the quantization error, but we cannot get the activation tensor and quantize it without input. In order to collect the full-precision activation tensor, a number of unlabeled input samples (calibration dataset) are used as input. The size of the calibration dataset (e.g., 128 randomly selected images) is much smaller than the training dataset.

In general, PTQ quantizes a network in three steps: (i) Select which operations in the network should be quantized and leave the other operations in full-precision. For example, some special functions such as softmax and GeLU often takes full-precision .Quantizing these operations will significantly increase the quantization error and they are not very computationally intensive; (ii) Collect the calibration samples. The distribution of the calibration samples should be as close as possible to the distribution of the real data to avoid over-fitting of quantization parameters on calibration samples; (iii) Use the proper method to select quantization parameters for weight tensors and activation tensors.

In the next sections, we will explore how to apply PTQ to the diffusion model step by step.

2 Exploration on Operation Selection

For the diffusion model, we will analyze the image generation process to determine which operations should be quantized. The diffusion model iteratively generate the xt−1\mathbf{x}_{t-1} from xt\mathbf{x}_{t}. At each timestep, the inputs of the network are xt\mathbf{x}_{t} and tt, and the outputs are the mean μ\boldsymbol{\mu} and variance Σ\boldsymbol{\Sigma}. Then xt−1\mathbf{x}_{t-1} is sampled from the distribution defined as Eq 5. As shown in Figure 2, the network in the diffusion model often takes UNet-like CNN architecture. The same as most previous PTQ methods, the computation-intensive convolution layers and fully-connected layers in the network should be quantized. The batch normalization can be folded into the convolution layer. The special functions such as SiLU and softmax are kept in full-precision.

There are two more questions for the diffusion model: 1. whether the network’s outputs, μ\boldsymbol{\mu} and Σ\boldsymbol{\Sigma}, can be quantized? 2. whether the sampled image xt−1\mathbf{x}_{t-1} can be quantized? To answer the two questions, we only quantize the operation generating μ\boldsymbol{\mu}, Σ\boldsymbol{\Sigma}, or xt−1\mathbf{x}_{t-1}. As shown in Table 1, we observe that they are not sensitive to quantization and we indicate that they can be quantized.

3 Exploration on Calibration Dataset

The second step is to collect the calibration samples for quantizing diffusion models. The calibration samples can be collected from the training dataset for quantizing other networks. However, the training dataset in the diffusion model is x0\mathbf{x}_{0}, which is not the network’s input. The real input is the generated samples xt\mathbf{x}_{t}. Should we use the generated samples in diffusion process or the generated samples in denoising process? At what time-step tt, should the generated samples be collected? This section will explore how to make a good calibration dataset.

By all-inclusively investigating several intuitive PTQ baselines, we obtain four meaningful observations (Sec. 3.3.1), which accordingly guide the design of our method (Sec. 3.3.2). Experimental results demonstrate that our method is efficient and effective. Through devised PTQ4DM calibration, the 8-bit post-training quantized diffusion model can perform at the same performance level as its full-precision counterpart, e.g., 8-bit diffusion model reaches 23.9 FID and 15.8 IS, while 32-bit one has 21.6 FID and 14.9 IS.

As discussed in Sec. 3.2, we desire the distribution of the collected calibration samples should be as close as possible to the distribution of the real data. In this way, the calibration set can supervise the quantization by minimizing the quantization error. Since previous works are implemented on single-time-step scenarios (e.g., CNN and ViT for image recognition and object detection) , they can directly collect samples from the real training dataset for quantizing networks. Due to the small size of the calibration dataset, its collection is extremely sensitive. If the distribution of the collected dataset is not representative of the real dataset, it can easily lead to overfitting for the calibration task.

We encounter more challenges when calibrating PTQ for DM. Since the inputs of the to-be-quantized network are the generated samples xt (t=0,1,⋯ ,T)\mathbf{x}_{t}~{}(t=0,1,\cdots,T), in which TT is a large number to maintain the diffusion process converging to isotropic Normal distribution. To quantize the diffusion model, we are required to design a novel and effective calibration dataset collection method in this particular multi-time-step scenario. We start by investigating both PTQ calibration and DMs, and then obtain the following instructive observations.

Observation 0: Distributions of activations changes along with time-step changing.

To understand the output distribution change of diffusion models, we investigate the activation distribution with respect to time-step. We would like to analyze the output distribution at different time-step, for example, given t1=0.1Tt_{1}=0.1T and t2=0.9Tt_{2}=0.9T, the output activation distributions of pθ(xt1−1∣xt1)p_{\theta}(\mathbf{x}_{{t_{1}}-1}|\mathbf{x}_{t_{1}}) and pθ(xt2−1∣xt2)p_{\theta}(\mathbf{x}_{{t_{2}}-1}|\mathbf{x}_{t_{2}}). Theoretically, if the distribution changes w.r.t.time-step, it would be difficult to implement previous PTQ calibration methods, as they are proposed for temporally-invariant calibration . We first analyze the overall activation distributions of the noise estimation network via boxplot as did, and then we take a closer look at the layer-wise distributions via histogram. The results are shown in Fig. 3. We can observe that at different time-steps, the corresponding activation distributions have large discrepancies, which makes previous PTQ calibration methods inapplicable for multi-time-step models (i.e., diffusion models).

Observation 1: Generated samples in the denoising process are more constructive for calibration.

In general, there are two directions to generate samples for PTQ calibration in diffusion: raw images as input for diffusion process, and noise as input for denoising process. Previous PTQ methods use raw images, as raw images can serve as ground truth, representing the training set’s distribution. We conduct a pair of comparison experiments, in which we separately collect two calibration sets with raw images for diffusion process and Gaussian noise for denoising process, and use these two sets to calibrate quantized models. Another similar intuitive baseline is to use the training samples in the diffusion process as calibration data. Specifically, we randomly generate a timestep tt for each image x0\mathbf{x}_{0}, and use Eq. 4 according to tt to generate xt\mathbf{x}_{t}. In other word, collect calibration samples in a “Image + Gaussian Noise” manner. We name this scheme as training-mimic baseline. The results are listed in Tab. 2. We find that the input noises for diffusion process are more constructive for calibrating quantized DMs.

Observation 2: Sample xt\mathbf{x}_{t} close to real image x0\mathbf{x}_{0} is more beneficial for calibration.

Based on the aforementioned observations, we establish a baseline of PTQ calibration for DM based on , in which the quantized diffusion models are calibrated with samples at time-step tt, i.e., a set of xt\mathbf{x}_{t}. We refer to this straight-forward approach as a naive PTQ-for-DM baseline. Specifically, given a set of Gaussion noise xT∼N(0,I)\mathbf{x}_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{I}), we use the diffusion model with the full-precision noise estimation network, pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) in Eq. 5 to generate a set of xt\mathbf{x}_{t} as calibration set. Then as described in Sec. 3.1, we use this collected set to calibrate our quantized noise estimation network, pθ′(xt−1∣xt)p_{\theta^{\prime}}(\mathbf{x}_{t-1}|\mathbf{x}_{t}), in which θ′\theta^{\prime} is the quantized parameters. We conduct a series of experiments with this calibration baseline in different time-steps, i.e., t=0.0T,0.2T,⋯ ,1.0Tt=0.0T,0.2T,\cdots,1.0T, where TT is the total denoising time-steps. The results are presented in Fig. 4. We can see that the 8-bit model calibrated by this naive baseline cannot synthesize satisfying images quantitatively and qualitatively.

Fortunately, there is a windfall from these experiments. The PTQ calibration helps more when the time-step tt approaches the real image x0\mathbf{x}_{0}. There is an intuitive explanation for this observation. In the denoising process, with tt decreasing, the distribution of outputs of network pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) is similar to real images’ distribution, which is a more significant phase in the image generation process.

Observation 3: Instead of a set of samples generated at the same time-step, calibration samples should be generated with varying time-steps.

Since our calibration dataset is collected for a multi-time-step scenario, while the common methods are proposed for single-time-step scenarios. We hypothesize that the calibration dataset for diffusion models should contain the samples with various time-steps, i.e., the calibration set should reflect the discrepancy of sample w.r.t.time-step. A straightforward way to test this hypothesis is to generate a set of uniformly sampled tt over the range of time-steps, i.e.,

where U(0,T)U(0,T) is a uniform distribution between and TT, NN is the size of calibration set, and TT is the number of time-steps in denoising process. Then given a Gaussion noise xT∼N(0,I)\mathbf{x}_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) and tit_{i}, we utilize the diffusion model with the full-precision noise estimation network, pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) in Eq. 5 to generate a xti\mathbf{x}_{t_{i}}. Finally, we get the calibration set, C={xti}i=1N\mathcal{C}=\{\mathbf{x}_{t_{i}}\}_{i=1}^{N}. Calibration samples can thus cover a wide range of time steps. We testify the effectiveness of this collection method, and present the results in Tab. 3. The result validates our hypothesis that calibration samples should reflect the time-step discrepancy.

3.2 Normally Distributed Time-step Calibration

Based on the above-demonstrated calibration baselines and observations, we desire the calibration samples: (1) generated by the denoising process (from noise xT\mathbf{x}_{T}) with the full-precision diffusion model; (2) relatively close to x0\mathbf{x}_{0}, far away from xT\mathbf{x}_{T}; (3) covered by various time-steps. Note that (2) and (3) are a pair of trade-off conditions, which can not be satisfied simultaneously.

Considering all the conditions, we propose a DM-specific calibration set collection method, termed as Normally Distributed Time-step Calibration (NDTC). In this method, the calibration set {xti}\{x_{t_{i}}\} are generated by the denoising process (for condition 1), where time-step tit_{i} are sampled from a skew Normal distribution (for balancing conditions 2 & 3). Specifically, we first generate a set of sampled {ti}\{t_{i}\} following skew normal distribution over the time-step range (satisfying condition 3), i.e.,

where N(μ,T2)\mathcal{N}(\mu,\frac{T}{2}) is a normal distribution with mean μ≤T2\mu\leq\frac{T}{2} and standard deviation T2\sqrt{\frac{T}{2}}, NN is the size of calibration set, and TT is the number of time-steps in denoising process. As μ\mu is less than or equal to the median of time-step, T2\frac{T}{2} (satisfying condition 2). Then given a Gaussian noise xT∼N(0,I)\mathbf{x}_{T}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) and tit_{i}, we utilize the diffusion model with the full-precision noise estimation network, pθ(xt−1∣xt)p_{\theta}(\mathbf{x}_{t-1}|\mathbf{x}_{t}) in Eq. 5 to generate a xti\mathbf{x}_{t_{i}} (satisfying condition 1). The above-mentioned process of sampling time-steps is presented in Fig. 5. Finally, we get the calibration set, C={xti}i=1N\mathcal{C}=\{\mathbf{x}_{t_{i}}\}_{i=1}^{N}. The detailed collection algorithm is presented in Alg. 1.

The effectiveness of NDTC is assessed by comparing it to the mentioned PTQ baselines and full-precision DMs. The results are presented in Tab. 3 and Fig. 6.

4 Exploration on Parameter Calibration

When the calibration samples are collected, the third step is selecting quantization parameters for tensors in the diffusion model. In this section, we explore the metric to calibrate the tensors. As shown in Table 4, the MSE is better than the L1 distance, cosine distance, and KL divergence. Therefore, we take MSE as the metric for quantizing the diffusion model.

More Experiments

We select the diffusion models that generating CIFAR10 32×3232\times 32 images or ImageNet down-sampled 64×6464\times 64 images. We experiment on both DDPM (4000 steps) and DDIM (100 and 250 steps) to generate the images.

We use the proposed method in Section 3.3 to generate 1024 calibration samples. And we quantize the network to 8-bit. Then we sample 10,000 images for evaluation. The results are listed in Table 5. Note that the number of samples that we generate is only 10,000 (50,000 in several papers) in order to efficiently compare our methods with other baselines quantitatively. Thus some reported results in this paper are slightly different from the results in the original papers. There is an exciting result in these experiments. In the setting of using DDPM to generate images with the size of 32×3232\times 32, the 8-bit DDPM quantized by our method outperforms the full-precision DDPM. As discussed in Sec. 1, there are two factors slowing down the denoising process: i) lengthy iterations for sampling images from noise, and ii) a cumbersome network for estimating noise in each iteration. The successes of previous DM acceleration methods validate the existence of model redundancy from the perspective of iteration length. With this exciting result, we uncover the redundancy from a previously unknown perspective, in which the noise estimation network is also redundant.

Conclusion

Two orthogonal factors slow down the denoising process: i) lengthy iterations for sampling images from noise, and ii) a cumbersome network for estimating noise in each iteration. Different from mainstream DM acceleration works focusing on the former, our work digs into the latter. In this paper, we propose Post-Training Quantization for Diffusion Models (PTQ4DM), in which a pre-trained diffusion model can be directly quantized into 8 bits without experiencing a significant degradation in performance. Importantly, our method can be added to other fast-sampling methods, such as DDIM .

References

Appendix

Notably, we have carefully chosen the nonuniform distribution for sampling timestep tt (Eq. 15). Specifically, except for the normal distribution in the paper, we also consider the Poisson and exponential distributions. Also, a series of hyperparameter selection experiments are conducted. More details are presented in Tab. 6.

2 Actual Acceleration

We test the latency(ms) of the original network (provided checkpoint) and the quantized network on Nvidia RTX A6000 GPU. The results in Table 7 show that the 8-bit quantization achieves about 2x speedup. The speedup can be more significant on NPU.