PTQ4ViT: Post-training quantization for vision transformers with twin uniform quantization

Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, Guangyu Sun

Introduction

The self-attention module is the basic building block of the transformer to capture global information . Inspired by the success of transformers on natural language processing (NLP) tasks, researchers have brought the self-attention module into computer vision . They replaced the convolution layers in convolutional neural networks (CNNs) with self-attention modules and they called these networks vision transformers. Vision transformers are comparable to CNNs on many computer vision tasks and have great potential to be deployed on various applications .

However, both the CNN and the vision transformer are computationally intensive and consume much energy. The larger and larger scales of neural networks block their deployment on various hardware devices, such as mobile phones and IoT devices, and increase carbon emissions. It is required to compress these neural networks. Quantization is one of the most effective ways to compress neural networks . The floating-point values are quantized to integers with a low bit-width, reducing the memory consumption and the computation cost.

There are two types of quantization methods, quantization-aware training (QAT) and post-training quantization (PTQ) . Although QAT can generate the quantized network with a lower accuracy drop, the training of the network requires a training dataset, a long optimization time, and the tuning of hyper-parameters. Therefore, QAT is impractical when the training dataset is not available or rapid deployment is required. While PTQ quantizes the network with unlabeled calibration images after training, which enables fast quantization and deployment.

Although PTQ has achieved great success on CNNs, directly bringing it to vision transformer results in more than 1% accuracy drop even with 8-bit quantization . Therefore, we analyze the problems of quantization on vision transformers. We collect the distribution of activation values in the vision transformer and observe there are some special distributions. 1) The values after softmax have a very unbalanced distribution in $$, where most of them are very close to zero. Although the number of large values is very small, they mean high attention between two patches, which is of vital importance in the attention mechanism. This requires a large scaling factor to make the quantization range cover the large value. However, a big scaling factor quantizes the small values to zero, resulting in a large quantization error. 2) The values after the GELU function have an asymmetrical distribution, where the positive values have a large distribution range while the negative values have a very small distribution range. It’s difficult to well quantify both the positive values and negative values with uniform quantization. Therefore, we propose the twin uniform quantization, which separately quantifies the values in two ranges. To enable its efficient processing on hardware devices, we design a data format and constrain the scaling factors of the two ranges.

The second problem is that the metric to determine the optimal scaling factor is not accurate on vision transformers. There are various metrics in previous PTQ methods, including MSE, cosine distance, and Pearson correlation coefficient between the layer outputs before and after quantization. However, we observe they are inaccurate to evaluate different scaling factor candidates because only the local information is used. Therefore, we propose to use the Hessian guided metric to determine the quantization parameters, which is more accurate. The proposed methods are demonstrated in Fig. 1.

We develop a post-training quantization framework for vision transformers using twin uniform quantization, PTQ4ViT. Code is in https://github.com/hahnyuan/PTQ4ViT. Experiments show the quantized vision transformers (ViT, DeiT, and Swin) achieve near-lossless prediction accuracy (less than 0.5% drop at 8-bit quantization) on the ImageNet classification task.

We find the problems in PTQ on vision transformers are special distributions of post-softmax and post-GELU activations and the inaccurate metric.

We propose the twin uniform quantization to handle the special distributions, which can be efficiently processed on existing hardware devices including CPU and GPU.

We propose to use the Hessian guided metric to determine the optimal scaling factors, which replaces the inaccurate metrics.

The quantized networks achieve near-lossless prediction accuracy, making PTQ acceptable on vision transformers.

Background and Related Work

In the last few years, convolution neural networks (CNNs) have achieved great success in computer vision. The convolution layer is a fundamental component of CNNs to extract features using local information. Recently, the position of CNNs in computer vision is challenged by vision transformers, which take the self-attention modules to make use of the global information. DETR is the first work to replace the object detection head with a transformer, which directly regresses the bounding boxes and achieves comparable results with the CNN-based head. ViT is the first architecture that replaces all convolution layers, which achieves better results on image classification tasks. Following ViT, various vision transformer architectures have been proposed to boost performance . Vision transformers have been successfully applied to downstream tasks . They have great potential for computer vision tasks .

The input of a transformer is a sequence of vectors. An image is divided into several patches and a linear projection layer is used to project each patch to a vector. These vectors form the input sequence of the vision transformer. We denote these vectors as X∈RN×DX\in R^{N\times D}, where NN is the number of patches and DD is the hidden size, which is the size of the vector after linear projection.

A vision transformer contains some blocks. As shown in Fig. 1, each block is composed of a multi-head self-attention module (MSA) and a multi-layer perceptron (MLP). MSA generates the attention between different patches to extract features with global information. Typical MLP contains two fully-connected layers (FC) and the GELU activation function is used after the first layer. The input sequence is first fed into each self-attention head of MSA. In each head, the sequence is linearly projected to three matrices, query Q=XWQQ=XW^{Q}, key K=XWKK=XW^{K}, and value V=XWVV=XW^{V}. Then, matrix multiplication QKTQK^{T} calculates the attention scores between patches. The softmax function is used to normalize these scores to attention probability PP. The output of the head is matrix multiplication PVPV. The process is formulated as Eq. 1:

where dd is the hidden size of head. The outputs of multiple heads are concatenated together as the output of MSA.

Vision transformers have a large amount of memory, computation, and energy consumption, which hinders their deployment in real-world applications. Researchers have proposed a lot of methods to compress vision transformers, such as patch pruning , knowledge distillation ,and quantization .

2 Quantization

Network quantization is one of the most effective methods to compress neural networks. The weight values and activation values are transformed from floating-point to integer with lower bit-width, which significantly decreases the memory consumption, data movement, and energy consumption. The uniform symmetric quantization is the most widely used method, which projects a floating-point value xx to a kk-bit integer value xqx_{q} with a scaling factor Δ\Delta:

where round projects a value to an integer and clamp constrains the output in the range that kk-bit integer can represent. We propose the twin uniform quantization, which separately quantifies the values in two ranges. also uses multiple quantization ranges. However, their method targets CNN and is not suitable for ViT. They use an extra bit to represent which range is used, taking 12.5% more storage than our method. Moreover, they use FP32 computation to align the two ranges, which is not efficient. Our method uses the shift operation, avoiding the format transformation and extra FP32 multiplication and FP32 addition.

There are two types of quantization methods, quantization-aware training (QAT) and post-training quantization (PTQ) . QAT methods combine quantization with network training. It optimizes the quantization parameters to minimize the task loss on a labeled training dataset. QAT can be used to quantize transformers . Q-BERT uses the Hessian spectrum to evaluate the sensitivity of the different tensors for mixed-precision, achieving 3-bit weight and 8-bit activation quantization. Although QAT achieves lower bit-width, it requires a training dataset, a long quantization time, and hyper-parameter tuning. PTQ methods quantize networks with a small number of unlabeled images, which is significantly faster than QAT and doesn’t require any labeled dataset. PTQ methods should determine the scaling factors Δ\Delta of activations and weights for each layer. Choukroun et al. proposed to minimize the mean square error (MSE) between the tensors before and after quantization. EasyQuant uses the cosine distance to improve the quantization performance on CNN. Recently, Liu et al. first proposed a PTQ method to quantize the vision transformer. Pearson correlation coefficient and ranking loss are used as the metrics to determine the scaling factors. However, these metrics are inaccurate to evaluate different scaling factor candidates because only the local information is used.

Method

In this section, we will first introduce a base PTQ method for vision transformers. Then, we will analyze the problems of quantization using the base PTQ and propose methods to address the problems. Finally, we will introduce our post-training quantization framework, PTQ4ViT.

Matrix multiplication is used in the fully-connected layer and the computation of QKTQK^{T} and PVPV, which is the main operation in vision transformers. In this paper, we formulate it as O=ABO=AB and we will focus on its quantization. AA and BB are quantized to kk-bit using the symmetric uniform quantization with scaling factors ΔA\Delta_{A} and ΔB\Delta_{B}. According to Eq. 2, we have Aq=Ψk(A,ΔA)A_{q}=\Psi_{k}(A,\Delta_{A}) and Bq=Ψk(B,ΔB)B_{q}=\Psi_{k}(B,\Delta_{B}). In base PTQ, the distance of the output before and after quantization is used as metric to determine the scaling factors, which is formulated as:

where O^\hat{O} is the output of the matrix multiplication after quantization O^=ΔAΔBAqBq\hat{O}=\Delta_{A}\Delta_{B}A_{q}B_{q}.

The same as , we use cosine distance as the metric to calculate the distance. We make the search spaces of ΔA{\Delta_{A}} and ΔB{\Delta_{B}} by linearly dividing [αAmax2k−1,βAmax2k−1][\alpha\frac{A_{max}}{2^{k-1}},\beta\frac{A_{max}}{2^{k-1}}] and [αBmax2k−1,βBmax2k−1][\alpha\frac{B_{max}}{2^{k-1}},\beta\frac{B_{max}}{2^{k-1}}] to nn candidates, respectively. AmaxA_{max} and BmaxB_{max} are the maximum absolute value of AA and BB. α\alpha and β\beta are two parameters to control the search range. We alternatively search for the optimal scaling factors ΔA∗{\Delta_{A}^{*}} and ΔB∗{\Delta_{B}^{*}} in the search space. Firstly, ΔB\Delta_{B} is fixed, and we search for the optimal ΔA\Delta_{A} to minimize distance(O,O^)\text{distance}(O,\hat{O}). Secondly, ΔA\Delta_{A} is fixed, and we search for the optimal ΔB\Delta_{B} to minimize distance(O,O^)\text{distance}(O,\hat{O}). ΔA\Delta_{A} and ΔB\Delta_{B} are alternately optimized for several rounds.

The values of AA and BB are collected using unlabeled calibration images. We search for the optimal scaling factors of activation or weight layer-by-layer. However, the base PTQ results in more than 1% accuracy drop on quantized vision transformer in our experiments.

2 Twin Uniform Quantization

The activation values in CNNs are usually considered Gaussian distributed. Therefore, most PTQ quantization methods are based on this assumption to determine the scaling factor. However, we observe the distributions of post-softmax values and post-GELU values are quite special as shown in Fig. 3. Specifically, (1) The distribution of activations after softmax is very unbalanced, in which most values are very close to zero and only a few values are close to one. (2) The values after the GELU function have a highly asymmetric distribution, in which the unbounded positive values are large while the negative values have a very small distribution range. As shown in Fig. 3, we demonstrate the quantization points of the uniform quantization using different scaling factors.

For the values after softmax, a large value means that there is a high correlation between the two patches, which is important in the self-attention mechanism. A larger scaling factor can reduce the quantization error of these large values, which causes smaller values to be quantized to zero. While a small scaling factor makes the large values quantized to small values, which significantly decreases the intensity of attention between two patches. For the values after GELU, it is difficult to quantify both positive and negative values well with symmetric uniform quantization. Non-uniform quantization can be used to solve the problem. It can set the quantization points according to the distribution, ensuring the overall quantization error is small. However, most hardware devices cannot efficiently process the non-uniform quantized values. Acceleration can be achieved only on specially designed hardware.

We propose the twin uniform quantization, which can be efficiently processed on existing hardware devices including CPUs and GPUs. As shown in Fig. 4, twin uniform quantization has two quantization ranges, R1 and R2, which are controlled by two scaling factors ΔR1\Delta_{\text{R1}} and ΔR2\Delta_{\text{R2}}, respectively. The kk-bit twin uniform quantization is be formulated as:

For values after softmax, the values in R1 = [0,2k−1ΔR1s)[0,2^{k-1}\Delta_{\text{R1}}^{s}) can be well quantified by using a small ΔR1s\Delta_{\text{R1}}^{s}. To avoid the effect of calibration dataset, we keeps ΔR2s\Delta_{\text{R2}}^{s} fixed to 1/2k−11/2^{k-1}. Therefore, R2 = $cancoverthewholerange,andlargevaluescanbewellquantifiedinR2.ForactivationvaluesafterGELU,negativevaluesarelocatedinR1=can cover the whole range, and large values can be well quantified in R2. For activation values after GELU, negative values are located in R1 =[-2^{k-1}\Delta_{\text{R1}}^{g},0]andpositivevaluesarelocatedinR2=and positive values are located in R2=[0,2^{k-1}\Delta_{\text{R2}}^{g}].Wealsokeep. We also keep\Delta_{\text{R1}}^{g}fixedtomakeR1justcovertheentirerangeofnegativenumbers.Sincedifferentquantizationparametersareusedforpositiveandnegativevaluesrespectively,thequantizationerrorcanbeeffectivelyreduced.Whencalibratingthenetwork,wesearchfortheoptimalfixed to make R1 just cover the entire range of negative numbers. Since different quantization parameters are used for positive and negative values respectively, the quantization error can be effectively reduced. When calibrating the network, we search for the optimal\Delta_{\text{R1}}^{s}andand\Delta_{\text{R2}}^{g}$.

The uniform symmetric quantization uses the kk bit signed integer data format. It consists of one sign bit and k−1k-1 bits representing the quantity. In order to efficiently store the twin-uniform-quantized values, we design a new data format. The most significant bit is the range flag to represent which range is used (0 for R1, 1 for R2). The other k−1k-1 bits compose an unsigned number to represent the quantity. Because the sign of values in the same range is the same, the sign bit is removed.

Data in different ranges need to be multiplied and accumulated in matrix multiplication. In order to efficiently process with the twin-uniform-quantized values on CPUs or GPUs, we constrain the two ranges with ΔR2=2mΔR1\Delta_{\text{R2}}=2^{m}\Delta_{\text{R1}}, where mm is an unsigned integer. Assuming aqa_{q} is quantized in R1 and bqb_{q} is quantized in R2, the two values can be aligned:

We left shift bqb_{q} by mm bits, which is the same as multiplying the value by 2m2^{m}. The shift operation is very efficient on CPUs or GPUs. Without this constraint, multiplication is required to align the scaling factor, which is much more expensive than shift operations.

3 Hessian Guided Metric

Next, we will analyze the metrics to determine the scaling factors of each layer. Previous works greedily determine the scaling factors of inputs and weights layer by layer. They use various kinds of metrics, such as MSE and cosine distance, to measure the distance between the original and the quantized outputs. The change in the internal output is considered positively correlated with the task loss, so it is used to calculate the distance.

We plot the performance of different metrics in Fig. 5. We observe that MSE, cosine distance, and Pearson correlation coefficient are inaccurate compared with task loss (cross-entropy) on vision transformers. The optimal scaling factors based on them are not consistent with that based on task loss. For instance, on blocks.6.mlp.fc1:activation, they indicate that a scaling factor around 0.4Amax2k−10.4\frac{A_{max}}{2^{k-1}} is the optimal one, while the scaling factor around 0.75Amax2k−10.75\frac{A_{max}}{2^{k-1}} is the optimal according to the task loss. Using these metrics, we get sub-optimal scaling factors, causing the accuracy degradation. The distance between the last layer’s output before and after quantization can be more accurate in PTQ. However, using it to determine the scaling factors of internal layers is impractical because it requires executing the network many times to calculate the last layer’s output, which consumes too much time.

where OlO^{l} and Ol^\hat{O^{l}} are the outputs of the ll-th layer before and after quantization, respectively. As shown in Fig. 5, the optimal scaling factor indicated by Hessian guided metric is closer to that indicated by task loss (CE). Although it still has a gap with the task loss, Hessian guided metric significantly improves the performance. For instance, on blocks.6.mlp.fc1:activation, the optimal scaling factor indicated by Hessian guided metric has less influence on task loss than other metrics.

4 PTQ4ViT Framework

To achieve fast quantization and deployment, we develop an efficient post-training quantization framework for vision transformers, PTQ4ViT. Its flow is described in Algorithm 1. It supports the twin uniform quantization and Hessian guided metric. There are two quantization phases. 1) The first phase is to collect the output and the gradient of the output in each layer before quantization. The outputs of the ll-th layer OlO^{l} are calculated through forward propagation on the calibration dataset. The gradients ∂L∂O1l,…,∂L∂Oal\frac{\partial L}{\partial O^{l}_{1}},\dots,\frac{\partial L}{\partial O^{l}_{a}} are calculated through backward propagation. 2) The second phase is to search for the optimal scaling factors layer by layer. Different scaling factors in the search space are used to quantize the activation values and weight values in the ll-th layer. Then the output of the layer Ol^\hat{O^{l}} is calculated. We search for the optimal scaling factor Δ∗\Delta^{*} that minimizes Eq. 7.

In the first phase, we need to store OlO^{l} and ∂L∂Ol\frac{\partial L}{\partial O^{l}}, which consumes a lot of GPU memory. Therefore, we transfer these data to the main memory when they are generated. In the second phase, we transfer OlO^{l} and ∂L∂Ol\frac{\partial L}{\partial O^{l}} back to GPU memory and destroy them when the quantization of ll-th layer is finished. To make full use of the GPU parallelism, we calculate Ol^\hat{O^{l}} and the influence on loss for different scaling factors in batches.

Experiments

In this section, we first introduce the experimental settings. Then we will evaluate the proposed methods on different vision transformer architectures. At last, we will take an ablation study on the proposed methods.

For post-softmax quantization, the search space of ΔR1s\Delta_{\text{R1}}^{s} is [12k,12k+1,...,12k+10][\frac{1}{2^{k}},\frac{1}{2^{k+1}},...,\frac{1}{2^{k+10}}]. The search spaces of scaling factors for weight and other activations are the same as that of base PTQ (Sec. 3.1). We set  alpha=0\ alpha=0,  beta=1.2\ beta=1.2, and n=100n=100. The search round #Round\#Round is set to 33. We experiment on the ImageNet classification task . We randomly select 32 images from the training dataset as calibration images. The ViT models are provided by timm .

We quantize all the weights and inputs for the fully-connect layers including the first projection layer and the last prediction layer. We also quantize the two input matrices for the matrix multiplications in self-attention modules. We use different quantization parameters for different self-attention heads. The scaling factors for WQW^{Q}, WKW^{K}, and WVW^{V} are different. The same as , we don’t quantize softmax and normalization layers in vision transformers.

2 Results on ImageNet Classification Task

We choose different vision transformer architectures, including ViT , DeiT , and Swin . The results are demonstrated in Tab. 1. From this table, we observe that base PTQ results in more than 1% accuracy drop on some vision transformers even at the 8-bit quantization. PTQ4ViT achieves less than 0.5% accuracy drop with 8-bit quantization. For 6-bit quantization, base PTQ results in high accuracy drop (9.8% on average) while PTQ4ViT achieves a much smaller accuracy drop (2.1% on average).

We observe that the accuracy drop on Swin is not as significant as ViT and DeiT. The prediction accuracy drops are less than 0.15% on the four Swin transformers at 8-bit quantization. The reason may be that Swin computes the self-attention locally within non-overlapping windows. It uses a smaller number of patches to calculate the self-attention, reducing the unbalance after post-softmax values. We also observe that larger vision transformers are less sensitive to quantization. For instance, the accuracy drops of ViT-S/224/32, ViT-S/224, ViT-B/224, and ViT-B/384 are 0.41, 0.38, 0.29, and 0.17 at 8-bit quantization and 4.08, 2.75, 2.89, and 2.65 at 6-bit quantization, respectively. The reason may be that the larger networks have more weights and generate more activations, making them more robust to the perturbation caused by quantization.

Tab. 2 demonstrates the results of different PTQ methods. EasyQuant is a popular post-training method that alternatively searches for the optimal scaling factors of weight and activation. However, the accuracy drop is more than 3% at 8-bit quantization. Liu et al. proposed using the Pearson correlation coefficient and ranking loss are used as the metrics to determine the scaling factors, which increases the Top-1 accuracy. Since the sensitivity of different layers to quantization is not the same, they also use the mixed-precision technique, achieving good results at 4-bit quantization. At 8-bit quantization and 6-bit quantization, PTQ4ViT outperforms other methods, achieving more than 1% improvement in prediction accuracy on average. At 4-bit quantization, the performance of PTQ4ViT is not good. Although bias correction can improve the performance of PTQ4ViT, the result at 4-bit quantization is lower than the mixed-precision of Liu et al. This indicates that mixed-precision is important for quantization with lower bit-width.

3 Ablation Study

Next, we take ablation study on the effect of the proposed twin uniform quantization and Hessian guided metric. The experimental results are shown in Tab. 3. As we can see, the proposed methods improve the top-1 accuracy of quantized vision transformers. Specifically, using the Hessian guided metric alone can slightly improve the accuracy at 8-bit quantization, and it significantly improves the accuracy at 6-bit quantization. For instance, on ViT-S/224, the accuracy improvement is 0.46% at 8-bit while it is 6.96% at 6-bit. And using them together can further improve the accuracy.

Based on the Hessian guided metric, using the twin uniform quantization on post-softmax activation or post-GELU activation can improve the performance. We observe that using the twin uniform quantization without the Hessian guided metric significantly decreases the top-1 accuracy. For instance, the top-1 accuracy on ViT-S/224 achieves 81.00% with both Hessian guided metric and twin uniform quantization at 8-bit quantization, while it decreases to 79.25% without Hessian guided metric, which is even lower than basic PTQ with 80.47% top-1 accuracy. This is also evidence that the metric considering only the local information is inaccurate.

Conclusion

In this paper, we analyzed the problems of post-training quantization for vision transformers. We observed both the post-softmax activations and the post-GELU activations have special distributions. We also found that the common quantization metrics are inaccurate to determine the optimal scaling factor. To solve these problems, we proposed the twin uniform quantization and a Hessian-guided metric. They can decrease the quantization error and improve the prediction accuracy at a small cost. To enable the fast quantization of vision transformers, we developed an efficient framework, PTQ4ViT. The experiments demonstrated that we achieved near-lossless prediction accuracy on the ImageNet classification task, making PTQ acceptable for vision transformers.

This work is supported by National Key R&D Program of China (2020AAA0105200), NSF of China (61832020, 62032001, 92064006), Beijing Academy of Artificial Intelligence (BAAI), and 111 Project (B18001).

Appendix

One of the targets of PTQ4ViT is to quickly quantize a vision transformer. We have proposed to pre-compute the output and gradient of each layer and compute the influence of scaling factor candidates in batches to reduce the quantization time. As demonstrated in Tab. 4, PTQ4ViT can quantize most vision transformers in several minutes using 32 calibration images. Using #ims=128\#ims=128 significantly increases the quantization time. We observe the Top-1 accuracy varies slightly, demonstrating PTQ4ViT is not very sensitive to #ims\#ims.

2 Base PTQ

Base PTQ is a simple quantization strategy and serves as a benchmark for our experiments. Like PTQ4ViT, we quantize all weights and inputs for fully-connect layers (including the first projection layer and the last prediction layer), as well as all input matrices of matrix multiplication operations. For fully-connected layers, we use layerwise scaling factors ΔW\Delta_{W} for weight quantization and ΔX\Delta_{X} for input quantization; while for matrix multiplication operations, we use ΔA\Delta_{A} and ΔB\Delta_{B} for A’s quantization and B’s quantization respectively.

To get the best scaling factors, we apply a linear grid search on the search space. The same as EasyQuant and Liu et al. , we take hyper-parameters α=0.5\alpha=0.5, β=1.2\beta=1.2, search round #Round=1\#Round=1 and use cosine distance as the metric. Note that in PTQ4ViT, we change the hyper-parameters to α=0\alpha=0, β=1.2\beta=1.2 and search round #Round=3\#Round=3, which slightly improves the performance.

It should be noticed that Base PTQ adopts a parallel quantization paradigm, which makes it essentially different from sequential quantization paradigms such as EasyQuant . In sequential quantization, the input data of the current quantizing layer is generated with all previous layers quantizing weights and activations. While in parallel quantization, the input data of the current quantizing layer is simply the raw output of the previous layer.

In practice, we found sequential quantization on vision transformers suffers from significant accuracy degradation on small calibration datasets. While parallel quantization shows robustness on small calibration datasets. Therefore, we choose parallel quantization for both Base PTQ and PTQ4ViT.

3 Derivation of Hessian guided metric

Our goal is to introduce as small an increment on task loss L=CE(y^,y)L=CE(\hat{y},y) as possible, in which y^\hat{y} is the prediction of the quantized model and yy is the ground truth. In PTQ, we don’t have labels of input data yy, so we make a fair assumption that the prediction of floating-point network yFPy_{FP} is close to the ground truth yy. Therefore, we use the CE(y^,yFP)CE(\hat{y},y_{FP}) as a substitution of the task loss LL, which, for convenience, is still denoted as LL in the following.

where gˉ(W)\bar{g}^{(W)} is the gradients and Hˉ(W)\bar{H}^{(W)} is the Hessian matrix. Since the pretrained model has converged to a local optimum, the gradients gˉ(W)\bar{g}^{(W)} is close to zero, thus the first-order term could be ignored and we only consider the second-order term.

where O=WTX∈RmO=W^{T}X\in\mathit{R}^{m} is the output of the layer. Note that ∂2Ok∂wi∂wj=0\dfrac{\partial^{2}O_{k}}{\partial w_{i}\partial w_{j}}=0, the first term of Eq. 9 is zero. We denote JO(W)J_{O}(W) as the Jacobian matrix of OO w.r.t. weight WW. Then we have

Since weight’s perturbation ϵ\epsilon is relatively small, we have a first-order Taylor expansion that (O^−O)≈JO(W)ϵ(\hat{O}-O)\approx J_{O}(W)\epsilon, where O^=(W+ϵ)TX\hat{O}=(W+\epsilon)^{T}X. The second-order term in Eq. 8 could be written as

Next, we introduce how to compute the Hessian matrix of the layer’s output OO. Following Liu et al., we use the Fisher Information Matrix II of OO to substitute Hˉ(O)\bar{H}^{(O)}. For our probabilistic model p(Y;θ)p(Y;\theta) where θ\theta is the model’s parameters and YY is the random variable of predicted probability, we have:

Notice that II would equal to the expected Hessian Hˉ(O)\bar{H}^{(O)} if the model’s distribution matches the true data distribution, i.e. y=yFPy=y_{FP}. We have assumed y≈yFPy\approx y_{FP} in the above derivation, and therefore it is reasonable to replace Hˉ(O)\bar{H}^{(O)} with II.

Using the original Fisher Information Matrix, however, still requires an unrealistic amount of computation. So we only consider elements on the diagonal, which is diag((∂L∂O1)2,⋯ ,(∂L∂Om)2)\text{diag}((\dfrac{\partial L}{\partial O_{1}})^{2},\cdots,(\dfrac{\partial L}{\partial O_{m}})^{2}). This only requires the first-order gradient on output OO, which introduces a relatively small computation overhead. Therefore, the Hessian guided metric is:

Using this metric, we can search for the optimal scaling factor for weight. The optimization is formulated as:

In PTQ4ViT, we make a search space for scaling factors. Then we compute the influence on the output of the layer O^−O\hat{O}-O for each scaling factor. The optimal scaling factor can be selected according to Eq. 14. We assume that ∂L∂O\frac{\partial L}{\partial O} doesn’t change when the weight is quantized. This assumption enables the pre-computation of ∂L∂O\frac{\partial L}{\partial O}, significantly improving the quantization efficiency.

4 More Ablation Study

We supply more ablation studies for the hyper-parameters. It is enough to set the number of quantization intervals ≥\geq 20 (accuracy change <0.3%<0.3\%). It is enough to set the upper bound of m ≥\geq 15 (no accuracy change). The best settings of alpha and beta vary from different layers. It is appropriate to set α=0\alpha=0 and β=1/2k−1\beta=1/2^{k-1}, which has little impact on search efficiency. We observe that #Round\#Round has little impact on the prediction accuracy (accuracy change << 0.05% when #Round>1\#Round>1).

We randomly take 32 calibration images to quantize different models 20 times and we observe the fluctuation is not significant. The mean/std of accuracies are: ViT-S/32 75.55%/0.055%75.55\%/0.055\% , ViT-S 80.96%/0.046%80.96\%/0.046\%, ViT-B 84.12%/0.068%84.12\%/0.068\%, DeiT-S 79.45%/0.094%79.45\%/0.094\% , and Swin-S 83.11%/0.035%83.11\%/0.035\%.

References