Q-ViT: Accurate and Fully Quantized Low-bit Vision Transformer

Yanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao, Peng Gao, Guodong Guo

Introduction

Inspired by the success in natural language processing (NLP), transformer-based models have shown great power in various computer vision (CV) tasks, such as image classification and object detection . Pre-trained with large-scale data, these models usually have a tremendous number of parameters. For example, there are 632M parameters taking up 2528MB memory usage and 162G FLOPs in the ViT-H model, which is both memory and computation expensive during inference. This limits these models for the deployment on resource-limited platforms. Therefore, compressed transformers are urgently needed for real applications.

Substantial efforts have been made to compress and accelerate neural networks for efficient online inference. Methods include compact network design , network pruning , low-rank decomposition , quantization , and knowledge distillation . Quantization is particularly suitable for deployment on AI chips because it reduces the bit-width of network parameters and activations for efficient inference. Prior post-training quantization (PTQ) methods on ViTs directly compute quantized parameters based on pre-trained full-precision models, which constrains the model performance to a sub-optimized level without fine-tuning. Furthermore, quantizing these models based on PTQ methods to ultra-low bits (e.g., 4 bits or lower) is ineffective and suffers from a significant performance reduction.

Differently, quantization-aware training (QAT) methods perform quantization during back propagation and achieve much less performance drop with a higher compression rate generally. QAT is shown to be effective for CNN models for CV tasks. However, QAT methods remain largely unexplored for low-bit quantization of vision transformers. Therefore, we first build a fully quantized ViT baseline, a straightforward yet effective solution based on common techniques. Our study discovers that the performance drop of fully quantized ViT lies in the information distortion among the attention mechanism in the forward process, and the ineffective optimization for eliminating the distribution difference through distillation in the backward propagation. First, the attention mechanism of ViT aims at modeling long-distance dependencies . However, our analysis shows that a direct quantization method leads to the information distortion, i.e., significant distribution variation for the query module between quantized ViT and full-precision counterpart. For example, as shown in Fig. 1, the variance difference is 0.4409 (1.2124 v.s. 1.6533) for the first block . This inevitably deteriorates the representation capability of the attention module on capturing the global dependency for the input. Second, the distillation for the fully quantized ViT baseline utilizes distillation token (following ) to directly supervise the classification output of the quantized ViT. However, we found that such a simple supervision is ineffective, which is coarse-grained for the large gap between the quantized attention scores and their full-precision counterparts.

To address the aforementioned issues, a fully quantized ViT (Q-ViT) is developed by retaining the distribution of quantized attention modules as that of full-precision counterparts (see the overview in Fig. 2). Accordingly, we propose to modify the distorted distribution over quantized attention modules through an Information Rectification Module (IRM) based on information entropy maximization, in the forward process. While in the backward process, we present a Distribution Guided Distillation (DGD) scheme to eliminate the distribution variation through attention similarity loss between quantized ViT and full-precision counterpart. The contributions of our work include:

We propose an Information Rectification Module (IRM) based on the information theory to address the information distortion problem. IRM applies quantized representations in the attention module with a maximized information entropy, allowing the quantized model to restore the representation of input images.

We develope a Distribution Guided Distillation (DGD) scheme to eliminate the distribution mismatch in distillation. DGD takes appropriate activations and utilizes knowledge from the similarity matrices in distillation to perform optimization accurately.

Our Q-ViT, for the first time, explores a promising way towards accurate and low-bit ViT. Extensive experiments on the ImageNet benchmark show that our Q-ViT outperforms the baseline by a large margin, and achieves comparable performances with the full-precision counterparts.

Related Work

Vision transformer. Motivated by the great success of Transformer in natural language processing, researchers are trying to apply Transformer architecture to Computer Vision tasks. Unlike mainstream CNN-based models, Transformer is capable of capturing long-distance visual relations by its self-attention module and provides the paradigm without image-specific inductive bias. ViT views 16 ×\times 16 image patches as token sequence and predicts classification via a unique class token, which shows promising results. Subsequently, many works, such as DeiT and PVT achieve further improvement on ViT, making it more efficient and applicable in downstream tasks. CONTAINER fully utilizes a hybrid ViT to aggregate dynamic and static information, exploring a new framework for visual tasks. However, these high-performing vision transformers are attributed to the large number of parameters and high computational overhead, limiting their adoption. Therefore, innovating a smaller and faster vision transformer becomes a new trend. DynamicViT presents a dynamic token sparsification framework to prune redundant tokens progressively and dynamically, achieving competitive complexity and accuracy trade-off. Evo-ViT proposes a slow-fast updating mechanism that guarantees information flow and spatial structure, trimming down both the training and inference complexity. While the above works focus on efficient model designing, this paper boosts the compression and acceleration in the track of quantization.

Quantization. Quantizing neural networks (QNNs) often possess low-bit (1∼\sim 4-bit) weights and activations to accelerate the model inference and save the memory usage. Specifically, ternary weights are introduced to reduce the quantization error in TWN . DoReFa-Net exploits convolution kernels with low bit-width parameters and gradients to accelerate both the training and inference. TTQ uses two full-precision scaling coefficients to quantize the weights to ternary values. presented a 2 ⁣∼ ⁣42\!\sim\!4-bit quantization scheme using a two-stage approach to alternately quantize the weights and activations, which provides an optimal trade-off among memory, efficiency, and performance. parameterizes the quantization intervals and obtain their optimal values by directly minimizing the task loss of the network and also the accuracy degeneration with further bit-width reduction. introduces transfer learning into network quantization to obtain an accurate low-precision model by utilizing the Kullback-Leibler (KL) divergence. enables accurate approximation for tensor values that have bell-shaped distributions with long tails and finds the entire range by minimizing the quantization error. In our Q-ViT, we aim to implement an accurate, fully quantized vision transformer under the QAT paradigm.

Baseline of Fully Quantized ViT

First of all, we build a baseline to study the fully quantized ViT since it has never been proposed in previous works. A straightforward solution is to quantize the representations (weights and activations) in ViT architecture in the forward propagation and apply distillation to the optimization in the backward propagation.

Quantized ViT architecture. We briefly introduce the technology of neural network quantization. We first introduce a general asymmetric activation quantization and symmetric weight quantization scheme as

Here, clip⁡{y,r1,r2}\operatorname{clip}\{y,r_{1},r_{2}\} returns yy with values below r1r_{1} set as r1r_{1} and values above r2r_{2} set as r2r_{2}, and ⌊y⌉\lfloor y\rceil rounds yy to the nearest integer. With quantizing activations to signed aa bits and weights to signed bb bits, Qnx=2a−1,Qpx=2a−1−1Q_{n}^{x}=2^{a-1},Q_{p}^{x}=2^{a-1}-1 and Qnw=2b−1,Qpw=2b−1−1Q_{n}^{\bf w}=2^{b-1},Q_{p}^{\bf w}=2^{b-1}-1. In general, the forward and backward propagation of quantization function in quantized network is formulated as

where J\mathcal{J} is loss function, Q(⋅)Q(\cdot) is applied in the forward propagation while the straight-through estimator (STE) is used to retain the derivation of gradient in backward propagation. ⊗\otimes denotes the matrix multiplication with efficient bit-wise operations.

The input images are first encoded as patches and passes through several transformer blocks. Such transformer block consists of two components: Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP). The computation of attention weight depends on the corresponding query q{\bf q}, key k{\bf k} and value v{\bf v}, and the quantized computation in one attention head is

where Q-Linear⁡q,Q-Linear⁡k,Q-Linear⁡v\operatorname{Q-Linear}_{q},\operatorname{Q-Linear}_{k},\operatorname{Q-Linear}_{v} denote the three quantized linear layers for q,k,v{\bf q},{\bf k},{\bf v}, respectively. Thus, the attention weight is formulated as

Training for Quantized ViT. Knowledge distillation is an essential supervision approach for training quantized neural networks, which bridges the performance gap between quantized models and full-precision counterparts. The usual practice is using distillation through attention as described in

Proposed Q-ViT

With aforementioned quantizing and training pipeline, a fully quantized ViT baseline is built, which however has a large performance gap with full-precision counterparts. Our study shows that the baseline suffers an severe information distortion among quantized attention scores in the forward propagation and distillation direction misleading in the backward propagation.

Intuitively, in the fully quantized ViT baseline, the information representation capability largely depends on the transformer-based architecture, such as the attention weight in MHSA module. However, the performance improvement brought by such architecture is severely limited by the quantized parameters, while the rounded and discrete quantization also significantly affect the optimization. The phenomenon identifies the bottleneck of the fully quantized ViT baseline comes from architecture and optimization for the forward and backward propagation, respectively.

Architecture bottleneck. We replace each module with the full-precision counterpart respectively and compare the accuracy drop as shown in Fig. 3. These ablation experiments are conducted the same as experiments in Tab. 2. We find that quantizing query, key, value and attention weight (i.e., softmax⁡(A)\operatorname{softmax}({\bf A}) in Eq. (4) to 2 bits brings the most significant drop of accuracy among all parts of the ViT, up to 10.03%. While quantized MLP layers and quantized weights of linear layers in MHSA brings only 1.78% and 4.26% drop, respectively. And once query, key, value and attention weight are quantized, even keep all weights of linear layers in MHSA module full-precision, the performance drops (10.57%) are still significant. Thus, improving the attention structure is pivotal to solve the performance drop problem of quantized ViT.

Optimization bottleneck. We calculate l2-norm distances between each attention weight among different blocks in DeiT-S architecture as shown in Fig. 4. The MHSA modules in full-precision ViT with different depth learn different representations from images. As mentioned in , lower ViT layers attend both locally and globally while higher ViT layers pay most attention to global representations. However, the fully quantized ViT (blue lines in Fig. 4) fails to learn accurate attention map distances. Thus, it requires a new design that could utilize the full-precision teacher’s information better.

2 Information Rectification in Q-Attention

To address the information distortion of quantized representations in the forward propagation, we propose an efficient Q-Attention structure based on information theory, which statistically maximizes the entropy of representation and revives the attention mechanism in the fully quantized ViT. Since the representations with extremely compressed bit-width in fully quantized ViT have limited capabilities, the ideal quantized representation should preserve the given full-precision counterparts as much as possible, which means the mutual information between quantized and full-precision representations should be maximized as mentioned in .

We further show the statistical results that the distribution of query and key values in ViT architectures intending to follow Gaussian distributions under the distilling supervision, whose histograms are in bell-shape . For example, in Fig. 1 and Fig. 5, we have shown the query and key distributions and their corresponding Probability Density Function (PDF) using the calculated mean and standard deviation for each MHSA layer. Therefore, the distributions of query and key in the MHSA modules of full-precision counterparts are formulated as

Since the weight and the activation with extremely compressed bit-width in fully quantized ViT have limited capabilities, the ideal quantization process should preserve the corresponding full-precision counterparts as much as possible, thus the mutual information between quantized and full-precision representations should be maximized . As shown in , for Gaussian distribution, the quantizers with maximum output entropy (MOE) and minimum average error (MAE) are approximately the same within a multiplicative constant. Thus the process of minimizing the error between full-precision values and quantized values is equivalent to maximizing the information entropy of the quantized values. Thus, when the deterministic quantization function is applied to quantized ViT, such objective is equivalent to maximizing the information entropy H(Qx)\mathcal{H}(Q_{\bf x}) of quantized representation QxQ_{\bf x} in Eq. (4), which is defined as

where qxq_{\bf x} is the random quantized variables in Qa(x)Q_{a}({\bf x}) (which is Qa(q)Q_{a}({\bf q}) or Qa(k)Q_{a}({\bf k}) in different conditions) with probability mass function p(⋅)p(\cdot). For better retaining the information contained in the MHSA modules from the full-precision counterparts, the information entropy in the quantization process should be maximized.

However, direct application of quantization function converting the values into finite fixed points brings irreversible disturbance to the distributions and the information entropy H(Qa(q))\mathcal{H}(Q_{a}({\bf q})) and H(Qa(k))\mathcal{H}(Q_{a}({\bf k})) degenerates to a much lower level than the full-precision counterparts. To mitigate the information degradation from the quantization process in the attention mechanism, a Information Rectification Module (IRM) is proposed for effectively maximizing the information entropy of quantized attention weights

Then to revive the attention mechanism to capture critic elements by information entropy maximization, the learnable parameters γq,βq\gamma_{\bf q},\beta_{\bf q} and γk,βk\gamma_{\bf k},\beta_{\bf k} reshape the distributions of the query and key values to achieve the state of information maximization. In a nutshell, in our IRM-Attention structure, the information entropy of quantized attention weight is maximized to alleviate its severe information distortion and revive the attention mechanism.

3 Distribution Guided Distillation through Attention

To address the attention distribution mismatch occurred in fully quantized ViT baseline in the backward propagation, we further propose a Distribution Guided Distillation (DGD) scheme with apposite distilled activations and the well-designed similarity matrices to effectively utilize knowledge from the teacher, which optimizes the fully quantized ViT more accurately.

As an optimization technique based on element-level comparison of activation, distillation allows the quantized ViT to mimic the full-precision teacher model about output logits. However, we find that the distillation procedure used in previous ViT and fully quantized ViT baseline (Sec. 3) unable to deliver meticulous supervision to attention weights (shown in Fig. 4), leading to insufficient optimization. To solve the optimization insufficiency in the distillation of the fully quantized ViT, we propose the Distribution-Guided Distillation (DGD) method in Q-ViT. We first build patch-based similarity pattern matrices for distilling the upstream query and key instead of attention following , which is formulated as

where LL and HH denote the number of ViT layers and heads. With the proposed Distribution Guided Distillation, the Q-ViT retains the distribution over query and key from the full-precision counterparts (as shown in Fig. 5).

Our DGD scheme first provides the distribution-aware optimization direction together with processing appropriate distilled parameters and then constructs similarity matrices to eliminate scale differences and numerical instability, thereby improves fully quantized ViT by accurate optimization.

Experiments

In this section, we evaluate the performance of the proposed Q-ViT model for image classification task using popular DeiT and Swin backbones. To the best of our knowledge, there is no publicly available source code on quantization-aware training of vision transformer at this point, so we implement the baseline and LSQ methods by ourselves.

Datasets. The experiments are carried out on the ILSVRC12 ImageNet classification dataset . The ImageNet dataset is more challenging due to its large scale and greater diversity. There are 1000 classes and 1.2 million training images, and 50k validation images in it. In our experiments, we use the classic data augmentation method described in .

Experimental settings. In our experiments, we initialize the weights of quantized model with the corresponding pretrained full-precision model. The quantized model is trained for 300 epochs with batch-size 512 and the base learning rate 2e ⁣− ⁣42e\!-\!4. We do not use warm-up scheme. For all the experiments, we apply LAMB optimizer with weight decay set as 0. Other training settings follow DeiT or Swin Transformer . Note that we use 8-bit for the patch embedding (first) layer and the classification (last) layer following .

Bakcbone. We evaluate our quantization method on two popular vision transformer implementation: DeiT and Swin Transformer . The DeiT-S, DeiT-B, Swin-T and Swin-S are adopted as the backbone models, whose Top-1 accuracy on ImageNet dataset are 79.9%, 81.8%, 81.2%, and 83.2% respectively. For a fair comparison, we utilize the official implementation of DeiT and Swin Transformer.

2 Ablation Study

We give quantitative results of the proposed IRM and DGD in Tab. 1 As shown in Tab. 1, the fully quantized ViT baseline suffers a severe performance drop on classification task (0.2%, 2.1% and 11.7% with 2/3/4-bit, respectively). IRM and DGD improve the performance when used alone, and the two techniques further boost the performance considerably when combined together. For example, the IRM improve the 2-bit Baseline by 1.7% and the DGD achieves 2.3% performance improvement. While combining the IRM and DGD together, the performance improvement achieves 3.8%.

To conclude, the two techniques can promote each other to improve Q-ViT and close the performance gap between fully quantized ViT and full-precision counterpart.

3 Main Results

The experimental results are shown in Tab. 2. We compare our method with 2/3/4-bit baseline and LSQ based on the same frameworks for the task of image classification with the ImageNet dataset. We also report the classification performance of the 8-bit post-training quantization networks percentile VT-PTQ . We firstly evaluate the proposed method on DeiT-S and DeiT-B models.

For DeiT-S backbone, compared with 8-bit VT-PTQ method, our 4bit Q-ViT achieves a much larger compression ratio than 8-bit VT-PTQ, but with significant performance improvement (78.1% vs. 80.9%). And it is worth noting that the proposed 2-bit model significantly compresses the DeiT-S by 21.5×\times on FLOPs. The proposed method boosts the performance of 2/3/4-bit Baseline by 3.9%, 1.5% and 1.2% with the same architecture and bit-width, which is significant on the ImageNet dataset. For larger DeiT-B, as shown in Tab. 2, the performance of the proposed method outperforms the 2/3/4-bit Baseline by 3.8%, 1.7% and 1.9%, a large margin. Also note that the proposed 2/3/4-bit model significantly compresses the DeiT-B by 21×\times, 12×\times and 7.6×\times on FLOPs. Compared with 8-bit post-training quantization methods, our method achieves significantly higher compression rate, and the performance improvement is significant.

Also, our method generates convincing results on Swin-transformers. As shown in Tab. 2, the performance of the proposed method with Swin-T outperforms the 2/3/4-bit Baseline method by 4.1% , 2.1% and 2.0%, a large margin. Compared with 8-bit post-training quantization methods, our method achieves significantly higher compression rate, and comparable performance. Note that our 4-bit Q-ViT surpasses the full-precision by 1.3% counterpart using Swin-T, which demonstrates the significance of our Q-ViT. For larger Swin-S, the performance of the proposed method outperforms the 2/3/4-bit Baseline by 4.3%, 1.8% and 1.5%. Also note that our 4-bit Q-ViT surpasses the full-precision by 1.1% counterpart using Swin-S and significantly compresses the Swin-S by 7.9×\times , which demonstrates the effectiveness and efficiency of our Q-ViT.

Conclusion

In this paper, we introduce Q-ViT to improve the fully quantized ViTs with high compression ratio and competitive performance. We first build a theoretical framework of fully quantized ViT and analysis the bottlenecks of the fully quantized ViT baseline. Then we introduce Information Rectification Module and Distribution Guided Distillation to Q-ViT for performance improvement. Our proposed Q-ViTs achieve comparable performance with full-precision counterparts with ultra-low bit weights and activations. Our work gives an insightful analysis and effective solution about the crucial issues in ViT full quantization, which blazes a promising path for the extreme compression of ViT.

Acknowledgement

This work was supported in part by the National Natural Science Foundation of China under Grant 62076016, under Grant 62206272 and 61827901, Beijing Natural Science Foundation-Xiaomi Innovation Joint Fund L223024, Foundation of China Energy Project GJNY-19-90.

References