Low-bit Quantization of Neural Networks for Efficient Inference

Yoni Choukroun, Eli Kravchik, Fan Yang, Pavel Kisilev

Introduction

Neural networks (NNs) proved to be extremely effective in solving a broad variety of problems in computer vision, speech recognition and natural language processing . Deep learning methods are usually evaluated only according to their accuracy over a given task. This criterion leads to the development of architectures with constantly increasing computational complexity and memory requirements. Thus, performing inference on low power System on a Chip (SoCs) used in smartphones or IoT devices is a significant challenge, due to the limited available memory and computational resources.

Several approaches have been proposed in order to make deep NNs less resource demanding. Network pruning of redundant and non-informative weights allows significant reduction of the network size . Matrix factorization via low-rank approximation exploits the redundancy property of the NN parameters in order to increase speed up . Distillation of NNs aims to transfer the knowledge contained in a pretrained large network to a compressed model via adapted training process . Also, new architectures (i.e.) with more efficient operations such as point/depth-wise or grouped convolutions allow the reduction of the model size, compared to existing over-parameterized architectures.

Another popular direction, which we focus on in this paper, is the quantization of NN. Quantization methods attempt to reduce the precision of the NN parameters and/or activations from single precision (32 bit floating point, or FP32) to lower bit representations. Several benefits of low-bit precision can be exploited by deep learning accelerators. The storage requirement for a low-bit precision model can be diminished substantially, as well as the power consumption. Similarly, the memory bandwidth requirements can be significantly reduced. Since the multiply accumulate (MAC) operations are performed on low-bit processing engines, the computational complexity can be reduced as well. Perhaps the most important benefit of low bit representation is the saving of chip area. For instance, 8 bits integer (INT8) operations can save up to 30x energy and up to 116x area compared to FP32 operations , allowing significantly better computational throughput. However, low-bit precision inference often causes loss of the task accuracy, which is usually compensated with the help of heavy full retraining, mixed precision or non-uniform quantization

In this paper, we address the quantization problem for weights and/or activations of a pretrained NN on highly constrained hardware, wherein complete retraining or mixed precision calculations cannot be tolerated. To the best of our knowledge, this is the first time INT4 only deployment of a pretrained NN via efficient linear quantization is performed with minimal loss of accuracy and data requirement. We propose a simple yet efficient optimization framework to find the optimal quantization parameters in the MMSE sense at each layer separately. The proposed MMSE reconstruction gives better accuracy than other existing MSE based optimization methods. NN parameters are quantized in a kernel-wise fashion that does not violate the linearity of the dot product operations, enabling efficient deployment on any common deep learning hardware. We identify key layers that are most sensitive to quantization errors, and provide an adaptive quantization scheme that makes use of multiple low precision tensors. Finally, we propose a refinement procedure for the scaling factors of the quantized tensors, which is performed on a small unlabelled calibration set. In a similar manner, the NN activations quantization coefficients are obtained offline via MMSE criterion. More sensitive activations are better approximated according to their reconstruction residual. The main contributions of this paper can be summarized as follows:

We propose a low-bit precision linear quantization framework for fast deployment of pretrained NNs on hardware which does not allow mixed precision operations.

We achieve minimal loss of task accuracy, while remaining compliant with modern deep learning hardware, by using fine grained partitioning of the NN weights. We propose optimal solution to the MMSE quantization problem and we deploy multiple tensors for key layers. Also, we refine the quantization factors of the network parameters for fast differentiable optimization.

Extensive experiments on ImageNet with various popular architectures demonstrate that our INT4 linear quantization method for both weights and activations, performs inference with only 3%3\% top-1 and 1.7%1.7\% top-5 mean accuracy degradation, as compared to the FP32 models, reaching, to the best of our knowledge, state-of-art results. The above degradation can be further reduced according to the complexity-accuracy trade-off inherent to the proposed multiple kernel method.

The remainder of the paper is organized as follows. Section 2 reviews related works. In section 3, after analyzing the quantization challenges, we develop our MMSE based quantization for accurate approximation of original models. In this section we also provide experimental results to demonstrate the usefulness of the steps in the proposed quantization pipeline. At the end of the section, we describe the modification of the presented algorithm components for the problem of activation quantization. Finally, quantization results from our experiments on several popular NNs are presented in Section 4.

Related Work

Neural network acceleration received increasing attention in the deep learning community, where the need for accurate yet fast and efficient frameworks is crucial for real-world applications. A powerful approach is quantization of NNs to low-bit representation. There are two main quantization scenarios. The first one is the full training of a given model to a desired lower bit precision. With this approach, the weights, the activations and even the gradients can be quantized to very low precision, enabling potentially fast training and inference . The major problem with the training approach above arises from the discritness of the parameters, wherein the backpropagation approach is not well defined. The ”straight-through estimator” has been used in in order to estimate the gradient of a stochastic neuron. proposed to use stochastic quantization of the weights via random rounding, in order to inject regularizing noise to the training process. suggests to approximate solution using variational Bayes method where the weights can be restricted to discrete values assuming Gaussian distribution. Instead of seeking for appropriate derivatives, assumed smooth approximation of parameters with defined gradients. Non-uniform quantization of NN parameters has been proposed in where the parameters are approximated using k-means algorithm. proposed a high order quantization scheme of weights where the approximation residual is further processed allowing better refinement of full precision input. Estimation of the quantization parameters by solving constrained optimization problem has been proposed for binary and ternary weights .

We focus on the second quantization scenario that targets direct quantization of a pretrained FP32 network to a lower bit-depth precision without full training. INT8 quantization of parameters has been proven to be relatively robust to quantization noise even with simple uniform quantization of weights . proposed L2L_{2} error minimization of weights via alternating optimization in order to obtain a generalizing ability during the training. Nevertheless, INT8 quantization of the network activations is more challenging because of real time constraints. Nvidia proposed in TensorRT a quantization framework that searches for saturation threshold of the activations, based on the Kullback-Leibler divergence measure between the quantized activations and their full precision counterpart. Recently, proposed to approximate activations, as if they were sampled from a known distribution in order to obtain, under some assumptions, analytically optimal threshold in the L2L_{2} sense. However, quantization of full precision weights and activations to less than 8-bits usually causes significant loss of accuracy, a problem that has not been solved yet. In order to overcome such degradation in performance, quantization frameworks resort to retraining procedures, mixed precision solutions or non-uniform quantization. These solutions make fast and easy deployment of quantized NNs impossible, especially on highly constrained HW such as mobiles or IoT devices.

Proposed MMSE Quantization Method

Conversion of a full precision NN into its fixed point version introduces quantization noise (error) into the network. Since significant noise may deteriorate the model performance, in the absence of training capabilities, efforts should be invested in minimizing the noise power, in order to approximate the original model as accurately as possible. Minimization of the noise power of weights and of activations via MSE optimization is a natural criterion for quantization quality, even though no direct relationship can be easily established between the noise of the output and the model accuracy. In the following, we investigate how quantization affects the NN output, and propose optimal solution to the quantization problem, via MSE optimization with fixed precision constraints.

A popular uniform quantization scheme is given by

2 Mean Squared Error Analysis of Quantization

The relation between the full precision tensor weights WW and activations XX and their respective approximations W^\hat{W} and X^\hat{X} can be obtained as follows

where nW,nXn_{W},n_{X} denote the quantization noise of the weights and activations respectively and where the approximation is obtained by neglecting second order noise term. Let us consider the case where the NN is composed of linear layers only. In such setting the NN output YY is defined as

For ease of notation we will omit the mean factor of the MSE. Defining eL2e_{L}^{2}, the MSE between the original model output and the quantized model output, we obtain in expectation that

where the approximation is obtained by assuming zero mean noise and where ∥⋅∥F\|\cdot\|_{F} denotes the Frobenius norm. The inequality is obtained using Cauchy Schwarz inequality and by assuming the weights/activations and the noises are statistically independent. We obtain here a recursive expression of the network output MSE. It is obvious that first layers may have significant impact on the output quality because of the recursion factor. Also, special attention should be given to the minimization of the weights noise which is coupled with possibly unbounded activations (ReLU). In real configurations, the non linear functions that play major role in the discriminative power of deep networks are much harder to model and analyze. Nevertheless, the linear analysis presented above, is supported empirically in our experiments, wherein the first layers always impact model accuracy the most. Also, in our setting, the approximation quality of the weights has much more influence on the performance than the activations, in contrast to the analysis of .

3 Kernel-Wise Quantization

4 Minimum MSE Quantization

For any arbitrary precision pp this optimization problem does not have an analytical solution. Assuming optimal α≠0\alpha\neq 0 is given, and denoting TiT_{i} the ithi^{th} element of the tensor TT, we remain with a one constraint optimization problem where we can rewrite eq. (7) objective as

5 Multiple Tensor Quantizations

Here we opt for a nested optimization approach, where the solution is obtained iteratively until convergence. Then, the optimization can be written in the nested form as

Alternating optimization is used until convergence, while at each iteration the quantized tensor and its scaling factor are obtained using the MSE quantization mapping suggested in Section 3.4. The proposed quantization approach for layers with high MSE is summarized in Algorithm 1. In the algorithm below, input TT refers to one convolutional kernel of a given key layer.

Because of its high computational cost, this method should be reserved to small tensors only (convolutional kernels). To emphasize the rational of the proposed line search optimization procedure over the alternating approach, we show the highly non-convex function of α1\alpha_{1} and α2\alpha_{2} for the INT4 setting in Figure 2. In our experiments, dual line-search approach improves the MSE by 5x in average over the tested dual layers.

Layers with high MSE that are approximated using multiple quantized tensors obviously require more parameters and computations, so a trade-off between accuracy and performance can be established. Defining speedup is not trivial since it is highly dependent on the hardware design. Modern GPUs can double the TOPS performance at INT4 precision, since multiple small low precision matrix multiplications can be performed in a distributed fashion. In order to measure the increase in storage and computations we define the compression ratio for analysis of the framework. The compression ratio 0<CR≤10<\text{CR}\leq 1 is defined as the ratio of the quantized model (weights and scaling factors) size to the FP32 model size. In this work, difficult layers in the INT4 setting are approximated with the dual method only (n=2n=2). Performance of the proposed algorithm as well as the corresponding compression ratios are summarized in Table 3.

6 Scaling Factors Refinement

Assuming f(X,{Wl}l)f(X,\{W_{l}\}_{l}) is the NN mapping function, we seek to minimize

where MM is the size of the calibration set. The advantage of this approach is that the number of optimized values is small and equal to the number of convolutional kernels. At contrary to training methods, this approach is fully differentiable avoiding sub-gradient definitions. Thus, the optimization can be conducted very efficiently using popular stochastic gradient descent methods on a small calibration set. Also, only a few optimization steps are required for both fast deployment and better generalization. Results of the refinement procedure, and its influence on accuracy improvement are presented in Table 4. These experiments show that the refinement procedure can improve the accuracy by up to to 23 percent. We also provide ablation study to demonstrate the usefulness of the refinement stage in low compression ratio tasks where no dual layers are allowed.

7 Activations Quantization

For key layers, optimal approximation using multiple quantizations as described in Section 3.5 suffers from strong overfitting. In order to approximate better, and to generalize at the same time, we obtain the optimal parameters by quantizing the residual from the first approximation, such that for a given activation XX we have

We obtain β1\beta_{1} from eq.(15), followed by obtaining β2\beta_{2} from eq.(16). This procedure is similar to the alternating approach described in Algorithm 1 with one iteration only.

Similar to the network weight compression ratio, we define the activations compression ratio as the ratio between the size of all the compressed activations and the size of all the original activations.

Table 5 presents the comparison of the variants of the proposed activation quantization method with several existing methods. The line-search method iterates over 50 samples only. Statistics of thresholding factors for different models are presented in Figure 3 where we can observe the severe saturation of the MMSE method.

Experiments

We evaluate the proposed framework on several popular architectures in order to demonstrate its robustness to aggressive quantization. We also evaluate quantization of non over-parameterized models such as SqueezeNet or DenseNet that are much more sensitive to quantization and are usually not analyzed in the NN quantization literature. We consider different INT4 quantization scenarios such as unsigned (with offset) and signed representations. We present also the accuracy-compression ratio trade off analysis. The framework has been implemented using the Pytorch library with its pretrained models.

The MSE grid search iterates over 500 sampling points for the weights and 50 for the activations. The scaling factors refinement stage requires 25 epochs over 500 randomly sampled images from the validation set. The calibration set used for the quantization of activations is comprised of 250 images. All the layers are quantized, including the network input. Also, it is important to notice that there is currently no existing framework for linear INT4 only quantization of weights and activations for comparison in the inference setting. The summary of the accuracy loss and the compression ratios is presented in Table 6. In Table 7 we show how modification of the MSE threshold τ\tau influences the accuracy-compression ratio trade-off.

Conclusion

In this paper we introduced an efficient and accurate MSE-based low-bit precision quantization framework for neural networks. Our approach deploys hardware-aware partitioning of the network parameters, and the refinement of high MSE layers, using quantized filter banks. Given a small calibration set, we further refine the quantization scaling factors for better approximation of the original model. We also provide a framework for the quantization of the network activations, wherein we propose a method of residual quantization for improved approximation of the most sensitive layers. The proposed approach can be adjusted to any desired precision for constrained hardware deployment, according to the inherent compression-complexity trade-off of the method. The framework allows fast and efficient deployment of pretrained models, producing a new state-of-the-art INT4 inference quantization results.

References