Data-Free Quantization Through Weight Equalization and Bias Correction
Markus Nagel, Mart van Baalen, Tijmen Blankevoort, Max Welling
Introduction
In recent years, deep learning based computer vision models have moved from research labs into the cloud and onto edge devices. As a result, power consumption and latency of deep learning inference have become an important concern. For this reason fixed-point quantization is often employed to make inference more efficient. By quantizing floating point values onto a regularly spaced grid, the original floating point values can be approximated by a set of integers, a scaling factor, and an optional zero point offset . This allows for the use of faster and more power-efficient integer operations in matrix multiplication and convolution computations, at the expense of lower representational power. We refer the reader to for details on commonly used, hardware-friendly quantization methods for deep learning models.
Quantization of 32-bit full precision (FP32) models into 8-bit fixed point (INT8) introduces quantization noise on the weights and activations, which often leads to reduced model performance. This performance degradation ranges from very minor to catastrophic. To minimize the quantization noise, a wide range of different methods have been introduced in the literature (see Section 2). A major drawback of these quantization methods is their reliance on data and fine-tuning. As an example, consider real-world actors that manage hardware for quantized models, such as cloud-based deep learning inference providers or cellphone manufacturers. To provide a general use quantization service they would have to receive data from the customers to fine-tune the models, or rely on their customers to do the quantization. In either case, this can add a difficult step to the process. For such stakeholders it would be preferable if FP32 models could be converted directly to INT8, without needing the know-how, data or compute necessary for running traditional quantization methods. Even for model developers that have the capability to quantize their own models, automation would save significant time.
In this paper, we introduce a quantization approach that does not require data, fine-tuning or hyperparameter tuning, resulting in accuracy improvement with a simple API call. Despite these restrictions we achieve near-original model performance when quantizing FP32 models to INT8. This is achieved by adapting the weight tensors of pre-trained models such that they are more amenable to quantization, and by correcting for the bias of the error that is introduced when quantizing models. We show significant improvements in quantization performance on a wide range of computer vision models previously thought to be difficult to quantize without fine-tuning.
In literature the practical application of proposed quantization methods is rarely discussed. To distinguish between the differences in applicability of quantization methods, we introduce four levels of quantization solutions, in decreasing order of practical applicability. Our hope is that this will enable other authors to explore solutions for each level, and makes the comparison between methods more fair. The axes for comparison are whether or not a method requires data, whether or not a method requires error backpropagation on the quantized model, and whether or not a method is generally applicable for any architecture or requires significant model re-working. We use the following definitions throughout the paper:
No data and no backpropagation required. Method works for any model. As simple as an API call that only looks at the model definition and weights.
Requires data but no backpropagation. Works for any model. The data is used e.g. to re-calibrate batch normalization statistics or to compute layer-wise loss functions to improve quantization performance. However, no fine-tuning pipeline is required.
Requires data and backpropagation. Works for any model. Models can be quantized but need fine-tuning to reach acceptable performance. Often requires hyperparameter tuning for optimal performance. These methods require a full training pipeline (e.g. ).
Requires data and backpropagation. Only works for specific models. In this case, the network architecture needs non-trivial reworking, and/or the architecture needs to be trained from scratch with quantization in mind (e.g. ). Takes significant extra training-time and hyperparameter tuning to work.
Background and related work
There are several works that describe quantization and improving networks for lower bit inference and deployment . These methods all rely on fine-tuning, making them level 3 methods, whereas data-free quantization improves performance similarly without that requirement. Our method is complementary to these and can be applied as a pre-processing before quantization aware fine-tuning.
In a whitepaper, Krishnamoorthi , introduces a level 1 ‘per-channel’ quantization scheme, in which the weights of a convolutional weight tensor are quantized per output channel. A major drawback of this method is that it is not supported on all hardware, and that it creates unnecessary overhead in the computation due to the necessity of scale and offset values for each channel individually. We show that our method improves on per-channel quantization, while keeping a single set of scale and offset values for the whole weight tensor instead.
Other methods to improve quantization need architecture changes or training with quantization in mind from the start . These methods are even more involved than doing quantization and fine-tuning. They also incur a relatively large overhead during training because of sampling and noisy optimization, and introduce extra hyperparameters to optimize. This makes them level 4 methods.
Methods that binarize or ternarize networks result in models with great inference efficiency as expensive multiplications and additions are replaced by bit-shift operations. However, quantizing models to binary often leads to strong performance degradation. Generally they need to be trained from scratch, making them level 4 methods.
Other approaches use low-bit floating point operations instead of integer operations, or other custom quantization implementations . We do not consider such approaches as the hardware implementation is less efficient.
In concurrent work, Meller et al. also exploits the scale equivariance of the ReLU function to rescale weight channels and notice the biased error introduced by weight quantization , leading to a method that resembles our data-free quantization approach. Stock et al. also use the scale equivariance property of the ReLU function, but use it for network optimization instead.
Motivation
While many trained FP32 models can be quantized to INT8 without much loss in performance, some models exhibit a significant drop in performance after quantization (). For example, when quantizing a trained MobileNetV2 model, Krishnamoorthi reports a drop in top-1 accuracy from 70.9% to 0.1% on the ImageNet validation set. The author restores near original model performance by either applying per-channel quantization, fine-tuning or both.
The fact that per-channel quantization yields much better performance on MobileNetV2 than per-tensor quantization suggests that, in some layers, the weight distributions differ so strongly between output channels that the same set of quantization parameters cannot be used to quantize the full weight tensor effectively. For example, in the case where one channel has weights in the range $(-0.5,0.5)$, the weights in the latter channel will all be quantized to when quantizing to 8-bits.
Figure 2 shows that large differences in output channel weight ranges do indeed occur in a (trained) MobileNetV2 model. This figure shows the weight distribution of the output channel weights of the depthwise-separable layer in the model’s first inverted residual block. Due to the strong differences between channel weight ranges that this layer exhibits, it cannot be quantized with reasonable accuracy for each channel. Several layers in the network suffer from this problem, making the overall model difficult to quantize.
We conjecture that performance of trained models after quantization can be improved by adjusting the weights for each output channel such that their ranges are more similar. We provide a level 1 method to achieve this without changing the FP32 model output in section 4.1.
2 Biased quantization error
A common assumption in literature (e.g. ) is that quantization error is unbiased and thus cancels out in a layer’s output, ensuring that the mean of a layer’s output does not change as a result of quantization. However, as we will show in this section, the quantization error on the weights might introduce biased error on the corresponding outputs. This shifts the input distribution of the next layer, which may cause unpredictable effects.
The biased error in a quantized layer’s output unit can be computed empirically using input data points as:
where and are the original outputs and the outputs generated using the quantized weight matrix, respectively.
Figure 3 shows the biased error per channel of a depthwise-separable convolution layer in a trained MobileNetV2 model. From this plot it is clear that for many channels in the layer’s output, the error introduced by weight quantization is biased, and influences the output statistics. Depthwise-separable layers are especially susceptible to this biased error effect as each output channel has only 9 corresponding weights.
Such a biased error on the outputs can be introduced in many settings, e.g. when weights or activations are clipped , or in non-quantization approaches, such as weight tensor factorization or channel pruning .
In section 4.2 we introduce a method to correct for this bias. Furthermore, we show that a model’s batch normalization parameters can be used to compute the expected biased error on the output, yielding a level 1 method to fix the biased error introduced by quantization.
Method
Our proposed data-free quantization method (DFQ) consists of three steps, on top of the normal quantization. The overall flow of the algorithm is shown in Figure 4.
We observe that for a ReLU activation function the following scaling equivariance property holds:
for any non-negative real number . This follows from the definition of the ReLU:
This equivariance also holds for the PreLU activation function. More generally, the positive scaling equivariance can be relaxed to for any piece-wise linear activation functions:
where is parameterized as , and . Note that contrary to equivariance defined in eq. 2 we now also change the function into .
1.1 Scaling equivariance in neural networks
The positive scaling equivariance can be exploited in consecutive layers in neural networks. Given two layers, and , through scaling equivariance we have that:
where is a diagonal matrix with value denoting the scaling factor for neuron . This allows us to reparameterize our model with , and . In case of CNNs the scaling will be per channel and broadcast accordingly over the spatial dimensions. The rescaling procedure is illustrated in Figure 5.
1.2 Equalizing ranges over multiple layers
We can exploit the rescaling and reparameterization of the model to make the model more robust to quantization. Ideally the ranges of each channel are equal to the total range of the weight tensor, meaning we use the best possible representative power per channel. We define the precision of a channel as:
where is the quantization range of channel in and is the total range of . We want to find such that the total precision per channel is maximized:
In the case of symmetric quantization we have and . Solving eq. 9 (see appendix A) leads to the necessary condition:
meaning the limiting channel defining the quantization range is given by . We can satisfy this condition by setting such that:
which results in . Thus the channel’s ranges between both tensors are matched as closely as possible.
When equalizing multiple layers at the same time, we iterate this process for pairs of layers that are connected to each other without input or output splits in between, until convergence.
1.3 Absorbing high biases
In case the equalization procedure increases bias . This could in turn increase the range of the activation quantization. In order to avoid big differences between per-channel ranges in the activations we introduce a procedure that absorbs high biases into the subsequent layer.
For a layer with ReLU function , there is a non-negative vector such that . The trivial solution holds for all . However, depending on the distribution of and the values of and , there can be some values for which this equality holds for (almost) all . Following the previous two layer example, these can be absorbed from layer into layer as:
where , , and .
To find without violating our data-free assumption we assume that the pre-bias activations are distributed normally with the batch normalization shift and scale parameters and as its mean and standard deviation. We set . If , the equality introduced above will hold for the of values of (those greater than ) under the Gaussian assumption. As we will show in section 5.1.1, this approximation does not harm the full precision performance significantly but helps for activation quantization. Note that, in case data is available, the pre-bias distribution of can be found empirically and used to set .
2 Quantization bias correction
As shown empirically in the motivation, quantization can introduce a biased error in the activations. In this section we show how to correct for the bias in the error on the layer’s output, and how we can use the network’s batch normalization parameters to compute this bias without using data.
For a fully connected layer with weight tensor , quantized weights , and input activations , we have and therefore , where we define the quantization error , as the layer pre-activations of the FP32 model, and that layer with quantization error added.
For implementation, the expected error can be subtracted from the layer’s bias parameter, since the expected error vector has the same shape as the layer’s output. This method easily extends to convolutional layers as described in Appendix B.
Due to the centralization and normalization applied by batch normalization, the mean and standard deviation of the pre-activations are known: these are the batch normalization scale and shift parameters (henceforth referred to as and respectively).
where is the pre-activation output for channel , which is assumed to be normally distributed with mean and variance , is the normal CDF, and the notation is used to denote the normal PDF.
Experiments
In this section we present two sets of experiments to validate the performance of data-free quantization (DFQ). We first show in section 5.1 the effect of the different aspects of DFQ and how they solve the problems observed earlier. Then we show in section 5.2 how DFQ generalizes to other models and tasks, and sets a new state-of-the-art for level 1 quantization.
To allow comparison to previously published results, we use both weights and activations are quantized using 8-bit asymmetric, per-tensor quantization in all experiments. Batch normalization is folded in the adjacent layer before quantization. Weight quantization ranges are the min and max of the weight tensor. Activation quantization ranges are set without data, by using the learned batch normalization shift and scale parameter vectors and as follows: We compute the activation range for channel as (with ), with the minimum clipped to 0 in case of ReLU activation. We observed a wide range of can be used without significant performance difference. All experiments are done in Pytorch . In appendix 8 we show additional experiments using short-term fine-tuning, symmetric quantization and per-channel quantization.
In this section we investigate the effect of our methods on a pre-trained MobileNetV2 modelWe use the Pytorch implementation of MobileNetV2 provided by https://github.com/tonylins/pytorch-mobilenet-v2.. We validate the performance of the model on the ImageNet validation set. We first investigate the effects of different parts of our approach through a set of ablation studies.
In this section we investigate the effects of cross-layer equalization and high-bias folding. We compare these methods to two baselines: the original quantized model and the less hardware friendly per-channel quantization scheme.
The models considered in this section employ residual connections . For these networks we apply cross-layer equalization only to the layers within each residual block. MobileNetV2 uses ReLU6 activation functions, which clips activation ranges to $$. To avoid ReLU6 requiring a different cut off per channel after applying the equalization procedure, we replace ReLU6 with regular ReLU.
The results of the equalization experiments are shown in Table 1. Similar to , we observe that the model performance is close to random when quantizing the original model to INT8. Further we note that replacing ReLU6 by ReLU does not significantly degrade the model performance. Applying equalization brings us to within 2% of FP32 performance, close to the performance of per-channel quantization. We note that absorbing high biases results in a small drop in FP32 performance, but it boosts quantized performance by 1% due to more precise activation quantization. Combining both methods improves performance over per-channel quantization, indicating the more efficient per-tensor quantization could be used instead.
To illustrate the effect of cross-layer equalization, we show the weight distributions per output channel of the depthwise-separable layer in the model’s first inverted residual block after applying the equalization in Figure 6. We observe that most channels ranges are now similar and that the strong outliers from Figure 2 have been equalized. Note, there are still several channels which have all weight values close to zero. These channels convey little information and can be pruned from the network with hardly any loss in accuracy.
1.2 Bias correction
In this section we present results on bias correction for a quantized MobileNetV2 model. We furthermore present results of bias correction in combination with a naive weight-clipping baseline, and combined with the cross-layer equalization approach.
To illustrate the effect of bias correction, Figure 3 shows the per output channel biased error introduced by weight quantization. The per-channel biases are obtained as described in eq. 1. This figure shows that applying bias correction reduces the bias in the error on the output of a layer to very close to 0 for most output channels.
Results for the experiments described above for MobileNet V2 on the ImageNet validation set are shown in Table 2. Applying bias correction improves quantized model performance, indicating that a part of the problem of quantizing this model lies in the biased error that is introduced. However, bias correction on its own does not achieve near-floating point performance. The reason for this is most likely that the problem described in 3.1 is more severe for this model. The experiments on weight-clipping show that bias correction can mitigate performance degradation due to biased error in non-quantized models as well as quantized models. Clipping without correction in the FP32 model introduces a 4.66% loss in accuracy; bias correction reduces that loss to a mere 0.57%. Furthermore, it shows that weight clipping combined with bias correction is a fairly strong baseline for quantizing MobileNet V2. Lastly, we show that bias correction improves results when combined with the cross-layer equalization and bias folding procedures. The combination of all methods is our data-free quantization (DFQ) method. The full DFQ approach achieves near-floating point performance with a reduction of 0.53% top 1 accuracy relative to the FP32 baseline.
2 Comparison to other methods and models
In this section we show how DFQ generalizes to other popular computer vision tasks, namely semantic segmentation and object detection, and other model architectures such as MobileNetV1 and Resnet18 . Afterwards we compare DFQ to methods in the literature, including more complex level 3 and 4 approaches. This set of models was chosen as they are efficient and likely to be used in mobile applications where 8-bit quantization is frequently used for power efficiency.
To demonstrate the generalization of our method to semantic segmentation we apply DFQ for DeeplabV3+ with a MobileNetV2 backend , performance is evaluated on the Pascal VOC segmentation challenge . For our experiments we use the publicly available Pytorch implementationhttps://github.com/jfzhang95/pytorch-deeplab-xception.
We show the results of this experiment in Table 3. As observed earlier for classification we notice a significant drop in performance when quantizing the original model which makes it almost unusable in practice. Applying DFQ recovers almost all performance degradation and achieves less than 1% drop in mIOU compared to the full precision model. DFQ also outperforms the less hardware friendly per-channel quantization. To the best of our knowledge we are the first to publish quantization results on DeeplabV3+ as well as for semantic segmentation.
To demonstrate the applicability of our method to object detection we apply DFQ for MobileNetV2 SSDLite , evaluated on the Pascal VOC object detection challenge . In our experiments we use the publicly available Pytorch implementation of SSDhttps://github.com/qfgaohao/pytorch-ssd.
The results are listed in Table 4. Similar to semantic segmentation we observe a significant drop in performance when quantizing the SSDLite model. Applying DFQ recovers almost all performance drop and achieves less than 1% drop in mAP compared to the full precision model, again outperforming per-channel quantization.
2.2 Comparison to other approaches
In this section we compare DFQ to other approaches in literature. We compare our results to two other level 1 approaches, direct per-layer quantization as well as per-channel quantization . In addition we also compare to multiple higher level approaches, namely quantization aware training as well as stochastic rounding and dynamic ranges , which are both level 3 approaches. We also compare to two level 4 approaches based on relaxed quantization , which involve training a model from scratch and to quantization friendly separable convolutions that require a rework of the original MobileNet architecture. The results are summarized in Table 5.
For both MobileNetV1 and MobileNetV2 per-layer quantization results in an unusable model whereas DFQ stays close to full precision performance. DFQ also outperforms per-channel quantization as well as most level 3 and 4 approaches which require significant fine-tuning, training or even architecture changes.
On Resnet18 we maintain full precision performance for 8-bit fixed point quantization using DFQ. Some higher level approaches report slightly higher results than our baseline model, likely due to a better training procedure than used in the standard Pytorch Resnet18 model. Since 8-bit quantization is lossless we also compare 6-bit results. DFQ clearly outperforms traditional per-layer quantization but stays slightly below per-channel quantization and higher level approaches such as QT and RQ .
Overall DFQ sets a new state-of-the-art for 8-bit fixed point quantization on several models and computer vision tasks. It is especially strong for mobile friendly architectures such as MobileNetV1 and MobileNetV2 which were previously hard to quantize. Even though DFQ is an easy to use level 1 approach, we generally show competitive performance when comparing to more complex level 2-4 approaches.
Conclusion
In this work, we introduced DFQ, a data-free quantization method that significantly helps quantized model performance without the need for data, fine-tuning or hyper-parameter optimization. The method can be applied to many common computer vision architectures with a straight-forward API call. This is crucial for many practical applications where engineers want to deploy deep learning models trained in FP32 to INT8 hardware without much effort. Results are presented for common computer vision tasks like image classification, semantic segmentation and object detection. We show that our method compares favorably to per-channel quantization , meaning that instead the more efficient per-tensor quantization can be employed in practice. DFQ achieves near original model accuracy for almost every model we tested, and even competes with more complicated training based methods.
Further we introduced a set of quantization levels to facilitate the discussion on the applicability of quantization methods. There is a difference in how easy a method is to use for generating a quantized model, which is a significant part of the impact potential of a quantization method in real world applications. We hope that the quantization levels and methods introduced in this paper will contribute to both future research and practical deployment of quantized deep learning models.
Acknowledgments
We would like to thank Christos Louizos, Harris Teague, Jakub Tomczak, Mihir Jain and Pim de Haan for their helpful discussions and valuable feedback.
References
Appendix A Optimal range equalization of two layers
Consider two fully-connected layers with weight matrices and , that we scale as in 4.1. We investigate the problem of optimizing the quantization ranges by rescaling the weight matrices by , where , such that and the weight matrices after rescaling. We investigate the case of symmetric quantization, which also gives good results in practice for asymmetric quantization. We denote
where are the per-channel weight ranges that are scaled by , the range for the scaled weight matrix and are the original unscaled ranges.
Using this in our optimization goal of eq. 9 leads to
We observe that the specific scaling of each channel cancels out as long as they do not increase , the range of the full weight matrix. We can reformulate the above to
By contradiction, if there is a small positive such that which will decrease by without affecting . Therefore such a solution would not be optimal for eq. 26.
The condition from eq. 27 implies there is a limiting channel which defines the quantization range of both weight matrices and . However, our optimization goal is not affected by the choice of the other given the resulting and are smaller than or equal to and , respectively. To break the ties of solutions we decide to set . Thus the channel’s ranges between both tensors are matched as closely as possible and the introduced quantization error is spread equally among both weight tensors. This results in our final rescaling factor
which satisfies our necessary condition from eq. 27 and ensures that .
Appendix B Bias correction for convolutional layers
Appendix C Clipped normal distribution
Given a normally distributed random variable with mean and variance , and a clipped-linear function that clips its argument to the range , s.t. , the mean and variance of can be determined using the standard rules of computing the mean and variance of a function:
Using the fact that is constant if we have that:
The first and last term can be computed as and respectively, where we define , , and , the normal CDF with zero mean and unit variance.
The integral over the linear part of can be computed as:
where we define , i.e. the standard normal pdf and is the normalization constant for a normal distribution with variance , thus
C.2 Variance of Clipped Normal Distribution
We again exploit the fact that is constant if :
The first and last term can be solved as and respectively.
The second term can be decomposed as follows:
where we use the result from the previous subsection and define , and where is the mean of the truncated normal distribution.
Appendix D Empirical quantization bias correction
If a network does not use batch normalization, or does not use batch normalization in all layers, a representative dataset can be used to compute the difference between pre-activation means before and after quantization. We then subtract this difference from the quantized model’s pre-activations. This procedure can be run with unlabeled data. The procedure should be run after BatchNorm folding and cross-layer range equalization. Clipping should be applied in the quantized network, but not in the floating point network. Since the activation function and the quantization operation are fused, this procedure is run on a network with quantized weights only. However, after this procedure is applied activations can be quantized as well. We bias correct a layer only after all the layers feeding into it have been bias-corrected. The procedure is as follows:
For each layer in the quantized model:
In Table 6 we compare this empirical bias correction procedure with the analytic bias correction introduced in section 4.2. We observe that both approaches leads to similar results.
Appendix E Additional experiments
The focus of our method is data-free quantization (level 1). However, our method can also be used as a pre-processing before quantization aware fine-tuning. To demonstrate this we used DFQ together with short-term quantization aware fine-tuning . After just 1 epoch of quantization aware fine-tuning MobileNet V2, accuracy increases from 71.19% to 71.42%, almost recovering the FP32 performance (71.72%).
In our experimental section all our experiments were performed with asymmetric quantization, since this is commonly used in literature. Here we also compare to symmetric quantization. Symmetric quantization does not use an offset, which eliminates several cross terms in the calculations done on hardware compared to asymmetric quantization. This makes symmetric quantization more efficient on some hardware at the expense of losing some expressive power.
In Table 7 we compare symmetric and asymmetric quantization in combination with DFQ. For all three models the advantage of asymmetric quantization is almost negligible. We noticed that cross-layer equalization is effective at removing outliers, resulting in weight distributions are often close to symmetric.
In our experiments we focused on per-tensor quantization since the more recent per-channel quantization is not efficiently supported on all hardware. For hardware that does support it, we analyze the effect of DFQ in combination with per-channel quantization.
In Table 8 we show the results of the different components of DFQ in combination with per-channel quantization. We notice that each individual component, cross-layer equalization, bias absorption and bias correction, incremental improve over per-channel quantization and reduce the total quantization error from 1.07% to only 0.39%.