Loss Aware Post-training Quantization

Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Alex M. Bronstein, Avi Mendelson

Introduction

Deep neural networks (DNNs) are a powerful tool that have shown unmatched performance in various tasks in computer vision, natural language processing and optimal control, to mention only a few. The high computational resource requirements, however, constitute one of the main drawbacks of DNNs, hindering their massive adoption on edge devices. With the growing number of tasks performed on edge devices, e.g., smartphones or embedded systems, and the availability of dedicated custom hardware for DNN inference, the subject of DNN compression has gained popularity.

One way to improve DNN computational efficiency is to use lower-precision representation of the network, also known as quantization. Most of the literature on neural network quantization involves training either from scratch or performing fine-tuning on a pre-trained full-precision model . While training is a powerful method to compensate for accuracy loss due to quantization, it is often desirable to be able to quantize the model without training since this is both resource consuming and requires access to the data the model was trained on, which is not always available. These methods are commonly referred to as post-training quantization and usually require only a small calibration dataset. Unfortunately, current post-training methods are not very efficient, and most existing works only manage to quantize parameters to the 8-bit integer representation (INT8).

In the absence of a training set, these methods typically aim at minimizing the local error introduced during the quantization process (e.g., round-off errors). Recently, a popular approach to minimizing this error has been to clip the tensor outliers. This means that peak values will incur a larger error, but, in total, this will reduce the distortion introduced by the limited resolution . Unfortunately, these schemes suffer from two fundamental drawbacks.

Firstly, it is hard or even impossible to choose an optimal metric for the network performance based on the tensor level quantization error. In particular, even for the same task, similar architectures may favor different objectives. Secondly, the noise in earlier layers might be amplified by successive layers, creating a dependency between quantizaton errors of different layers. This cross-layer dependency makes it necessary to jointly optimize the quantization parameters across all network layers; however, current methods optimize them separately for each layer. Figure 1 plots the challenging loss surface of this optimization process.

Below, we outline the main contributions of the present work along with the organization of the remaining sections.

First we consider current layer-by-layer quantization methods where the quantization step size within each layer is optimized to accommodate the dynamic range of the tensor while keeping it small enough to minimize quantization noise. Although these methods optimize the quantization step size of each layer independently of the other layers, we observe strong interactions between the layers, explaining their suboptimal performance at the network level.

Accordingly, we consider network quantization as a multivariate optimization problem where the layer quantization step sizes are jointly optimized to minimize the neural network cross-entropy loss. We observe that layer-by-layer quantization identifies solutions in a small region around the optimum, where degradation is quadratic in the distance from it. We provide analytical justification as well as empirical evidence showing this effect.

Finally, we propose to combine layer-by-layer quantization with multivariate quadratic optimization. Our method is shown to significantly outperform state-of-the-art methods on two different challenging tasks and six DNN architectures.

The rest of the paper is organized as follows: Section 2 reviews the related work, Section 3 studies the properties of the loss function of the quantized network, Section 4 describes a proposed method, Section 5 provides the experimental results, and Section 6 concludes the paper.

Related work

The approaches to neural network quantization can be divided roughly into two major categories : quantization-aware training, which introduces the quantization at some point during the training, and post-training quantization, where the network weights are not optimized during the quantization. The recent progress in quantization-aware training has allowed the realization of results comparable to a baseline for as low as 2–4 bits per parameter and show decent performance even for single bit (binary) parameters . The major drawback of quantization-aware methods is the necessity for a vast amount of labeled data and high computational power. Consequently, post-training quantization is widely used in existing embedded hardware solutions. Most work that proposes hardware-friendly post-training schemes has only managed to get to 8-bit quantization without significant degradation in performance or requiring substantial modifications in exisiting hardware.

Another efficient approach quantization is to change the distribution of the tensors, making it more suitable for quantization, such as equalizing the weight ranges or splitting outlier channels .

Recent works noted that the quantization process introduces a bias into distributions of the parameters and focused on correction of this bias. Finkelstein et al. addressed a problem of MobileNet quantization. They claimed that the source of degradation was shifting in the mean activation value caused by inherent bias in the quantization process and proposed a scheme for fixing this bias. Multiple alternative schemes for the correction of quantization bias were proposed, and those techniques are widely applied in state-of-the-art quantization approaches . In our work, we utilize the bias correction method by developed Banner et al. .

While the simplest form of quantization applies a single quantizer to the whole tensor (weights or activations), finer quantization allows reduction of performance degradation. Even though this approach usually boosts performance significantly , it requires more parameters and special hardware support, which makes it unfavorable for real-life deployment.

One more source of performance improvement is using more sophisticated ways to map values to a particular bin. This includes clustering of the tensor entry values and non-uniform quantization . By allowing a more general quantizer, these approaches provide better performance than uniform quantization, but also require hardware support for efficient inference, and thus, similarly to fine-grained quantization, are less suitable for deployment on consumer-grade hardware.

To the best of our knowledge, previous works did not take into account the fact that the loss might be not separable during optimization, usually performed per layer or even per channel. Notable exceptions are Gong et al. and Zhao et al. , who did not show any kind of optimization. Nagel et al. partially addressed the lack of separability by treating pairs of consecutive layers together.

Loss landscape of quantized DNNs

In this section, we introduce the notion of separability of the loss function. We study the separability and the curvature of the loss function and show how quantization of DNNs affect these properties. Finally, we show that during aggressive quantization, the loss function becomes highly non-separable with steep curvature, which is unfavorable for existing post-training quantization methods. Our method addresses these properties and makes post-training quantization possible at low bit quantization.

By constraining the range of xx to [−c,c][-c,c], the connection between cc and Δ\Delta is given by:

In the case of activations, we limit ourselves to the ReLU function, which allows us to choose a quantization range of [0,c][0,c]. In such cases, the quantization step Δ\Delta is given by:

Suppose the loss function of the network L\mathcal{L} depends on a certain set of variables (weights, activations, etc.), which we denote by a vector v\mathbf{v}. We would like to measure the effect of adding quantization noise to this set of vectors. In the following we show that for sufficiently small quantization noise, we can treat it as an additive noise vector ε\bm{\varepsilon}, allowing coordinate-wise optimization. However, when quantization noise is increased, the degradation in one layer is associated with other layers, calling for more laborious non-separable optimization techniques. From the Taylor expansion:

When the quantization error ε\bm{\varepsilon} is sufficiently small, higher-order terms can be neglected so that degradation ΔL\Delta\mathcal{L} can be approximated as a sum of the quadratic functions,

One can see from Eq. 5 that when quantization error \normε2\norm{\bm{\varepsilon}}^{2} is sufficiently small, the overall degradation ΔL\Delta\mathcal{L} can be approximated as a sum of NN independent separable degradation processes as follows:

On the other hand, when \normε2\norm{\bm{\varepsilon}}^{2} is larger, one needs to take into account the interactions between different layers, corresponding to the second term in Eq. 5 as follows:

where QIT refers to the quantization interaction term.

In Fig. 2 we provide a visualization of these interactions at 2, 3 and 4 bitwidth representations.

2 Curvature

We now analyze how the steepness of the curvature of the loss function with respect to the quantization step changes as the quantization error increases. We will show that at aggressive quantization, the curvature of the loss becomes steep, which is unfavorable for the methods that aim to minimize quantization error on the tensor level.

We start by defining a measure to quantify curvature of the loss function with respect to quantization step size Δ\Delta. Given a quantized neural network, we denote L(Δ1,Δ2,…,Δn)\mathcal{L}(\Delta_{1},\Delta_{2},\dots,\Delta_{n}) the loss with respect to quantization step size Δi\Delta_{i} of each individual layer. Since L\mathcal{L} is twice differentiable with respect to Δi\Delta_{i} we can calculate the Hessian matrix:

To quantify the curvature, we use Gaussian curvature , which is given by:

where Δ=\quantity(Δ1,Δ2,…,Δn)\Delta=\quantity(\Delta_{1},\Delta_{2},\dots,\Delta_{n}). We calculated the Gaussian curvature at the point that minimizes the L2L_{2} norm of the quantization error and acquired the following values:

This means that the flat surface for 4 bits, shown in Fig. 2, is a generic property of the fine-grained quantization loss and not of the specific layer choice. Similarly, we conclude that coarser quantization generally has steeper curvature than more fine-grained quantization.

Moreover, the Hessian matrix provides additional information regarding the coupling between different layers. As could be expected, adjacent off-diagonal terms have higher values than distant elements, corresponding to higher dependencies between clipping parameters of adjacent layers (the Hessian matrix is presented in Fig. 0.A.1 in the Appendix).

When the curvature is steep, even small changes in quantization step size may change the results drastically. To justify this hypothesis experimentally, in Fig. 3, we evaluate the accuracy of ResNet-50 for five different quantization steps. We choose quantization steps which minimize the LpL_{p} norm of the quantization error for different values of pp.

While at 4-bit quantization, the accuracy is almost not affected by small changes in the quantization step size, at 2-bit quantization, the same changes shift the accuracy by more than 20%. Moreover, the best accuracy is obtained with a quantization step that minimizes L3.5L_{3.5} and not the MSE, which corresponds to L2L_{2} norm minimization.

Loss Aware Post-training Quantization (LAPQ)

In the previous section, we showed that the loss function L(Δ)\mathcal{L}(\Delta), with respect to the quantization step size Δ\Delta, has a complex, non-separable landscape that is hard to optimize. In this section, we suggest a method to overcome this intrinsic difficulty. Our optimization process involves three consecutive steps.

In the first phase, we find the quantization step Δp\Delta_{p} that minimizes the LpL_{p} norm of the quantization error of the individual layers for several different values of pp. Then, we perform quadratic interpolation to approximate an optimum of the loss with respect to pp. Finally, we jointly optimize the parameters of all layers acquired on the previous step by applying a gradient-free optimization method . The pseudo-code of the whole algorithm is presented in Algorithm 1. Fig. 6 provides the algorithm visualization.

Our method starts by minimizing the LpL_{p} norm of the quantization error of weights and activations in each layer with respect to clipping values:

Given a real number p>0p>0, the set of optimal quantization steps Δp={Δ1p,Δ2p,…,Δnp}\Delta_{p}=\{\Delta_{1_{p}},\Delta_{2_{p}},\dots,\Delta_{n_{p}}\}, according to Eq. 12, minimizes the quantization error within each layer. In addition, optimizing the quantization error allows us to get Δp\Delta_{p} in the vicinity of the optimum Δ∗\Delta^{*} . Different values of pp result in different quantization step sizes Δp\Delta_{p}, which are still optimal under some metric LpL_{p} due to the trade-off between clipping and quantization error (Fig. 4).

2 Quadratic approximation

Assuming a quantization step size Δ\Delta in the vicinity of the optimal quantization step Δ∗\Delta^{*}, the loss function can be approximated with a Taylor series as follows:

where H(Δ∗)\mathbf{H}(\Delta^{*}) is the Hessian matrix with respect to Δ\Delta. Since Δ∗\Delta^{*} is a minimum, the first derivative vanishes and we acquire a quadratic approximation of L\mathcal{L},

Our method exploits this quadratic property for optimization. Fig. 5(a) demonstrates empirical evidence of such a quadratic relationship for ResNet-18 around the optimal quantization step Δ∗\Delta^{*} obtained by our method. First, we sample a few data points {Δp}\{\Delta_{p}\} to build a trajectory on the graph of L(Δ)\mathcal{L}(\Delta) (orange points in Fig. 6). Then, we use the prior quadratic assumption to approximate the minimum of the L\mathcal{L} on that trajectory by fitting a quadratic function f(p)f(p) to the sampled Δp\Delta_{p}. Finally, we minimize f(p)f(p) and use the optimal quantization step size Δp∗\Delta_{p^{*}} as a starting point for a gradient-free joint optimization algorithm, such as Powell’s method , to minimize the loss and find Δ∗\Delta^{*}.

3 Joint optimization

By minimizing both the quantization error and the loss using quadratic interpolation, we get a decent approximation of the global minimum Δ∗\Delta^{*}. Due to steep curvature of the minimum, however, for a low bitwidth quantization, even a small error in the value of Δ\Delta leads to performance degradation. Thus, we use a gradient-free joint optimization, specifically an iterative algorithm (Powell’s method ), to further optimize Δp∗\Delta_{p^{*}}.

At every iteration, we optimize the set of parameters, initialized by Δp∗\Delta_{p^{*}}. Given a set of linear search directions D=\quantityd1,d2,…,dND=\quantity{d_{1},d_{2},\dots,d_{N}}, the new position Δt+1\Delta_{t+1} is expressed by the linear combination of the search directions as following Δt+∑iλidi\Delta_{t}+\sum_{i}{\lambda_{i}d_{i}}. The new displacement vector ∑iλidi\sum_{i}{\lambda_{i}d_{i}} becomes part of the search directions set, and the search vector, which contributed most to the new direction, is deleted from the search directions set. For further details, see Algorithm 1.

Experimental Results

In this section we conduct extensive evaluations and offer a comparison to prior art of the proposed method on two challenging benchmarks, image classification on ImageNet and a recommendation system on NCF-1B. In addition, we examine the impact of each part of the proposed method on the final accuracy. In all experiments we first calibrate the optimal clipping values on a small held-back calibration set using our method and then evaluate the validation set.

We evaluate our method on several CNN architectures on ImageNet. We select a calibration set of 512 random images for the optimization step. The size of the calibration set defines the trade-off between generalization and running time (the analysis is given in Appendix 0.B). Following the convention , we do not quantize the first and last layers.

Many successful methods of post-training quantization perform finer parameter assignment, such as group-wise , channel-wise , pixel-wise or filter-wise quantization, which require special hardware support and additional computational resources. Finer parameter assignment appears to provide unconditional improvement, independently of the underlying methods used. In contrast with those approaches, our method performs layer-wise quantization, which is simple to implement on any existing hardware that supports low precision integer operations. Thus, we do not include the above-mentioned methods in our comparison study. We apply bias correction, as proposed by Banner et al. , on top of the proposed method. In Table 1 (more results are available in Table 0.C.1 in the Appendix) we compare our method with several other layer-wise quantization methods, as well as the minimal MSE baseline. In most cases, our method significantly outperforms all the competing methods, showing acceptable performance even for 4-bit quantization.

2 NCF-1B

In addition to the vision models, we evaluated our method on a recommendation system task, specifically on a Neural Collaborative Filtering (NCF) model. We use mlperfhttps://github.com/mlperf/training/tree/master/recommendation/pytorch implementation to train the model on the MovieLens-1B dataset. Similarly to the ImageNet, the calibration set of 50k random user/item pairs is significantly smaller than both the training and validation sets.

In Table 2 we present results for the NCF-1B model compared to the MMSE method. Even at 8-bit quantization, NCF-1B suffers from significant degradation when using the naive MMSE method. In contrast, LAPQ achieves near baseline accuracy with 0.5% degradation from FP32 results.

3 Ablation study

The proposed method comprises of two steps: layer-wise optimization and quadratic approximation as the initialization for the joint optimization. In Table 3, we show the results for ResNet-18 under different initializations and their results after adding the joint optimization. LAPQ suggest a better initialization for the joint optimization.

Prior research has shown that CNNs are sensitive to quantization bias. To address this issue, we perform bias correction of the weights as proposed by Banner et al. , which can easily be combined with LAPQ, in all our CNN experiments. In Table 4 we show the effect of adding bias correction to the proposed method and compare it with MMSE. We see that bias correction is especially important in compact models, such as MobileNet.

Conclusion

We have analyzed the loss function of quantized neural networks. At low precision, the function is non-separable, with steep curvature, which is unfavorable for existing post-training quantization methods. Accordingly, we have introduced Loss Aware Post-training Quantization (LAPQ), which jointly optimizes all quantization parameters by minimizing the loss function directly. We have shown that our method outperforms current post-training quantization methods. Also, our method does not require special hardware support such as channel-wise or filter-wise quantization. In some models, LAPQ is the first to show compatible performance in 4-bit post-training quantization regime with layer-wise quantization, almost achieving a the full-precision baseline accuracy.

The research was funded by National Cyber Security Authority and the Hiroshi Fujiwara Technion Cyber Security Research Center.

References

Appendix 0.A Hessian of the loss function

To estimate dependencies between clipping parameters of different layers, we analyze the structure of the Hessian matrix of the loss function. The Hessian matrix contains the second-order partial derivatives of the loss L(Δ)\mathcal{L}(\Delta), where Δ\bm{\Delta} is a vector of quantization steps:

In the case of separable functions, the Hessian is a diagonal matrix. This means that the magnitude of the off-diagonal elements can be used as a measure of separability. In Fig. 0.A.1 we show the Hessian matrix of the loss function under quantization of 4 and 2 bits. As expected, higher dependencies between quantization steps emerge under more aggressive quantization.

Appendix 0.B Calibration set size

Calibration set size reflects the balance between the running time and the generalization. To determine the required size, we ran the proposed method on ResNet-18 for various calibration set sizes and different bitwidths. As shown in Fig. 0.B.2, a calibration set size of 512 is a good choice to balance this trade-off.

Appendix 0.C ImageNet additional results

In Table 0.C.1 we show additional results for the proposed method for the image classification task on ImageNet dataset.