HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision

Zhen Dong, Zhewei Yao, Amir Gholami, Michael Mahoney, Kurt Keutzer

I Introduction

There has been a significant increase in the computational resources required for Neural Network (NN) training and inference. This is mainly due to larger input sizes (e.g., higher image resolution) as well as larger NN models requiring more FLOPs and significantly larger memory footprint. For example, in 1998 the state-of-the-art NN was LeNet-5 applied to MNIST dataset with an input image size of 1×28×281\times 28\times 28. Twenty years later, a common benchmark dataset is ImageNet, with an input resolution that is 200×200\times larger than MNIST, and with NN models that have orders of magnitude higher memory footprint.

In fact, ImageNet resolution is now considered “small” for many applications such as autonomous driving where input resolutions are significantly larger (more than 40×40\times in certain cases).

This combination of larger models and higher resolution images has created a major challenge in the deployment of NNs in application environments with computationally constrained resources such as surveillance systems or ADAS systems in passenger cars. This trend is going to accelerate further in the near future.

There has been a significant effort taken by many researchers to address these issues. These could be broadly categorized as follows. (i) Finding NNs that provide the required accuracy, while remaining compact by design (i.e., with small memory footprint) and which require relatively small FLOPs. SqueezeNet was an early effort here, followed by more efficient NNs such as . (ii) Co-designing NN architecture and hardware together. This can allow significant speed ups and savings in power consumption of the hardware without losing accuracy. SqueezeNext is an example work here where the neural network and associated accelerator are co-designed. (iii) Pruning redundant filters of NN layers. Seminal works here are . (iv) Applying AutoML for both hardware aware NN design as well as quantization. Notable works here are DNAS and HAQ . (v) Using quantization (reduced precision) instead of float or double precision, which can significantly speed up inference time and reduce power consumption. This paper exclusively focuses on quantization, but other approaches could be used in conjunction of our method to allow for further possible reduction on the model size.

Quantization needs to be performed for both NN parameters (i.e., weights) as well as the activations to reduce the total memory footprint of the model during inference. However, the main challenge here is that a naïve quantization can lead to significant loss in accuracy. In particular, it is not possible to reduce the number of bits of all weights/activations of a general convolutional network to ultra low-precision without significant accuracy loss. This is because not all the layers of a convolutional network allow the same quantization level. A possible approach to address this is to use mixed-precision quantization, where higher precision is used for certain “sensitive” layers of the network, and lower precision for “non-sensitive” layers. However, the search space for finding the right precision for each layer is exponential in the number of layers. Moreover, to avoid accuracy loss we need to perform fine-tuning (i.e. re-training) of the model. As we will discuss below, quantizing the whole model at once and then fine-tuning is not optimal. Instead, we need to perform multi-stage quantization, where at each stage parts of the network are quantized to low-precision followed by quantization-aware fine-tuning to recover accuracy. However, the search space to determine which layers to quantize first is factorial in the number of layers. In this paper, we propose a Hessian guided approach to address these challenges. In particular, our contributions are the following.

The search space for choosing mixed-precision quantization is exponential in the number of layers. Thus, we present a novel, deterministic method for determining the relative quantization level of layers based on the Hessian spectrum of each layer.

The search space for quantization-aware fine-tuning of the model is factorial in the number of blocks/layers. Thus, we propose a Hessian based method to determine fine-tuning order for different NN blocks.

We perform ablation study of HAWQ, and we present novel quantization results using ResNet20 on Cifar10, as well as Inception-V3/ResNet50/SqueezeNext on ImageNet. Comparison with state-of-the-art shows that our method achieves higher precision (up to 1%), smaller model size (up to 20%20\%), and smaller activation size (up to 8×8\times).

The paper is organized as follows. First, in § II, we will discuss related works on model compression. This is followed by describing our method in § III, and our results in § IV. Finally, we present ablation study in § V, followed by conclusions.

II Related work

Recently, significant efforts have been spent on developing new model compression solutions to reduce the parameter size as well as computational complexity of NNs . In , pruning is used to reduce the number of non-zero weights in NN models. This approach is very useful for models that have very large fully connected layers (such as AlexNet or VGG ).

For instance, the first fully-connected layer in VGG-16 occupies 408MB alone, which is 77.3% of total model size. Large fully-connected layers have been removed in other fully convolutional networks such as ResNet , or Inception family .

Knowledge distillation introduced in is another direction for compressing NNs. The main idea is to distill information from a pre-trained, large model into a smaller model. For instance, it was shown that with knowledge distillation it is possible to reduce model size by a factor of 3.63.6 with an accuracy of 91.61%91.61\% on Cifar-10 .

Another fundamental approach has been to architect models which are, by design, both small and hardware-efficient. An initial effort here was SqueezeNet which could achieve AlexNet level accuracy with 50×50\times smaller footprint through network design, and additional 10×10\times reduction through quantization , resulting in a NN with 500×500\times smaller memory footprint. Other notable works here are , where more accurate networks are presented. Another work here is SqueezeNext , where a similar approach is taken, but with co-design of both hardware architecture along with a compact NN model.

Quantization is another orthogonal approach for model compression, where lower bit representation are used instead of redesigning the NN. One of the major benefits of quantization is that it increases a NN’s arithmetic intensity (which is the ratio of FLOPs to memory accesses). This is particularly helpful for layers that are memory bound and have low arithmetic intensity. After quantization, the volume of memory accesses reduces, which can alleviate/remove the memory bottleneck.

However, directly quantizing NNs to ultra low precision may cause significant accuracy degradation.

One possibility to address this is to use Mixed-Precision quantization (MP) . A second possibility, Multi-Stage Quantization (MSQ), is proposed by . MP and MSQ can improve the accuracy of quantized NNs, but face an exponentially large search space. This is a major problem that has not been addressed in existing literature for quantization. Applying existing methods require often ad-hoc rules to choose precision of different layers which are problem/model specific and do not generalize. The goal of our work here is to address this challenge using second-order information.

III Methodology

Assume that the NN is partitioned into bb blocks denoted by {B1, B2 …, Bb}\{B_{1},~{}B_{2}~{}\ldots,~{}B_{b}\}, with learnable parameters {W1, W2, …, Wb}\{W_{1},~{}W_{2},~{}\ldots,~{}W_{b}\}. A block can be a single/multiple layer(s) (or a single/multiple residual block(s) for the case of residual networks). For a supervised learning framework, the loss function L(θ)L(\theta) is:

The training is performed by solving an Empirical Risk Minimization problem, to find the optimal model parameters. This process is typically performed in single precision, where both the weights and activations are stored with 32-bit precision.

After the training is finished, each of these blocks will have a specific distribution of floating point numbers for both the parameters, θ\theta, as well as input/output activations. For quantization, we need to restrict these floating numbers to a finite set of values, defined by the following function:

where (tj,tj+1](t_{j},t_{j+1}] denotes an interval in the real numbers (j=0, … ,2k−1j=0,~{}\ldots~{},2^{k}-1), kk is the quantization bits, and zz is either an activation or the weights. This means that all the values in the range of (tj,tj+1](t_{j},t_{j+1}] are mapped to qjq_{j}. In the extreme case of binary quantization (k=1k=1), Q(z)Q(z) is basically the sign function. For cases other than binary quantization, the choice of these intervals can be important. One popular option is to use a uniform quantization function, where the above range is equally split . However, it has been argued that (i) not all layers have the same distribution of floating point values, and (ii) the network can have significantly different sensitivity to quantization of each layer. To address the first issue, different quantization schemes such as uniformly discretizing logarithmic-domain have been proposed . However, this does not completely address the sensitivity problem. A sensitive layer cannot be quantized to the same level as a non-sensitive layer.

One possible approach that can be used to measure quantization sensitivity is to use first-order information, based on the gradient vector. However, the gradient can be very misleading. This can be easily illustrated by considering a simple 1-d parabolic function of the form y=12ax2y=\frac{1}{2}ax^{2} at origin (i.e., x=0x=0). The gradient signal at the origin is zero, irrespective of the value of aa. However, this does not mean that the function is not sensitive to perturbation in xx. We can get a better metrics for sensitivity by using second-order information, based on the Hessian matrix. This clearly shows that higher values of aa result in more sensitivity to input perturbations.

For the case of high dimensions, the second order information is stored in the Hessian matrix, of size ni×nin_{i}\times n_{i} for each block. For this case, we can compute the eigenvalues of the Hessian to measure sensitivity, as described next.

We compute the eigenvalues of the Hessian (i.e., the second-order operator) of each block in the network. Note that it is not possible to explicitly form the Hessian since the size of a block (denoted by nin_{i} for ith block) can be quite large. However, it is possible to compute the Hessian eigenvalues without explicitly forming it, using a matrix-free power iteration algorithm . This method requires computation of the so-called Hessian matvec, which is the result of multiplication of the Hessian matrix with a given (possibly random) vector vv. To illustrate how this can be done for a deep network, let us first denote gig_{i} as the gradient of loss LL with respect to the ithi^{th} block parameters,

For a random vector vv (which has the same dimension as gig_{i}), we have:

where HiH_{i} is the Hessian matrix of LL with respect to WiW_{i}. Note that the second equality above, comes from the fact that vv is independent of WiW_{i}. We can then use power-iteration method to compute the top eigenvalue of HiH_{i}, as shown in Algorithm 1. Intuitively the algorithm requires multiple evaluations of the Hessian matvec, which can be computed using Eq. 4.

It is well known, based on the theory of Minimum Description Length (MDL), that fewer bits are required to specify a flat region up to a given threshold, and vice versa for a region with sharp curvature . The intuition for this is that the noise created by imprecise location of a flat region is not magnified for a flat region, making it more amenable to aggressive quantization. The opposite is true for sharp regions, in that even small round off errors may be amplified. Therefore, it is expected that layers with higher Hessian spectrum (i.e., larger eigenvalues) are more sensitive to quantization. The distribution of these eigenvalues for different blocks are shown in Figure 1 for ResNet20 on CIFAR-10 and Inception-V3 on ImageNet. As one can see, different blocks exhibit orders of magnitude difference in the Hessian spectrum. For instance, ResNet20 is an order of magnitude more sensitive to perturbations to its 9th9^{th} block, than its last block.

To further illustrate this, we provide 1D visualizations of the loss landscape as well. To this end, we first compute the Hessian eigenvector of each block, and we perturb each block individually along the eigenvector and compute how the loss changes. This is illustrated in Figure 2 and 3 for ResNet20 (on Cifar-10) and Inception-V3 (on ImageNet), respectively. It can be clearly seen that blocks with larger Hessian eigenvalue (i.e., sharper curvature) exhibit larger fluctuations in the loss, as compared to those with smaller Hessian eigenvalue (i.e., flatter curvature). A corresponding 3D plot is also shown in Figure 1, where instead of just considering the top eigenvector, we also compute the second top eigenvector and visualize the loss by perturbing the weights along these two directions. These surface plots are computed for the 9th9^{th} and last blocks of ResNet20, as well as 2nd and last blocks of Inception-V3 (the loss landscape for other blocks is shown in Figure 6 and Figure 7).

III-B Algorithm

We approximate the Hessian as a block diagonal matrix, scaled by its top eigenvalue, λ\lambda, as {Hi≈λiI}i=1b\{H_{i}\approx\lambda_{i}I\}_{i=1}^{b}. Based on the MDL theory, layers with large λ\lambda cannot be quantized to ultra low precision without significant perturbation to the model. Thus we can use the Hessian spectrum of each block to sort the different blocks and perform less aggressive quantization to layers with large spectrum. However, some of these blocks may contain very large number of parameters, and using higher bits here would lead to large memory footprint of the quantized network. Therefore, as a compromise, we weight the spectrum with block’s memory footprint and use the following metric for sorting the blocks:

where λi\lambda_{i} is the top eigenvalue of HiH_{i}. Based on this sorting, layers that have large number of parameters and have small eigenvalue would be quantized to lower bits, and vice versa. That is, after SiS_{i} is computed, we sort SiS_{i} in descending order and use it as a metric to determine the quantization precision.Note that, as mentioned in the limitations section, SiS_{i} does not give us the exact bit precision but a relative ordering for the bits of different blocks.

Quantization-aware re-training of the neural network is necessary to recover performance which can sharply drop due to ultra-low precision quantization. A straightforward way to do this is to re-train (hereafter referred to as fine-tune) the whole quantized network at once. However, as we will discuss in §IV, this can lead to sub-optimal results. A better strategy is to perform multi-stage fine-tuning. However, the order in multi-stage tuning is important and different ordering could lead to very different accuracies.

We sort different blocks for fine-tuning based on the following metric:

where ii refers to ithi^{th} block, λi\lambda_{i} is the Hessian eigenvalue, and ∥Q(Wi)−Wi∥2\|Q(W_{i})-W_{i}\|_{2} is the L2L_{2} norm of quantization perturbation. The intuition here is to first fine-tune layers that have high curvature as well as large number of parameters which cause more perturbations after quantization. Note that the latter metric depends on the bits used for quantization and is not a fixed metric. (See Table V in appendix, where we show how this metric changes for different quantization precision.) The motivation for choosing this order is that fine-tuning blocks with large Ωi\Omega_{i} can significantly affect other blocks, thus making prior fine-tuning of layers with small Ωi\Omega_{i} futile.

IV Results

In this section, we first present our quantization results for ResNet20 on Cifar-10, and then we present our results for Inception-V3, ResNet50, and SqueezeNext quantization on ImageNet. See appendix for details regarding the training procedure and hyper-parameters used.

After computing the eigenvalues of block Hessian (shown in Figure 1), we compute the weighted sensitivity metric of Eq. 5, along with Ωi\Omega_{i} based on Eq. 6. We then perform the quantization based on HAWQ algorithm. Results are shown in Table I.

For comparison, we test the quantization performance without using the Hessian information, which we refer to as “Direct” method, as well as other methods in the literature including Dorefa , PACT , LQ-Net , and DNAS , as shown in Table I.

For methods that use Mixed-Precision (MP), we report the lowest bits used for weights (“w-bits”), and activations (“a-bits”).

The Direct method achieves good compression, but it results in 2.03%2.03\% accuracy drop, as shown in Table I.Furthermore, comparison with other state-of-the-art shows a similar trend. There have been several methods proposed in the literature to address this reduction, with the latest method introduced in , where a learnable quantization method is used. As one can see, LQ-Nets results in 0.77%0.77\% accuracy degradation with 10.67×10.67\times compression ratio, whereas HAWQ has only 0.15%0.15\% accuracy drop with 13.11×13.11\times compression. Moreover, HAWQ achieves similar accuracy as compared to DNAS but with 8×8\times higher compression ratio for activations.

ImageNet

Here, we test the HAWQ method for quantizing Inception-V3 on ImageNet. Inception-V3 is appealing for efficient hardware implementation, as it does not use any residual connections. Such non-linear structures create dependencies that may be very difficult to optimize for fast inference . As before, we first compute the block Hessian eigenvalues, which are reported in Figure 1, and then compute the corresponding weighted sensitivity metric. We also plot the 1D loss landscape of all Inception-V3 blocks in Figure 3.

We report the quantization results in Table II, where as before we compare with a direct quantization, as well as recently proposed “Integer-Only” , and RVQuant methods . Direct quantization of Inception-V3 (i.e., without use of second-order information), results in 7.69%7.69\% accuracy degradation. Using the approach proposed in results in more than 2%2\% accuracy drop, even though it uses higher bit precision. However, HAWQ results in a generalization gap of 2%2\% with a compression ratio of 12.04×12.04\times, both of which are better than previous work .We should emphasize here that the work of uses integer arithmetic, and it is not completely fair to compare their results with ours.

We also compare with Deep Compression and the AutoML based method of HAQ, which has been recently introduced . We compare our HAWQ results with their ResNet50 quantization, as shown in Table III. HAWQ achieves higher top-1 accuracy of 75.48% with a model size of 7.96MB, whereas the AutoML based HAQ method has a top-1 of 75.30% even with 16% larger model size of 9.22MB.

Furthermore, we apply HAWQ to quantize SqueezeNext on ImageNet. We choose the wider SqueezeNext model which has a baseline accuracy of 69.38%69.38\% with 2.5 million parameters (10.1MB in single precision). We are able to quantize this model to uniform 8-bit precision, with just 0.04% top-1 accuracy drop. Direct quantization of SqueezeNext (i.e., without use of second-order information), results in 3.98%3.98\% accuracy degradation. HAWQ results in an unprecedented 1MB model size, with only 1.36% top-1 accuracy drop. The significance of this result is that it allows deployment of the whole model on-chip or on hardwares with very limited memory and power constraints.

V Ablation Study

Here we discuss the ablation study for the HAWQ. The HAWQ method has two main steps: (i) relative precision order for different blocks using second-order information, and (ii) relative order for fine-tuning these blocks. Below we discuss the ablation study for each step separately.

We first discuss the ablation study for step (i), where the quantization precision is chosen based on Eq. 5. As discussed above, blocks with higher values of SiS_{i} are assigned higher quantization precision, and vice versa for layers with relatively lower values of SiS_{i}. For the ablation study we reverse this order and avoid performing the block-wise fine-tuning of step (ii) so we can isolate step (i). Instead of the fine-tuning phase, we re-train the whole network at once after the quantization is performed. The results are shown in Figure 4, where we perform 50 epochs of fine-tuning using Inception-V3 on ImageNet. As one can see, HAWQ results in significantly better accuracy (74.26% as compared to 66.72%) than the reverse method (labeled as “HAWQ-Reverse-Precision”). This is despite the fact that the latter approach only has a compression ratio of 7.2×7.2\times, whereas HAWQ has a compression ratio of 12.0×12.0\times.

Another interesting observation is that the convergence speed of the Hessian aware approach is significantly faster than the reverse method. Here, HAWQ converges in about 30 epochs, whereas the HAWQ-Reverse-Precision case takes 50 epochs before converging to a sub-optimal value (Figure 4).

V-B Block-Wise Fine-Tuning

Here we perform the ablation study for the Hessian based fine-tuning part of HAWQ. The block-wise fine tuning is performed based on Ωi\Omega_{i} (Eq. 6) of each block. The blocks are fine-tuned based on the descending order of Ωi\Omega_{i}. Similar to the above, we compare the quantization performance when a reverse ordering is used (i.e., we use the ascending order of Ωi\Omega_{i} and refer to this as “HAWQ-Reverse-Tuning”).

We test this ablation study using Inception-V3 on ImageNet, as shown in Figure 5. As one can see, the fine-tuning for HAWQ method quickly converges in just 25 epochs, allowing it to switch to fine-tuning the next block. However, “HAWQ-Reverse-Tuning” takes more than 50 epochs to converge for this block.

VI Conclusions

We have introduced HAWQ, a new quantization method for neural network training. Our method is based on exploiting second-order (Hessian) information to systematically select both quantization precision as well as the order for block-wise fine-tuning. We performed an ablation study for both the relative quantization bit-order for different blocks, as well as the fine-tuning order. We showed that HAWQ can achieve good testing performance with high compression-ratio, as compared to state-of-the-art. In particular, we showed results for ResNet20 on Cifar-10, where we can achieve similar testing performance as , but with 8×8\times higher compression ratio for activations. We also showed results for Inception-V3 on ImageNet, for which we showed ultra low precision quantization results with 2-bit for weights and 4-bit for activations, with only 1.93%1.93\% accuracy drop. For ResNet50 model, our approach results in higher accuracy of 75.48% with smaller model size of 7.96MB, as compared to HAQ method with top-1 of 75.30% and 9.22MB . Furthermore, our method applied to SqueezeNext can result in an unprecedented 1MB model size with 68.02 top-1 accuracy on ImageNet.

Limitations and Future Work. We believe it is critical for every work to clearly state its limitations, especially in this area. An important limitation is that computing the second-order information adds some computational overhead. However, we only need to compute the top eigenvalue of the Hessian, which can be found using the matrix-free method presented in Algorithm 1. (The total computational overhead is equivalent to about 20 gradient back-propogations to compute top Hessian eigenvalue of each block). Another limitation is that in this work we solely focused on image classification, but it would be interesting to see how HAWQ would perform for more complex tasks such as segmentation, object detection, or natural language processing. Furthermore, one has to consider that implementation of a NN with mixed-precision inference for embedded processors is not as straightforward as the case with uniform quantization precision. Practical solutions have been proposed in recent works . Another limitation is that we can only determine the relative ordering for quantization precision, and not the absolute value of the bits. However, the search space for this is significantly smaller than the original exponential complexity. Finally, even though we showed benefits of HAWQ as compared to DNAS or HAQ , it may be possible to combine these methods for more efficient AutoML search. We leave this as part of future work.

Acknowledgments

This work was supported by a gracious fund from Intel corporation, Berkeley Deep Drive (BDD), and Berkeley AI Research (BAIR) sponsors. We would like to thank the Intel VLAB team for providing us with access to their computing cluster. We also gratefully acknowledge the support of NVIDIA Corporation for their donation of two Titan Xp GPU used for this research. We would also like to acknowledge ARO, DARPA, NSF, and ONR for providing partial support of this work.

References

VII Appendix

Here, we provide additional experimental results as well as quantization details for the neural networks that we tested.

In § VII-A we discuss the fine-tuning details.

In § VII-B we present extra results for 3D plots for loss landscape of different blocks of ResNet20 and Inception-V3 as well as exemplary results showing distribution of Ωi\Omega_{i} in Eq. 6.

In § VII-C we show the exact bit-precision used for different blocks of ResNet20 on Cifar-10 as well as Inception-V3 on ImageNet.

The results were tested on two classification datasets of Cifar-10 and ImageNet:

This is a classification dataset with 10 classes consisting of 50,000 training images and 10,000 test images of size 3×32×323\times 32\times 32. We used pre-trained ResNet20 model and performed quantization on this model in PyTorch framework. We follow the same learning rate policy as the baseline (i.e., decaying learning rate from 0.1 to 0.0001).

ImageNet

This is a classification problem with 1000 classes consisting of more than 1.2 million training images and 50,000 validation images of size 3×224×2243\times 224\times 224 on SqueezeNext and ResNet50 , and 3×299×2993\times 299\times 299 on Inception-V3. (i) We used pre-trained Inception-V3 model and used a fixed learning rate of 0.0002 for fine-tuning of each block. (ii) We used pre-trained ResNet50 model and used a fixed learning rate of 0.0001 for fine-tuning of each block. (iii) We used pre-trained SqueezeNext and used a fixed learning rate of 0.0001 for fine-tuning of each block. All experiments were performed on PyTorch framework. As for data augmentation, we used standard random crop, resizing and horizontal flip in all experiments.

VII-B Extra results

In Table V, we show how Ωi\Omega_{i} changes as a function of quantization precision. In Figure 6, we plot the rest surface visualization of ResNet20 on Cifar-10. And in Figure 7, we plot the rest surface visualization of Inception-V3 on ImageNet.

VII-C Mixed-precision details

In this section, we give the details about how we separate blocks and details about weight/activation precision of each individual block. We show the exact bit-precision used for different blocks of ResNet20 (Table VI) on Cifar-10 as well as Inception-V3 (Table VII).