Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Christopher De Sa, Zhiru Zhang
Introduction
Over the past few years, deep neural networks (DNNs) have become the state-of-the-art approach for many large-scale computer vision and sequence modeling problems. Deep convolutional networks dominate the leaderboards for popular image classification and object detection datasets such as ImageNet (Deng et al., 2009) and Microsoft COCO (Lin et al., 2014). However, the significant compute and memory requirements of running DNNs impedes the adoption of neural nets in application domains such as edge computing or latency-critical services (Xu et al., 2018). One approach to reducing the costs of DNN execution is to quantize the floating-point weights and activations into low-precision fixed-point numbers. This reduces the model size as well as the complexity of multiply-accumulate (MAC) operations in hardware, enabling better throughput and energy efficiency. DNN quantization is an active area of research (Wu et al., 2018; Jacob et al., 2018; Choi et al., 2018b; Banner et al., 2018) and sees deployment in commercial systems such as Google’s TPU (Jouppi et al., 2017), NVIDIA’s TensorRT (Migacz, 2017), and Microsoft’s Brainwave (Chung et al., 2018).
The majority of literature on DNN quantization involves training — either from scratch (Courbariaux et al., 2015; Wu et al., 2018; Jacob et al., 2018) or retraining/fine-tuning from a floating-point model. (Han et al., 2016; Zhou et al., 2017). Although such techniques are valuable, there are important real-world scenarios in which (re)training is not applicable. Consider an ML service provider (e.g. Amazon, Microsoft, Google) which wants to run a black-box floating-point client model in low-precision. The service provider does not have the training data, and the client may not be able to train for quantization because: (1) it lacks the expertise or manpower; (2) it is using an off-the-shelf or legacy model for which training data is not available. The importance of post-training quantization can be seen from NVIDIA’s TensorRT, a product specifically designed to perform 8-bit integer quantization without (re)training. This paper focuses on post-training DNN quantization.
DNN weights and activations follow a bell-shaped distribution after training. However, commodity hardware uses a linear number representation with evenly-spaced grid points. The naïve approach is to linearly map the entire range of the distribution to the range of the quantization grid (Figure 1(a)). Here the grid points extend to the maximum value in the distribution (Hubara et al., 2017). Clearly, this method over-provisions grid points for the rarely-occurring outliers. A better approach is to make the grid narrower than the distribution — this is known as clipping, as it is equivalent to thresholding the outliers before applying linear quantization (Figure 1(b)). Empirically, clipping can improve the accuracy of quantized DNNs, and many techniques exist to choose the optimal clip threshold (Sung et al., 2015; Zhuang et al., 2018; Migacz, 2017). Unfortunately, clipping can only reduce overall quantization error by increasing the distortion on the outliers — it is constrained by this tradeoff.
Another approach to handling outliers is to quantize them separately from the central values. Such outlier-aware quantization (Park et al., 2018a, b) is highly effective, but involves the use of dedicated non-commodity hardware.
In this paper, we propose outlier channel splitting (OCS). OCS identifies a small number of channels containing outliers, duplicates them, then halves the values in those channels. This creates a functionally identical network, but moves the affected outliers towards the center of the distribution (Figure 1 (c)). OCS takes inspiration from Net2Net (Chen et al., 2016); it does not require retraining and can be used on commodity CPUs and GPUs. OCS introduces a new tradeoff: it reduces quantization error at the expense of making the neural network larger. Experimental evaluation shows that for practical CNN and RNN models, OCS can significantly improve post-training quantization accuracy over state-of-the-art clipping methods with just a few percent overhead.
To present a comprehensive study of post-training quantization, we also evaluate different techniques for optimizing the clip threshold on both weights and activations. To our best knowledge, we are the first to perform a detailed literature comparison. Code for both OCS and clipping is available in open source https://github.com/cornell-zhang/dnn-quant-ocs. Our specific contributions are as follows:
We propose outlier channel splitting, a technique to improve DNN model quantization that does not require retraining and works with commodity hardware.
We present a comprehensive evaluation of post-training clipping techniques found in literature. To our best knowledge this is the first such study.
We demonstrate that OCS can outperform state-of-the-art clipping techniques on weight quantization, while incurring negligible overheads.
Related Work
Clipping is the state-of-the-art for DNN quantization without training. Clipping can be applied to both weights and activations — for the latter the activation distributions are sampled from a small number of inputs. (Sung et al., 2015) and (Shin et al., 2016) examined post-training quantization for CNNs and RNNs, respectively. They adopt a clip threshold that minimizes the L2-norm of the quantization error. ACIQ (Banner et al., 2018) fits a Gaussian and Laplacian to the sampled distribution, then uses the better-fitting curve to analytically compute the optimal clip threshold. In a similar vein, SAWB (Choi et al., 2018a) linearly extrapolates the clip threshold using statistics from fitting six different distributions. (McKinstry et al., 2018) clips using a percentile of the sampled values, the exact percentile depends on the quantization bitwidth. Different from the others, (Settle et al., 2018) tunes the bitwidth and floating-point format, achieving 32-bit accuracy performance with only 8-6 bits.
NVIDIA’s TensorRT (Migacz, 2017) is a commercial library that quantizes floating-point models to 8-bit for GPU inference. Clipping is used for the activations to control the effect of outliers. TensorRT profiles the activation distributions using a small number (1000s) of user-provided training samples, then computes a clipping threshold by minimizing the KL divergence between the original and quantized distributions.
OCS is different from these works as it leverages model expansion to improve quantization.
2 Outlier-Aware Quantization
Park et al. propose outlier-aware quantization (Park et al., 2018b, a), which uses a low-precision grid for the center values and a high-precision grid for the outliers. Placing 3% of values on the high-precision grid enabled post-training quantization of many popular CNN models to 4-bit without accuracy loss. This technique requires a specialized outlier-aware DNN accelerator; our approach is very different as it is designed to be applicable on commodity hardware.
3 Net2Net
OCS is inspired by Net2Net (Chen et al., 2016), which presents transformations to make a neural network wider or deeper while preserving functional equivalence. The goal of Net2Net was to speed up training by expanding a smaller DNN model into a larger one; the larger model inherits knowledge from the smaller and does not need to be trained from scratch. In this work we apply the Net2WiderNet transform to reduce outliers and improve quantization.
4 Cell Division
Cell division (Park & Choi, 2019) examines the same idea as OCS, and was published concurrently with our work. Though conceptually very similar, there are some technical differences between their work and ours: (1) they apply OCS on weights only while we examine both weights and activations; (2) they first tune the fixed-point bitwidths per layer with 50K training images while we use no data for weight OCS; (3) they do not compare against clipping as a baseline; (4) they evaluate on MNIST, CIFAR-10, and AlexNet while we evaluate on more modern (i.e. post-ResNet) ImageNet CNNs and language models.
Outlier Channel Splitting
The simplest form of linear quantization maps the inputs to a set of discrete, evenly-spaced grid points which span the entire dynamic range of the inputs. The maximum quantization error for any single value is one-half of the increment. For symmetric -bit quantization, we have grid points:
Because each value is scaled by , LinearQuant is sensitive to the largest inputs, i.e. the outliers. Many existing works first clip the range of prior to linear quantization; a survey of such clipping techniques can be found in Section 4.
2 Improving Quantization with Net2WiderNet
The core idea of OCS is to reduce the magnitude of outlier weights and/or activations in a DNN layer by duplicating a neuron, then either (1) halving its output; (2) halving the outgoing weight connections. This leaves the layer functionally equivalent but makes the weight/activation distribution narrower and thus more suitable for linear quantization. Such layer transformations were originally proposed as Net2WiderNet in Net2Net (Chen et al., 2016); we leverage them to improve quantization.
More formally, consider a linear layer in a DNN which takes as input the -channel activation vector , where each can be a single value (FC layer) or a 2D feature map (conv layer). Let be the -channel output. We can define a linear layer as follows:
where represents the weight(s) connecting and and represents multiplication or 2D convolution over a single channel. Without loss of generality, consider using OCS to split the last channel . This equates to rewriting Equation 2 as follows:
In both cases, we split channel into 2 channels. To preserve equivalence, we can halve the weights (Equation 3) or halve the input activations (Equation 4). Figure 2(a) taken from the Net2Net paper illustrates weight OCS visually: by duplicating we can cut its outgoing weight in half.
OCS is an alternative to clipping for reducing the dynamic range of DNN values without retraining. Compared to clipping, OCS preserves the outliers but incurs additional network overhead. The outlier values are the largest values in a layer and contribute the most to the outputs. We expect OCS to outperform clipping in neural network accuracy — the question is whether it can do so with low overhead.
Figure 2(b) shows some additional caveats of OCS in a layer with 2 inputs and 2 outputs. The top equation describes the original layer; the next two equations illustrate OCS to split the activations and the weights, respectively. One caveat is that to split any weight value, an entire row must be added to the weight matrix. For a conv layer, OCS requires duplicating an entire 2D activation channel and all 2D weight filters connected to that channel. A second caveat is that not all values need to be split. At the bottom of Figure 2(b), is split in half while is not split.
3 Quantization-Aware Splitting
It is clear that , i.e. naïve OCS does not preserve the quantized value. The maximum total quantization error is doubled as both halves may be rounded in the same direction (e.g., and each half is ).
To address this, we propose the following quantization-aware (QA) splitting function:
Intuitively, this forces to round in different directions when is close to the midpoint between grid points. More formally, we can prove the following:
The last line is simply , showing that QA OCS preserves the original quantization result. To derive the last line, we apply Hermite’s Identity (Savchev & Andreescu, 2003) with :
We can further show that QA splitting is optimal, i.e. there exists no way to split which results in lower quantization error. This proof is omitted due to length.
4 Channel Selection
As stated earlier, OCS cannot target individual weights or activations and must duplicate entire channels. OCS performs splits one at a time, and always splits the channel containing the largest absolute value in the layer. By prioritizing channels containing the largest values, OCS seeks to minimize distortion caused by any subsequent clipping.
We use a simple method to determine how many splits to perform in each layer (i.e. how many extra channels are created). For a layer containing channels, OCS splits channels, where is the expansion ratio, a hyperparameter that determines approximately the level of tolerable overhead in the network. This method allocates extra channels without considering each layer’s weight or activation distributions. We also tried a more intelligent approach which formulates extra channel allocation as a knapsack problem. The reward function is the percentage reduction in the dynamic range of the distribution, and the cost is the increase in memory size. We optimize the number of extra channels for all layers simultaneously subject to a constraint on the memory overhead. Unfortunately, the knapsack approach is experimentally not better than the simple method described above, and for space reasons we do not show results with knapsack.
Channel selection on DNN weights is straightforward to implement as the weights are known and fixed post-training. For the activations, we take an approach similar to TensorRT (Migacz, 2017): we use a small number of training images to sample the activations in each layer. The sampled distributions in each layer are then used for OCS.
5 Implementation on Commodity Hardware
A key strength of OCS is simplicity, allowing it to be used in practical scenarios with either commodity hardware or emerging deep learning accelerators. Figure 2(b) shows the network modifications needed to implement OCS — we need to duplicate and possibly scale certain channels in the weights and activations. The weight modifications can be done off-line prior to serving the model. For the activations, a custom layer can be inserted which simply copies and scales the appropriate channels.
Clipping
Clipping represents the state-of-the-art in post-training quantization. This section gives a brief overview of different methods for optimizing the clip threshold in literature; we present an evaluation of these methods in Section 5.
This method chooses a clip threshold which minimizes the mean squared error (MSE) or L2-norm between the floating-point and quantized values (Sung et al., 2015; Shin et al., 2016). It first constructs a histogram of the floating-point values. Let and be the bin values and frequencies, and denote the bins. The MSE is defined as:
where is the quantization function. In our experiments, we generate a large number of candidate clip thresholds evenly spaced between and the max absolute value, and choose the one with minimal MSE.
2 ACIQ
Proposed by (Banner et al., 2018), ACIQ first determines whether distribution is closer to a Gaussian or a Laplacian. Using statistics from the appropriate distribution, it uses an (approximate) closed-form solution for the clip threshold which minimizes MSE. Compared to the MSE method above, ACIQ avoids sweeping candidate thresholds and is much faster — this allows the clip threshold to be adjusted between input batches for activation quantization.
We used open-source code from the authors https://github.com/submission2019/AnalyticalScaleForInteger Quantization. Banner et al. assumed that an -bit fixed-point format contains grid points; this representation lacks a grid point at zero for signed values. We use grid points instead (i.e. sign-magnitude) as it is the default in our framework, and slightly adjusted the formulas from the paper to suit.
3 KL Divergence
This method chooses a clip threshold which (approximately) minimizes the KL divergence between the floating-point and quantized. Similar to the MMSE method, it works on the histogram of values and selects the optimal clip threshold from a set of candidates. The method was first proposed in a set of slides on NVIDIA’s TensorRT (Migacz, 2017), which unfortunately does not contain enough technical detail for replication. Instead, we adapted an open-source implementation from Apache MXNet (Chen et al., 2015).
In general, floating-point and quantized distributions do not have the same support and the KL divergence is thus undefined. To get around this, the MXNet implementation smooths the quantized histogram slightly by moving some of the probability mass into zero-frequency bins.
Experimental Evaluation on CNNs
This section reports experiments on CNN models for ImageNet classification (Deng et al., 2009) conducted using PyTorch (Paszke et al., 2017) and Intel’s open-source Distiller https://github.com/NervanaSystems/distiller quantization library. Post-training quantization was performed using Distiller’s symmetric linear quantizer, which scales the quantization grid based on the maximum absolute value following Equation 1. For activation quantization, we first sampled the activation distributions using 512 training images (i.e. images not part of the validation/test set) to determine the quantization grid points, then use this grid during testing. This profiling took between 40 and 200 seconds on our machine using an NVIDIA GTX 1080 Ti. Weight clipping and OCS does not require profiling and was performed without any input data.
The chosen CNN benchmarks are four popular ImageNet classification models: VGG16 (Simonyan & Zisserman, 2015) with batch normalization added, ResNet-50 (He et al., 2015), DenseNet-121 (Huang et al., 2017), and Inception-V3 (Szegedy et al., 2015). Pre-trained weights were obtained from the PyTorch model zoo and we ran inference only. The first layer was not quantized as it generally requires more bits than the others, and contains only 3 input channels meaning OCS would incur a large overhead.
The first experiment compares our proposed quantization-aware (QA) splitting (Section 3.3) against simply dividing by two as per Net2Net. Table 1 displays results from ResNet-20 for CIFAR-10 (Krizhevsky & Hinton, 2009). Although the difference is negligible until 4 bits (at which point there is significant accuracy degradation), QA splitting is clearly better than the naive method. This validates our mathematical ideas in Section 3.3, and we use QA splitting in all ensuing experiments.
2 Weight Quantization
Table 2 compares different clipping methods and OCS on weight quantization. The weights were quantized to 8-4 bits, while the activations were quantized to 8 bits. Floating-point accuracy is displayed under the model name. Linear quantization without clipping or OCS is shown in the Clip - None column. For ease of comparison we copy the best clipping result to the Clip - Best column. A range of small expand ratios was chosen for OCS.
Our results indicate that for large bitwidths, there is no advantage to doing weight clipping. This is in line with (Migacz, 2017), which reported the same at 8 bits. Clipping becomes beneficial at 6 bits or fewer, improving accuracy by up to for Resnet-50 and for Inception-V3. Interestingly, the best-performing clipping technique depends on bitwidth and follows a consistent pattern across network architectures. As we go from high to low bitwidth, the winning technique goes from no clipping, to MSE/ACIQ, to KL at 4 bits.
Weight OCS with an expansion ratio of only outperforms our benchmarked clipping methods at 8-5 bits. At 8 and 7 bits the difference between OCS and clipping is small and there isn’t a clear trend of improvement for higher expand ratios; in this regime OCS is not especially effective and the accuracy differences between expand ratios are mostly noise. At 6 and 5 bits, OCS with outperforms clipping by for all models except ResNet-50, and up to for Inception-V3. This demonstrates that the basic idea of OCS works — by splitting the outliers to preserve their values instead of clipping them, OCS can improve the accuracy of post-training quantization. Another trend is that OCS gets most of its gains from small expansion ratios. The gain from (no OCS) to is always larger than the gain from moving to higher values. Note that our benchmark networks have channel widths in the tens to hundreds so equates to a single channel split in many layers. This again makes intuitive sense: the first channel split will target the unique largest outlier, guaranteeing a narrower weight distribution. Further splits target smaller values which occur with higher frequency, making OCS less effective at reducing the distribution width.
Given this intuition, we expect that a combination of OCS (to remove the largest outliers) followed by clipping (to further shrink the quantization grid) might surpass either method alone. The rightmost columns of Table 2 show results for OCS + Best Clip (i.e. applying OCS followed by the best performing clip method at each bitwidth). At 5 and 4 bits, OCS plus clipping cleanly outperforms OCS alone. OCS and clipping both seek to shrink the dynamic range of the quantized values, and thus there is some level of overlap between them. At high precision, OCS with our chosen expand ratios eliminates enough outliers such that additional clipping is unnecessary. At low precision, we believe that OCS would require huge expand ratios to fully address the outlier problem — in this regime OCS can combine with clipping to produce the best quantization results.
3 Activation Quantization
The same benchmarks and setup were used for activation quantization, except weights were kept at 8 bits while the bitwidth was varied for activations. To select the channels to split, we sampled activation distributions and counted the number of extreme values (we used values greater than the 99’th percentile) in each channel. Channels with the highest counts were split.
Table 4 shows activation quantization results. Unlike the weights, clipping is effective at all bitwidths tested. This is again in agreement with (Migacz, 2017), which applied clipping to 8-bit activations. MSE clipping outperforms the other clip threshold techniques in nearly all cases. The gap between MSE and KL divergence is very small for large bitwidths, but at fewer bits MSE is clearly better. ACIQ performs worse than the other two methods with the exception of ResNet-50, where it showed good performance.
Activation OCS provides some improvement over simple linear quantization, but performs worse than clipping. This is likely because OCS relies on being able to identify the exact channel containing the largest outlier. With activations, profiling can only indicate which channels are likely to contain outliers, the best channel to split varies from input to input. To test our explanation, we experiment with Oracle OCS, which is simply OCS with exact knowledge of the activations generated by the network during testing — Oracle OCS chooses different channels to split in each input batch. Table 4 displays the results for Oracle OCS with different batch size on two models with 6 bit activations. Even at batch size 32, the oracle can already match or surpass the best clipping result. Further reducing the batch size (allowing channel selection at a finer granularity) leads to even better accuracy. These results show that OCS and our channel selection strategy can be effective for activations. However, channel selection must be done dynamically, requiring additional run-time analysis which is difficult to implement and likely inefficient in commodity systems.
4 OCS Memory Overhead
Because OCS increases the input channels by a factor of , rounded up, the expand ratio is a lower bound for the model size overhead. Table 5 shows both weight and activation overhead for ResNet-50 with different values of , show that the true overhead matches very closely.
Experimental Evaluation on RNNs
This section reports experiments on an RNN model with two stacked LSTM layers for language modeling (Zaremba et al., 2014). The corpus is the WikiText-2 dataset (Merity et al., 2016) with a vocabulary of 33,278 words. Each LSTM layer has a hidden size of 650, and the dimension of the word embedding in the input layer is 650. As the CNN results have shown that activation OCS is not effective, we focused on OCS and clipping on the weights. Activations and the hidden state are kept in floating-point for this experiment.
Table 6 compares the effects of OCS combined with different clipping methods on weight quantization. Lower perplexity is better, and the baseline floating-point model achieves a perplexity of . The best result on each row (i.e. the best clipping method at each OCS expand ratio) is bolded. Clipping is not effective on this model — none of the clipping techniques achieve any perplexity improvement. OCS achieves a much better result. At 6 bits, OCS begins to outperform the baseline with . At 5 bits, OCS sees steady perplexity decrease with successively larger expand ratios, clearly outperforming the best clipping result past . This is strong evidence that OCS can effectively improve post-training quantization beyond what can be achieved via clipping.
Conclusions and Future Work
We propose outlier channel splitting, a method to improve DNN quantization without retraining which can be applied on commodity hardware. OCS splits channels in a layer to reduce the magnitude of outliers. Unlike the existing clip-based methods, OCS introduces a new tradeoff by reducing quantization error at the cost of network size overhead. Experimental results demonstrate that OCS on weights outperforms state-of-the-art clipping techniques with minimal overhead on deep CNN and RNN benchmarks. At very low precision, OCS in conjunction with clipping outperforms either method alone. Because clipping is used in NVIDIA TensorRT — a commercial post-training quantization flow — we believe that OCS has potential applicability in real-life systems.
Future work includes a more in-depth study into different channel selection methods, as well as applying OCS quantization during training. Specifically, we believe that OCS can help shape weight distributions during training to obtain better results than training for quantization alone.
This work was supported in part by the Semiconductor Research Corporation (SRC) and DARPA. One of the Titan Xp GPUs used for this research was donated by NVIDIA.