Convolutional Neural Networks using Logarithmic Data Representation

Daisuke Miyashita, Edward H. Lee, Boris Murmann

Introduction

Deep convolutional neural networks (CNN) have demonstrated state-of-the-art performance in image classification Krizhevsky et al. (2012); Simonyan & Zisserman (2014); He et al. (2015) but have steadily grown in computational complexity. For example, the Deep Residual Learning He et al. (2015) set a new record in image classification accuracy at the expense of 11.311.3 billion floating-point multiply-and-add operations per forward-pass of an image and 230230 MB of memory to store the weights in its 152152-layer network.

In order for these large networks to run in real-time applications such as for mobile or embedded platforms, it is often necessary to use low-precision arithmetic and apply compression techniques. Recently, many researchers have successfully deployed networks that compute using 88-bit fixed-point representation Vanhoucke et al. (2011); Abadi et al. (2015) and have successfully trained networks with 1616-bit fixed point Gupta et al. (2015). This work in particular is built upon the idea that algorithm-level noise tolerance of the network can motivate simplifications in hardware complexity.

Interesting directions point towards matrix factorization Denton et al. (2014) and tensorification Novikov et al. (2015) by leveraging structure of the fully-connected (FC) layers. Another promising area is to prune the FC layer before mapping this to sparse matrix-matrix routines in GPUs Han et al. (2015b). However, many of these inventions aim at systems that meet some required and specific criteria such as networks that have many, large FC layers or accelerators that handle efficient sparse matrix-matrix arithmetic. And with network architectures currently pushing towards increasing the depth of convolutional layers by settling for fewer dense FC layers He et al. (2015); Szegedy et al. (2015), there are potential problems in motivating a one-size-fits-all solution to handle these computational and memory demands.

We propose a general method of representing and computing the dot products in a network that can allow networks with minimal constraint on the layer properties to run more efficiently in digital hardware. In this paper we explore the use of communicating activations, storing weights, and computing the atomic dot-products in the binary logarithmic (base-2 logarithmic) domain for both inference and training. The motivations for moving to this domain are the following:

Training networks with weight decay leads to final weights that are distributed non-uniformly around .

Similarly, activations are also highly concentrated near . Our work uses rectified Linear Units (ReLU) as the non-linearity.

Logarithmic representations can encode data with very large dynamic range in fewer bits than can fixed-point representation Gautschi et al. (2016).

Data representation in log⁡\log-domain is naturally encoded in digital hardware (as shown in Section 4.3).

we show that networks obtain higher classification accuracies with logarithmic quantization than linear quantization using traditional fixed-point at equivalent resolutions.

we show that activations are more robust to quantization than weights. This is because the number of activations tend to be larger than the number of weights which are reused during convolutions.

we apply our logarithmic data representation on state-of-the-art networks, allowing activations and weights to use only 33b with almost no loss in classification performance.

we generalize base-22 arithmetic to handle different base. In particular, we show that a base-2\sqrt{2} enables the ability to capture large dynamic ranges of weights and activations but also finer precisions across the encoded range of values as well.

we develop logarithmic backpropagation for efficient training.

Related work

Reduced-precision computation. Shin et al. (2016); Sung et al. (2015); Vanhoucke et al. (2011); Han et al. (2015a) analyzed the effects of quantizing the trained weights for inference. For example, Han et al. (2015b) shows that convolutional layers in AlexNet Krizhevsky et al. (2012) can be encoded to as little as 5 bits without a significant accuracy penalty. There has also been recent work in training using low precision arithmetic. Gupta et al. (2015) propose a stochastic rounding scheme to help train networks using 16-bit fixed-point. Lin et al. (2015) propose quantized back-propagation and ternary connect. This method reduces the number of floating-point multiplications by casting these operations into powers-of-two multiplies, which are easily realized with bitshifts in digital hardware. They apply this technique on MNIST and CIFAR10 with little loss in performance. However, their method does not completely eliminate all multiplications end-to-end. During test-time the network uses the learned full resolution weights for forward propagation. Training with reduced precision is motivated by the idea that high-precision gradient updates is unnecessary for the stochastic optimization of networks Bottou & Bousquet (2007); Bishop (1995); Audhkhasi et al. (2013). In fact, there are some studies that show that gradient noise helps convergence. For example, Neelakantan et al. (2015) empirically finds that gradient noise can also encourage faster exploration and annealing of optimization space, which can help network generalization performance.

Hardware implementations. There have been a few but significant advances in the development of specialized hardware of large networks. For example Farabet et al. (2010) developed Field-Programmable Gate Arrays (FPGA) to perform real-time forward propagation. These groups have also performed a comprehensive study of classification performance and energy efficiency as function of resolution. Zhang et al. (2015) have also explored the design of convolutions in the context of memory versus compute management under the RoofLine model. Other works focus on specialized, optimized kernels for general purpose GPUs Chetlur et al. (2014).

Concept and Motivation

The first proposed method as shown in Figure 1(b) is to transform one operand to its log⁡\log representation, convert the resulting transformation back to the linear domain, and multiply this by the other operand. This is simply

Quantizing the activations and weights in the log⁡\log-domain (log⁡2(x)\log_{2}(x) and log⁡2(w)\log_{2}(w)) instead of xx and ww is also motivated by leveraging structure of the non-uniform distributions of xx and ww. A detailed treatment is shown in the next section. In order to quantize, we propose two hardware-friendly flavors. The first option is to simply floor the input. This method computes ⌊log⁡2(w)⌋\lfloor\log_{2}(w)\rfloor by returning the position of the first 11 bit seen from the most significant bit (MSB). The second option is to round to the nearest integer, which is more precise than the first option. With the latter option, after computing the integer part, the fractional part is computed in order to assert the rounding direction. This method of rounding is summarized as follows. Pick mm bits followed by the leftmost 11 and consider it as a fixed point number FF with 0 integer bit and mm fractional bits. Then, if F≥2−1F\geq\sqrt{2}-1, round FF up to the nearest integer and otherwise round it down to the nearest integer.

2 Proposed Method 2.

The second proposed method as shown in Figure 1(c) is to extend the first method to compute dot products in the log⁡\log-domain for both operands. Additions in linear-domain map to sums of exponentials in the log⁡\log-domain and multiplications in linear become log⁡\log-addition. The resulting dot-product is

3 Accumulation in log\log domain

Experiments of Proposed Methods

Here we evaluate our methods as detailed in Sections 3.1 and 3.2 on the classification task of ILSVRC-2012 Deng et al. (2009) using Chainer Tokui et al. (2015). We evaluate method 1 (Section 3.1) on inference (forward pass) in Section 4.1. Similarly, we evaluate method 2 (Section 3.2) on inference in Sections 4.2 and 4.3. For those experiments, we use published models (AlexNet Krizhevsky et al. (2012), VGG16 Simonyan & Zisserman (2014)) from the caffe model zoo (Jia et al. (2014)) without any fine tuning (or extra retraining). Finally, we evaluate method 2 on training in Section 4.4.

This experiment evaluates the classification accuracy using logarithmic activations and floating point 32b for the weights. In similar spirit to that of Gupta et al. (2015), we describe the logarithmic quantization layer LogQuant that performs the element-wise operation as follows:

In order to evaluate our logarithmic representation, we detail an equivalent linear quantization layer described as

2 Logarithmic Representation of Weights of Fully Connected Layers

The FC weights are quantized using the same strategies as those in Section 4.1, except that they have sign bit. We evaluate the classification performance using log⁡\log data representation for both FC weights and activations jointly using method 2 in Section 3.2. For comparison, we use linear for FC weights and log⁡\log for activations as reference. For both methods, we use optimal 44b log⁡\log for activations that were computed in Section 4.1.

Table 4 compares the mentioned approaches along with floating point. We observe a small 0.4%0.4\% win for log⁡\log over linear for AlexNet but a 0.2%0.2\% decrease for VGG16. Nonetheless, log⁡\log computation is performed without the use of multipliers.

An added benefit to quantization is a reduction of the model size. By quantizing down to 44b log⁡\log including sign bit, we compress the FC weights for free significantly from 1.91.9 Gb to 0.270.27 Gb for AlexNet and 4.44.4 Gb to 0.970.97 Gb for VGG16. This is because the dense FC layers occupy 98.2%98.2\% and 89.4%89.4\% of the total model size for AlexNet and VGG16 respectively.

3 Logarithmic Representation of Weights of Convolutional Layers

We now represent the convolutional layers using the same procedure. We keep the representation of activations at 44b log⁡\log and the representation of weights of FC layers at 44b log⁡\log, and compare our log⁡\log method with the linear reference and ideal floating point. We also perform the dot products using two different bases: 2,22,\sqrt{2}. Note that there is no additional overhead for log⁡\log base-2\sqrt{2} as it is computed with the same equation shown in Equation 4.

Table 5 shows the classification results. The results illustrate an approximate 6%6\% drop in performance from floating point down to 5b base-2 but a relatively minor 1.7%1.7\% drop for 5b base-2\sqrt{2}. They includes sign bit. There are also some important observations here.

We first observe that the weights of the convolutional layers for AlexNet and VGG16 are more sensitive to quantization than are FC weights. Each FC weight is used only once per image (batch size of 1) whereas convolutional weights are reused many times across the layer’s input activation map. Because of this, the quantization error of each weight now influences the dot products across the entire activation volume. Second, we observe that by moving from 55b base-22 to a finer granularity such as 55b base-2\sqrt{2}, we allow the network to 1) be robust to quantization errors and degradation in classification performance and 2) retain the practical features of log⁡\log-domain arithmetic.

4 Training with Logarithmic Representation

We incorporate log⁡\log representation during the training phase. This entire algorithm can be computed using Method 2 in Section 3.2. Table 6 illustrates the networks that we compare. The proposed log⁡\log and linear networks are trained at the same resolution using 4-bit unsigned activations and 55-bit signed weights and gradients using Algorithm 1 on the CIFAR10 dataset with simple data augmentation described in He et al. (2015). Note that unlike BinaryNet Courbariaux & Bengio (2016), we quantize the backpropagated gradients to train log⁡\log-net. This enables end-to-end training using logarithmic representation at the 55-bit level. For linear quantization however, we found it necessary to keep the gradients in its unquantized floating-point precision form in order to achieve good convergence. Furthermore, we include the training curve for BinaryNet, which uses unquantized gradients.

Fig. 7 illustrates the training results of log⁡\log, linear, and BinaryNet. Final test accuracies for log⁡\log-55b, linear-55b, and BinaryNet are 0.93790.9379, 0.92530.9253, 0.88620.8862 respectively where linear-55b and BinaryNet use unquantized gradients. The test results indicate that even with quantized gradients, our proposed network with log⁡\log representation still outperforms the others that use unquantized gradients.

Conclusion

In this paper, we describe a method to represent the weights and activations with low resolution in the log⁡\log-domain, which eliminates bulky digital multipliers. This method is also motivated by the non-uniform distributions of weights and activations, making log⁡\log representation more robust to quantization as compared to linear. We evaluate our methods on the classification task of ILSVRC-2012 using pretrained models (AlexNet and VGG16). We also offer extensions that incorporate end-to-end training using log⁡\log representation including gradients.

References