Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers

Yukuan Yang, Shuang Wu, Lei Deng, Tianyi Yan, Yuan Xie, Guoqi Li

I INTRODUCTION

Deep neural networks have achieved state-of-art results in many fields like image processing , object detection , natural language processing , and robotics through learning high-level features from a large amount of input data. However, due to the existence of a huge number of floating-point (FP) values and complex FP multiply-accumulate operations (MACs) in the process of network training and inference, the intensive memory overhead, large computational complexity, and high energy consumption impede the wide deployment of deep learning models. DNN quantization which converts FP MACs to bit-wise operations is an effective way to reduce the memory and computation costs and improve the speed of deep learning accelerators.

With the deepening of research, DNN quantization gradually transfers from inference quantization (BWN , XNOR-Net , ADMM ) to training quantization (DoReFa , GXNOR-Net , FP8 ). Usually, the inference quantization focuses on the forward pass; while the training quantization further quantizes the backward pass and weight updates. Recently, training quantization becomes a hot topic in the network compression community. Whereas, there are still two major issues in existing schemes. The first issue lies in the incomplete quantization, including two aspects: partial quantization and FP dependency. Partial quantization means that only parts of dataflows, not all of them, are quantized (e.g. DoReFa , GXNOR-Net and QBP2 ); FP dependency still remains FP values during the training process (e.g. MP and FP8 ). The second issue is that the quantization of batch normalization (BN) is ignored by most schemes (e.g. MP-INT and FX Training ). BN is an essential layer for the training of DNNs by addressing the problem of the internal covariate shift of each layer’s inputs, especially as the network deepens, allowing a much higher learning rate and less careful weight initialization.

Compared with all the studies above, WAGE is the most thorough work of DNNs quantization, which quantizes the data including W (Weights), A (Activation), G (Gradient), E (Error), U (Update) and replacing each BN layer with a constant scaling factor. WAGE has achieved competitive results on LeNet , VGG , and AlexNet , providing a good inspiration for this work. However, we find that WAGE is difficult to be applied in large-scale DNNs due to the absence of BN layers. Besides, it is known that the gradient descent optimizer such as Momentum or Adam increases the stability and even helps get rid of the local optimum, thus the speed and final performance are significantly improved. A complete quantization should cover the entire training process, including W, A, G, E, U, BN, and the optimizer. Regretfully, up to now, there is still no such solution that can achieve this complete quantization, especially on large-scale DNNs.

To address the issues of incomplete quantization and ignored BN quantization mentioned above and extend quantization framework to large-scale datasets and networks with high performance, we propose a unified complete quantization framework termed “WAGEUBN” to constrain W, A, G, E, U, BN, and the optimizer in the low-bit integer (INT) space. To the best of our knowledge, WAGEUBN is the first complete quantization framework achieving high performance in large-scale datasets, where all computation steps and operands in DNNs are decomposed and quantized.

We mainly make the following efforts to create the complete quantization framework. Firstly, according to the various data distributions and the role the quantized data plays in DNN training, we fuse three quantization functions to satisfy the different precision requirements. Furthermore, we propose a new storage and computing method by introducing a flag bit to expand the data coverage and solve the non-convergence problem caused by insufficient data representation. Last but not least, we quantize BN and Momentum optimizer for the first time, converting all FP operations in DNNs to bit-wise operations. Compared with the full precision DNNs, DNNs under the full 8-bit WAGEUBN framework can achieve about 4×4\times memory saving. More importantly, the multiplication and accumulation operations of WAGEUBN, which are the main operations in DNNs, can perform >>3×\times and 9×\times faster in speed, 10×\times and >>30×\times lower in power, 9×\times and >>30×\times smaller in circuit area, respectively. Besides, the efficient INT8 multiplication and accumulation operations also make WAGEUBN a big step ahead of most existing quantization schemes in computational costs, whether it is FP8, INT16, FP16 or INT32. In addition to the huge advantages in memory cost, computing speed, energy consumption and circuit area, WAGEUBN also shows competitive accuracy on large-scale networks (ResNet18/34/50) and dataset (ImageNet ). What’s more, the hardware design of WAGEUBN is much simpler and more efficient because of the complete quantization. Due to the improvement of computing speed and saving in hardware resources and energy consumption, WAGEUBN provides a feasible idea for the architecture design of future efficient online learning chips used in portable devices with limited computational resources. The contributions of this work are twofold, which are summarized as follows:

We address two main issues existing in most quantization schemes via fully quantizing all the data paths, including W, A, G, E, U, BN, and the optimizer, greatly reducing the memory and compute costs. What’s more, we constrain the data to INT8 for the first time, pushing the training quantization to a new bit level compared with the existing FP16, INT16, and FP8 solutions.

Our quantization framework is validated in large-scale DNN models (ResNet18/34/50) over ImageNet dataset and achieves competitive accuracy with much fewer overheads, indicating great potential for future portable devices with online learning ability.

The organization of this paper is as follows: Section II introduces the related work of DNN quantization; Section III details the WAGEUBN framework; Section IV presents the experiment results of WAGEUBN and the corresponding analyses; Section V summarizes this work and delivers the conclusion.

II Related Work

With the wide applications of DNNs, the related compression technologies have been proposed rapidly, among which the quantization plays an important role. The development of DNN quantization can be divided into two stages, inference quantization and training quantization, according to the different quantization objects.

Inference quantization: Inference quantization starts from constraining W into {−1, 1}\{-1,~{}1\} (BWN ), replacing complex FP MACs with simple accumulations. BNN and XNOR-Net further quantize both W and A, making the inference computation dominated by bit-wise operations. However, extremely low bit-width quantization usually leads to significant accuracy loss. For example, when the bit width comes to <<4 bits, the accuracy degradation becomes obvious, especially for large-scale DNNs. Instead, the bit width of W and A for inference quantization can be reduced to 8 bits with little accuracy degradation. The study of inference quantization is sufficient for the deep learning inference accelerators. Whereas, this is not enough for efficient online learning accelerators because only the data in the forward pass are considered.

Training quantization: To further extend the quantization towards the training stage, DoReFa trains DNNs with low bit-width W, A, and G, while leaving E and BN unprocessed. MP and MP-INT use FP16 and INT16 values, respectively, to constrain W, A, and G. Recently, FP8 further pushes W, A, G, E, and U to 8, 8, 8, 8, and 16-bit FP values, respectively, still leaving BN untouched. QBP2 replaces the conventional BN with range BN and constrains W, A, and E to INT8 values while calculating G with FP MACS. Recently, WAGE adopts a layer-wise scaling factor instead of using the BN layer and quantizes W, A, G, E, and U to 2, 8, 8, 8, and 8 bits, respectively. Despite its thorough quantization, WAGE is difficult to be applied to large-scale DNNs due to the absence of powerful BN layers. In summary, there still lacks a complete INT8 quantization framework for the training of large-scale DNNs with high accuracy.

III WAGEUBN Framework

The main idea of WAGEUBN is to quantize all the data in DNN training to INT8 values. In this section, we detail the WAGEUBN framework implemented in large-scale DNN models. The organization of this section is as follows: Subsection III-A introduces the straight-through estimator (STE) method which is accepted and used by most researchers to solve the non-differentiable problem of quantization; Subsection III-B and Subsection III-C describe the notations and quantization functions, respectively; Subsection III-D explains the specific quantization schemes for W, A, G, E, U, BN, and the Momentum optimizer, respectively; Subsection III-E goes through the overall implementation of WAGEUBN, including in both forward and backward passes; Subsection III-F summarizes the whole process and shows the pseudo codes.

In the early stage of the study, the non-differentiable problem in mathematical sense caused by quantization has been hindering the development of quantization research. However, since the straight-through estimator (STE) method was used to estimate the gradient of quantized data in BNN , almost all the works in the field of quantization have adopted this method to avoid the mathematical non-derivative problem .

The STE method used in WAGEUBN can be illustrated as the following

where x\bm{x} is the data to be quantized, Q(⋅)Q(\cdot) is the quantization function that may be non-differentiable, and LL denotes the objective function.

III-B Notations

Before introducing the WAGEUBN quantization framework formally, we need to define some notations. Considering the ll-th layer of DNNs, we divide the forward pass of DNNs into four steps as described in Figure 1 (BN is divided into two steps: Normalization & QBNQ_{BN} and Scale & Offset).

Different from most existing schemes, we define e\bm{e} and g\bm{g} respectively, where e\bm{e} represents the gradient of A (activation) which is used in the error backpropagation and g\bm{g} represents the gradient of W (weights) which is used in the weight update. Moreover, we quantize the BN layers, including both the forward and backward passes, which is not well touched in most prior work. Similar to the forward pass, we divide the backward pass of the ll-th layer into five steps as shown in Figure 2. According to the derivative chain rules, we have

where LL is the loss function, e4l+1\bm{e}_{4}^{l+1} represents the error from the (l+1)(l+1)-th layer, and ⊙\odot represents the Hadamard product. For vectors with the same dimension, such as a=(a1,a2,⋯ ,an)\bm{a}=(a_{1},a_{2},\cdots,a_{n}) and b=(b1,b2,⋯ ,bn)\bm{b}=(b_{1},b_{2},\cdots,b_{n}), we have: a⊙b=(a1b1,a2b2,⋯ ,anbn)\bm{a}\odot\bm{b}=(a_{1}b_{1},a_{2}b_{2},\cdots,a_{n}b_{n}). Two quantization functions are used here: QE1Q_{E_{1}} is the quantization function detailed as Equation (15) that converts high bit-width integers to low bit-width integers; QE2Q_{E_{2}} detailed as Equation (17) is trying to convert FP values to low bit-width integers. WqlT{W_{q}^{l}}^{T} is the transposed matrix of WqlW_{q}^{l}, and ∂x4l/∂x3l\partial\bm{x}_{4}^{l}/\partial\bm{x}_{3}^{l} represents the gradient of activation. When relurelu is used as the activation function, ∂x4l/∂x3l\partial\bm{x}_{4}^{l}/\partial\bm{x}_{3}^{l} is a tensor containing only 0 and 1 elements.

According to the definitions given above, the gradients of W, γ\bm{\gamma}, and β\bm{\beta} can be summarized as follows

To further reduce the bit width of G that will increase greatly after the multiplication, we have

where QGWQ_{G_{W}}, QGγQ_{G_{\bm{\gamma}}}, and QGβQ_{G_{\bm{\beta}}} are quantization functions for the gradient of W, γ{\bm{\gamma}}, and β{\bm{\beta}}, respectively, which will be shown in Equation (18).

Some notations to be used below are also explained here. kWk_{W}, kAk_{A}, kGWk_{G_{W}}, kEk_{E} (kE1k_{E_{1}} and kE2k_{E_{2}}), and kBNk_{BN} are the bit width of W, A, G, E, and BN, respectively. kWUk_{WU}, kγUk_{\bm{\gamma}U}, and kβUk_{\bm{\beta}U} are the bit width of W, γ\bm{\gamma}, and β\bm{\beta} update, which are also the bit width of data stored in memory. kγk_{\bm{\gamma}}, kβk_{\bm{\beta}}, kμk_{\mu}, and kσk_{\sigma} are the bit width of γ\bm{\gamma}, β\bm{\beta}, μ\mu, and σ\sigma, respectively, used in the BN layer. kGγk_{G\bm{\gamma}} and kGβk_{G\bm{\beta}} are the bit width of γ\bm{\gamma} and β\bm{\beta} gradient, respectively. kMomk_{Mom} and kAcck_{Acc} are the bit width of momentum coefficient (MomMom) and accumulation (AccAcc), respectively, used in the Momentum optimizer. At last, klrk_{lr} is the bit width of the learning rate.

III-C Quantization Functions

There are three quantization functions used in WAGEUBN. The direct-quantization function uses the nearest fixed-point values to represent the continuous values of W, A, and BN. The constant-quantization function for G is used to keep the bit width of U (update) fixed since G is directly related to U. Because U and the weights stored in memory have the same bit width, the bit width of weights stored in memory can be fixed, which is more hardware-friendly. The magnitude of E is very small, so the shift-quantization function reduces the bit width of E greatly compared with the direct-quantization function under the same precision.

The direct-quantization function simply approximates a continuous value to its nearest discrete state and is defined as

where kk is the bit width, and round(⋅)round(\cdot) rounds a number to its nearest INT value.

The intention of constant-quantization function is to normalize a tensor firstly, then limit it to INT, and finally maintain its magnitude. It is governed by

The illustration of constant-quantization function is described in Figure 3. Here, R(⋅)R(\cdot) is used to project the maximum value of x\bm{x} to its nearest fixed-point value, which is prepared for normalization; Sr(⋅)Sr(\cdot) is a stochastic rounding function used for converting a continuous float value to its nearby INT value in a probabilistic manner and PxP_{x} is the rounding probability; Norm(⋅)Norm(\cdot) denotes normalization and clip(⋅)clip(\cdot) is a saturation function limiting the data range;Sd(⋅)Sd(\cdot) is to shift the distribution of x\bm{x} and limit x\bm{x} to INT values between −dr+1-dr+1 and dr−1dr-1. Here dr∈[2k−1,2k−2,...,1]dr\in[2^{k-1},2^{k-2},...,1] limits the data range after mapping and decreases as the training goes on, presenting the same effect as reducing the learning rate. For example, Sd(⋅)Sd(\cdot) maps G to {-127, -126, ⋯\cdots, 126, 127} and {-63, -62, ⋯\cdots, 62, 63} in the early training stage (k=8k=8, dr=128dr=128, epoch in $)andlatertrainingstage() and later training stage (k=7,,dr=64,epochin, epoch in),respectively.), respectively.CQ(\cdot)isutilizedtomaintainthemagnitudeorderofdata,whereis utilized to maintain the magnitude order of data, where2^{k_{GC}-1}isaconstantscalingfactorandis a constant scaling factor andk_{GC}$ is its bit width.

The shift-quantization function serves for the quantization of E and is defined as

where d(⋅)d(\cdot) is the minimum interval for a k-bit INT and Q(⋅)Q(\cdot) is the direct-quantization function defined in Equation (6).

The shift-quantization function normalizes E first, then converts E to fixed-point values, and finally uses a layer-wise scaling factor (R(⋅)R(\cdot) defined in Equation (7)) to maintain the magnitude. The differences between the constant-quantization function and the shift-quantization function mainly exist in two points: First, the constant-quantization uses a constant to keep the magnitude for hardware friendliness while the shift-quantization uses a lay-wise scaling factor; Second, the constant-quantization contains a stochastic rounding process while the shift-quantization function does not.

III-D Quantization Schemes in WAGEUBN

After introducing the quantization functions used in our WAGEUBN framework, we provide detailed quantization schemes.

Since weights are stored and used as fixed-point values, weights should be also initialized discretely. An initialization method proposed by MSRA has been evidenced helpful for faster training. The initialization of weights can be formulated as follows

where ninn_{in} is the layer’s fan-in number, and kWUk_{WU} is the bit width of weight update and the memory storage.

Because of the different bit width for weight storage and computation, it should be quantized from kWUk_{WU} bits to kWk_{W} (the bit width of weights used for convolution) bits for convolution. In addition, we also limit the data range of W. Finally, the quantization function for W is

As aforementioned, BN plays an important role in training large-scale DNNs. WAGE has proved that simple scaling layers are not enough to replace BN layers. Conventional BN layer can be divided into two steps as

where μl\mu^{l} and σl\sigma^{l} are the mean and standard, respectively, deviation of x\bm{x} over one mini-batch in the ll-th layer; ϵ\epsilon is a small positive value added to σ\sigma to avoid the case of dividing by zero; γl\bm{\gamma^{l}} and βl\bm{\beta^{l}} are the scale and offset parameters, respectively.

Under the WAGEUBN framework, the BN layer is also quantized. Through the operations described in Equation (12), all operands are quantized and all operations are bit-wise. Specifically, the quantization follows

where Qμ,Qσ,Qγ,Qβ,QBNQ_{\mu},Q_{\sigma},Q_{\bm{\gamma}},Q_{\bm{\beta}},Q_{BN} are the quantization functions converting the operands to fixed-point values defined as

And ϵq\epsilon_{q} is a small fixed-point value, playing the same role as ϵ\epsilon in Equation (11); kμ,kσ,kγ,kβ,kBNk_{\mu},k_{\sigma},k_{\bm{\gamma}},k_{\bm{\beta}},k_{BN} are the bit width of μ,σ,γ,β\mu,\sigma,\bm{\gamma},\bm{\beta} and x^\hat{\bm{x}}, respectively.

After the convolution and BN layers in the forward pass, the bit width of operands increases due to the multiplication operation. To reduce the bit width and keep the input bit width of each layer consistent, we need to quantize the activations. Here, the quantization function for activations can be described as

where kAk_{A} is the bit width of activations.

In Equation (3), we have given the definition of E and quantized E. Through investigating the importance of error propagation in DNN training, we find that the quantization of E is very essential for the model convergence. If E is naively quantized using the direct-quantization function, it will require a large bit width of operands to realize the convergence of DNNs. Instead, we use the following shift-quantization function

where SQ(⋅)SQ(\cdot) is the shift-quantization function defined in Equation (8), and kE1k_{E_{1}} is the bit width of e0l\bm{e}_{0}^{l} defined in Equation (3).

As mentioned above, we use QE1Q_{E_{1}} and QE2Q_{E_{2}} for the error quantization. However, the precision requirements of QE1Q_{E_{1}} and QE2Q_{E_{2}} vary a lot. Experiments show that kE1=8k_{E_{1}}=8 affects little on accuracy while kE2≤8k_{E_{2}}\leq 8 will cause the non-convergence of large-scale DNNs when using SQ(⋅)SQ(\cdot) as the quantization function. kE2=16k_{E_{2}}=16 is a proper value for the training of DNNs with minimum accuracy degradation. More analyses will be given in Subsection IV-E. Here we will provide two versions of QE2Q_{E_{2}}, the 16-bit and 8-bit versions. The 16-bit QE2Q_{E_{2}} is defined as

where kE2k_{E_{2}} is the bit width of e3l\bm{e}_{3}^{l} defined in Equation (3).

Experiments have proved the data range covered by 8-bit QE2Q_{E_{2}} (kE2=8k_{E_{2}}=8) is not sufficient to train DNNs. In order to expand the coverage of quantization function while still maintaining a low bit width, we introduce a layer-wise scaling factor ScSc and a flag bit. Then, to distinguish it from QE2Q_{E_{2}} defined in Equation (16), we name the quantization function Flag QE2Q_{E_{2}} and the quantization process is governed by

where kE2=8k_{E_{2}}=8, min=−2kE2+1min=-2^{k_{E_{2}}}+1, and max=2kE2−1max=2^{k_{E_{2}}}-1.

By introducing a layer-wise scaling factor and a flag bit, we can expand the data coverage greatly. Details can be found in Figure 4. The flag bit is used to indicate whether the absolute value of xx stored in memory is less than the layer-wise scaling factor (e.g., 0 represents ∣x∣\textlessSc|x|\textless Sc and 1 represents ∣x∣≥Sc|x|\geq Sc). The sign bit is used to denote the positive or negative direction of the value. The data bit follows the conventional binary format. According to the definition, the values stored in Figure 4 and 4 are +Sc/128+Sc/128 and −127×Sc-127\times Sc when kE2=8k_{E_{2}}=8, respectively. Therefore, the 9-bit data format can cover almost the same data range as the direct 15-bit quantization described in Equation (16). Since the flag bit is just used for judgment, the effective value for computation is still INT8.

The gradient is another important part in DNN training because it is directly related to the weight update. The rules for calculating and quantizing the gradients of W, γ\bm{\gamma}, and β\bm{\beta} are described as Equation (4) and (5). Since e1l\bm{e_{1}}^{l}, e3l\bm{e_{3}}^{l}, x0l\bm{x_{0}}^{l}, and x2l\bm{x_{2}}^{l} are all fixed-point values, the conventional FP MACs operations can be replaced with bit-wise operations during the process of calculating gWl\bm{g}_{W}^{l}, gγl\bm{g}_{\bm{\gamma}}^{l}, and gβl\bm{g}_{\bm{\beta}}^{l}. The quantization functions are defined to further reduce the bit width of gradients and prepare for the next step of the optimizer. Specifically, we have

where CQ(⋅)CQ(\cdot) is the constant-quantization function defined in Equation (7); kGWk_{G_{W}}, kGγk_{G_{\bm{\gamma}}}, and kGβk_{G_{\bm{\beta}}} are the bit width of the gradient of W, γ\bm{\gamma}, and β\bm{\beta}, respectively.

Momentum optimizer is one of the most common optimizers used in DNN training, especially for classification tasks. For the ii-th training step of the ll-th layer, the conventional Momentum optimizer works as follows

where AccilAcc_{i}^{l} and Acci−1lAcc_{i-1}^{l} are the accumulation in the ii-th and (i−1)(i-1)-th training step, respectively; MomMom is a constant value used as a coefficient; gil\bm{g}_{i}^{l} is the gradient of W, γ\bm{\gamma}, or β\bm{\beta}.

Momentum optimizer under the WAGEUBN framework is trying to constrain all operands to fixed-point values. The process can be formulated as

where Acc(i−1)qlAcc_{(i-1)q}^{l} is the quantized accumulation in the (i−1)(i-1)-th training step; giqlg^{l}_{iq} is the quantized gradient of W, γ\bm{\gamma}, or β\bm{\beta}; QAcc(⋅)Q_{Acc}(\cdot) is the quantization function defined as

To guarantee the consistency of bit width, we further set

The parameter update is the last step in the training of each mini-batch. Different from conventional DNNs where the learning rate can take any FP value, the learning rate under WAGEUBN must also be a fixed-point value and the bit width of update is directly related to the bit width of learning rate. The update under quantized Momentum optimizer can be described as

where ΔW\Delta W is the update of W with kWUk_{WU} bits, and lrlr is the fixed-point learning rate with klrk_{lr} bits. The updates of γ\bm{\gamma} and β\bm{\beta} are the same as in Equation (23). According to Equation (20), (22), and (23), we have

Through our evaluations, the precision of the update has the greatest impact on the accuracy of DNNs because it is the last step to constrain the parameters. Thus, we need to set a reasonable bit width for update to balance the model accuracy and memory cost.

III-E Quantization Framework

Given the quantization details of W, A, G, E, U, BN, and the Momentum optimizer, the overall quantization framework is depicted in Figure 5. Under this framework, conventional FP MACs can be replaced with bit-wise operations. Here, the forward pass of the ll-th layer in DNNs is divided into three parts: Conv (convolution), BN, and activation. x0l\bm{x}_{0}^{l}, x1l\bm{x}_{1}^{l}, x2l\bm{x}_{2}^{l}, x3l\bm{x}_{3}^{l}, and x4l\bm{x}_{4}^{l} are defined in Equation (2). The weights (WlW^{l}) are stored as kWUk_{WU}-bit integers and then QWQ_{W} maps WlW^{l} to kWk_{W}-bit INT values (WqlW^{l}_{q}) before convolution. After convolution, MEAN&QμMEAN\&Q_{\mu} and STD&QσSTD\&Q_{\sigma} operations are used to calculate the mean and standard deviation of x1l\bm{x}_{1}^{l} in one mini-batch and then quantize them to kμk_{\mu} and kσk_{\sigma} bits, respectively. The BW&QBNBW\&Q_{BN} operation constrains x2l\bm{x}_{2}^{l} to kBNk_{BN} bits in BN. Similar to W, γ\bm{\gamma} and β\bm{\beta} are stored as kγUk_{\bm{\gamma}U} and kβUk_{\bm{\beta}U}-bit integers and used in kγk_{\bm{\gamma}} and kβk_{\bm{\beta}} bits (γql\bm{\gamma}^{l}_{q} and βql\bm{\beta}^{l}_{q}) after the QγQ_{\bm{\gamma}} and QβQ_{\bm{\beta}} quantization, respectively. After the second step of BN, activation and quantization are implemented with the ACT&QAACT\&Q_{A} operation, reducing the increased bit width to kAk_{A} bits again and preparing inputs for the next layer.

The backward pass of the ll-th layer is much more complicated than the forward pass, including error propagation, gradient of weight, gradient of BN, Momentum optimizer, and weight update. In the process of error propagation, e0l\bm{e}_{0}^{l}, e1l\bm{e}_{1}^{l}, e2l\bm{e}_{2}^{l}, e3l\bm{e}_{3}^{l}, and e4l\bm{e}_{4}^{l} are defined in Equation (3) and there are two locations needing quantization using QE1Q_{E_{1}} and QE2Q_{E_{2}}. QE1Q_{E_{1}} reduces the bit width of e4l+1\bm{e}_{4}^{l+1} from kE2+kW−1k_{E_{2}}+k_{W}-1 to kE1k_{E_{1}}. ACT′ACT^{\prime} is the derivative of activation function (relurelu) and QE2Q_{E_{2}} is used to constrain e3l\bm{e}_{3}^{l} to kE2k_{E_{2}} bits. In the phase of calculating the gradients of weights and BN, QGWQ_{G_{W}}, QGγQ_{G_{\bm{\gamma}}}, and QGβQ_{G_{\bm{\beta}}} are leveraged to reduce the increased bit width caused by the multiplication operations.

All parameters of the Momentum optimizer in the ii-th training step are quantized. Different from the conventional learning rate with FP value, WAGEUBN requires a discrete learning rate so that the bit width of weight updates can be controlled. The updates of γ\bm{\gamma} and β\bm{\beta} in BN layers are also similar to the weight update, which are omitted in Figure 5 for simplicity.

III-F Overall Algorithm

Given the framework of WAGEUBN, we summarize the entire quantization process and present the pseudo codes for both the forward and backward passes as shown in Algorithm 1 and 2, respectively.

IV Results

To verify the effectiveness of the proposed quantization framework, we apply WAGEUBN on ResNet18/34/50 on ImageNet dataset. We provide two versions of WAGEUBN, one with full 8-bit INT where kWk_{W}, kAk_{A}, kGWk_{G_{W}}, kE1k_{E_{1}}, kE2k_{E_{2}}, kγk_{\bm{\gamma}}, and kβk_{\bm{\beta}} are equal to 8. The other version has 16-bit kE2k_{E_{2}}. The only difference between the 16-bit E2E_{2} version and the full 8-bit version exists in the quantization function QE2Q_{E_{2}} (see Equation (16) and (17), respectively). kGγk_{G_{\bm{\gamma}}}, kGβk_{G_{\bm{\beta}}}, and kGCk_{GC} are 15 and we set kmomk_{mom}, kAcck_{Acc}, klrk_{lr}, and kWUk_{WU} to 3, 13, 10, and 24 respectively to satisfy Equation (22) and (24). In addition, we set kBNk_{BN}, kμk_{\mu}, and kσk_{\sigma} to 16. Since W, A, G, and E occupy the majority of memory and compute costs, their bit-width values are reduced as much as possible. Other parameters occupying much less resources can increase the bit width to maintain the accuracy, e.g., μql\mu_{q}^{l}, σql\sigma_{q}^{l}, gγqlg_{\bm{\gamma q}}^{l}, and gβqlg_{\bm{\beta q}}^{l} in BN layers.

The first and last layers are believed to differ from the rest because of their interface with network inputs and outputs. The quantization of these two layers will cause significant accuracy degradation compared to hidden layers and they just consume few overheads due to the small number of neurons. Therefore, we do not quantize the first and last layers, as previous work did .

For the quantization frameworks, it’s a common issue of saturated learning and vanishing gradients because the data in DNNs may be clipped very small after quantization and the model may be stuck on a local optimum because of the few updates. And this situation can be more serious when the quantization framework is used in large-scale complex datasets. WAGEUBN has made many efforts to alleviate this problem. Firstly, the newly designed shift-quantization function ensures that the quantized error in WAGEUBN framework is on the same magnitude as that of a traditional FP network, making the errors (gradients of activations) not vanishing in the process of back propagation. Secondly, based on the observation that it is the orientation rather than the magnitude of gradients (gradients of weights) that guides deep neural networks (DNNs) to converge, it’s more important to keep the orientation rather than the magnitude. What’s more, the updates (ΔW\Delta W) jointly determined by the learning rate and gradients are the data that ultimately determines the final changes of the network. The much bigger learning rate setting in WAGEUBN (minimum learning rate: WAGEUBN 1.95×10−31.95\times 10^{-3} (2−92^{-9}) VS floating-point network 5×10−65\times 10^{-6} ) can make up for the problem of update magnitude caused by the smaller quantized gradients. Through the technics above, the problems of saturated learning and dead gradients can be solved. In addition to the above measures which can help the model get rid of local optimum, the proposed quantized Momentum optimizer, fixed-point updates which is precise enough and a proper batch size selection can also reduce the probability of falling into a local optimum greatly.

IV-B Training Curve

Figure 6 illustrates the accuracy comparison between vanilla DNNs (FP32), DNNs with the full 8-bit version of WAGEUBN, and DNNs with the 16-bit E2E_{2} version of WAGEUBN. The initial learning rate and momentum coefficient under WAGEUBN and full precision are slightly different. To ensure the best recognition accuracy, we use the official parameter settings of TensorFlow , where the initial learning rate and momentum coefficient are set to 0.05 (batch size is 128) and 0.9 while those of WAGEUBN framework are set to 0.05078125 (26×2−926\times 2^{-9}, 10-bit integer) and 0.75 (3×2−23\times 2^{-2}, 3-bit integer). And this may cause the phenomenon that WAGEUBN converges faster at the beginning of training. During epoch 30 and 60, we have reduced the learning rate, which is a general practice in training process . The training curves show that there is little difference between vanilla DNNs and the ones under the WAGEUBN framework when the training epoch is less than 60, which reflects the effectiveness of our approach. As the epoch evolves, the accuracy gap begins to grow because the learning rate in vanilla DNNs is much lower than that in WAGEUBN, such as 5×10−65\times 10^{-6} v.s. 1.95×10−31.95\times 10^{-3} (klr=10k_{lr}=10), thus the update of vanilla DNNs is more precise than that under the WAGEUBN framework. We can further improve the accuracy by reducing the learning rate, while the bit width values of learning rate and update need to increase accordingly at the expense of more overheads.

Table I quantitatively presents the accuracy comparison between vanilla DNNs and WAGEUBN DNNs on the ImageNet dataset. We have achieved the state-of-the-art accuracy on large-scale DNNs with full 8-bit INT quantization. The 16-bit E2E_{2} WAGEUBN only loses 3.46% mean accuracy compared with the vanilla DNNs. Because the bit width of most data keeps the same between the full 8-bit WAGEUBN and the 16-bit E2E_{2} WGAEUBN, the overhead difference between them is negligible. The DNNs under WAGEUBN framework have achieved an accuracy that is comparable to FP8 and QBP2 . And although MP and MP-INT can achieve the accuracy close to the full precision networks, the computational cost is much higher than WAGEUBN because of the floating-point data type and higher bit width, which will be detailed in Section IV-F. Compared with the vanilla DNNs, about 4×4\times memory size shrink, much faster processing speed, and much less energy and circuit area can be achieved under the proposed WAGEUBN framework.

IV-C Quantization Strategies for W, A, G, E, and BN

In our WAGEUBN framework, we use different quantization strategies for W, A, G, E, and BN, i.e. Q(⋅)Q(\cdot) for W, A, and BN; CQ(⋅)CQ(\cdot) for G, and SQ(⋅)SQ(\cdot) for E. Different quantization strategies are based on the data distribution, data sensitivity, and hardware friendliness. Figure 7 shows the distribution comparison between W, BN (x2l\bm{x}_{2}^{l} defined in Equation (2)), A, G (weight gradient), and E (e0l\bm{e}_{0}^{l}, e3l\bm{e}_{3}^{l} defined in Equation (3)) before and after quantization.

According to the definition, the resolution of the direct-quantization function is 2−72^{-7} when the bit width equals 8 and there is no limitation on the data range. Because W, BN, and A in the inference stage directly affect the loss function and further influence the backpropagation, the quantization of W, BN, and A should be as precise as possible to avoid the loss fluctuation. This is guaranteed for the reason that the resolution of the direct-quantization function is enough for W, BN, and A, which indicates that the direct-quantization function barely changes their data distributions.

The constant-quantization function has a resolution of 2−142^{-14} and the data range after quantization is about [−2−7,2−7][-2^{-7},2^{-7}] in the case of kGC=15k_{GC}=15 and k=8k=8. kk will decrease as the training epoch goes on, causing the data range reduction. Figure 7 reveals that the constant-quantization function changes the data distribution of G greatly while the network accuracy has not declined much as a result. The reason behind this phenomenon is that it is the orientation rather than the magnitude of gradients that guides DNNs to converge. In the meantime, it is easy to ensure that the bit width of updates can be fixed when kGCk_{GC} is fixed, which is more hardware-friendly since the bit width of weights stored in memory can also be fixed during training.

The shift-quantization retains the magnitude order and omits the general values whose absolute value is less than 2−72^{-7} when k=8k=8. The 8-bit shift-quantization works well for the quantization of error after activation (e1l\bm{e}_{1}^{l} defined in Equation (3)). However, we find the shift-quantization is not enough for the quantization of errors between Conv and BN (e3l\bm{e}_{3}^{l} defined in Equation (3)). Therefore, the newly designed quantization function in Equation (17) named 8-bit Flag QE2Q_{E_{2}} is utilized. Figure 7 shows that the distribution of E (e3l\bm{e}_{3}^{l}) is almost the same before and after quantization, revealing the validity of the 8-bit Flag QE2Q_{E_{2}} quantization function.

IV-D Accuracy Sensitivity Analysis

To compare the influences of W, A, G, E, and BN quantization individually, we quantize them to 8-bit INT separately with the FP32 update. Taking kW=8k_{W}=8 as an example, we quantize only W to 8-bit INT and leaving others (A, BN, G, E, and U) still kept in FP32. The quantization function for single data used here is the same as what Section III-D describes (Equation (17) is used for the error quantization when kE2=8k_{E_{2}}=8).

The results of ResNet18 under the WAGEUBN framework with single data quantization is shown in Table II. The accuracy of single data quantization reflects the difficulty degree when quantizing W, A, G, E, and BN, separately. From the table, we can see that the quantization of E, especially e3l\bm{e}_{3}^{l} defined in Equation (3), makes the most impacts on accuracy. In addition, we find that the accuracy heavily fluctuates during training when e3l\bm{e}_{3}^{l} is constrained to 8-bit INT, which does not appear in the quantization of other data. To sum up, the E data, especially e3l\bm{e}_{3}^{l}, demands the highest precision and is the most sensitive component under our WAGEUBN framework.

Because the BN process in the forward pass and the average gradients in the backward pass all involve the batch size, the sensibility of batch size to network performance under WAGEUBN is also explored. We have done comparative experiments and the results are illustrated in Figure 8. When the batch size changes between 128 and 16, the accuracy of the full precision DNNs won’t drop significantly, while that of the full 8-bit WAGEUBN DNNs has a relatively large decline in accuracy when the batch size equals 16. One reason for this lies in that there is a momentum used to update the moving means and standard deviations in a batch for full precision DNNs while WAGEUBN abandons this considering the computational cost. Besides, we think it is normal because quantization will inevitably result in the loss of some precise information. And the loss of information will bring a slightly higher batch size sensitivity. However, the robustness of WAGEUBN is still quite good, because there is a significant decline in accuracy only when the batch size is less than a very small number 16.

IV-E Analysis of the Error Quantization between Conv and BN

Error backpropagation is the foundation of DNN training. If the error of E caused by quantization is too large, the convergence of DNNs will be degraded. Especially, because the error quantization between convolution and BN (e3l\bm{e}_{3}^{l}) is directly related to the weight update of the ll-th layer, the impact of e3l\bm{e}_{3}^{l} quantization on the model accuracy is critical.

To further analyze the reason why 8-bit QE2Q_{E_{2}} (defined in Equation (16) where kE2=8k_{E_{2}}=8) causes the non-convergence of DNNs and compare the distributions of e3l\bm{e}_{3}^{l} under 8-bit QE2Q_{E_{2}}, 8-bit Flag QE2Q_{E_{2}} (defined in Equation (17) where kE2=8k_{E_{2}}=8), and full precision, the data distribution of e3l\bm{e}_{3}^{l} of the first quantized layer on ResNet18 is shown in Figure 9. From the figure, we can see that the e31\bm{e}_{3}^{1} distributions under 8-bit QE2Q_{E_{2}} quantization and full precision differs a lot and those of 8-bit Flag QE2Q_{E_{2}} quantization and full precision are almost the same. The major difference between the 8-bit QE2Q_{E_{2}} quantization and full precision lies in the interval of [−2−8R(e31),2−8R(e31)][-{2^{-8}}R(\bm{e}_{3}^{1}),{2^{-8}}R(\bm{e}_{3}^{1})] (R(⋅)R(\cdot) is defined in Equation (7)), where the 8-bit QE2Q_{E_{2}} quantization forces the data in this range to zero.

The only difference between 8-bit QE2Q_{E_{2}} and 8-bit Flag QE2Q_{E_{2}} quantization functions lies in the data range. Theoretically, the covered data range of 8-bit QE2Q_{E_{2}} and 8-bit Flag QE2Q_{E_{2}} are about [−R(e3l),−2−8R(e3l)]∪{0}∪[2−8R(e3l),R(e3l)][-R(\bm{e}_{3}^{l}),-{2^{-8}}R(\bm{e}_{3}^{l})]\cup\{0\}\cup[{2^{-8}}R(\bm{e}_{3}^{l}),R(\bm{e}_{3}^{l})] and [−R(e3l),−2−15R(e3l)][-R(\bm{e}_{3}^{l}),-{2^{-15}}R(\bm{e}_{3}^{l})] ∪\cup {0}\{0\} ∪\cup [2−15R(e3l),R(e3l)][{2^{-15}}R(\bm{e}_{3}^{l}),R(\bm{e}_{3}^{l})], respectively. Because the distribution of e3l\bm{e}_{3}^{l} is not uniform, the range covered by different quantization methods varies a lot. The data ratios (Proportion of non-zero values after quantization.) of 8-bit QE2Q_{E_{2}} and 8-bit Flag QE2Q_{E_{2}} quantization functions covered by each layer of ResNet18 are illustrated as Figure 10. Although the larger values take greater impacts on the model accuracy in the process of error propagation, the smaller values also contain useful information and occupy the majority. Compared with 8-bit Flag QE2Q_{E_{2}}, the data ratio 8-bit QE2Q_{E_{2}} covers is too little because of the smaller data range. That is to say, although the most important information contained by the larger values is retained, the information contained by the smaller values is ignored, resulting in the non-convergence of DNNs. In addition, there is also a rough trend that the data ratio decreases as the network becomes shallower, either in the 8-bit QE2Q_{E_{2}} or 8-bit Flag QE2Q_{E_{2}} quantization.

IV-F Cost Discussion

Although it is recognized that DNN quantization can greatly reduce memory and compute costs, resulting in lower energy consumption, quantitative analysis is rarely seen in recent research. In order to compare the full INT8 quantization with other precision solutions (FP32, INT32, FP16, INT16, and FP8) more clearly, we have simulated the processing speed, power consumption, and circuit area for single multiplication and accumulation operation on FPGA platform. Figure 11 shows the results. With FP32 as the baseline, taking the multiplication operation as an example, INT8 can perform >>3×\times faster in speed, 10×\times lower in power, and 9×\times smaller in circuit area. Similarly, compared with FP32, the speed of INT8 accumulation is about 9×\times faster, and the energy consumption and circuit area are reduced by >>30×\times. In addition, the INT8 multiplication and accumulation operations are more advantageous than other data type operations, whether it is FP8, INT16, FP16 or INT32. In conclusion, the proposed full INT8 quantization has great advantages in hardware overheads, whether in terms of memory cost, processing speed, power consumption, and circuit area. Given the huge advantages in hardware resources and computing speed of quantization, the idea of low-precision computing and the quantization functions in WAGEUBN can also extend to other research fields which involve large-scale matrix operations, such as large-scale control systems , weather forecasting models , etc.

V CONCLUSIONS

We propose a unified framework termed as “WAGEUBN” to achieve a complete quantization of large-scale DNNs in both training and inference with competitive accuracy. We are the first to quantize DNNs over all data paths and promote DNN quantization to the full INT8 level. In this way, all the operations can be replaced with bit-wise operations, causing significant improvements in memory overhead, processing speed, circuit area, and energy consumption. Extensive experiments evidence the effectiveness and efficiency of WAGEUBN. This work provides a feasible solution for the online training acceleration of large-scale and high-performance DNNs and further shows the great potential for the applications in future efficient portable devices with online learning ability. Although great progress has been made, WAGEUBN still needs to build a special hardware architecture design and develop more applications. Future works could transfer to the design of computing architecture, memory hierarchy, interconnection infrastructure, and mapping tool to enable the specialized machine learning chips and more applications.

ACKNOWLEDGMENT

The work was supported partially by the National Natural Science Foundation of China (No. 61603209, 61876215).

References