Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers
Yukuan Yang, Shuang Wu, Lei Deng, Tianyi Yan, Yuan Xie, Guoqi Li
I INTRODUCTION
Deep neural networks have achieved state-of-art results in many fields like image processing , object detection , natural language processing , and robotics through learning high-level features from a large amount of input data. However, due to the existence of a huge number of floating-point (FP) values and complex FP multiply-accumulate operations (MACs) in the process of network training and inference, the intensive memory overhead, large computational complexity, and high energy consumption impede the wide deployment of deep learning models. DNN quantization which converts FP MACs to bit-wise operations is an effective way to reduce the memory and computation costs and improve the speed of deep learning accelerators.
With the deepening of research, DNN quantization gradually transfers from inference quantization (BWN , XNOR-Net , ADMM ) to training quantization (DoReFa , GXNOR-Net , FP8 ). Usually, the inference quantization focuses on the forward pass; while the training quantization further quantizes the backward pass and weight updates. Recently, training quantization becomes a hot topic in the network compression community. Whereas, there are still two major issues in existing schemes. The first issue lies in the incomplete quantization, including two aspects: partial quantization and FP dependency. Partial quantization means that only parts of dataflows, not all of them, are quantized (e.g. DoReFa , GXNOR-Net and QBP2 ); FP dependency still remains FP values during the training process (e.g. MP and FP8 ). The second issue is that the quantization of batch normalization (BN) is ignored by most schemes (e.g. MP-INT and FX Training ). BN is an essential layer for the training of DNNs by addressing the problem of the internal covariate shift of each layer’s inputs, especially as the network deepens, allowing a much higher learning rate and less careful weight initialization.
Compared with all the studies above, WAGE is the most thorough work of DNNs quantization, which quantizes the data including W (Weights), A (Activation), G (Gradient), E (Error), U (Update) and replacing each BN layer with a constant scaling factor. WAGE has achieved competitive results on LeNet , VGG , and AlexNet , providing a good inspiration for this work. However, we find that WAGE is difficult to be applied in large-scale DNNs due to the absence of BN layers. Besides, it is known that the gradient descent optimizer such as Momentum or Adam increases the stability and even helps get rid of the local optimum, thus the speed and final performance are significantly improved. A complete quantization should cover the entire training process, including W, A, G, E, U, BN, and the optimizer. Regretfully, up to now, there is still no such solution that can achieve this complete quantization, especially on large-scale DNNs.
To address the issues of incomplete quantization and ignored BN quantization mentioned above and extend quantization framework to large-scale datasets and networks with high performance, we propose a unified complete quantization framework termed “WAGEUBN” to constrain W, A, G, E, U, BN, and the optimizer in the low-bit integer (INT) space. To the best of our knowledge, WAGEUBN is the first complete quantization framework achieving high performance in large-scale datasets, where all computation steps and operands in DNNs are decomposed and quantized.
We mainly make the following efforts to create the complete quantization framework. Firstly, according to the various data distributions and the role the quantized data plays in DNN training, we fuse three quantization functions to satisfy the different precision requirements. Furthermore, we propose a new storage and computing method by introducing a flag bit to expand the data coverage and solve the non-convergence problem caused by insufficient data representation. Last but not least, we quantize BN and Momentum optimizer for the first time, converting all FP operations in DNNs to bit-wise operations. Compared with the full precision DNNs, DNNs under the full 8-bit WAGEUBN framework can achieve about memory saving. More importantly, the multiplication and accumulation operations of WAGEUBN, which are the main operations in DNNs, can perform 3 and 9 faster in speed, 10 and 30 lower in power, 9 and 30 smaller in circuit area, respectively. Besides, the efficient INT8 multiplication and accumulation operations also make WAGEUBN a big step ahead of most existing quantization schemes in computational costs, whether it is FP8, INT16, FP16 or INT32. In addition to the huge advantages in memory cost, computing speed, energy consumption and circuit area, WAGEUBN also shows competitive accuracy on large-scale networks (ResNet18/34/50) and dataset (ImageNet ). What’s more, the hardware design of WAGEUBN is much simpler and more efficient because of the complete quantization. Due to the improvement of computing speed and saving in hardware resources and energy consumption, WAGEUBN provides a feasible idea for the architecture design of future efficient online learning chips used in portable devices with limited computational resources. The contributions of this work are twofold, which are summarized as follows:
We address two main issues existing in most quantization schemes via fully quantizing all the data paths, including W, A, G, E, U, BN, and the optimizer, greatly reducing the memory and compute costs. What’s more, we constrain the data to INT8 for the first time, pushing the training quantization to a new bit level compared with the existing FP16, INT16, and FP8 solutions.
Our quantization framework is validated in large-scale DNN models (ResNet18/34/50) over ImageNet dataset and achieves competitive accuracy with much fewer overheads, indicating great potential for future portable devices with online learning ability.
The organization of this paper is as follows: Section II introduces the related work of DNN quantization; Section III details the WAGEUBN framework; Section IV presents the experiment results of WAGEUBN and the corresponding analyses; Section V summarizes this work and delivers the conclusion.
II Related Work
With the wide applications of DNNs, the related compression technologies have been proposed rapidly, among which the quantization plays an important role. The development of DNN quantization can be divided into two stages, inference quantization and training quantization, according to the different quantization objects.
Inference quantization: Inference quantization starts from constraining W into (BWN ), replacing complex FP MACs with simple accumulations. BNN and XNOR-Net further quantize both W and A, making the inference computation dominated by bit-wise operations. However, extremely low bit-width quantization usually leads to significant accuracy loss. For example, when the bit width comes to 4 bits, the accuracy degradation becomes obvious, especially for large-scale DNNs. Instead, the bit width of W and A for inference quantization can be reduced to 8 bits with little accuracy degradation. The study of inference quantization is sufficient for the deep learning inference accelerators. Whereas, this is not enough for efficient online learning accelerators because only the data in the forward pass are considered.
Training quantization: To further extend the quantization towards the training stage, DoReFa trains DNNs with low bit-width W, A, and G, while leaving E and BN unprocessed. MP and MP-INT use FP16 and INT16 values, respectively, to constrain W, A, and G. Recently, FP8 further pushes W, A, G, E, and U to 8, 8, 8, 8, and 16-bit FP values, respectively, still leaving BN untouched. QBP2 replaces the conventional BN with range BN and constrains W, A, and E to INT8 values while calculating G with FP MACS. Recently, WAGE adopts a layer-wise scaling factor instead of using the BN layer and quantizes W, A, G, E, and U to 2, 8, 8, 8, and 8 bits, respectively. Despite its thorough quantization, WAGE is difficult to be applied to large-scale DNNs due to the absence of powerful BN layers. In summary, there still lacks a complete INT8 quantization framework for the training of large-scale DNNs with high accuracy.
III WAGEUBN Framework
The main idea of WAGEUBN is to quantize all the data in DNN training to INT8 values. In this section, we detail the WAGEUBN framework implemented in large-scale DNN models. The organization of this section is as follows: Subsection III-A introduces the straight-through estimator (STE) method which is accepted and used by most researchers to solve the non-differentiable problem of quantization; Subsection III-B and Subsection III-C describe the notations and quantization functions, respectively; Subsection III-D explains the specific quantization schemes for W, A, G, E, U, BN, and the Momentum optimizer, respectively; Subsection III-E goes through the overall implementation of WAGEUBN, including in both forward and backward passes; Subsection III-F summarizes the whole process and shows the pseudo codes.
In the early stage of the study, the non-differentiable problem in mathematical sense caused by quantization has been hindering the development of quantization research. However, since the straight-through estimator (STE) method was used to estimate the gradient of quantized data in BNN , almost all the works in the field of quantization have adopted this method to avoid the mathematical non-derivative problem .
The STE method used in WAGEUBN can be illustrated as the following
where is the data to be quantized, is the quantization function that may be non-differentiable, and denotes the objective function.
III-B Notations
Before introducing the WAGEUBN quantization framework formally, we need to define some notations. Considering the -th layer of DNNs, we divide the forward pass of DNNs into four steps as described in Figure 1 (BN is divided into two steps: Normalization & and Scale & Offset).
Different from most existing schemes, we define and respectively, where represents the gradient of A (activation) which is used in the error backpropagation and represents the gradient of W (weights) which is used in the weight update. Moreover, we quantize the BN layers, including both the forward and backward passes, which is not well touched in most prior work. Similar to the forward pass, we divide the backward pass of the -th layer into five steps as shown in Figure 2. According to the derivative chain rules, we have
where is the loss function, represents the error from the -th layer, and represents the Hadamard product. For vectors with the same dimension, such as and , we have: . Two quantization functions are used here: is the quantization function detailed as Equation (15) that converts high bit-width integers to low bit-width integers; detailed as Equation (17) is trying to convert FP values to low bit-width integers. is the transposed matrix of , and represents the gradient of activation. When is used as the activation function, is a tensor containing only 0 and 1 elements.
According to the definitions given above, the gradients of W, , and can be summarized as follows
To further reduce the bit width of G that will increase greatly after the multiplication, we have
where , , and are quantization functions for the gradient of W, , and , respectively, which will be shown in Equation (18).
Some notations to be used below are also explained here. , , , ( and ), and are the bit width of W, A, G, E, and BN, respectively. , , and are the bit width of W, , and update, which are also the bit width of data stored in memory. , , , and are the bit width of , , , and , respectively, used in the BN layer. and are the bit width of and gradient, respectively. and are the bit width of momentum coefficient () and accumulation (), respectively, used in the Momentum optimizer. At last, is the bit width of the learning rate.
III-C Quantization Functions
There are three quantization functions used in WAGEUBN. The direct-quantization function uses the nearest fixed-point values to represent the continuous values of W, A, and BN. The constant-quantization function for G is used to keep the bit width of U (update) fixed since G is directly related to U. Because U and the weights stored in memory have the same bit width, the bit width of weights stored in memory can be fixed, which is more hardware-friendly. The magnitude of E is very small, so the shift-quantization function reduces the bit width of E greatly compared with the direct-quantization function under the same precision.
The direct-quantization function simply approximates a continuous value to its nearest discrete state and is defined as
where is the bit width, and rounds a number to its nearest INT value.
The intention of constant-quantization function is to normalize a tensor firstly, then limit it to INT, and finally maintain its magnitude. It is governed by
The illustration of constant-quantization function is described in Figure 3. Here, is used to project the maximum value of to its nearest fixed-point value, which is prepared for normalization; is a stochastic rounding function used for converting a continuous float value to its nearby INT value in a probabilistic manner and is the rounding probability; denotes normalization and is a saturation function limiting the data range; is to shift the distribution of and limit to INT values between and . Here limits the data range after mapping and decreases as the training goes on, presenting the same effect as reducing the learning rate. For example, maps G to {-127, -126, , 126, 127} and {-63, -62, , 62, 63} in the early training stage (, , epoch in $k=7dr=64CQ(\cdot)2^{k_{GC}-1}k_{GC}$ is its bit width.
The shift-quantization function serves for the quantization of E and is defined as
where is the minimum interval for a k-bit INT and is the direct-quantization function defined in Equation (6).
The shift-quantization function normalizes E first, then converts E to fixed-point values, and finally uses a layer-wise scaling factor ( defined in Equation (7)) to maintain the magnitude. The differences between the constant-quantization function and the shift-quantization function mainly exist in two points: First, the constant-quantization uses a constant to keep the magnitude for hardware friendliness while the shift-quantization uses a lay-wise scaling factor; Second, the constant-quantization contains a stochastic rounding process while the shift-quantization function does not.
III-D Quantization Schemes in WAGEUBN
After introducing the quantization functions used in our WAGEUBN framework, we provide detailed quantization schemes.
Since weights are stored and used as fixed-point values, weights should be also initialized discretely. An initialization method proposed by MSRA has been evidenced helpful for faster training. The initialization of weights can be formulated as follows
where is the layer’s fan-in number, and is the bit width of weight update and the memory storage.
Because of the different bit width for weight storage and computation, it should be quantized from bits to (the bit width of weights used for convolution) bits for convolution. In addition, we also limit the data range of W. Finally, the quantization function for W is
As aforementioned, BN plays an important role in training large-scale DNNs. WAGE has proved that simple scaling layers are not enough to replace BN layers. Conventional BN layer can be divided into two steps as
where and are the mean and standard, respectively, deviation of over one mini-batch in the -th layer; is a small positive value added to to avoid the case of dividing by zero; and are the scale and offset parameters, respectively.
Under the WAGEUBN framework, the BN layer is also quantized. Through the operations described in Equation (12), all operands are quantized and all operations are bit-wise. Specifically, the quantization follows
where are the quantization functions converting the operands to fixed-point values defined as
And is a small fixed-point value, playing the same role as in Equation (11); are the bit width of and , respectively.
After the convolution and BN layers in the forward pass, the bit width of operands increases due to the multiplication operation. To reduce the bit width and keep the input bit width of each layer consistent, we need to quantize the activations. Here, the quantization function for activations can be described as
where is the bit width of activations.
In Equation (3), we have given the definition of E and quantized E. Through investigating the importance of error propagation in DNN training, we find that the quantization of E is very essential for the model convergence. If E is naively quantized using the direct-quantization function, it will require a large bit width of operands to realize the convergence of DNNs. Instead, we use the following shift-quantization function
where is the shift-quantization function defined in Equation (8), and is the bit width of defined in Equation (3).
As mentioned above, we use and for the error quantization. However, the precision requirements of and vary a lot. Experiments show that affects little on accuracy while will cause the non-convergence of large-scale DNNs when using as the quantization function. is a proper value for the training of DNNs with minimum accuracy degradation. More analyses will be given in Subsection IV-E. Here we will provide two versions of , the 16-bit and 8-bit versions. The 16-bit is defined as
where is the bit width of defined in Equation (3).
Experiments have proved the data range covered by 8-bit () is not sufficient to train DNNs. In order to expand the coverage of quantization function while still maintaining a low bit width, we introduce a layer-wise scaling factor and a flag bit. Then, to distinguish it from defined in Equation (16), we name the quantization function Flag and the quantization process is governed by
where , , and .
By introducing a layer-wise scaling factor and a flag bit, we can expand the data coverage greatly. Details can be found in Figure 4. The flag bit is used to indicate whether the absolute value of stored in memory is less than the layer-wise scaling factor (e.g., 0 represents and 1 represents ). The sign bit is used to denote the positive or negative direction of the value. The data bit follows the conventional binary format. According to the definition, the values stored in Figure 4 and 4 are and when , respectively. Therefore, the 9-bit data format can cover almost the same data range as the direct 15-bit quantization described in Equation (16). Since the flag bit is just used for judgment, the effective value for computation is still INT8.
The gradient is another important part in DNN training because it is directly related to the weight update. The rules for calculating and quantizing the gradients of W, , and are described as Equation (4) and (5). Since , , , and are all fixed-point values, the conventional FP MACs operations can be replaced with bit-wise operations during the process of calculating , , and . The quantization functions are defined to further reduce the bit width of gradients and prepare for the next step of the optimizer. Specifically, we have
where is the constant-quantization function defined in Equation (7); , , and are the bit width of the gradient of W, , and , respectively.
Momentum optimizer is one of the most common optimizers used in DNN training, especially for classification tasks. For the -th training step of the -th layer, the conventional Momentum optimizer works as follows
where and are the accumulation in the -th and -th training step, respectively; is a constant value used as a coefficient; is the gradient of W, , or .
Momentum optimizer under the WAGEUBN framework is trying to constrain all operands to fixed-point values. The process can be formulated as
where is the quantized accumulation in the -th training step; is the quantized gradient of W, , or ; is the quantization function defined as
To guarantee the consistency of bit width, we further set
The parameter update is the last step in the training of each mini-batch. Different from conventional DNNs where the learning rate can take any FP value, the learning rate under WAGEUBN must also be a fixed-point value and the bit width of update is directly related to the bit width of learning rate. The update under quantized Momentum optimizer can be described as
where is the update of W with bits, and is the fixed-point learning rate with bits. The updates of and are the same as in Equation (23). According to Equation (20), (22), and (23), we have
Through our evaluations, the precision of the update has the greatest impact on the accuracy of DNNs because it is the last step to constrain the parameters. Thus, we need to set a reasonable bit width for update to balance the model accuracy and memory cost.
III-E Quantization Framework
Given the quantization details of W, A, G, E, U, BN, and the Momentum optimizer, the overall quantization framework is depicted in Figure 5. Under this framework, conventional FP MACs can be replaced with bit-wise operations. Here, the forward pass of the -th layer in DNNs is divided into three parts: Conv (convolution), BN, and activation. , , , , and are defined in Equation (2). The weights () are stored as -bit integers and then maps to -bit INT values () before convolution. After convolution, and operations are used to calculate the mean and standard deviation of in one mini-batch and then quantize them to and bits, respectively. The operation constrains to bits in BN. Similar to W, and are stored as and -bit integers and used in and bits ( and ) after the and quantization, respectively. After the second step of BN, activation and quantization are implemented with the operation, reducing the increased bit width to bits again and preparing inputs for the next layer.
The backward pass of the -th layer is much more complicated than the forward pass, including error propagation, gradient of weight, gradient of BN, Momentum optimizer, and weight update. In the process of error propagation, , , , , and are defined in Equation (3) and there are two locations needing quantization using and . reduces the bit width of from to . is the derivative of activation function () and is used to constrain to bits. In the phase of calculating the gradients of weights and BN, , , and are leveraged to reduce the increased bit width caused by the multiplication operations.
All parameters of the Momentum optimizer in the -th training step are quantized. Different from the conventional learning rate with FP value, WAGEUBN requires a discrete learning rate so that the bit width of weight updates can be controlled. The updates of and in BN layers are also similar to the weight update, which are omitted in Figure 5 for simplicity.
III-F Overall Algorithm
Given the framework of WAGEUBN, we summarize the entire quantization process and present the pseudo codes for both the forward and backward passes as shown in Algorithm 1 and 2, respectively.
IV Results
To verify the effectiveness of the proposed quantization framework, we apply WAGEUBN on ResNet18/34/50 on ImageNet dataset. We provide two versions of WAGEUBN, one with full 8-bit INT where , , , , , , and are equal to 8. The other version has 16-bit . The only difference between the 16-bit version and the full 8-bit version exists in the quantization function (see Equation (16) and (17), respectively). , , and are 15 and we set , , , and to 3, 13, 10, and 24 respectively to satisfy Equation (22) and (24). In addition, we set , , and to 16. Since W, A, G, and E occupy the majority of memory and compute costs, their bit-width values are reduced as much as possible. Other parameters occupying much less resources can increase the bit width to maintain the accuracy, e.g., , , , and in BN layers.
The first and last layers are believed to differ from the rest because of their interface with network inputs and outputs. The quantization of these two layers will cause significant accuracy degradation compared to hidden layers and they just consume few overheads due to the small number of neurons. Therefore, we do not quantize the first and last layers, as previous work did .
For the quantization frameworks, it’s a common issue of saturated learning and vanishing gradients because the data in DNNs may be clipped very small after quantization and the model may be stuck on a local optimum because of the few updates. And this situation can be more serious when the quantization framework is used in large-scale complex datasets. WAGEUBN has made many efforts to alleviate this problem. Firstly, the newly designed shift-quantization function ensures that the quantized error in WAGEUBN framework is on the same magnitude as that of a traditional FP network, making the errors (gradients of activations) not vanishing in the process of back propagation. Secondly, based on the observation that it is the orientation rather than the magnitude of gradients (gradients of weights) that guides deep neural networks (DNNs) to converge, it’s more important to keep the orientation rather than the magnitude. What’s more, the updates () jointly determined by the learning rate and gradients are the data that ultimately determines the final changes of the network. The much bigger learning rate setting in WAGEUBN (minimum learning rate: WAGEUBN () VS floating-point network ) can make up for the problem of update magnitude caused by the smaller quantized gradients. Through the technics above, the problems of saturated learning and dead gradients can be solved. In addition to the above measures which can help the model get rid of local optimum, the proposed quantized Momentum optimizer, fixed-point updates which is precise enough and a proper batch size selection can also reduce the probability of falling into a local optimum greatly.
IV-B Training Curve
Figure 6 illustrates the accuracy comparison between vanilla DNNs (FP32), DNNs with the full 8-bit version of WAGEUBN, and DNNs with the 16-bit version of WAGEUBN. The initial learning rate and momentum coefficient under WAGEUBN and full precision are slightly different. To ensure the best recognition accuracy, we use the official parameter settings of TensorFlow , where the initial learning rate and momentum coefficient are set to 0.05 (batch size is 128) and 0.9 while those of WAGEUBN framework are set to 0.05078125 (, 10-bit integer) and 0.75 (, 3-bit integer). And this may cause the phenomenon that WAGEUBN converges faster at the beginning of training. During epoch 30 and 60, we have reduced the learning rate, which is a general practice in training process . The training curves show that there is little difference between vanilla DNNs and the ones under the WAGEUBN framework when the training epoch is less than 60, which reflects the effectiveness of our approach. As the epoch evolves, the accuracy gap begins to grow because the learning rate in vanilla DNNs is much lower than that in WAGEUBN, such as v.s. (), thus the update of vanilla DNNs is more precise than that under the WAGEUBN framework. We can further improve the accuracy by reducing the learning rate, while the bit width values of learning rate and update need to increase accordingly at the expense of more overheads.
Table I quantitatively presents the accuracy comparison between vanilla DNNs and WAGEUBN DNNs on the ImageNet dataset. We have achieved the state-of-the-art accuracy on large-scale DNNs with full 8-bit INT quantization. The 16-bit WAGEUBN only loses 3.46% mean accuracy compared with the vanilla DNNs. Because the bit width of most data keeps the same between the full 8-bit WAGEUBN and the 16-bit WGAEUBN, the overhead difference between them is negligible. The DNNs under WAGEUBN framework have achieved an accuracy that is comparable to FP8 and QBP2 . And although MP and MP-INT can achieve the accuracy close to the full precision networks, the computational cost is much higher than WAGEUBN because of the floating-point data type and higher bit width, which will be detailed in Section IV-F. Compared with the vanilla DNNs, about memory size shrink, much faster processing speed, and much less energy and circuit area can be achieved under the proposed WAGEUBN framework.
IV-C Quantization Strategies for W, A, G, E, and BN
In our WAGEUBN framework, we use different quantization strategies for W, A, G, E, and BN, i.e. for W, A, and BN; for G, and for E. Different quantization strategies are based on the data distribution, data sensitivity, and hardware friendliness. Figure 7 shows the distribution comparison between W, BN ( defined in Equation (2)), A, G (weight gradient), and E (, defined in Equation (3)) before and after quantization.
According to the definition, the resolution of the direct-quantization function is when the bit width equals 8 and there is no limitation on the data range. Because W, BN, and A in the inference stage directly affect the loss function and further influence the backpropagation, the quantization of W, BN, and A should be as precise as possible to avoid the loss fluctuation. This is guaranteed for the reason that the resolution of the direct-quantization function is enough for W, BN, and A, which indicates that the direct-quantization function barely changes their data distributions.
The constant-quantization function has a resolution of and the data range after quantization is about in the case of and . will decrease as the training epoch goes on, causing the data range reduction. Figure 7 reveals that the constant-quantization function changes the data distribution of G greatly while the network accuracy has not declined much as a result. The reason behind this phenomenon is that it is the orientation rather than the magnitude of gradients that guides DNNs to converge. In the meantime, it is easy to ensure that the bit width of updates can be fixed when is fixed, which is more hardware-friendly since the bit width of weights stored in memory can also be fixed during training.
The shift-quantization retains the magnitude order and omits the general values whose absolute value is less than when . The 8-bit shift-quantization works well for the quantization of error after activation ( defined in Equation (3)). However, we find the shift-quantization is not enough for the quantization of errors between Conv and BN ( defined in Equation (3)). Therefore, the newly designed quantization function in Equation (17) named 8-bit Flag is utilized. Figure 7 shows that the distribution of E () is almost the same before and after quantization, revealing the validity of the 8-bit Flag quantization function.
IV-D Accuracy Sensitivity Analysis
To compare the influences of W, A, G, E, and BN quantization individually, we quantize them to 8-bit INT separately with the FP32 update. Taking as an example, we quantize only W to 8-bit INT and leaving others (A, BN, G, E, and U) still kept in FP32. The quantization function for single data used here is the same as what Section III-D describes (Equation (17) is used for the error quantization when ).
The results of ResNet18 under the WAGEUBN framework with single data quantization is shown in Table II. The accuracy of single data quantization reflects the difficulty degree when quantizing W, A, G, E, and BN, separately. From the table, we can see that the quantization of E, especially defined in Equation (3), makes the most impacts on accuracy. In addition, we find that the accuracy heavily fluctuates during training when is constrained to 8-bit INT, which does not appear in the quantization of other data. To sum up, the E data, especially , demands the highest precision and is the most sensitive component under our WAGEUBN framework.
Because the BN process in the forward pass and the average gradients in the backward pass all involve the batch size, the sensibility of batch size to network performance under WAGEUBN is also explored. We have done comparative experiments and the results are illustrated in Figure 8. When the batch size changes between 128 and 16, the accuracy of the full precision DNNs won’t drop significantly, while that of the full 8-bit WAGEUBN DNNs has a relatively large decline in accuracy when the batch size equals 16. One reason for this lies in that there is a momentum used to update the moving means and standard deviations in a batch for full precision DNNs while WAGEUBN abandons this considering the computational cost. Besides, we think it is normal because quantization will inevitably result in the loss of some precise information. And the loss of information will bring a slightly higher batch size sensitivity. However, the robustness of WAGEUBN is still quite good, because there is a significant decline in accuracy only when the batch size is less than a very small number 16.
IV-E Analysis of the Error Quantization between Conv and BN
Error backpropagation is the foundation of DNN training. If the error of E caused by quantization is too large, the convergence of DNNs will be degraded. Especially, because the error quantization between convolution and BN () is directly related to the weight update of the -th layer, the impact of quantization on the model accuracy is critical.
To further analyze the reason why 8-bit (defined in Equation (16) where ) causes the non-convergence of DNNs and compare the distributions of under 8-bit , 8-bit Flag (defined in Equation (17) where ), and full precision, the data distribution of of the first quantized layer on ResNet18 is shown in Figure 9. From the figure, we can see that the distributions under 8-bit quantization and full precision differs a lot and those of 8-bit Flag quantization and full precision are almost the same. The major difference between the 8-bit quantization and full precision lies in the interval of ( is defined in Equation (7)), where the 8-bit quantization forces the data in this range to zero.
The only difference between 8-bit and 8-bit Flag quantization functions lies in the data range. Theoretically, the covered data range of 8-bit and 8-bit Flag are about and , respectively. Because the distribution of is not uniform, the range covered by different quantization methods varies a lot. The data ratios (Proportion of non-zero values after quantization.) of 8-bit and 8-bit Flag quantization functions covered by each layer of ResNet18 are illustrated as Figure 10. Although the larger values take greater impacts on the model accuracy in the process of error propagation, the smaller values also contain useful information and occupy the majority. Compared with 8-bit Flag , the data ratio 8-bit covers is too little because of the smaller data range. That is to say, although the most important information contained by the larger values is retained, the information contained by the smaller values is ignored, resulting in the non-convergence of DNNs. In addition, there is also a rough trend that the data ratio decreases as the network becomes shallower, either in the 8-bit or 8-bit Flag quantization.
IV-F Cost Discussion
Although it is recognized that DNN quantization can greatly reduce memory and compute costs, resulting in lower energy consumption, quantitative analysis is rarely seen in recent research. In order to compare the full INT8 quantization with other precision solutions (FP32, INT32, FP16, INT16, and FP8) more clearly, we have simulated the processing speed, power consumption, and circuit area for single multiplication and accumulation operation on FPGA platform. Figure 11 shows the results. With FP32 as the baseline, taking the multiplication operation as an example, INT8 can perform 3 faster in speed, 10 lower in power, and 9 smaller in circuit area. Similarly, compared with FP32, the speed of INT8 accumulation is about 9 faster, and the energy consumption and circuit area are reduced by 30. In addition, the INT8 multiplication and accumulation operations are more advantageous than other data type operations, whether it is FP8, INT16, FP16 or INT32. In conclusion, the proposed full INT8 quantization has great advantages in hardware overheads, whether in terms of memory cost, processing speed, power consumption, and circuit area. Given the huge advantages in hardware resources and computing speed of quantization, the idea of low-precision computing and the quantization functions in WAGEUBN can also extend to other research fields which involve large-scale matrix operations, such as large-scale control systems , weather forecasting models , etc.
V CONCLUSIONS
We propose a unified framework termed as “WAGEUBN” to achieve a complete quantization of large-scale DNNs in both training and inference with competitive accuracy. We are the first to quantize DNNs over all data paths and promote DNN quantization to the full INT8 level. In this way, all the operations can be replaced with bit-wise operations, causing significant improvements in memory overhead, processing speed, circuit area, and energy consumption. Extensive experiments evidence the effectiveness and efficiency of WAGEUBN. This work provides a feasible solution for the online training acceleration of large-scale and high-performance DNNs and further shows the great potential for the applications in future efficient portable devices with online learning ability. Although great progress has been made, WAGEUBN still needs to build a special hardware architecture design and develop more applications. Future works could transfer to the design of computing architecture, memory hierarchy, interconnection infrastructure, and mapping tool to enable the specialized machine learning chips and more applications.
ACKNOWLEDGMENT
The work was supported partially by the National Natural Science Foundation of China (No. 61603209, 61876215).