Forward and Backward Information Retention for Accurate Binary Neural Networks
Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, Jingkuan Song
Introduction
Deep neural networks (DNNs), especially convolutional neural networks (CNNs), have been well demonstrated in a wide variety of computer vision applications such as image classification , object detection and semantic segmentation . Traditional CNNs are usually with massive parameters and high computational complexity for the requirement of high accuracy. Consequently, deploying the most advanced deep CNN models requires expensive storage and computing resources, which largely limits the applications of DNNs on portable devices such as mobile phones and cameras. Binary neural networks are appealing to the community for their tiny storage usage and efficient inference , which results from the binarization of both weights and activations and the efficient convolution implemented by bitwise operations. Although much progress has been made on binarizing DNNs, the existing quantization methods remains a significant drop of accuracy compared with the full-precision counterparts .
The performance degradation of binary neural networks is mainly caused by the limited representation ability and discreteness of binarization, which results in severe information loss in both forward and backward propagation. In the forward propagation, when the activations and weights are restricted to two values, the model’s diversity sharply decreases, while the diversity is proved to be the key of pursuing high accuracy of neural networks . Two approaches are widely used to increase the diversity of neural networks: increasing the number of neurons or increasing the diversity of feature maps. For example, Bi-Real Net is targeted at the latter by adding on a full-precision shortcut to the quantized activations, which achieves significant performance improvement. However, with the additional floating-point add operation, Bi-Real Net inevitably faces worse efficiency than vanilla binary neural networks.
The diversity implies the ability to carry enough information during the forward propagation, meanwhile, accurate gradients in the backward propagation provide correct information for optimization. However, during the training process of binary neural networks, the discrete binarization always leads to inaccurate gradients and the wrong optimization direction. To better deal with the discreteness, different approximations of binarization for backward propagation have been studied , mainly categorized into either improving the updating ability or reducing the mismatching area between function and the approximation one. Unfortunately, the difference between the early and later training stages is always be ignored, where in practice strong updating ability is usually highly required when the training process starts and small gradient error becomes more important at the end of the training. It is insufficient to acquire as much information reflected from the loss function as possible only focusing on one point.
To solve the above-mentioned problems, this paper is the first to study the model binarization from the view of information flow and proposes a novel Information Retention Network (IR-Net) (see the overview in Fig 1). Our goal is to train highly accurate binarized models by retaining the information in the forward and backward propagation: (1) IR-Net introduces a balanced and standardized quantization method called Libra Parameter Binarization (Libra-PB) in the forward propagation. With Libra-PB, we can minimize the information loss in forward propagation by maximizing the information entropy of the quantized parameters and minimizing the quantization error, which ensures a high diversity. (2) In backward propagation, IR-Net adopts the Error Decay Estimator (EDE) to calculate gradients and minimizes the information loss by better approximating the function, which ensures sufficient updating at the beginning and accurate gradients at the end of the training.
Our IR-Net presents a new and practical perspective to understand how the binarized network works. Besides the strong capability of preserving the information forward/backward in the deep network, it also enjoys good versatility and can be optimized in a standard network training pipeline. We evaluate our IR-Net with image classification tasks on the CIFAR-10 and ImageNet datasets. The experimental results show that our method performs remarkably well across various network structures such as ResNet-20, VGG-Small, ResNet-18, and ResNet-34, surpassing previous quantization methods by a wide margin. Our code is released at GitHub.
Related Work
Network binarization aims to accelerate the inference of neural networks and save memory occupancy without much accuracy degradation. One approach to speed up low-precision networks is to utilize bitwise operations. By directly binarizing the 32-bit parameters in DNNs including weights and activations, we can achieve significant accelerations and memory reductions. XNOR-Net utilizes a deterministic binarization scheme and minimizes the quantization error of the output matrix by employing some scalars in each layer. TWN and TTQ enhance the representation ability of neural networks with more available quantization points. ABC-Net recommends using more binary bases for weights and activations to improve accuracy, while compression and acceleration ratios are reduced accordingly. proposed HWGQ considering the quantization error from the aspect of activation function. further proposed LQ-Net with more training parameters, which achieved comparable results on the ImageNet benchmark but increased the memory overhead.
Compared with other model compression methods, e.g., pruning and matrix decomposition , network binarization can greatly reduce the memory consumption of the model, and make the model fully compatible with bitwise operations to get good acceleration. Although much progress has been made on network binarization, the existing quantization methods still cause a significant drop of accuracy compared with the full-precision models, since great information loss still exists in the training of binary neural networks. Therefore, to retain the information and ensure a correct information flow during the forward and backward propagation of binarized training, IR-Net is designed.
Preliminaries
The main operation in deep neural networks is expressed as:
The goal of network binarization is to represent the floating-point weights and/or activations with 1-bit. In general, the quantization can be formulated as:
where denotes floating-point parameters including floating-point weights and activations , and denotes binary values including binary weights and activations . denotes scalars for binary values including for weights and for activations. And we usually use function to get :
With the quantized weights and activations, the vector multiplications in the forward propagation can be reformulated as
where denotes the inner product for vectors with bitwise operations XNOR and Bitcount.
In the backward propagation, the derivative of the function is zero almost everywhere, which makes it incompatible with backward propagation, since exact gradients for the original values before the discretization (pre-activations or weights) would be zeroed. So "Straight-Through Estimator (STE) " is generally used to train binary models, which propagates the gradient through or function.
Information Retention Network
In the paper, we point out that the bottleneck of training highly accurate binary neural networks mainly lies in the severe information loss of the training process. Information loss caused by the forward function and the backward approximation for gradient greatly harms the accuracy of binary neural networks. In this paper, we propose a novel model, Information Retention Network (IR-Net), which retains the information in the training process and acquires highly accurate binarized models.
In the forward propagation, the quantization operation brings information loss. Many quantized convolutional neural networks, including binarized models , find the optimal quantizer by minimizing the quantization error :
where indicates the full-precision parameters, denotes the quantized parameters and denotes the quantization error between full-precision and binary parameters. The objective function (Eq. (5)) assumes that quantized models should completely follow the pattern of full-precision models. However, this is not always true especially when extremely low bit-width is applied. For binary models, the representation ability of their parameters is limited to two values, which makes the information carried by neurons easy to lose. The solution space of binary neural networks also quite differs from that of full-precision neural networks. Therefore, without retaining the information through networks, it is insufficient and difficult to promise a good binarized network only by minimizing the quantization error.
To retain the information and minimize the information loss in forward propagation, we propose Libra Parameter Binarization (Libra-PB) that jointly considers both quantization error and information loss. For a random variable obeying Bernoulli distribution, whose probability mass function is
where is the probability of taking the value +1, , and each element in can be viewed as a sample of . The entropy of in Eq. (2) can be calculated by:
If we only pursue the goal of minimizing quantization error, the information entropy of the quantized parameters can be close to zero in extreme case. Therefore, Libra-PB combines the quantization error and information entropy of quantized values as objective function, which is defined as
Under the Bernoulli distribution assumption, when , the information entropy of the quantized values takes the maximum value. This means the quantized values should be evenly distributed. Therefore, we balance weights with zero-mean attribute by subtracting the mean of full-precision weights. Moreover, to make the training more stable without negative effect deriving from the weight magnitude and thus the gradient, we further normalize the balanced weight. The standardized balanced weights are obtained through standardization and balance operations as follows:
where the denotes the standard deviation. has two characteristics: (1) , which maximizes the obtained binary weights’ information entropy. (2) , which makes the full-precision weights involved in binarization more dispersed. Therefore, compared with the direct use of the balanced progress, the use of standardized balanced progress makes the weights steadily updated, and makes the binary weights more stable during the training.
Because of using Libra-PB for weights in each layer, we have , the mean of output is zero. Therefore, the information entropy of activations in each layer can be maximized, which means that the information in activations can be retained.
To further minimize the quantization error and avoid extra expensive floating-point calculation in previous binarization methods, Libra-PB introduces an integer bit-shift scalar to expand the representation ability of binary weights. The optimal bit-shift scalar can be solved by:
where stands for left or right bit-shift. is calculated by , thus can be solved as:
where and denote the dimension and L1-norm of the vector, respectively.
Therefore, our Libra Parameter Binarization for the forward propagation can be presented as below:
The main operations in IR-Net can be expressed as:
As shown in Fig 2, the parameters quantized by Libra-PB have the maximum information entropy under the Bernoulli distribution. We call our binarization method "Libra Parameter Binarization" because the parameters are balanced before the binarization to retain information.
Note that Libra-PB serves an implicit rectifier that reshapes the data distribution before binarization. In the literature, a few studies also realized this positive effect on the performance of BNNs and adopted empirical settings to redistribute parameters . For example, proposed the specific degeneration problem of binarization and solved it using a specially designed additional regularization loss. Different from these works, we first straightforwardly present the information view to rethink the impact of parameter distribution before binarization, and promise the optimal solution by maximizing the information entropy. Moreover, in this framework, Libra-PB can accomplish the distribution adjustment by simply balancing and standardizing the weights before the binarization. This means that our method can be easily and widely applied to various neural network architectures and be directly plugged into the standard training pipeline with a very limited extra computation cost.
2 Error Decay Estimator in Backward Propagation
Limited by the discontinuity of binarization, approximation of gradients is inevitable for the backward propagation. Thus, huge information loss is caused because the influence of quantization cannot be exactly modeled with the approximation. The approximation can be formulated as:
where represents the loss function, denotes the approximation of the function and is the derivative of . There are two common practices for approximation used in previous works:
The function straightly passes the gradient information of output values to input values and completely neglects the effect of binarization. As shown in the shaded area of Fig 3(a), the gradient error is huge and will accumulate during the backward propagation. It is highly required to retain correct gradient information to avoid unstable training rather than ignore the noise caused by for utilizing the Stochastic Gradient Descent Algorithm.
The function takes the clipping attribute of binarization into account to reduce the gradient error. But it can only pass the gradient information inside the clipping interval. It can be seen in Fig 3(b) that for parameters outside [-1, +1], the gradient is clamped to 0. This means once the value jumps outside the clamping interval, it cannot be updated anymore. This characteristic greatly harms the updating ability of backward propagation, which can be proved by the fact that ReLU is a superior activation function compared with Tanh. Thus the approximation increases the difficulty for optimization and decreases the accuracy in practice. It is crucial to ensure enough updating possibility, especially during the beginning of the training process.
function loses the gradient information of quantization while function loses the gradient information outside the clipping interval. There is a contradiction between these two kinds of gradient information loss. To make a balance and acquire the optimal approximation of backward gradient, we design Error Decay Estimator:
where is the backward approximation substitute for the forward function, and are control variables varying during the training process:
where is the current epoch and is the number of epochs, and .
To retain the information deriving from loss function in the backward propagation, EDE introduces a progressive two-stage approach to approximate gradients.
Stage 1: Retain the updating ability of the backward propagation algorithm. We keep the gradient estimation function’s derivative value close to one, and then progressively reduce the clipping value from a large number to one. With this rule, our estimation function evolves from to approximation, which ensures the update ability at the early stage of training.
Stage 2: Retain the accurate gradients for parameters around zero. We keep the clipping value as one and gradually push the derivative curse to the shape of the staircase function. With this rule, our estimation function evolves from approximation to function, which ensures the consistency of forward and backward propagation.
The shape change of EDE for each stage is shown in Fig 3(c). Our EDE updates all parameters in the first stage, and further makes the parameters more accurate in the second stage. Based on the two-stage estimation, EDE reduces the gap between the forward binarization function and the backward approximation function and meanwhile all parameters can be reasonably updated.
3 Analysis and Discussions
The training process of our IR-Net is summarized in Algorithm 1. In this section, we will analyze IR-Net from different aspects.
Since Libra-PB is applied on weights, there is extra operation for binarizing activations in IR-Net. And in Libra-PB, with the novel bit-shift scales, the computation costs are reduced compared with the existing solutions with floating-point scalars (e.g., XNOR-Net, and LQ-Net), as shown in Table 1. Later, we further test the real speed of deployment on hardware and the results can be seen in the Deployment Efficiency Section.
3.2 Stabilize Training
In Libra Parameter Binarization, weight standardization is introduced for reducing the gap between full-precision weights and the binarized ones, avoiding the noise caused by binarization. Fig 4 shows the data distribution of weights without standardization, obviously more concentrated around 0. This phenomenon means the signs of most weights are easy to change during the process of optimization, which directly causes unstable training of binary neural networks. By redistributing the data, weight standardization implicitly sets up a bridge between the forward Libra-PB and backward EDE, contributing to a more stable training of binary neural networks.
Experiments
In this section, we conduct experiments on two benchmark datasets: CIFAR-10 and ImageNet (ILSVRC12) to verify the effectiveness of the proposed IR-Net and compare it with other state-of-the-art (SOTA) methods.
IR-Net: We implement our IR-Net using PyTorch because of its high flexibility and powerful automatic differentiation mechanism. When constructing a binarized model, we simply replace the convolutional layers in the origin models with the binary convolutional layer binarized by our method.
Network Structures: We employ the widely-used network structures including VGG-Small , ResNet-18, ResNet-20 for CIFAR-10, and ResNet-18, ResNet-34 for ImageNet. To prove the versatility of our IR-Net, we evaluate it on both the normal structure and the Bi-Real structure of ResNet. All convolutional and fully-connected layers except the first and last one are binarized, and we select as our activation function instead of ReLU when we binarize the activation.
Initialization: Our IR-Net is trained from scratch (random initialization) without leveraging any pre-trained model. To evaluate our IR-Net on various network architectures, we mostly follow the hyper-parameter settings of their original papers . Among the experiments, we apply SGD as our optimization algorithm.
In this part, we investigate the behaviors and effects of the proposed Libra-PB and EDE techniques on BNN performance.
Our Libra-PB can maximize the information entropy of binary weights and binary activations in IR-Net by adjusting the distribution of weights in the network. Since an explicit balance operation is used before the binarization, the binary weights of each layer in the network have maximum information entropy. Binary activations affected by binary weights in IR-Nets also have maximum information entropy. To demonstrate the information retention of Libra-PB in IR-Net, in Fig 5, we show the information loss of each layer’s binary activations in the networks quantized by vanilla binarization and Libra-PB. Vanilla binarization suffers a large reduction in the information entropy of the binary activations. In the network quantized by Libra-PB, the activation of each layer is close to the maximum information entropy under the Bernoulli distribution. And in the forward propagation, the information loss of binary activations caused by vanilla binarization accumulates layer by layer. Fortunately, the results in Fig 5 show that Libra-PB can retain the information in the binary activation of each layer.
1.2 Effect of EDE
To demonstrate the necessity and effect of our well-designed EDE, we show the data distribution of weights in different stages of training, as shown in Fig 6. The figures in the first row show the distribution and the figures in the second row show the corresponding derivative curve. Among the derivative curves, the blue one represents EDE and the yellow one represents the common STE (with clipping). It can be seen that during the first stage of EDE (epoch 10 to epoch 200 in Fig 6), there is much data outside the range [-1, +1], thus there should not be much clipping which will be harmful to update ability. Besides, the peakedness of weight distribution is high and much data gather around zero at the beginning of training. EDE keeps the derivative similar to function at this stage to ensure the derivative around zero not too large, and thus avoids severely unstable training. Fortunately, with the binarization introduced into training, the weights will gradually approach -1/+1 in the later stages of training. Thus we can slowly increase the value of derivative and approximate a standard function to reduce the gradient mismatch. The visualized results prove that our EDE approximation for backward propagation agrees with the real data distribution, which is the key to improving accuracy.
1.3 Ablation Performance
We further investigate the performance using different parts of IR-Net with the ResNet-20 model on CIFAR-10, which helps understand how our IR-Net works in practice. Table 2 shows the performance in different settings. From the table, we can see that using Libra-PB or EDE alone can improve the accuracy, and the weight standardization in Libra-PB also plays an important role. Moreover, the improvements brought by these parts together can be superimposed, that is why our method can train highly accurate binarized models.
2 Comparison with SOTA methods
We further comprehensively evaluate IR-Net by comparing it with the existing SOTA methods.
Table 3 lists the performance using different methods on CIFAR-10, including RAD over ResNet-18 (based on ), DoReFa-Net , LQ-Net , DSQ over ResNet-20 (based on ) and BNN , LAB , RAD, XNOR-Net over VGG-Small. In all cases, our IR-Net obtains the best performance. More importantly, our method gets a significant improvement over the SOTA methods when using 1-bit weights and 1-bit activations (1W/1A), whether we use the original ResNet structure or the Bi-Real structure. For example, in the 1W/1A bit-width setting, compared with SOTA on ResNet-20, the absolute accuracy increase is as high as 2.4%, and the gap to its full-precision (FP) counterpart is reduced to 4.3%.
2.2 ImageNet
For the large-scale ImageNet dataset, we study the performance of IR-Net over ResNet-18 and ResNet-34. Table 4 shows a number of SOTA quantization methods over ResNet-18 and ResNet-34, including BWN , HWGQ , TWN , LQ-Net , DoReFa-Net , ABC-Net , Bi-Real , XNOR++ , BWHN , SQ-BWN and SQ-TWN . We can observe that when only quantizing weights over ResNet-18, IR-Net using 1-bit outperforms most other methods by large margins, and even surpasses TWN using 2-bit weights. And in the 1W/1A setting, the Top-1 accuracy of IR-Net is also significantly better than that of the SOTA methods (e.g., 58.1% vs. 56.4% for ResNet-18). The experimental results prove that our IR-Net is more competitive than the existed methods.
3 Deployment Efficiency
To further validate the efficiency of IR-Net when deployed into the real-world mobile devices, we further implement our IR-Net on Raspberry Pi 3B, which has a 1.2 GHz 64-bit quad-core ARM Cortex-A53, and test its real speed in practice. We utilize the SIMD instruction SSHL on ARM NEON to make inference framework daBNN compatible with our IR-Net. We have to point out that till now there have been very few studies that reported their inference speed in real-world devices, especially when using 1-bit binarization. In Table 5, we compare our IR-Net with the existing high-performance inference implementation including NCNN and DSQ . From the table, we can easily find that the inference speed of IR-Net is much faster, the model size of IR-Net can be greatly reduced and the bit-shift scales in IR-Net bring almost no extra inference time and storage consumption.
Conclusion
In this paper, we propose IR-Net to retain the information propagated in binary neural networks, mainly consisting of two novel techniques: Libra-PB for keeping diversity in forward propagation and EDE for reducing the gradient error in backward propagation. Libra-PB conducts a simple yet effective transformation on weights from the view of information entropy, which simultaneously reduces the information loss of both weights and activations, without additional operation on activations. Thus, the diversity of binary neural networks can be kept as much as possible and meanwhile the efficiency will not be harmed. Besides, the well-designed gradient estimator EDE retains the gradient information during backward propagation. Owing to the sufficient updating ability and accurate gradients, the performance with EDE surpasses that with STE by a large margin. Extensive experiments prove that the IR-Net consistently outperforms the existed state-of-the-art binary neural networks.
Acknowledgement This work was supported by National Natural Science Foundation of China (61872021, 61690202), and Beijing Nova Program of Science and Technology (Z191100001119050).