RBCN: Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs

Chunlei Liu, Wenrui Ding, Xin Xia, Yuan Hu, Baochang Zhang, Jianzhuang Liu, Bohan Zhuang, Guodong Guo

Introduction

Deep convolutional neural networks (DCNNs) have been successfully demonstrated on many computer vision tasks such as object detection and image classification. DCNNs deployed in practical environments, however, still face many challenges. They usually involve millions of parameters and billions of FLOPs during computation. This is critical because models of vision applications may consume very large amounts of memory and computation, making them impractical for most embedded platforms.

Binary filters instead of using full-precision filter weights have been investigated in DCNNs to compress the deep models to handle the aforementioned problems. Many works attempt to quantize the weights of a network while keeping the activations (feature maps) to 32-bit floating points Zhou et al. (2017); Zhu et al. (2016); Wang et al. (2018). Although this scheme leads to less performance decrease compared to its full-precision counterpart, it still needs a substantial amount of computational resource to handle the full-precision activations. Therefore, the so-called 1-bit DCNNs, which target the problem of training the networks with both 1-bit quantized weights and 1-bit activations, become more promising and significant in the field of DCNNs compression. As presented in Rastegari et al. (2016), by reconstructing the full-precision filters with a single scaling factor, XNOR provides an efficient implementation of convolutional operations. More recently, Bi-Real Net Liu et al. (2018) explores a new variant of residual structure to preserve the real activations before the sign function. And the researchers in Hou et al. (2016) propose a new value approximation method that considers the effect of binarization on the loss to further obtain binarized weights. PCNN Gu et al. (2019) learns a set of diverse quantized kernels by exploiting multiple projections with discrete back propagation.

The investigation into prior arts reveals that how to use the full-precision models is the key issue to obtain the optimized BCNNs. Most existing methods use the full-precision models as an initialization Rastegari et al. (2016) Liu et al. (2018), or for kernel approximation Gu et al. (2019) Rastegari et al. (2016). Besides, knowledge distillation uses a teacher model (e.g., a full-precision model) to provide a guidance to quantize the network Polino et al. (2018); Zhuang et al. (2018); Mishra and Marr (2017). While these methods generally use a regularization term to minimize the difference between the student’s and teacher’s posterior probabilities or intermediate feature representations, they fail to consider the full-precision feature maps (activations) in a comprehensive way. This might be the reason why the knowledge distillation methods have not been employed to obtain the extreme 1-bit CNNs yet. To narrow down the performance gap between a BCNN and its full-precision model, we propose that the full-precision kernels and feature maps should be considered in a more comprehensive way, in order to fully exploit the multi-cue information.

In this paper, we introduce a rectified binary convolutional network (RBCN) to calculate an optimized BCNN in which a novel learning architecture is introduced to combine the full-precision feature maps and the kernels approximation in an end-to-end manner. Based on the powerful probability fitting ability of generative adversarial network (GAN), we discover that training a BCNN network with GAN, a better performance can be obtained by fitting the distribution of feature maps between full-precision and 1-bit binary networks. By doing so, GAN is introduced to distill RBCN from full-precision network by exploiting their full-precision feature maps. To the best of our knowledge, we are the first to use a GAN to do binary approximation of the full-precision model. The whole process is illustrated in Fig. 1, where the full-precision model and the 1-bit binary model (generator) respectively provide “real” and “fake” feature maps to the discriminators. The discriminators try to distinguish the “real” from the “fake”, and the generator tries to make the discriminators unable to work well. By repeating this process, the multi-cue information (full-precision kernels and feature maps) is sufficiently employed in the training process to enhance the representational ability of the 1-bit binary model. Besides, kernel (filter) approximation (RBConv in Fig. 1) is integrated in the framework. Also, multiple discriminators are used to further improve the performance of RBCN. This process involving the GAN and the kernel approximation is a rectified process, which can lead to a unique architecture with more precise estimation of the full-precision model. The contributions of this paper are summarized as follows.

(1) A novel BCNN learning architecture, referred to as rectified binary convolutional network (RBCN), is proposed, which employs the full-precision kernels and feature maps to rectify the binarization process in a comprehensive framework.

(2) To the best of our knowledge, we are the first to use a GAN to calculate a BCNN. Besides, we discover that using multiple discriminators in the GAN can significantly improve the performance of the 1-bit binary model.

(3) Extensive experiments demonstrate the superior performance of the proposed RBCNs over state-of-the-art BCNNs on the object classification and tracking tasks.

Rectified Binary Convolutional Networks (RBCNs)

We design RBCNs via kernel approximation and training with GANs to rectify BCNNs in a unified framework. During this process, the multi-cue information of the full-precision feature maps and kernelsIn this paper, the terms “filter” and “kernel” are exchangeable. is exploited to improve the performance degraded by binarization. The rectified convolutional layers are generic and flexible, which can be easily incorporated into existing CNNs, such as WideResNets and ResNets. First of all, Table 1 gives the main notation used in this paper.

The rectified process combines the full-precision kernels and feature maps to rectify the binarization process. It includes kernel approximation and adversarial learning. This learnable kernel approximation can lead to an unique architecture with more precise estimation of the original convolutional filters through minimizing a kernel loss. The discriminators D(⋅)D(\cdot) with filters YY are introduced to distinguish the feature maps RR of the full-precision model from those TT of RBCN. The generator (RBCN) with filters WW and learnable matrixs CC is learned together with YY by using the knowledge from the supervised feature maps RR. Therefore, WW, CC and YY are learned by solving the following optimization problem:

where LAdv(W,W^,C,Y)\mathscr{L}_{Adv}(W,\hat{W},C,Y) is the adversarial loss:

where D(⋅)D(\cdot) consists of four basic blocks, each of which has a linear layer and a LeakyRelu layer.

In addition, LKernel(W,W^,C)\mathscr{L}_{Kernel}(W,\hat{W},C) is the kernel loss between the learned full-precision filters WW and the binarized filters W^\hat{W}, which is expressed by MSE:

Finally, LS(W,W^,C)\mathscr{L}_{S}(W,\hat{W},C) is a traditional problem-dependent loss such as the softmax loss.

For simplicity, the update of the discriminators is omitted in the following description until Algorithm 1. Besides, we find that the loglog in Equ. 2 has little effect during training and so it is omitted too. Then, based on the Lagrangian method, the optimization problem in Equ. 1 is rewritten as:

In Equ. 4, the target is to obtain WW, W^\hat{W} and CC with YY fixed, which is why the term D(R;Y)D(R;Y) in Equ. 2 can be ignored. The update of YY can be found in Algorithm 1. The advantage of our formulation in Equ. 4 lies in that the loss function is trainable, meaning that it can be easily incorporated into existing learning frameworks.

2 Forward Propagation in RBCNs

In RBCNs, a binary filter W^il\hat{W}_{i}^{l} is calculated as:

where WilW_{i}^{l} is the corresponding full-precision filter, and the values of W^il\hat{W}_{i}^{l} are 11 or −1-1. Both WilW_{i}^{l} and W^il\hat{W}_{i}^{l} are jointly obtained in the end-to-end learning.

In RBCNs, the convolution is implemented based on ClC^{l} and FinlF_{in}^{l} to calculate the feature maps FoutlF_{out}^{l}:

where RBConvRBConv denotes the convolution operation implemented as a new module, FinlF_{in}^{l} and FoutlF_{out}^{l} are the feature maps before and after the convolution, respectively, and ⊙\odot is the element-by-element product. Note that FinlF_{in}^{l} is binary after the sign operation (see Fig. 1), and CC is actually C∗C^{*}, which will be elaborated at the end of section 3.3.

3 Backward Propagation in RBCNs

In RBCNs, what need to be learned and updated are the full-precision filters WW and the learnable matrixs CC. These two sets of parameters are jointly learned. In each convolutional layer, an RBCN updates WW first and then CC.

Let δWil{{\delta}_{W_{i}^{l}}} be the gradient of the full-precision filter WilW_{i}^{l}. During backpropagation, the gradients pass to W^il\hat{W}_{i}^{l} first and then to WilW_{i}^{l}. Thus:

which is an approximation of the 2×2\timesdirac-delta function Liu et al. (2018). Furthermore,

where η1\eta_{1} is a learning rate. Then:

3.2 Updating C𝐶C

We further update the learnable matrix ClC^{l} with WlW^{l} fixed. Let δCl{{\delta}_{C^{l}}} be the gradient of ClC^{l}. Then we have:

where η2\eta_{2} is another learning rate. Further,

The above derivations show that the rectified process is trainable in an end-to-end manner. The complete training process is summarized in Algorithm 1, including the update of the discriminators. Besides, in the implementation, the batch normalization (BN) layers are updated with WW and CC fixed after each epoch.

We note that in our implementation, the value of CC will be replaced by its average during the forward process, resulting into a new matrix denoted by C∗C^{*}its elements are equal. By doing so, only a scalar instead of a matrix involve into the convolution which thus speed up the calculation.

Experiments

Our RBCNs are evaluated first on object classification using MNIST Lecun et al. (1998), CIFAR10/100 Krizhevsky and Hinton (2009) and ILSVRC12 ImageNet datasets Russakovsky et al. (2015), and then on object tracking. For object classification, WideResNet (WRN) Zagoruyko and Komodakis (2016) and ResNet He et al. (2016) are employed as the backbone networks to build our RBCNs. Also, binarizing the neuron activations is carried out in all of our experiments.

Datasets: The MINIST Lecun et al. (1998) dataset is composed of a training set of 60,000 and a testing set of 10,000 32×3232\times 32 grayscale images of hand-written digits from 0 to 9.

CIFAR10 Krizhevsky and Hinton (2009) is a natural image classification dataset containing a training set of 50,00050,000 and a testing set of 10,00010,000 32×3232\times 32 color images across the following 10 classes: airplanes, automobiles, birds, cats, deers, dogs, frogs, horses, ships, and trucks, while CIFAR100 consists of 100 classes.

ImageNet object classification dataset Russakovsky et al. (2015) is more challenging due to its large scale and greater diversity. There are 1000 classes and 1.2 million training images and 50k validation images in it. We compare our method with the state-of-the-art on the ImageNet dataset, and we adopt ResNet18 to validate the superiority and effectiveness of RBCNs.

WRN Backbone: WRN is a network structure similar to ResNet with a depth factor kk to control the feature map depth dimension expansion through 3 stages, within which the dimensions remain unchanged. For simplicity we fix the depth factor to 1. Each WRN has a parameter ii which indicates the channel dimension of the first stage, and we set it to 16, leading to a network structures 1616-1616-3232-6464. The training details are the same as in Zagoruyko and Komodakis (2016). λ1\lambda_{1} and λ2\lambda_{2} are set as 0.01 with a degradation of 10% for every 60 epochs before reaching the maximum epoch of 200 for CIFAR10/100. For example, WRN22 is a network with 22 convolutional layers and similarly for WRN40.

ResNet18 Backbone: SGD is used as the optimization algorithm with a momentum of 0.90.9 and a weight decay 1e-4. λ1\lambda_{1} and λ2\lambda_{2} are set as 0.1 with a degradation of 10% for every 20 epochs before reaching the maximum epoch of 70 on ImageNet, while on CIFAR10/100, λ1\lambda_{1} and λ2\lambda_{2} are set as 0.01 with a degradation of 10% for every 60 epochs before reaching the maximum epoch of 200.

2 Ablation Study

In this section, we study the performance contributions of the components in RBCNs, which include kernel approximation, GAN, and the update of the BN layers. CIFAR100 and ResNet18 with different kernel stages are used in this experiment. The details are given below.

1) We only replace the convolution in Bi-Real Net with our kernel approximation (RBConvRBConv) and compare the results. As shown in the R column in Table LABEL:ablation, RBCN achieves 1.62% accuracy improvement over Bi-Real Net (56.54% vs. 54.92%) using the same network structure as in ResNet18 with 32-32-64-128. This significant improvement verifies the effectiveness of the learnable matrixs.

2) In RBCNs, if we use the GAN to help binarization, we can find a more significant improvement from 56.54% to 59.13% with the kernel stage of 32-32-64-128, which shows that the GAN can really enhance the binarized networks.

3) We find that a training trick can also improve RBCNs, which is to update the BN layers with WW and CC fixed after each epoch (line 17 in Algorithm 1). This trick makes RBCN boost 2.51% (61.64% vs. 59.13%) in CIFAR100 with 32-32-64-128.

3 Accuracy Comparison with State-of-the-Art

CIFAR10/100: The same parameter settings are used in RBCNs on both CIFAR10 and CIFAR100. We first compare our RBCNs with the original ResNet18 with different stage kernels, followed by a comparison with the original WRNs with the initial channel dimension 6464 in Table 3. Thanks to the rectified process, our results on both the datasets are close to the full-precision networks ResNe18 and WRN40. Then, we compare our results with other state-of-the-arts such as Bi-Real Net Liu et al. (2018), PCNN Gu et al. (2019), Scheme-A Mishra and Marr (2017) and XNOR Rastegari et al. (2016). All these BCNNs have both binary filters and binary activations. It is observed that at most 6.17% (== 61.09%−-54.92%) accuracy improvement is gained with our RBCN, and in other cases, larger margins are achieved.

ImageNet: Five state-of-the-art methods on ImageNet are chosen for comparison: Bi-Real Net Liu et al. (2018), BinaryNet Courbariaux et al. (2016), XNOR Rastegari et al. (2016), PCNN Gu et al. (2019) and ABC-Net Lin et al. (2017). Again, these networks are representative methods of binarizing both network weights and activations and achieve state-of-the-art results. All the methods in Table 4 perform the binarization of ResNet18. The results in Table 4 are quoted directly from their papers, except that the result of BinaryNet is from Lin et al. (2017). The comparison clearly indicates that the proposed RBCN outperforms the five binary networks by a considerable margin in terms of both the top-1 and top-5 accuracies. Specifically, for top-1 accuracy, RBCN outperforms BinaryNet and ABC-Net with a gap over 16%, achieves 7.9% improvement over XNOR, 3.1% over the very recent Bi-Real Net, and 2.2% over the latest PCNN. In Fig. 2, we plot the training and testing loss curves of XNOR and RBCN. It clearly shows that using our rectified process, RBCN converges faster than XNOR.

4 Experiments on object tracking

The key message conveyed in the proposed method is that although the conventional binary training method has a limited model capability, the proposed rectified process can lead to a powerful model. In this section, we show that this framework can also be used in object tracking. In particular, we consider the problem of tracking an arbitrary object in videos, where the object is identified solely by a rectangle in the first frame. For object tracking, it is necessary to update the weights of the network online, severely compromising the speed of the system. To directly apply the proposed framework to this application, we can construct a binary convolution with the same structure to reduce the convolution time. In this way, RBCN can be used to binarize the network further to guarantee the tracking performance.

In this paper, we use SiamFC Network as the backbone for object tracking. We binarize SiamFC as Rectified Binary Convolutional SiamFC Network (RB-SF). We evaluate RB-SF on four datasets, GOT-10K Huang et al. (2018), OTB50 Wu et al. (2013), OTB100 Wu et al. (2015), and UAV123 Mueller et al. (2016), using accuracy occupy (AO) and success rate (SR). The results are shown in Table LABEL:tracking. Intriguingly, our framework achieves about 7% AO improvement over XNOR, both using the same network architecture as in SiamFC Network on GOT-10k. Further, our framework brings so much benefit that Bi-SF performs almost as well as the full-precision SiamFC Network.

5 Efficiency Analysis

The memory usage is computed as the summation of 32 bits times the number of real-valued parameters and 1 bit times the number of binary parameters in the network. Further, we use FLOPs to measure the speed. The results are given in Table 6. The FLOPs are calculated as the amount of real-valued floating point multiplications plus 1/64 of the amount of 1-bit multiplications Liu et al. (2018). As shown in Table 6, the proposed RBCN, along with XNOR, reduces the memory usage of the full-precision ResNet18 by 11.10 times. For efficiency, both RBCN and XNOR gain 10.86×10.86\times speedup over ResNet18. Note the computational and storage costs brought by learnable scalar C∗C^{*} can be negligible.

Conclusion

In this paper, we introduce rectified binary convolutional networks (RBCNs), towards optimized BCNNs, by exploiting the full-precision kernels and feature maps in an end-to-end manner. In particular, we use a GAN to train the 1-bit binary network with the guidance of its corresponding full-precision model, which significantly improves the performance of the BCNN. Furthermore, as a general model, RBCNs can be used not only in object classification but also in other tasks such as object tracking. The experiments on both object classification and object tracking demonstrate the superior performance of the proposed RBCNs over state-of-the-art binary models.

Acknowledgment

The work was supported by the National Key Research and Development Program of China (Grant No. 2016YFB0502602) and the Natural Science Foundation of China under Contract 61672079. Also, it is in part supported by the Fundamental Research Funds for the Central Universities. Baochang Zhang is the corresponding author.

References