IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks

Ke Sun, Mingjie Li, Dong Liu, Jingdong Wang

Introduction

There have been imperative demands for portable and efficient deep convolutional neural networks with high accuracies in vision applications. The recent efforts include two paths, network compression: approximate pre-trained models by pruning superfluous connections and channels or decomposing convolutional matrices; and lightweight network design: design less-redundant kernels for constructing networks trained from scratch.

We are interested in designing lightweight networks using less-redundant kernels to form a dense convolutional kernel. There are two main representative schemes, either of which has independently shown the success in building small networks. One scheme is to use a sequence of low-rank convolutional kernels to compose a linear kernel or nonlinear kernel with intermediate nonlinear activations included in the sequence of low-rank kernels, e.g., bottleneck [He et al.(2016c)He, Zhang, Ren, and Sun] and inverted residual block [Sandler et al.(2018)Sandler, Howard, Zhu, Zhmoginov, and Chen]. The other scheme is to use a sequence of sparse (and possible dense) kernels to compose a kernel, e.g., interleaved group convolutions, MobileNetV1 and Xception, and deep roots [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang, Howard et al.(2017)Howard, Zhu, Chen, Kalenichenko, Wang, Weyand, Andreetto, and Adam, Chollet(2016), Ioannou et al.(2016)Ioannou, Robertson, Cipolla, and Criminisi].

In this paper, we design a modularized convolutional block, which simultaneously explores both low-rank and sparse kernels to compose a dense convolutional kernel. We start from IGCV22 [Xie et al.(2018)Xie, wang, Zhang, Lai, Hong, and Qi], which is composed of a channel-wise spatial convolution, group point-wise convolutions, and intermediate permutation operations. We introduce the low-rank design pattern into group point-wise convolutions, where one group convolution expands the feature dimension, and the other group convolution projects the feature back. We introduce a loose complementary condition over channels for constructing dense composed kernels, which is equivalent to the strict complementary condition, presented in IGCV22 [Xie et al.(2018)Xie, wang, Zhang, Lai, Hong, and Qi], imposed over super-channels (a set of channels), so as to handle the case that the number of input and output channels are different and avoid over-wide convolutions. The resulting network is called IGCV33. We empirically demonstrate that the combination of low-rank and sparse kernels boosts the performance and the superiority of our proposed approach to the state-of-the-arts, IGCV22 and MobileNetV22, over image classification on CIFAR and ImageNet and object detection on COCO.

Related Work

To obtain portable and efficient models, most of works devote great efforts to reduce the redundancy in the convolutional kernels, which mainly lie in two extents: high numerical precision and superfluous weights.

Low-precision kernels. Approaches in this field quantize the weights of CNNs from floating point into lower bit-depth representations. These methodologies reduce the redundancies in the numerical precision, such as binarization [Courbariaux et al.(2016)Courbariaux, Hubara, Soudry, El-Yaniv, and Bengio] constraining the weights to either +1+1 or −1-1, trinarization [Li et al.(2016a)Li, Zhang, and Liu, Zhou et al.(2016)Zhou, Wu, Ni, Zhou, Wen, and Zou, Zhu et al.(2016)Zhu, Han, Mao, and Dally] making weights to be ternary-valued and quantization [Han et al.(2015a)Han, Mao, and Dally, Zhou et al.(2017)Zhou, Yao, Guo, Xu, and Chen] a general form to convert weights into a low-precision version.

Sparse or Low-rank kernels. In this field, the methods adopt one sparse or low-rank kernel as an approximation of the original dense or high-rank kernel. (i) Sparse kernels:Sparse\ kernels: Many regularizations are designed to adding the sparse constraint on the convolutional kernels, such as L1L1 or L2L2 regularization [Han et al.(2015a)Han, Mao, and Dally, Han et al.(2015b)Han, Pool, Tran, and Dally] and structured sparsity regularizer [Li et al.(2016b)Li, Kadav, Durdanovic, Samet, and Graf, Wen et al.(2016)Wen, Wu, Wang, Chen, and Li, Alvarez and Salzmann(2016)]. Group convolution is a pre-defined structured-sparse matrix, which is widely used in the mobile models [Xie et al.(2016)Xie, Girshick, Dollár, Tu, and He, Zhao et al.(2016)Zhao, Wang, Li, and Tu]. (ii) LowLow-rank kernels.rank\ kernels. Pruning redundant weights or channels from pre-trained models [Park et al.(2016)Park, Li, Wen, Tang, Li, Chen, and Dubey, Li et al.(2016b)Li, Kadav, Durdanovic, Samet, and Graf, He et al.(2017)He, Zhang, and Sun, Luo et al.(2017)Luo, Wu, and Lin] is a main branch in network compression. What’s more, many works [Simonyan and Zisserman(2014), Szegedy et al.(2016)Szegedy, Vanhoucke, Ioffe, Shlens, and Wojna, Denton et al.(2014)Denton, Zaremba, Bruna, LeCun, and Fergus] replace a large kernel with small kernels is to reduce the ranks in the spatial domain.

Composition from multiple sparse or low-rank kernels. Design the composition of sparse or low-rank kernels to keep the original dense connections. (i) Multiplying the sparse kernels: IGCV11 [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang] is the product of two structured-sparse matrices and IGCV22 [Xie et al.(2018)Xie, wang, Zhang, Lai, Hong, and Qi] decomposes the dense kernel into more sparse kernels. And the permutation operations [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang] are adopted to keep dense connectivity between input and output channels, which can increase the expressive capacity of networks [Sharir and Shashua(2017)]. Xception [Chollet(2016)] is an extreme case of IGCV11: a pointwise convolution followed by a channel-wise convolution. (ii) Composition from lowComposition\ from\ low-rank kernels.rank\ kernels. In a clear case, a 3×33\times 3 kernel can be decomposed into a 3×13\times 1 convolution followed by a 1×31\times 3 convolution [Ioannou et al.(2015)Ioannou, Robertson, Shotton, Cipolla, and Criminisi, Jaderberg et al.(2014)Jaderberg, Vedaldi, and Zisserman, Mamalet and Garcia(2012)], which approximates the dense kernel along the spatial domain. Bottleneck [He et al.(2016c)He, Zhang, Ren, and Sun, Iandola et al.(2016)Iandola, Han, Moskewicz, Ashraf, Dally, and Keutzer, Sandler et al.(2018)Sandler, Howard, Zhu, Zhmoginov, and Chen], if neglecting the intermediate ReLUs, can be viewed as the low-rank approximation along the output channel domain.

Our Approach

In this section, we firstly review prior works, IGC and MobileNets, and then describe our proposed IGCV33. Finally, we analyze the loose complementary condition and give a discussion on the architecture of IGCV33 blocks.

Interleaved Group Convolution (IGCV11). The IGCV11 block consists of primary and secondary group convolutions, which is mathematically formulated as follows:

where xc=[x1c,x2c,...,xKc]\mathbf{x}^{c}=[x_{1}^{c},x_{2}^{c},...,x_{K}^{c}] is a column vector. Wgi\mathbf{W}_{g}^{i} (i=1i=1 or 22) is the kernel matrix over the corresponding channels in the ggth branch, GiG_{i} is the number of branches in the iith group convolution. In the case suggested in [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang], the primary group convolution is a group 3×33\times 3 convolution with G1=C2G_{1}=\frac{C}{2} groups, and Wg1\mathbf{W}^{1}_{g} is a matrix of size 2×(2K)2\times(2K). The secondary group convolution is a group 1×11\times 1 convolution with G2=2G_{2}=2 groups, where W12\mathbf{W}^{2}_{1} and W22\mathbf{W}^{2}_{2} are both dense matrices of size C2×C2\frac{C}{2}\times\frac{C}{2}.

Interleaved Structured Sparse Convolution (IGCV22). IGCV22 extends IGCV11 by decomposing the convolution matrix into more structured sparse matrices:

Here W1W_{1} corresponds to a channel-wise spatial convolution, and Wl,l∈{1,2,...,L}W_{l},l\in\{1,2,...,L\} corresponds to group point-wise convolutions.

MobileNetV11. A MobileNetV11 block consists of a channel-wise spatial convolution and a point-wise convolution. The mathematical formulation is given as follows:

where W1\mathbf{W^{1}} and W2\mathbf{W^{2}} corresponds to the channel-wise and point-wise convolution respectively. It is an extreme case of IGCV11 [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang]: both channel-wise and point-wise convolutions are extreme group convolutions.

MobileNetV22. A MobileNetV22 block consists of a dense pointwise convolution, a channel-wise spatial convolution, and a dense pointwise convolution. It uses an inverted bottleneck: the first pointwise convolution increases the width and the second one reduces the width.

For convenience, we convert the input x\mathbf{x} to x^\mathbf{\hat{x}}, and accordingly the bottleneck is mathematically formulated:

2 Interleaved Low-Rank Group Convolutions

The proposed Interleaved Low-Rank Group Convolutions, named IGCV33, extend IGCV22 by using low-rank group convolutions to replace group convolutions in IGCV22. It consists of a channel-wise spatial convolution, a low-rank group point-wise convolution with G1G_{1} groups that reduces the width and a low-rank group point-wise convolution with G2G_{2} groups which expands the width back. The mathematical formulation is given as follows,

Here, P1\mathbf{P}^{1} and P2\mathbf{P}^{2} are permutation matrices similar to permutation matrices given in [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang]. W1\mathbf{W}^{1} corresponds to the channel-wise 3×33\times 3 convolution. W^0\mathbf{\hat{W}}^{0} and W2\mathbf{W}^{2} are low-rank structured sparse matrices. The two low-rank sparse matrices are mathematically formulated as follows,

Construct a dense composed kernel. Similar to IGCV22, we also aim to design the two group convolutions such that the composed pointwise convolutional kernel, P2W2P1W1\mathbf{P}^{2}\mathbf{W}^{2}\mathbf{P}^{1}\mathbf{W}^{1}, is dense. Different from IGCV22, the number of input and output channels in the low-rank group convolution are not the same, so the permutation operation cannot work as before.

We divide input, output and intermediate channels into a set of (e.g., CsC_{s}) super-channels, where the super-channel for input and output response maps contains CCs\frac{C}{C_{s}} channels and for intermediate response maps the super-channel contains CintCs\frac{C_{int}}{C_{s}} channels. The sketch is shown in Fig 2, where super-channels are represented as 3-D boxes in different colors. Then, we review the strict complementary condition proposed in [Xie et al.(2018)Xie, wang, Zhang, Lai, Hong, and Qi] and describe the loose complementary condition based on the concept of super-channel.

The two group convolutions are thought complementary if the channels in the input and output response maps and in the intermediate response maps lying in the same branch in one group convolution lie in different branches and come from all the branches in the other group convolution.

Loose complimentary condition. We empirically observed that over-sparse convolutions lead to over-wide response maps and thus too large memory cost, but do not lead to higher accuracy, which is consistent to the empirical results in IGCV11 [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang] and IGCV22 [Xie et al.(2018)Xie, wang, Zhang, Lai, Hong, and Qi]. Therefore, we relax the strict complementary condition to a loose version.

The two group convolutions are thought complementary if the super-channels lying in the same branch in one group convolution lie in different branches and come from all the branches in the other group convolution.

3 Discussions and Analysis

Linear, nonlinear, and inverted IGCV33. It is shown in IGCV11 [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang] that the linear version, i.e., no intermediate ReLU is included in the block, performs better than the very deep networks which include lots of ReLU activations. Similar observations are also obtained in the bottleneck block [He et al.(2016c)He, Zhang, Ren, and Sun] and Xception [Chollet(2016)]. On the other hand, for not very deep networks, e.g., there are only 10+10+ IGCV33 blocks, the non-linear version, i.e., there are one or more intermediate ReLUs in the block, performs better. We provide the ablation study in Sec. 5.

In our experiments, we do not see difference between the normal version and the inverted version of bottleneck. We adopt the inverted IGCV33 block to save the memory footprints during the training and inference process (shown in Fig 1(c)): low-rank group convolution (width increasing), channel-wise spatial convolution, and low-rank group convolution (width reduction) and accordingly the identity connection is built between two low-dimensional representations, which is similar to inverted bottleneck presented in MobileNetV22 [Sandler et al.(2018)Sandler, Howard, Zhu, Zhmoginov, and Chen].

Wider and deeper IGCV33 networks. We stack our IGCV33 blocks to form a deep network and provide two versions for fair comparisons with MobileNetV22. One is to let the depth (the number of blocks) be the same and to widen the network. In our implementation, Cs=4C_{s}=4 super-channels is used to form group convolutions (G1=4G_{1}=4 and G2=4G_{2}=4). The other one is to let the width (of the corresponding block) be the same and to deepen the network. In our implementation, Cs=2C_{s}=2 super-channels is used (G1=2G_{1}=2 and G2=2G_{2}=2). The empirical study is given in Sec. 5.

Experiments

CIFAR. The CIFAR datasets, CIFAR-1010 [Krizhevsky(2009)] and CIFAR-100100 [Torralba et al.(2008)Torralba, Fergus, and Freeman], are subsets of the 8080 million tiny images. Both datasets contain 6000060000 32×3232\times 32 color images with 5000050000 images for training and 1000010000 images for test. The standard data augmentation scheme we adopt is widely used for these datasets [He et al.(2016a)He, Zhang, Ren, and Sun, Lee et al.(2015)Lee, Xie, Gallagher, Zhang, and Tu, Huang et al.(2016)Huang, Liu, and Weinberger, Larsson et al.(2016)Larsson, Maire, and Shakhnarovich, Lin et al.(2013)Lin, Chen, and Yan, Romero et al.(2014)Romero, Ballas, Kahou, Chassang, Gatta, and Bengio, Springenberg et al.(2014)Springenberg, Dosovitskiy, Brox, and Riedmiller, Srivastava et al.(2015)Srivastava, Greff, and Schmidhuber]: we zero-pad the images with 44 pixels on each side, and then randomly crop them to produce 32×3232\times 32 images, followed by horizontally mirroring half of the images and normalizing them by using the channel means.

For training, we use the SGD algorithm and train all networks from scratch. We initialize the weights similar to [He et al.(2016a)He, Zhang, Ren, and Sun, He et al.(2016b)He, Zhang, Ren, and Sun], and set the weight decay as 0.0001 and the momentum as 0.9. The initial learning rate is 0.1 and is reduced by a factor 10 at the 200, 300 and 350 training epochs. Each training batch consists of 64 images on four asynchronous GPUs.

ImageNet. The ILSVRC 2012 classification dataset [Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei] contains over 1.2 million training images and 50,000 validation images, and each image is labeled from 1000 categories. We adopt the same data augmentation scheme as in [He et al.(2016a)He, Zhang, Ren, and Sun, He et al.(2016b)He, Zhang, Ren, and Sun]. We use SGD to train the networks with the same hyperparameters (batch size =96=96, weight decay =0.00004=0.00004 and momentum =0.9=0.9) on 4 GPUs. We train the models for 480480 epochs with extra 5050 epochs for retraining. We start from a learning rate of 0.0450.045, and then scale it by 0.980.98 every epochs.

MSCOCO. We evaluate and compare the performance of IGCV3 and MobileNetV2 for object detection on COCO dataset [Lin et al.(2014)Lin, Maire, Belongie, Hays, Perona, Ramanan, Dollár, and Zitnick], which are used as a backbone for SSDLite and SSDLite2 respectively. As in previous work [Sandler et al.(2018)Sandler, Howard, Zhu, Zhmoginov, and Chen], we train using the union of 80k training images and a 35k subset of val images (trainval35k) and evaluate on test-dev. We train the models for 240240 epochs on MXNet. We start from a learning rate of 0.0040.004, and then divide it by 10 every 60 epochs. The input resolution is shown in Table 4.

2 Comparisons with IGCV111 and IGCV222.

We compare IGCV33 with IGCV11 and IGCV22 on CIFAR and ImageNet. The models of IGCV11 and IGCV22 replace the blocks in MobileNetV11 with their proposed blocks. Please check the details in [Zhang et al.(2017a)Zhang, Qi, Xiao, and Wang, Xie et al.(2018)Xie, wang, Zhang, Lai, Hong, and Qi]. IGCV33 adopts two group convolutions with G1=2G_{1}=2 and G2=2G_{2}=2 respectively for the deeper version (IGCV33-D).

The comparisons of IGC blocks are shown in Table 1. IGCV33 outperforms the prior works slightly on CIFAR datasets, and achieves significant improvement about 1.5% on ImageNet. By introducing low-rank design into the sparse kernels, IGCV33 reduces the redundancies in feature dimensions and increases the depth of network to improve the generalization and the overall performance.

3 Comparisons with Other Mobile Networks.

We compare IGCV33 with MobileNetV22 and other mobile networks for classification and detection on several benchmarks. IGCV33 adopts two group convolutions with G1=2G_{1}=2 and G2=2G_{2}=2 for the deeper version (IGCV33-D) and with G1=4G_{1}=4 and G2=4G_{2}=4 for the wider version (IGCV33-W). Due to the limitation, we reproduce MobileNetV22 on 4 GPUs not proposed 16 GPUs. We try our best to train MobileNetV22 and provide the reproduced results as (our impl.). We adopt the settings of the best reproduced MobileNetV22 to train our models.

CIFAR classification. The classification accuracy comparisons of MobileNetV22 and IGCV33 are presented in Table 2. Introducing sparse design into low-rank kernels, IGC V3 further reduces the superfluous weights in convolutional kernels, and also makes the best use of parameters to improve the performance. IGCV33 outperforms MobileNetV22 a lot with the similar number of parameters. Moreover, IGCV33 with 50% parameters still achieves a better performance, which has the same depth as MobileNetV2. The reason may be that the number of ReLU is half of MobileNet V2.

ImageNet classification. We compare IGCV33 with other mobile models on ImageNet, shown in Table 3. With the similar computation complexity, IGCV33 always outperforms other networks. IGCV33-D (1.4) achieves a superior performance under the same training settings (74.55% vs 73.76%), though marginally worse than the original MobileNetV22 (1.4). Due to the difference of the final FC-layer, IGCV33-D has the different number of parameters in Table 3 from Table 1 and 2 (1.31.3M (1000 classes), 1313k (100 classes) and 1.31.3k (10 classes)).

Detection on COCO. We extend IGCV33 to be a backbone for detection networks. SSDLite is proposed in [Sandler et al.(2018)Sandler, Howard, Zhu, Zhmoginov, and Chen], we follow the original framework, but replace all the feature extraction blocks with IGCV3, denoted by "SSDLite2". The comparisons of different methods for detection are shown in Table 4. We adopt IGCV33-W as our backbone for extracting the features from the layers at the same depth as MobileNetV2. IGCV3 is slightly better than MobileNetV22 with fewer parameters, and outperforms YOLOV22 0.6% mAP with much fewer number of parameters.

Ablation Study

We explore three main design choices: (i) deeper and wider networks; (ii) intermediate ReLUs; (iii) #branches in group convolutions. These three elements determine the architecture of the IGCV3 block and the whole network.

Deeper and wider networks. Most of related works enlarge the width of networks, which output high-dimensional features capturing more information.

Comparisons between deeper and wider networks are shown in Table 5. It can be observed that the deeper version significantly outperforms the wider version, which is as expected: (i) there are redundancies in feature dimensions, so further enlarging the width cannot bring about gains; (ii) the networks built by stacking bottlenecks improve the final performance with the increasing of depth, consistent to the observations in [He et al.(2016a)He, Zhang, Ren, and Sun, Huang et al.(2016)Huang, Liu, and Weinberger].

Intermediate ReLUs. We explore the positions of non-linearities and the results are shown in Table 6. The second block (IGCV33 block) has obvious advantages over other blocks. The first block includes two ReLUs, and thus the resulting network has two times ReLUs more than others, which may be a reason for its worse performance. The third block is equivalent to a convolutional kernel followed by a ReLU, which is worse than IGC blocks.

#Branches in group convolutions. It’s easy to find that there are many possible combinations of G1G_{1} and G2G_{2} for two group convolutions satisfying the loose complementary condition, only if the product of G1G_{1} and G2G_{2} is the common divisor of the number of input and output channels. Table 7 provides another two cases.

From the results, shown in Table 7, we can find that the first group convolution prefers to be denser. The third group convolution projects the high-dimensional features back to the low-dimensional space, which results in information loss. Therefore, reducing more kernels may have little effect on its performance. In these cases, the IGCV33 blocks in the resulting networks share the same settings, e.g. G1=2G_{1}=2, G2=2G_{2}=2. Actually, IGCV33 blocks can adopt the different settings independently, which leads to a deeper network and achieves a better performance. In our experiments, we adopt G1=2G_{1}=2, G2=2G_{2}=2 to reduce the memory cost, which also achieves a good performance.

Conclusion

In this paper, we propose a novel block, interleaved low-rank group convolutions, which enjoys the benefits of both low-rank and sparse convolutions. We introduce super-channels and the loose complementary condition to guide the design. Empirical results demonstrate the efficient of our models and the superiority over the state-of-the-arts, MobileNetV22 and IGCV22, on several benchmarks.

References