Ultimate tensorization: compressing convolutional and FC layers alike

Timur Garipov, Dmitry Podoprikhin, Alexander Novikov, Dmitry Vetrov

Introduction

Convolutional Neural Networks (CNNs) show state-of-the-art performance on many problems in computer vision, natural language processing and other fields . At the same time, CNNs require millions of floating point operations to process an image and therefore real-time applications need powerful CPU or GPU devices. Moreover, these networks contain millions of trainable parameters and consume hundreds of megabytes of storage and memory bandwidth . Thus, CNNs are forced to use RAM instead of solely relying on the processor cache – orders of magnitude more energy efficient memory device – which increases the energy consumption even more. These reasons restrain the spread of CNNs on mobile devices.

To address the storage and memory requirements of neural networks, used tensor decomposition techniques to compress fully-connected layers. They represented the parameters of the layers in the Tensor Train format and learned the network from scratch in this representation. This approach provided enough compression to move the storage bottleneck of VGG-16 from the fully-connected layers to convolutional layers. For a more detailed literature overview, see Sec. 5.

In this paper, we propose a tensor factorization based method to compress convolutional layers. Our contributions are:

We experimentally show that applying the Tensor Train decomposition – the compression technique used in – directly to the tensor of a convolution yields poor results (see Sec. 6). We explain this behavior and propose a way to reshape the 4-dimensional kernel of a convolution into a multidimensional tensor to fully utilize the compression power of the Tensor Train decomposition (see Sec. 4).

We experimentally show that the proposed approach allows compressing a network that consists only of convolutions up to 4×4\times times with 2%2\% accuracy decrease (Sec. 6).

We combine the proposed approach with the fully-connected layers compression of . Compressing both convolutional and fully-connected layers of a network yields 82×82\times network compression with 1%1\% accuracy drop, see Sec. 6.

Convolutional Layer

To improve the computational performance, many deep learning frameworks reduce the convolution (1) to a matrix-by-matrix multiplication (see Fig. 1). We exploit this matrix formulation to motivate a particular way of applying the Tensor Train format to the convolutional kernel (see Sec. 4). In the rest of this section, we introduce the notation needed to reformulate convolution (1) as a matrix-by-matrix multiplication \mathboldY=\mathboldX\mathboldK\mathbold{Y}=\mathbold{X}\mathbold{K}.

Using the matrices defined above, we can rewrite the convolution definition (1) as \mathboldY=\mathboldX\mathboldK\mathbold{Y}=\mathbold{X}\mathbold{K}.

Note that the compression approach presented in the rest of the paper works with other types of convolutions, such as convolutions with padding, stride larger than 11, or rectangular filters. But for clarity, we illustrate the proposed idea on the basic convolution (1).

Tensor Train Decomposition

The elements of the collection {rk}k=0d\{r_{k}\}_{k=0}^{d} are called TT-ranks. The collections of matrices {{\mathboldGk[jk]}jk=1nk}k=1d\{\{\mathbold{G}_{k}[j_{k}]\}_{j_{k}=1}^{n_{k}}\}_{k=1}^{d} are called TT-cores .

TT-convolutional Layer

In this section, we propose two ways to represent a convolutional kernel K\boldsymbol{\mathcal{K}} in the TT-format. One way is to apply the TT-decomposition to the tensor K\boldsymbol{\mathcal{K}} directly. To see the drawbacks of this approach, consider a 1×11\times 1 convolution, which is a small fully-connected layer applied to the channels of the input image in each pixel location. The kernel of such convolution is essentially a 22-dimensional array, and the TT-decomposition of 22-dimensional arrays coincides with the matrix low-rank format. But for fully-connected layers, the matrix TT-format proved to be more efficient than the matrix low-rank format . Thus, we seek for a decomposition that would coincide with the matrix TT-format on 1×11\times 1 convolutions.

where c′=c1+∑i=2d(ci−1)∏j=1i−1Cjc^{\prime}=c_{1}+\sum_{i=2}^{d}(c_{i}-1)\prod_{j=1}^{i-1}C_{j} and s′=s1+∑i=2d(si−1)∏j=1i−1Sjs^{\prime}=s_{1}+\sum_{i=2}^{d}(s_{i}-1)\prod_{j=1}^{i-1}S_{j}.

To train a network with TT-conv layers, we treat the elements of the TT-cores as the parameters of the layer and apply stochastic gradient descent with momentum to them. To compute the necessary gradients we use automatic differentiation implemented in TensorFlow .

Related Work

Fully-connected layers of neural networks are traditionally considered as the memory bottleneck and numerous works focused on compressing these layers . However, several state-of-the-art neural networks are either bottlenecked by convolutional layers , or their fully-connected layers can be compressed to move the bottleneck to the convolutional layers . This leads to a number of works focusing on compressing and speeding up the convolutional layers .

One approach to compressing a convolutional layer is based on either pruning less important weights from the convolutional kernel, or restricting possible variation of the weights (quantization), or both . Our approach is compatible with the quantization technique: one can quantize the elements of the TT cores of the decomposition. Some works also add Huffman coding on top of other compression techniques , which is also compatible with the proposed method.

Another approach is to use tensor or matrix decompositions. CP-decomposition and Kronecker product factorization allow to speed up the inference time of convolutions and compress the network as a side effect.

Experiments

We evaluated the compressing strength of the proposed approach on CIFAR-10 dataset , which has 50 00050\,000 train images and 10 00010\,000 test images. In all the experiments, we used stochastic gradient descent with momentum with coefficient 0.90.9, trained for 100100 epochs starting from the learning rate of 0.10.1 and decreased it 10×10\times after each 3030 epochs. To make the experiments reproducible, we released the codebasehttps://github.com/timgaripov/TensorNet-TF.

We used two architectures as references: the first one is dominated by the convolutions (they occupy 99.54%99.54\% parameters of the network), and the second one is dominated by the fully-connected layers (they occupy 95.98%95.98\% parameters of the network).

The first network has the following architecture: conv (6464 output channels); BN; ReLU; conv (6464 output channels); BN; ReLU; max-pool (3×33\times 3 with stride 22); conv (128128 output channels); BN; ReLU; conv (128128 output channels); BN; ReLU; max-pool (3×33\times 3 with stride 22); conv (128128 output channels); BN; ReLU; conv (128128 output channels); avg-pool (4×44\times 4); fc (128×10128\times 10), where ’BN’ stands for batch normalization and all convolutional filters are of size 3×33\times 3. To compress the network we replace each convolutional layer excluding the first one (it contains less than 1%1\% of the network parameters) with the TT-conv layer (see Sec. 4). For training, we initialize the TT-cores of the TT-conv layers with random noise and train the whole network from scratch.

We compare the proposed TT-convolution against the naive approach – directly applying the TT-decomposition to the 44-dimensional convolutional kernel (see Sec.4). We report that on the 2×2\times compression level the proposed approach (0.8%0.8\% loss of accuracy) outperforms the naive baseline (2.4%2.4\% loss of accuracy), for details see Tbl. 1a.

Network with convolutions and fully-connected layers.

The second reference network was obtained from the first one by replacing the average pooling with two fully-connected layers of size 8192×15368192\times 1536 and 1536×5121536\times 512.

To compress the second network, we replace all layers excluding the first and the last one (they occupy less than 1%1\% of parameters) with TT-conv and TT-fc layers. To speed up the convergence, we trained the network in two stages: first we replaced only the convolutional layers with TT-conv layers and trained the network; then we replaced the fully-connected layers with randomly initialized TT-fc layers and fine-tuned the whole model.

To compare against we include the results of compressing only fully-connected layers (Tbl. 1b). Initially, the fully-connected part was the memory bottleneck and it was more fruitful to compress it while leaving the convolutions untouched: we obtained 10.72×10.72\times network compression with 0.2%0.2\% accuracy drop by compressing only fully-connected layers, and 9.61×9.61\times compression with 0.4%0.4\% accuracy drop by compressing both fully-connected and convolutional layers. But after the first gains, the bottleneck moved to the convolutional part and the fully-connected layers compression capped at about 21×21\times network compression. At this point, by additionally factorizing the convolutions we raised the network compression up to 80×80\times while losing 1.1%1.1\% of accuracy (Tbl. 1b).

Conclusion

In this paper, we proposed a tensor decomposition approach to compressing convolutional layers of a neural network. By combing this convolutional approach with the work for fully-connected layers, we compressed a convolutional network 80×80\times times. These results make a step towards the era of embedding compressed models into smartphones to allow them constantly look and listen to its surroundings.

In the future work, we will experiment with the proposed approach on the ILSVRC-2012 dataset on state-of-the-art neural architectures.

References