Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets

Kai Han, Yunhe Wang, Qiulin Zhang, Wei Zhang, Chunjing Xu, Tong Zhang

Introduction

Deep convolutional neural networks (CNNs) have achieved great success in many visual tasks, such as image recognition , object detection , and super-resolution . In the past few decades, the evolution of neural architectures has greatly increased the performance of deep learning models. From LeNet and AlexNet to modern ResNet and EfficientNet , there are a number of novel components including shortcuts and depth-wise convolution. Neural architecture search also provides more possibility of network architectures. These various architectures have provided candidates for a large variety of real-world applications.

To deploy the networks on mobile devices, the depth, the width and the image resolution are continuously adjusted to reduce memory and latency. For example, ResNet provides models with different number of layers, and MobileNet changes the number of channels (i.e. the width of neural network) and image resolution for different FLOPs. Most of existing works only scale one of the three dimensions – resolution, depth, and width (denoted as rr, dd, and ww). Tan and Le explore the EfficientNet , which enlarges CNNs with a compound scaling method. The great success made by EfficientNets bring a Rubik’s cube to the deep learning community, i.e. we can twist it for better neural architectures using some pre-defined formulas. For example, the EfficientNet-B7 is derivate from the B0 version by uniformly increasing these three dimensions. Nevertheless, the original EfficientNet and some improved versions only discuss the giant formula, the rules for effectively downsize the baseline model has not been fully investigated.

The straightforward way for designing tiny networks is to apply the experience used in EfficientNet . For example, we can obtain an EfficientNet-B-1 with a 200M FLOPs (floating-point operations). Since the giant formula is explored for enlarging networks, this naive strategy could not perfectly find a network with the highest performance. To this end, we randomly generate 100 models by twisting the three dimensions (r,d,w)(r,d,w) from the baseline EfficientNet-B0. FLOPs of these models are less than or equal to that of the baseline. It can be found in Figure 1, the performance of best models is about 2.5% higher than that of models obtained using the inversed giant formula of EfficientNet (green line) with different FLOPs.

In this paper, we study the relationship between the accuracy and the three dimensions (r,d,w)(r,d,w) and explore a tiny formula for the model Rubik’s cube. Firstly, we find that resolution and depth are more important than width for retaining the performance of a smaller neural architecture. We then point out that the inversed giant formula, i.e. the compound scaling method in EfficientNets is no longer suitable for designing portable networks for mobile devices, due to the reduction on the resolution is relatively large. Therefore, we explore a tiny formula for the cube through massive experiments and observations. In contrast to the giant formula in EfficientNet that is handcrafted, the proposed scheme twists the three dimensions based on the observation of frontier models. Specifically, for the given upper limit of FLOPs, we calculate the optimal resolution and depth exploiting the tiny formula, i.e. the Gaussian process regression on frontier models. The width of the resulting model is then determined according to the FLOPs constraint and previously obtained rr and dd. The proposed tiny formula for establishing TinyNets is simple yet effective. For instance, TinyNet-A achieves a 76.8% Top-1 accuracy with about 339M FLOPs but the EfficientNet-B0 with the similar performance needs about 387M FLOPs. In addition, TinyNet-E achieves a 59.9% Top-1 accuracy with only 24M FLOPs, being 1.9% higher than the previous best MobileNetV3 with similar FLOPs. To our best knowledge, we are the first to study how to generate tiny neural networks via simultaneously twisting resolution, depth and width. Besides the validations on EfficientNet, our tiny formula can be directly applied on ResNet architectures to obtain small but effective neural networks.

Related Work

Here we revisit the existing model compression methods for shrinking neural networks, and discuss about resolution, depth and width of CNNs.

Resolution, Depth and Width of CNNs.

The three dimensions including resolution, depth and width of convolutional neural networks have much impact on the performance and have been explored for scaling the CNNs. ResNet proposes models of different depth, from ResNet-18 to ResNet-152, to provide choices between model size and model performance. WideResNet propose to decrease the depth and increase the width of residual networks, demonstrating that a wider network is superior to a deep and thin counterpart. The input images of higher resolution provides more information that is helpful to model performance but also leads to higher computation cost . Considering all the three CNN dimensions into account, EffectiveNet proposes a compound scaling method to scale up networks with a handcrafted formula. However, it is still an open problem of how to shrink a given model to small and compact versions.

Approach

In this section, we first rethink the importance of resolution, depth and width, and find original EfficientNet rule lose its efficiency for smaller models. Based on the observation, we propose a new tiny formula for model Rubik’s cube to generate smaller neural networks.

Given a baseline CNN, we aim to find the smaller versions of it for deployment on low-resource devices. Resolution, depth and width are three key factors that affect the performance of CNNs as discussed in EfficientNet . However, which of them has more impact on the performance has not been well investigated in the previous works. Here we propose to evaluate the impact of (r,d,w)(r,d,w) under the fixed FLOPs or memory constraint. In practice, the FLOPs constraint is more common so we explore under FLOPs constraint and the methods can also be applied for memory constraint. To be specific, the FLOPs of the given baseline CNN are C0\mathcal{C}_{0}, the resolution of the input image is R0×R0\mathcal{R}_{0}\times\mathcal{R}_{0}, the width is W0\mathcal{W}_{0} and the depth is D0\mathcal{D}_{0}.

Here we start from EfficientNet-B0 with C0\mathcal{C}_{0} FLOPs, and sample models with FLOPs of around 0.5C00.5\mathcal{C}_{0}. In order to search around the target FLOPs, we randomly search the resolution and the depth, and tune the width around w=0.5/d/r2w=\sqrt{0.5/d/r^{2}} to make the resulted model has the target FLOPs (the difference ≤3%\leq 3\%). These random searched models are fully trained for 100 epochs on ImageNet-100 dataset. As shown in Figure 2, the accuracy is more related to the resolution, compared with the depth and the width. We find that the top accuracies are obtained around the range from 0.8 to 1.4. When r<0.8r<0.8, the accuracy is higher if the resolution is larger, while the accuracy drops slightly when r>1.4r>1.4. As for the depth, the models with high performance may have various depth from 0.5 to 2, that is to say, we may miss some good models if we narrowly restrict the depth. When fixing FLOPs, the width has roughly negative correlation to the accuracy. The good models are mostly distributed at w<1w<1.

If we follow the EfficientNet rule to obtain a model with 0.5C00.5\mathcal{C}_{0} FLOPs, namely EfficientNet-B-1, whose three dimensions are calculated and tuned as r=0.86,d=0.8,w=0.89r=0.86,d=0.8,w=0.89. Its accuracy on ImageNet-100 is only 75.8%, which is far from the optimal combination for 0.5C00.5\mathcal{C}_{0} FLOPs. It can be found in Figure 2, there is a number of models with higher performance even though they are randomly generated. This observation motivates us to explore a new model twisting formula that can obtain better models under a fixed FLOPs constraint.

2 Tiny Formula for Model Rubik’s Cube

For a given arbitrary baseline neural network, and with a FLOPs constraint of c⋅C0c\cdot\mathcal{C}_{0}, where 0<c<10<c<1 is the reduction factor, our goal is to provide the optimal values of the three dimensions (r,d,w)(r,d,w) for shrinking the model. Basically, we assume the optimal coefficients rr, ww, dd for shrinking resolution, width and depth are

where f1(⋅)f_{1}(\cdot), f2(⋅)f_{2}(\cdot) and f3(⋅)f_{3}(\cdot) are the functions for calculating the three dimensions. We will give the formulation of the equations in the following.

Then, we randomly sample a number of models with different coefficients and verify them to explore the relationship between the performance and the three dimensions. The coefficients are randomly sampled from a given range. We preserve the models whose FLOPs are between 0.03⋅C00.03\cdot\mathcal{C}_{0} and 1.05⋅C01.05\cdot\mathcal{C}_{0}. After fully training and testing these models, we can obtain their accuracies on validation set. We plot scatter diagram of accuracy v.s. FLOPs as shown in Figure 1. Obviously, there are a number of models whose performance is better than the vanilla EfficientNet-B0 and its shrunken versions obtained by exploiting the inversed compound scaling scheme.

To further explore the property of the best models, we select the models on the Accuracy-FLOPs Pareto front. Pareto front is a set of nondominated solutions, being chosen as optimal, if no objective can be improved without sacrificing at least one other objective . In particular, the top 20% models with higher performance and lower computational complexities (i.e. FLOPs) are selected using NSGA-III nondominated sorting strategy . We show the relation between depth/width/resolution and FLOPs of these selected models in Figure 3. Spearman correlation coefficient (Spearmanr) are calculated to measure the correlation between depth/width/resolution and FLOPs. From the results in Figure 3, the rank of correlations between the three dimensions with FLOPs is r>d>wr>d>w. Wherein, the Spearmanr score for resolution is 0.81, which is much higher than that of width.

where μ∗=K(c∗,c⃗)(K(c⃗,c⃗)+σ2I)−1r⃗\mu_{*}=K(c_{*},\vec{c})(K(\vec{c},\vec{c})+\sigma^{2}I)^{-1}\vec{r} and Σ∗=k(c∗,c∗)+σ2−K(c∗,c⃗)(K(c⃗,c⃗)+σ2I)−1K(c⃗,c∗)\Sigma_{*}=k(c_{*},c_{*})+\sigma^{2}-K(c_{*},\vec{c})(K(\vec{c},\vec{c})+\sigma^{2}I)^{-1}K(\vec{c},c_{*}) are the mean and variance for the test point c∗c_{*}. The formula for depth can be obtained similarly. Then, the last dimension, i.e. width, can be determined by FLOPs constraint:

In contrast to the handcrafted compound scaling method in EfficientNets, the proposed model shrinking rule is designed based on the observation of frontier small models, which are more effective for producing tiny networks with higher performance.

Our tiny formula for model Rubik’s cube can be applied to any network architecture. Here we start from the excellent baseline network, EfficientNet-B0 , and apply our shrinking method to obtain smaller networks. With about 5.3M parameters and 390M FLOPs, EfficientNet-B0 consists of 16 mobile inverted residual bottlenecks , in addition to the normal stem layer and classification head layers. To apply our tiny formula, we first construct a number of networks whose (r,w,d)(r,w,d) are randomly sampled as shown in Figure 1. After obtaining the accuracy on ImageNet-100, we can train the Gaussian process regression model for resolution and depth. Here we adopt the widely used RBF kernel as the covariance function. Then, given the desired FLOPs constraint cc, we can determine the three dimensions, aka (r,w,d)(r,w,d), by the above equations. We set cc in {0.9,0.5,0.25,0.13,0.06}\{0.9,0.5,0.25,0.13,0.06\}, and obtain a series of smaller EfficientNet-B0, namely, TinyNet-A to E.

Experiments

In this section, we apply our tiny formula for model Rubik’s cube to shrink EfficientNet-B0 and ResNet-50. The effectiveness of our method is verified on the visual recognition benchmarks.

ImageNet ILSVRC2012 dataset is a large-scale image classification dataset containing 1.2 million images for training and 50,000 validation images belonging to 1,000 categories. We use the common data augmentation strategy including random crop, random flip and color jitter. The base input resolution is 224 for r=1r=1.

ImageNet-100.

ImageNet-100 is the subset of ImageNet-1000 that contains randomly sampled 100 classes. 500 training images are randomly sampled for each class, and the corresponding 5,000 images are used as validation set. The data augmentation strategy is the same as that in ImageNet-1000.

Implementation details.

All the models are implemented using PyTorch and trained on NVIDIA Tesla V100 GPUs. The EfficientNet-B0 based models are trained using similar settings as . We train the models for 450 epochs using the RMSProp optimizer with momentum 0.9 and decay 0.9. The weight decay is 1e-5 and batch normalization momentum is set as 0.99. The initial learning rate is 0.048 and decays by 0.97 every 2.4 epochs. Learning rate warmup is applied for the first 3 epochs. The batch size is 1024 for 8 GPUs with 128 images per chip. The dropout of 0.2 is applied on the last fully-connected layer for regularization. We also use exponential moving average (EMA) with decay 0.9999. For ResNets, the models are trained for 90 epochs with batch size of 1024. SGD optimizer with the momentum 0.9 and weight decay 1e-4 is used to update the weights. The learning rate starts from 0.4 and decays by 0.1 every 30 epochs.

In the original EfficientNet rule for giant models , the FLOPs value of a model is calculated as 2−ϕ⋅C02^{-\phi}\cdot\mathcal{C}_{0}. We denote the models obtained from the inversed giant formula in original EfficientNet as EfficientNet-B-ϕ where ϕ=1,2,3,4\phi=1,2,3,4, with about 200M, 100M, 50M, 25M FLOPs, respectively.

2 Experiments on ImageNet-100

As stated in the above sections, we randomly sample a number of models with different resolution, depth and width. In particular, resolution, depth or width is randomly sampled from the range of 0.35≤r≤2.80.35\leq r\leq 2.8, 0.35≤d≤2.80.35\leq d\leq 2.8 and 0.35≤w≤2.80.35\leq w\leq 2.8. The sampled models are trained on ImageNet-100 dataset for 100 epochs. The other training hyperparameters are the same as those in implementation details for EfficientNet-B0 based models. 100 models are sampled in total and it takes about 2.5 GPU hours on average to train one model. The results of all the models are shown in Figure 1. Larger FLOPs lead to higher accuracy generally. Some of the sampled models perform better than the shrunken models using inversed giant formula of EfficientNet. For example, a sampled model with 318M FLOPs achieves 79.7% accuracy while EfficientNet-B0 with 387M FLOPs only achieves 78.8%. These observations indicate the necessity to design a more effective model shrinking method.

Comparison to EfficientNet Rule.

In order to verify the effectiveness of the proposed model shrinking method, we compare our method with the inversed giant formula of EfficientNet and separately changing rr, dd or ww. From Table 1, the proposed method outperforms both EfficientNet rule and separately adjusting resolution, depth or width, demonstrating the effectiveness of the proposed tiny formula for model shrinking.

Shrinking ResNet.

In addition to EfficientNet-B0, we also apply our method for shrinking the widely-used ResNet network architecture. ResNet-50 is adopted as the baseline model, and it is shrunken in different ways including reducing layers, EfficientNet rule and our method. The results on ImageNet-100 are shown in Table 2. Our models outperform other models generally, suggesting the effectiveness of the proposed model shrinking method for ResNet architecture.

3 Experiments on ImageNet-1000

The tiny formula obtained on ImageNet-100 can be well transferred to other datasets as demonstrated in NAS literature . We evaluate our tiny formula on the large-scale ImageNet-1000 dataset to verify its generalization.

We compare TinyNet models with other competitive small neural networks, including the models from original EfficientNet rule, i.e. EfficientNet-B-ϕ, and other state-of-the-art small CNNs such as MobileNet series , ShuffleNet series , and MnasNet , are compared here. Several competitive NAS-based models are also included. Table 3 shows the performance of all the compared models. Our TinyNet models generally outperform other CNNs. In particular, our TinyNet-E achieves 59.9% Top-1 accuracy with 24M FLOPs, being 1.9% higher than the previous best MobileNetV3 Small 0.5×\times with similar computational cost.

RandAugment is a practical automated data augmentation strategy to improve the generalization of deep learning models. We use RandAugment with magnitude 9 and standard deviation 0.5 to improve the performance of our TinyNet-A and EfficientNet-B0, and show the results in Table 3. For the TinyNet models, RandAugment is beneficial to the performance. In particular, TinyNet-A + RA achieves 77.7% Top-1 accuracy which is 0.9% higher than vanilla TinyNet-A.

Visualization of Learning Curves.

To better demonstrate the effect of our method, we plot the learning curves of EfficientNet-B-4 and our TinyNet-E in Figure 4. From Figure 4(a), the accuracy of TinyNet-E is higher than that of EfficientNet-B-4 by a large margin consistently during training. In the end of training, our TinyNet-E outperforms EfficientNet-B-4 by an accuracy gain of 3.2%. The train and validation loss curves in Figure 4(b) also show the superiority of our TinyNet.

Visualization of Class Activation Map.

We visualize the class activation map for EfficientNet-B-4 and TinyNet-E to better demonstrate the superiority of our TinyNet. The images are randomly picked from ImageNet-1000 validation set. As shown in Figure 5, TinyNet-E pays attention to the more relevant regions, while EfficientNet-B-4 sometimes only focuses on the unrelated objects or the local part of target objects.

Shrinking by r,d,w𝑟𝑑𝑤r,d,w Separately.

We also compare the proposed model shrinking rule with the naive method, i.e. changing resolution, depth and width separately. We tune resolution, width or depth separately to form models with 100M and 200M FLOPs. Note that the minimum viable depth for EfficientNet-B0 is reached with 7 inverted residual bottlenecks, and the corresponding FLOPs 174M. The results are shown in Figure 6. In general, all shrinking approaches lead to lower accuracy when the number of FLOPs decreases, but our model shrinking method can alleviate the accuracy drop, suggesting the effectiveness of the proposed method.

Inference Latency Comparison.

We also measure the inference latency of several representative CNNs on Huawei P40 smartphone. We test under single-threaded mode with batch size 1 using the MindSpore Lite tool . The results are listed in Table 4, where we run 1000 times and report average latency. Our TinyNet-A runs 15% faster than EfficientNet-B0 while their accuracies are similar. TinyNet-E can obtain 3.2% accuracy gain compared to EfficientNet-B-4 with similar latency.

Generalization on Object Detection.

To verify the generalization of our models, we apply tinynets on object detection task. We adopt SSDLite with 512×\times512 input as baseline network due to its efficiency and test on MS COCO dataset . The experimental setting is similar to that in . From results in Table 5, we can see that our TinyNet-D outperforms EfficientNet-B-3 by a large margin with comparable computational cost.

Conclusion and Discussion

In this paper, we study the model Rubik’s cube for shrinking deep neural networks. Based on a series of observations, we find that the original giant formula in EfficientNet is unsuitable for generating smaller neural architectures. To this end, we thoroughly analyze the importance of resolution, depth and width w.r.t. the performance of portable deep networks. Then, we suggest to pay more concentration on the resolution and depth and calculate the model width to satisfy FLOPs constraint. We explore a series of TinyNets by utilizing the tiny formula to twist the three dimensions. The experimental results for both EfficientNets and ResNets demonstrate effectiveness of the proposed simple but effective scheme for designing tiny networks. Moreover, the tiny formula in this work is summarized according to the observation on smaller models. These smaller models can also be further enlarged to obtain higher performance with some new rules beyond the giant formula in EfficientNets, which will be investigated in future works.

Broader Impact

The widely usage of deep neural networks which require large amount of computation resource is putting pressure on the energy source and the natural environment. The proposed model shrinking method for obtaining tiny neural networks is beneficial to energy conservation and environment protection.

Appendix A Appendix

We also verify the effectiveness of the proposed model Rubik’s cube on the state-of-the-art portal network GhostNet . We first build a baseline GhostNet with about 591M FLOPs, which is denoted as GhostNet-A (details in Table 7). We start from GhostNet-A and shrink the model by the proposed tiny formula, resulting in a serious of smaller models, i.e. GhostNet-B/C/D. The new models are trained using the similar setting as TinyNets. The comparison with the original GhostNets on ImageNet dataset is shown in Table 6. We can see that the new GhostNet models obtained by the proposed model Rubik’s cube outperform the original GhostNets which only change the width.

Note that GhostNets use ReLU as activation function, and more complex activations (e.g. HSwish ) may further improve the performance. We also show the results with automated data augmentation strategy, i.e. RandAugment in Table 6. Both fancy activation function and data augmentation could boost GhostNets to higher performance.

For better comparison, we plot the results in Fig. 7. The compared methods include recent state-of-the-art models, i.e. MobileNetV2 , MobileNetV3 , EfficientNet , ShuffleNetV2 , FBNet , MnasNet , ProxylessNAS and GreedyNAS . GhostNet models enhanced by TinyNet technique consistently outperform other models by a significant margin.

References