AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
Jiahui Yu, Thomas Huang
Introduction
The channel configuration (a.k.a. filter numbers or channel numbers) of a neural network plays a critical role in its affordability on resource constrained platforms, such as mobile phones, wearables and Internet of Things (IoT) devices. The most common constraints , i.e., latency, FLOPs and runtime memory footprint, are all bound to the number of channels. For example, in a single convolution or fully-connected layer, the FLOPs (number of Multiply-Adds) increases linearly by the output channels. The memory footprint can also be reduced by reducing the number of channels in bottleneck convolutions for most vision applications .
Despite its importance, the number of channels has been chosen mostly based on heuristics. LeNet-5 selected 6 channels in its first convolution layer, which is then projected to 16 channels after sub-sampling. AlexNet adopted five convolutions with channels equal to , , , and . A commonly used heuristic, the “half size, double channel” rule, was introduced in VGG nets , if not earlier. The rule is that when spatial size of feature map is halved, the number of filters is doubled. This heuristic has been more-or-less used in followup network architecture designs including ResNets , Inception nets , MobileNets and networks for many vision applications . Other heuristics have also been explored. For example, the pyramidal rule suggested to gradually increase the channels in all convolutions layer by layer, regardless of spatial size. Figure 1 visually summarizes these heuristics for setting channel numbers in a neural network.
Beyond the macro-level heuristics across entire network, recent works have also digged into channel configuration for micro-level building blocks (a network building block is usually composed of several and convolutions). These micro-level heuristics have led to better speed-accuracy trade-offs. The first of its kind, bottleneck residual block, was introduced in ResNet . It is composed of , , and convolutions, where the layers are responsible for reducing and then restoring dimensions, leaving the layer a bottleneck ( reduction). MobileNet v2 , however, argued that the bottleneck design is not efficient and proposed the inverted residual block where layers are used for expanding feature first ( expansion) and then projecting back after intermediate depthwise convolution. Furthermore, MNasNet and ProxylessNAS nets included expansion version of inverted residual block into search space, and achieved even better accuracy under similar runtime latency.
Apart from these human-designed heuristics, efforts on automatically optimizing channel configuration have been made explicitly or implicitly. A recent work suggested that many network pruning methods can be thought of as performing network architecture search for channel numbers. Liu et al. showed that training these pruned architectures from scratch leads to similar or even better performance than fine-tuning and pruning from a large model. More recently, MNasNet proposed to directly search network architectures, including filter sizes, using reinforcement learning algorithms . Although the search is performed on the factorized hierarchical search space, massive network samples and computational cost are required for an optimized network architecture.
In this work, we study how to set channel numbers in a neural network to achieve better accuracy under constrained resources. To start, the first and the most brute-force approach came in mind is the exhaustive search: training all possible channel configurations of a deep neural network for full epochs (e.g., MobileNets are trained for approximately 480 epochs on ImageNet). Then we can simply select the best performers that are qualified for efficiency constraints. However, it is undoubtedly impractical since the cost of this brute-force approach is too high. For example, we consider a -layer convolutional networks and a search space limited to 10 candidates of channel numbers (e.g., , , …, ) for each layer. As a result, there are totally candidate network architectures.
To address this challenge, we present a simple and one-shot solution AutoSlim. Our main idea lies in training a slimmable network to approximate the network accuracy of different channel configurations. Yu et al. introduced slimmable networks that can run at arbitrary width with equally or even better performance than same architecture trained individually. Although the original motivation is to provide instant and adaptive accuracy-efficiency trade-offs, we find slimmable networks are especially suitable as benchmark performance estimators for several reasons: (1) Training slimmable models (using the sandwich rule ) is much faster than the brute-force approach. (2) A trained slimmable model can execute at arbitrary width, which can be used to approximate relative performance among different channel configurations. (3) The same trained slimmable model can be applied on search of optimal channels for different resource constraints.
In AutoSlim, we first train a slimmable model for a few epochs (e.g., 10% to 20% of full training epochs) to quickly get a benchmark performance estimator. We then iteratively evaluate the trained slimmable model and greedily slim the layer with minimal accuracy drop on validation set (for ImageNet, we randomly hold out samples of training set as validation set). After this single pass, we can obtain the optimized channel configurations under different resource constraints (e.g., network FLOPs limited to 150M, 300M and 600M). Finally we train these optimized architectures individually or jointly (as a single slimmable network) for full training epochs. We experiment with various networks including MobileNet v1, MobileNet v2, ResNet-50 and RL-searched MNasNet on the challenging setting of 1000-class ImageNet classification. We compare our results with two baselines: (1) the default channel configuration of these networks, and (2) channel pruning methods on same network architectures .
Our contributions are summarized as follows:
We present the first one-shot approach on network architecture search for channel numbers with experiments on large-scale ImageNet classification.
We demonstrate the importance of channel configuration in neural networks and the effectiveness of our approach on addressing this challenging problem.
We achieve the state-of-the-art speed-accuracy trade-offs by setting the optimized channel configurations using AutoSlim.
Related Work
In this part, we mainly discuss previous methods on automatic architecture search for channel numbers. Human-designed heuristics have been introduced in Section 1 and visually summarized in Figure 1.
Neural Architecture Search (NAS). Recently there has been a growing interest in automating the neural network architecture design . Significant improvements have been achieved by these automatically searched architectures in many vision and language tasks . However, most neural architecture search methods did not include channel configuration into search space, and instead applied human-designed heuristics. More recently, the RL-based searching algorithms are also applied to prune channels or search for filter numbers directly. He et al. proposed AutoML for Model Compression (AMC) which leveraged reinforcement learning (deep deterministic policy gradient ) to provide the model compression policy. MNasNet proposed to directly search network architectures, including filter sizes, for mobile devices. In the search, each sampled model is trained on epochs using an aggressive learning rate schedule, and evaluated on a validation set. In total, Tan et al. sampled about models during architecture search. Further, ProxylessNAS proposed to directly learn the architectures for large-scale target tasks and target hardware platforms, based on DARTS . For each residual block, ProxylessNAS followed the channel configuration of MNasNet , while inside each block, the choices can be or version of inverted residual blocks. The memory consumption issue was addressed by binarizing the architecture parameters and forcing only one path to be active.
2 Slimmable Networks
Slimmable networks were firstly introduced in . A general slimmable training algorithm and the switchable batch normalization were introduced to train a single neural network executable at different widths, permitting instant and adaptive accuracy-efficiency trade-offs at runtime. However, one drawback of the switchable batch normalization is that the width can only be chosen from a predefined widths set. The drawback was addressed in , where the authors introduced universally slimmable networks, extending slimmable networks to execute at arbitrary width, and generalizing to networks both with and without batch normalization layers. Meanwhile, two improved training techniques, the sandwich rule and inplace distillation, were proposed to enhance training process and boost testing accuracy. Moreover, with the proposed methods, one can train nonuniform universally slimmable networks, where the width ratio is not uniformly applied to all layers. In other words, each layer in a nonuniform universally slimmable network can adjust its number of channels independently during inference. In this work, we simply refer to nonuniform universally slimmable networks as slimmable networks, if not explicitly noted. While the original motivation of slimmable networks is to provide instant and adaptive accuracy-efficiency trade-offs at runtime for different devices, we present an approach that uses slimmable networks for searching channel configurations of deep neural networks.
Network Slimming by Slimmable Networks
In this section, we first present an overview of our proposed approach for searching channel configuration of neural networks. We then discuss and analyze the difference of our approach compared with other baselines, i.e., network pruning methods and network architecture search methods. Afterwards we present each individual module in our proposed solution and discuss its non-trivial details.
The goal of channel configuration search is to optimize the number of channels in each layer, such that the network architecture with optimized channel configuration can achieve better accuracy under constrained resources. The constraints can be FLOPs, latency, memory footprint or model size. Our approach is conceptually simple, and it has two essential steps:
(1) Given a network architecture (e.g., MobileNets, ResNets), we first train a slimmable model for a few epochs (e.g., 10% to 20% of full training epochs). During the training, many different sub-networks with diverse channel configurations have been sampled and trained. Thus, after training one can directly sample its sub-network architectures for instant inference, using the correspondent computational graph and same trained weights.
(2) Next, we iteratively evaluate the trained slimmable model on the validation set. In each iteration, we decide which layer to slim by comparing their feed-forward evaluation accuracy on validation set. We greedily slim the layer with minimal accuracy drop, until reaching the efficiency constraints. No training is required in this step.
The flow diagram of our approach is shown in Figure 2. Our approach is also flexible for different resource constraints, since the FLOPs, latency, memory footprint and model size are all deterministic given a channel configuration and a runtime environment. By a single pass of greedy slimming in step (2), we can obtain the (FLOPs, latency, memory footprint, model size, accuracy) tuples of different channel configurations. It is noteworthy that the latency and accuracy are relative values, since the latency may be different across different hardware and the accuracy can be improved by training the network for full epochs. In the setting of optimizing channel numbers, we benefit from these relative values as performance estimators.
Discussion. We compare the flow diagram of our approach with the baselines, i.e., network pruning methods and network architecture search methods.
Network architecture search methods commonly consist of three major components: search space, search strategy, and performance estimation strategy. A typical pipeline is shown in Figure 4. First the search space is defined, based on which the search agent samples network architectures. The architecture is then passed to a performance estimator, which returns rewards (e.g., predictive accuracy after training and/or network runtime latency) to the search agent. In the process, the search agent learns from the repetitive loop to design better network architectures. One major drawback of network architecture search methods is their high computational cost and time cost . Although recently differentiable architecture search methods were proposed, they cannot be applied on search of channel numbers directly. Most of them were still using human-designed heuristics for setting channel numbers, which may introduce human bias.
2 Training Slimmable Networks
Warmup. We warmup by a brief review of training techniques for slimmable networks. More details can be found in . Slimmable networks were firstly introduced and trained with switchable batch normalization , which employed individual BNs for different sub-networks. During training, features are normalized with current mini-batch mean and variance, thus a simple modification to switchable batch normalization is introduced in : re-calibrating BN statistics after training. With this simple modification, one can train universally slimmable networks that can run with arbitrary channel numbers. Moreover, two improved training techniques the sandwich rule and inplace distillation were introduced to enhance training process and boost testing accuracy. We use all these techniques in training slimmable models by default.
Assumption. Our approach lies in the assumption that the slimmable model is a good accuracy estimator of individually trained models given same channel configuration. More specifically, we are interested in the relative ranking of accuracy among networks with different channel configurations. We use the instant inference accuracy of a slimmable model as the performance estimator. We note that assumptions and approximations commonly exist in other related methods. For example, in network channel pruning methods , one may assume that weights with smaller norm are less informative and can be pruned, which may not be the case as shown in . Recently the Lottery Ticket Hypothesis was also introduced. In network architecture search methods , one may believe the transferability among different datasets, accuracy approximations using aggressive learning rates and fewer training epochs, and approximation in runtime latency modeling.
The Search Space. The executable sub-networks in a slimmable model compose the search space of channel configurations given a network architecture. To train a slimmable model, we simply apply two width multipliers as the upper bound and lower bound of channel numbers. For example, for all mobile networks , we train a slimmable model that can execute between and . In each training iteration, we randomly and independently sample the number of channels in each layer. It is noteworthy that in residual networks, we first sample the channel number of residual identity pathway and then randomly and independently sample channel number inside each residual block. Moreover, we make all layers in a neural network slimmable, including the first convolution layer and last fully-connected layer. In each layer, we divide the channels into groups evenly (e.g., 10 groups) to reduce the search space. In other words, during training or slimming, we sample or remove an entire group, instead of an individual channel. We note that even with channel grouping, the search space is still large.
We implement a distributed training framework with synchronized stochastic gradient descent (SGD) on PyTorch . We set different random seeds in different processes such that each GPU samples diverse channel configurations in each SGD training step. All other techniques introduced in and distributed training techniques introduced in are used by default. All code will be released.
3 Greedy Slimming
After training a slimmable model, we evaluate it on the validation set (on ImageNet we randomly hold out images in training set as validation set). We start with the largest model (e.g., ) and compare the network accuracy among the architectures where each layer is slimmed by one channel group. We then greedily slim the layer with minimal accuracy drop. During the iterative slimming, we obtain optimized channel configurations under different resource constraints. We stop until reaching the strictest constraint (e.g., 50M FLOPs or 30ms CPU latency).
Large Batch Size. During greedy slimming, no training is involved. Thus we directly put the model in evaluation mode (no gradients are required), which enables us to use a larger batch size (for example during slimming we use mini-batch size for each GPU with totally V100 GPUs). Large batch size brings two benefits. First, previous work shows that BN statistics will be accurate if it is calibrated with the batch size larger than . Thus post-statistics of BN in our greedy slimming can be computed online without additional cost. Second, with large batch size we can simply use single feed-forward prediction accuracy as the performance estimator. In practice we find it speeds up greedy slimming and simplifies implementation without affecting final performance.
Training Optimized Networks. Similar to architecture search methods, after the search, we train these optimized network architectures from scratch. By default we search for the network FLOPs at approximately 200M, 300M and 500M, and train a slimmable model.
Experiments
Table 1 summarizes our results on ImageNet classification with various network architectures including MobileNet v1 , MobileNet v2 , MNasNet , and one large model ResNet-50 . We compare our results with their default channel configurations and recent channel pruning methods . The top-1 errors of our baselines are from corresponding works . To have a clear view, we divide the network architectures into four groups, namely, 200M FLOPs, 300M FLOPs, 500M FLOPs and heavy models (basically ResNet-50 based models). We evaluate their latency on same hardware environment with single-core CPU to ensure fairness. Device memory is reported as a summary of all feature maps and weights. We note that the memory footprint can be largely optimized by improving memory reusing and implementation of dedicated operators. For example, the inverted residual block can be optimized by splitting channels into groups and performing partial execution for multiple times . For all network architectures we train 50 epochs with squeezed learning rate schedule to obtain a slimmable model for greedy slimming. After search, we train the optimized network architectures for full epochs (300 epochs with linearly decaying learning rate for mobile networks, 100 epochs with step learning rate schedule for ResNet-50 based models) with other training settings following previous works (weight initialization, weight decay, data augmentation, training/testing image resolution, optimizer, hyper-parameters of batch normalization). We exclude the parameters and FLOPs of Batch Normalization layers following common practice since they can be fused into convolution layers.
As shown in Table 1, our models have better top-1 accuracy compared with the default channel configuration of MobileNet v1, MobileNet v2 and ResNet-50 across different computational budgets. We even have improvements over RL-searched MNasNet , where the filter numbers are already included in its search space. Notably, by setting optimized channel numbers, our AutoSlim-MobileNet-v2 at 305M FLOPs achieves 74.2% top-1 accuracy, 2.4% better than default MobileNet-v2 (301M FLOPs), and even 0.2% better than RL-searched MNasNet (317M FLOPs). Our AutoSlim-ResNet-50 at 570M FLOPs, without depthwise convolutions, achieves 1.3% better accuracy than MobileNet-v1 (569M FLOPs).
2 Visualization and Discussion
In this part, we visualize our optimized channel configurations and discuss some insights from the results.
Comparison with Default Channel Numbers. We first compare our results with default channels in MobileNet v2 . We show the optimized number of channels (left) and the percentage compared with default channels (right) in Figure 5. Compared with default MobileNet v2, our optimized configuration has fewer channels in shallow layers and more channels in deep ones.
Comparison with Width Multiplier Heuristic. Applying width multiplier , a global hyper-parameter across all layers, is a commonly used heuristic to trade off between model accuracy and efficiency . We search optimal channels at 207M, 305M and 505M FLOPs corresponding to MobileNet v2 , and . Figure 6 shows the pattern that under different budgets, AutoSlim applies different width scaling in each layer.
Comparison with Model Pruning Methods. Next, we compare our optimized channel configuration with model pruning method AMC . In Figure 6, we show the number of channels in all layers of optimized MobileNet v2. We observe several characteristics of our optimized channel configurations. First, AutoSlim-MobileNet-v2 has much more channels in deep layers, especially for deep depthwise convolutions. For example, AutoSlim-MobileNet-v2 has channels in the second last layer, compared with channels in AMC-MobileNet-v2. Second, AutoSlim-MobileNet-v2 has fewer channels in shallow layers. For example, AutoSlim-MobileNet-v2 has only channels in first convolution layer, while AMC-MobileNet-v2 has channels. It is noteworthy that although shallow layers have a small number of channels, the spatial size of feature maps is large. Thus overall these layers take up large computational overheads.
3 CIFAR10 Experiments
In addition to ImageNet dataset, we also conduct experiments on CIFAR10 dataset. We use same weight decay hyper-parameter, initial learning rate and learning rate schedule as ImageNet experiments. We note that these training settings may not be optimal for CIFAR10 dataset, nevertheless we report ablative study with same hyper-parameters and settings. We first report the performance of MobileNet v2 with the default channel configurations. We then search with proposed AutoSlim to obtain optimized channel configurations at same FLOPs (we hold out images from training set as validation set during the search). Finally we train the optimized architectures individually with same settings as the baselines. Table 2 shows that AutoSlim models have higher accuracy than baselines on CIFAR10 dataset.
We further study the transferability of the network architectures learned from ImageNet to CIFAR10 dataset, and compare it with the channel configuration searched on CIFAR10 directly. The results are shown in Table 3. It suggests that the optimized channel configuration on ImageNet cannot generalize to CIFAR10. Compared with the optimized architecture for ImageNet, we observed that the optimized architecture for CIFAR10 have much fewer channels in deep layers, which we guess may lead to better generalization on test set for small datasets like CIFAR10. It may also due to inconsistent image resolutions between ImageNet () and CIFAR10 ().
Conclusion
We presented the first one-shot approach on network architecture search for channel numbers, with extensive experiments on large-scale ImageNet classification. Our proposed solution AutoSlim automates the design of efficient network architectures for resource constrained devices.