MnasNet: Platform-Aware Neural Architecture Search for Mobile

Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, Quoc V. Le

Introduction

Convolutional neural networks (CNN) have made significant progress in image classification, object detection, and many other applications. As modern CNN models become increasingly deeper and larger , they also become slower, and require more computation. Such increases in computational demands make it difficult to deploy state-of-the-art CNN models on resource-constrained platforms such as mobile or embedded devices.

Given restricted computational resources available on mobile devices, much recent research has focused on designing and improving mobile CNN models by reducing the depth of the network and utilizing less expensive operations, such as depthwise convolution and group convolution . However, designing a resource-constrained mobile model is challenging: one has to carefully balance accuracy and resource-efficiency, resulting in a significantly large design space.

In this paper, we propose an automated neural architecture search approach for designing mobile CNN models. Figure 1 shows an overview of our approach, where the main differences from previous approaches are the latency aware multi-objective reward and the novel search space. Our approach is based on two main ideas. First, we formulate the design problem as a multi-objective optimization problem that considers both accuracy and inference latency of CNN models. Unlike in previous work that use FLOPS to approximate inference latency, we directly measure the real-world latency by executing the model on real mobile devices. Our idea is inspired by the observation that FLOPS is often an inaccurate proxy: for example, MobileNet and NASNet have similar FLOPS (575M vs. 564M), but their latencies are significantly different (113ms vs. 183ms, details in Table 1). Secondly, we observe that previous automated approaches mainly search for a few types of cells and then repeatedly stack the same cells through the network. This simplifies the search process, but also precludes layer diversity that is important for computational efficiency. To address this issue, we propose a novel factorized hierarchical search space, which allows layers to be architecturally different yet still strikes the right balance between flexibility and search space size.

We apply our proposed approach to ImageNet classification and COCO object detection . Figure 2 summarizes a comparison between our MnasNet models and other state-of-the-art mobile models. Compared to the MobileNetV2 , our model improves the ImageNet accuracy by 3.0% with similar latency on the Google Pixel phone. On the other hand, if we constrain the target accuracy, then our MnasNet models are 1.8×\mathbf{1.8\times} faster than MobileNetV2 and 2.3×\mathbf{2.3\times} faster thans NASNet with better accuracy. Compared to the widely used ResNet-50 , our MnasNet model achieves slightly higher (76.7%) accuracy with 4.8×\mathbf{4.8\times} fewer parameters and 10×\mathbf{10\times} fewer multiply-add operations. By plugging our model as a feature extractor into the SSD object detection framework, our model improves both the inference latency and the mAP quality on COCO dataset over MobileNetsV1 and MobileNetV2, and achieves comparable mAP quality (23.0 vs 23.2) as SSD300 with 42×\mathbf{42\times} less multiply-add operations.

To summarize, our main contributions are as follows:

We introduce a multi-objective neural architecture search approach that optimizes both accuracy and real-world latency on mobile devices.

We propose a novel factorized hierarchical search space to enable layer diversity yet still strike the right balance between flexibility and search space size.

We demonstrate new state-of-the-art accuracy on both ImageNet classification and COCO object detection under typical mobile latency constraints.

Related Work

Improving the resource efficiency of CNN models has been an active research topic during the last several years. Some commonly-used approaches include 1) quantizing the weights and/or activations of a baseline CNN model into lower-bit representations , or 2) pruning less important filters according to FLOPs , or to platform-aware metrics such as latency introduced in . However, these methods are tied to a baseline model and do not focus on learning novel compositions of CNN operations.

Another common approach is to directly hand-craft more efficient mobile architectures: SqueezeNet reduces the number of parameters and computation by using lower-cost 1x1 convolutions and reducing filter sizes; MobileNet extensively employs depthwise separable convolution to minimize computation density; ShuffleNets utilize low-cost group convolution and channel shuffle; Condensenet learns to connect group convolutions across layers; Recently, MobileNetV2 achieved state-of-the-art results among mobile-size models by using resource-efficient inverted residuals and linear bottlenecks. Unfortunately, given the potentially huge design space, these hand-crafted models usually take significant human efforts.

Recently, there has been growing interest in automating the model design process using neural architecture search. These approaches are mainly based on reinforcement learning , evolutionary search , differentiable search , or other learning algorithms . Although these methods can generate mobile-size models by repeatedly stacking a few searched cells, they do not incorporate mobile platform constraints into the search process or search space. Closely related to our work is MONAS , DPP-Net , RNAS and Pareto-NASH which attempt to optimize multiple objectives, such as model size and accuracy, while searching for CNNs, but their search process optimizes on small tasks like CIFAR. In contrast, this paper targets real-world mobile latency constraints and focuses on larger tasks like ImageNet classification and COCO object detection.

Problem Formulation

We formulate the design problem as a multi-objective search, aiming at finding CNN models with both high-accuracy and low inference latency. Unlike previous architecture search approaches that often optimize for indirect metrics, such as FLOPS, we consider direct real-world inference latency, by running CNN models on real mobile devices, and then incorporating the real-world inference latency into our objective. Doing so directly measures what is achievable in practice: our early experiments show it is challenging to approximate real-world latency due to the variety of mobile hardware/software idiosyncrasies.

Given a model mm, let ACC(m)ACC(m) denote its accuracy on the target task, LAT(m)LAT(m) denotes the inference latency on the target mobile platform, and TT is the target latency. A common method is to treat TT as a hard constraint and maximize accuracy under this constraint:

However, this approach only maximizes a single metric and does not provide multiple Pareto optimal solutions. Informally, a model is called Pareto optimal if either it has the highest accuracy without increasing latency or it has the lowest latency without decreasing accuracy. Given the computational cost of performing architecture search, we are more interested in finding multiple Pareto-optimal solutions in a single architecture search.

While there are many methods in the literature , we use a customized weighted product methodWe pick the weighted product method because it is easy to customize, but we expect methods like weighted sum should be also fine. to approximate Pareto optimal solutions, with optimization goal defined as:

where ww is the weight factor defined as:

where α\alpha and β\beta are application-specific constants. An empirical rule for picking α\alpha and β\beta is to ensure Pareto-optimal solutions have similar reward under different accuracy-latency trade-offs. For instance, we empirically observed doubling the latency usually brings about 5% relative accuracy gain. Given two models: (1) M1 has latency ll and accuracy aa; (2) M2 has latency 2l2l and 5% higher accuracy a⋅(1+5%)a\cdot(1+5\%), they should have similar reward: Reward(M2)=a⋅(1+5%)⋅(2l/T)β≈Reward(M1)=a⋅(l/T)βReward(M2)=a\cdot(1+5\%)\cdot(2l/T)^{\beta}\approx Reward(M1)=a\cdot(l/T)^{\beta}. Solving this gives β≈−0.07\beta\approx-0.07. Therefore, we use α=β=−0.07\alpha=\beta=-0.07 in our experiments unless explicitly stated.

Figure 3 shows the objective function with two typical values of (α,β)(\alpha,\beta). In the top figure with (α=0,β=−1\alpha=0,\beta=-1), we simply use accuracy as the objective value if measured latency is less than the target latency TT; otherwise, we sharply penalize the objective value to discourage models from violating latency constraints. The bottom figure (α=β=−0.07\alpha=\beta=-0.07) treats the target latency TT as a soft constraint, and smoothly adjusts the objective value based on the measured latency.

Mobile Neural Architecture Search

In this section, we will first discuss our proposed novel factorized hierarchical search space, and then summarize our reinforcement-learning based search algorithm.

As shown in recent studies , a well-defined search space is extremely important for neural architecture search. However, most previous approaches only search for a few complex cells and then repeatedly stack the same cells. These approaches don’t permit layer diversity, which we show is critical for achieving both high accuracy and lower latency.

In contrast to previous approaches, we introduce a novel factorized hierarchical search space that factorizes a CNN model into unique blocks and then searches for the operations and connections per block separately, thus allowing different layer architectures in different blocks. Our intuition is that we need to search for the best operations based on the input and output shapes to obtain better accurate-latency trade-offs. For example, earlier stages of CNNs usually process larger amounts of data and thus have much higher impact on inference latency than later stages. Formally, consider a widely-used depthwise separable convolution kernel denoted as the four-tuple (K,K,M,N)(K,K,M,N) that transforms an input of size (H,W,M)(H,W,M)We omit batch size dimension for simplicity. to an output of size (H,W,N)(H,W,N), where (H,W)(H,W) is the input resolution and M,NM,N are the input/output filter sizes. The total number of multiply-adds can be described as:

Here we need to carefully balance the kernel size KK and filter size NN if the total computation is constrained. For instance, increasing the receptive field with larger kernel size KK of a layer must be balanced with reducing either the filter size NN at the same layer, or compute from other layers.

Figure 4 shows the baseline structure of our search space. We partition a CNN model into a sequence of pre-defined blocks, gradually reducing input resolutions and increasing filter sizes as is common in many CNN models. Each block has a list of identical layers, whose operations and connections are determined by a per-block sub search space. Specifically, a sub search space for a block ii consists of the following choices:

Convolutional ops ConvOpConvOp: regular conv (conv), depthwise conv (dconv), and mobile inverted bottleneck conv .

Convolutional kernel size KernelSizeKernelSize: 3x3, 5x5.

Squeeze-and-excitation ratio SERatioSERatio: 0, 0.25.

Skip ops SkipOpSkipOp: pooling, identity residual, or no skip.

ConvOpConvOp, KernelSizeKernelSize, SERatioSERatio, SkipOpSkipOp, FiF_{i} determines the architecture of a layer, while NiN_{i} determines how many times the layer will be repeated for the block. For example, each layer of block 4 in Figure 4 has an inverted bottleneck 5x5 convolution and an identity residual skip path, and the same layer is repeated N4N_{4} times. We discretize all search options using MobileNetV2 as a reference: For #layers in each block, we search for {0, +1, -1} based on MobileNetV2; for filter size per layer, we search for its relative size in {0.75, 1.0, 1.25} to MobileNetV2 .

Our factorized hierarchical search space has a distinct advantage of balancing the diversity of layers and the size of total search space. Suppose we partition the network into BB blocks, and each block has a sub search space of size SS with average NN layers per block, then our total search space size would be SBS^{B}, versing the flat per-layer search space with size SB∗NS^{B*N}. A typical case is S=432,B=5,N=3S=432,B=5,N=3, where our search space size is about 101310^{13}, versing the per-layer approach with search space size 103910^{39}.

2 Search Algorithm

Inspired by recent work , we use a reinforcement learning approach to find Pareto optimal solutions for our multi-objective search problem. We choose reinforcement learning because it is convenient and the reward is easy to customize, but we expect other methods like evolution should also work.

Concretely, we follow the same idea as and map each CNN model in the search space to a list of tokens. These tokens are determined by a sequence of actions a1:Ta_{1:T} from the reinforcement learning agent based on its parameters θ\theta. Our goal is to maximize the expected reward:

where mm is a sampled model determined by action a1:Ta_{1:T}, and R(m)R(m) is the objective value defined by equation 2.

As shown in Figure 1, the search framework consists of three components: a recurrent neural network (RNN) based controller, a trainer to obtain the model accuracy, and a mobile phone based inference engine for measuring the latency. We follow the well known sample-eval-update loop to train the controller. At each step, the controller first samples a batch of models using its current parameters θ\theta, by predicting a sequence of tokens based on the softmax logits from its RNN. For each sampled model mm, we train it on the target task to get its accuracy ACC(m)ACC(m), and run it on real phones to get its inference latency LAT(m)LAT(m). We then calculate the reward value R(m)R(m) using equation 2. At the end of each step, the parameters θ\theta of the controller are updated by maximizing the expected reward defined by equation 5 using Proximal Policy Optimization . The sample-eval-update loop is repeated until it reaches the maximum number of steps or the parameters θ\theta converge.

Experimental Setup

Directly searching for CNN models on large tasks like ImageNet or COCO is expensive, as each model takes days to converge. While previous approaches mainly perform architecture search on smaller tasks such as CIFAR-10 , we find those small proxy tasks don’t work when model latency is taken into account, because one typically needs to scale up the model when applying to larger problems. In this paper, we directly perform our architecture search on the ImageNet training set but with fewer training steps (5 epochs). As a common practice, we reserve randomly selected 50K images from the training set as the fixed validation set. To ensure the accuracy improvements are from our search space, we use the same RNN controller as NASNet even though it is not efficient: each architecture search takes 4.5 days on 64 TPUv2 devices. During training, we measure the real-world latency of each sampled model by running it on the single-thread big CPU core of Pixel 1 phones. In total, our controller samples about 8K models during architecture search, but only 15 top-performing models are transferred to the full ImageNet and only 1 model is transferred to COCO.

For full ImageNet training, we use RMSProp optimizer with decay 0.9 and momentum 0.9. Batch norm is added after every convolution layer with momentum 0.99, and weight decay is 1e-5. Dropout rate 0.2 is applied to the last layer. Following , learning rate is increased from 0 to 0.256 in the first 5 epochs, and then decayed by 0.97 every 2.4 epochs. We use batch size 4K and Inception preprocessing with image size 224×224224\times 224. For COCO training, we plug our learned model into SSD detector and use the same settings as , including input size 320×320320\times 320.

Results

In this section, we study the performance of our models on ImageNet classification and COCO object detection, and compare them with other state-of-the-art mobile models.

Table 1 shows the performance of our models on ImageNet . We set our target latency as T=75msT=75ms, similar to MobileNetV2 , and use Equation 2 with α\alpha=β\beta=-0.07 as our reward function during architecture search. Afterwards, we pick three top-performing MnasNet models, with different latency-accuracy trade-offs from the same search experiment and compare them with existing mobile models.

As shown in the table, our MnasNet A1 model achieves 75.2% top-1 / 92.5% top-5 accuracy with 78ms latency and 3.9M parameters / 312M multiply-adds, achieving a new state-of-the-art accuracy for this typical mobile latency constraint. In particular, MnasNet runs 1.8×\mathbf{1.8\times} faster than MobileNetV2 (1.4) on the same Pixel phone with 0.5% higher accuracy. Compared with automatically searched CNN models, our MnasNet runs 2.3×\mathbf{2.3\times} faster than the mobile-size NASNet-A with 1.2% higher top-1 accuracy. Notably, our slightly larger MnasNet-A3 model achieves better accuracy than ResNet-50 , but with 4.8×\mathbf{4.8\times} fewer parameters and 10×\mathbf{10\times} fewer multiply-add cost.

Given that squeeze-and-excitation (SE ) is relatively new and many existing mobile models don’t have this extra optimization, we also show the search results without SE in the search space in Table 2; our automated approach still significantly outperforms both MobileNetV2 and NASNet.

2 Model Scaling Performance

Given the myriad application requirements and device heterogeneity present in the real world, developers often scale a model up or down to trade accuracy for latency or model size. One common scaling technique is to modify the filter size using a depth multiplier . For example, a depth multiplier of 0.5 halves the number of channels in each layer, thus reducing the latency and model size. Another common scaling technique is to reduce the input image size without changing the network.

Figure 5 compares the model scaling performance of MnasNet and MobileNetV2 by varying the depth multipliers and input image sizes. As we change the depth multiplier from 0.35 to 1.4, the inference latency also varies from 20ms to 160ms. As shown in Figure 5(a), our MnasNet model consistently achieves better accuracy than MobileNetV2 for each depth multiplier. Similarly, our model is also robust to input size changes and consistently outperforms MobileNetV2 (increaseing accuracy by up to 4.1%) across all input image sizes from 96 to 224, as shown in Figure 5(b).

In addition to model scaling, our approach also allows searching for a new architecture for any latency target. For example, some video applications may require latency as low as 25ms. We can either scale down a baseline model, or search for new models specifically targeted to this latency constraint. Table 4 compares these two approaches. For fair comparison, we use the same 224x224 image sizes for all models. Although our MnasNet already outperforms MobileNetV2 with the same scaling parameters, we can further improve the accuracy with a new architecture search targeting a 22ms latency constraint.

3 COCO Object Detection Performance

For COCO object detection , we pick the MnasNet models in Table 2 and use them as the feature extractor for SSDLite, a modified resource-efficient version of SSD . Similar to , we compare our models with other mobile-size SSD or YOLO models.

Table 3 shows the performance of our MnasNet models on COCO. Results for YOLO and SSD are from , while results for MobileNets are from . We train our models on COCO trainval35k and evaluate them on test-dev2017 by submitting the results to COCO server. As shown in the table, our approach significantly improve the accuracy over MobileNet V1 and V2. Compare to the standard SSD300 detector , our MnasNet model achieves comparable mAP quality (23.0 vs 23.2) as SSD300 with 7.4×\mathbf{7.4\times} fewer parameters and 42×\mathbf{42\times} fewer multiply-adds.

Ablation Study and Discussion

In this section, we study the impact of latency constraint and search space, and discuss MnasNet architecture details and the importance of layer diversity.

Our multi-objective search method allows us to deal with both hard and soft latency constraints by setting α\alpha and β\beta to different values in the reward equation 2. Figure 6 shows the multi-objective search results for typical α\alpha and β\beta. When α=0,β=−1\alpha=0,\beta=-1, the latency is treated as a hard constraint, so the controller tends to focus more on faster models to avoid the latency penalty. On the other hand, by setting α=β=−0.07\alpha=\beta=-0.07, the controller treats the target latency as a soft constraint and tries to search for models across a wider latency range. It samples more models around the target latency value at 75ms, but also explores models with latency smaller than 40ms or greater than 110ms. This allows us to pick multiple models from the Pareto curve in a single architecture search as shown in Table 1.

2 Disentangling Search Space and Reward

To disentangle the impact of our two key contributions: multi-objective reward and new search space, Figure 5 compares their performance. Starting from NASNet , we first employ the same cell-base search space and simply add the latency constraint using our proposed multiple-object reward. Results show it generates a much faster model by trading the accuracy to latency. Then, we apply both our multi-objective reward and our new factorized search space, and achieve both higher accuracy and lower latency, suggesting the effectiveness of our search space.

3 MnasNet Architecture and Layer Diversity

Figure 7(a) illustrates our MnasNet-A1 model found by our automated approach. As expected, it consists of a variety of layer architectures throughout the network. One interesting observation is that our MnasNet uses both 3x3 and 5x5 convolutions, which is different from previous mobile models that all only use 3x3 convolutions.

In order to study the impact of layer diversity, Table 6 compares MnasNet with its variants that only repeat a single type of layer (fixed kernel size and expansion ratio). Our MnasNet model has much better accuracy-latency trade-offs than those variants, highlighting the importance of layer diversity in resource-constrained CNN models.

Conclusion

This paper presents an automated neural architecture search approach for designing resource-efficient mobile CNN models using reinforcement learning. Our main ideas are incorporating platform-aware real-world latency information into the search process and utilizing a novel factorized hierarchical search space to search for mobile models with the best trade-offs between accuracy and latency. We demonstrate that our approach can automatically find significantly better mobile models than existing approaches, and achieve new state-of-the-art results on both ImageNet classification and COCO object detection under typical mobile inference latency constraints. The resulting MnasNet architecture also provides interesting findings on the importance of layer diversity, which will guide us in designing and improving future mobile CNN models.

Acknowledgments

We thank Barret Zoph, Dmitry Kalenichenko, Guiheng Zhou, Hongkun Yu, Jeff Dean, Megan Kacholia, Menglong Zhu, Nan Zhang, Shane Almeida, Sheng Li, Vishy Tirumalashetty, Wen Wang, Xiaoqiang Zheng, and the larger device automation platform team, TensorFlow Lite, and Google Brain team.

References