Latency-aware Spatial-wise Dynamic Networks

Yizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue, Shiji Song, Guangyu Sun, Gao Huang

Introduction

Dynamic neural networks have attracted great research interests in recent years. Compared to static models which treat different inputs equally during inference, dynamic networks can allocate the computation in a data-dependent manner. For example, they can conditionally skip the computation of network layers or convolutional channels , or perform spatially adaptive inference on the most informative image regions (e.g. the foreground areas) . Spatial-wise dynamic networks, which typically decide whether to compute each feature pixel with masker modules (Figure 1 (a)), have shown promising results in improving the inference efficiency of convolution neural networks (CNNs).

Despite the remarkable theoretical efficiency achieved by spatial-wise dynamic networks , researchers have found it challenging to translate the theoretical results into realistic speedup, especially on some multi-core processors, e.g., GPUs . The challenges are two-fold: 1) most previous approaches perform spatially adaptive inference at the finest granularity: every pixel is flexibly decided whether to be computed or not. Such flexibility induces non-contiguous memory access and requires specialized scheduling strategies (Figure 1 (b)); 2) the existing literature has only adopted the hardware-agnostic FLOPs (floating-point operations) as an inaccurate proxy for the efficiency, lacking latency-aware guidance on the algorithm design. For dynamic networks, the adaptive computation with sub-optimal scheduling strategies further enlarges the discrepancy between the theoretical FLOPs and the practical latency. Note that it has been validated by previous works that the latency on CPUs has a strong correlation with FLOPs . Therefore, we mainly focus on the GPU platform in this paper, which is more challenging and less explored.

We address the above challenges by proposing a latency-aware spatial-wise dynamic network (LASNet). Three key factors to the inference latency are considered: the algorithm, the scheduling strategy, and the hardware properties. Given a target hardware device, we directly use the latency, rather than the FLOPs, to guide our algorithm design and scheduling optimization (Figure 1 (c)).

Because the memory access pattern and the scheduling strategies in our dynamic operators differ from those in static networks, the libraries developed for static models (e.g. cuDNN) are sub-optimal for the acceleration of dynamic models. Without the support of libraries, each dynamic operator requires scheduling optimization, code optimization, compiling, and deployment for each device. Therefore, it is laborious to evaluate the network latency on different hardware platforms. To this end, we propose a novel latency prediction model to efficiently estimate the realistic latency of a network by simultaneously considering the aforementioned three factors. Compared to the hardware-agnostic FLOPs, our predicted latency can better reflect the practical efficiency of dynamic models.

Guided by this latency prediction model, we establish our latency-aware spatial-wise dynamic network (LASNet), which adaptively decides whether to allocate computation on feature patches instead of pixels (Figure 2 top). We name this paradigm as spatially adaptive inference at a coarse granularity. While less flexible than the pixel-level adaptive computation in previous works , it facilitates more contiguous memory access, benefiting the realistic speedup on hardware. The scheduling strategy and the implementation are further ameliorated for faster inference.

It is worth noting that LASNet is designed as a general framework in two aspects: 1) the coarse-grained spatially adaptive inference paradigm can be conveniently implemented in various CNN backbones, e.g., ResNets , DenseNets and RegNets ; and 2) the latency predictor is an off-the-shell tool which can be directly used for various computing platforms (e.g. server GPUs and edge devices).

We evaluate the performance of LASNet on multiple CNN architectures on image classification, object detection, and instance segmentation tasks. Experiment results show that our LASNet improves the efficiency of deep CNNs both theoretically and practically. For example, the inference latency of ResNet-101 on ImageNet is reduced by 36% and 46% on an Nvidia Tesla V100 GPU and an Nvidia Jetson TX2 GPU, respectively, without sacrificing the accuracy. Moreover, the proposed method outperforms various lightweight networks in a low-FLOPs regime.

Our main contributions are summarized as follows:

We propose LASNet, which performs coarse-grained spatially adaptive inference guided by the practical latency instead of the theoretical FLOPs. To the best of our knowledge, LASNet is the first framework that directly considers the real latency in the design phase of dynamic neural networks;

We propose a latency prediction model, which can efficiently and accurately estimate the latency of dynamic operators by simultaneously considering the algorithm, the scheduling strategy, and the hardware properties;

Experiments on image classification and downstream tasks verify that our proposed LASNet can effectively improve the practical efficiency of different CNN architectures.

Related works

Spatial-wise dynamic network is a common type of dynamic neural networks . Compared to static models which treat different feature locations evenly during inference, these networks perform spatially adaptive inference on the most informative regions (e.g., foregrounds), and reduce the unnecessary computation on less important areas (e.g., backgrounds). Existing works mainly include three levels of dynamic computation: resolution level , region level and pixel level . The former two generally manipulate the network inputs or require special architecture design . In contrast, pixel-level dynamic networks can flexibly skip the convolutions on certain feature pixels in arbitrary CNN backbones . Despite its remarkable theoretical efficiency, pixel-wise dynamic computation brings considerable difficulty to achieving realistic speedup on multi-core processors, e.g., GPUs. Compared to the previous approaches which only focus on reducing the theoretical computation, we propose to directly use the latency to guide our algorithm design and scheduling optimization.

Hardware-aware network design. To bridge the gap between theoretical and practical efficiency of deep models, researchers have started to consider the real latency in the network design phase. There are two lines of works in this direction. One directly performs speed tests on targeted devices, and summarizes some guidelines to facilitate hand-designing lightweight models . The other line of work searches for fast models using the neural architecture search (NAS) technique . However, all existing works try to build static models, which have intrinsic computational redundancy by treating different inputs in the same way. However, speed tests for dynamic operators on different hardware devices can be very laborious and impractical. In contrast, our proposed latency prediction model can efficiently estimate the inference latency on any given computing platforms by simultaneously considering algorithm design, scheduling strategies and hardware properties.

Methodology

In this section, we first introduce the preliminaries of spatially adaptive inference, and then demonstrate the architecture design of our LASNet. The latency prediction model is then explained, which guides the granularity settings and the scheduling optimization for LASNet. We further present the implementation improvements for faster inference, followed by the training strategies.

Scheduling strategy. During inference, the current scheduling strategy for spatial-wise dynamic convolutions generally involve three steps (Figure 1 (b)): 1) gathering, which re-organizes the selected pixels (if the convolution kernel size is greater than 1×11\times 1, the neighbors are also required) along the batch dimension; 2) computation, which performs convolution on the gathered input; and 3) scattering, which fills the computed pixels on their corresponding locations of the output feature.

Limitations. Compared to performing convolutions on the entire feature map, the aforementioned scheduling strategy reduces the computation while bringing considerable overhead to memory access due to the mask generation and the non-contiguous memory access. Such overhead would increase the overall latency, especially when the granularity of dynamic convolution is at the finest pixel level.

2 Architecture design

Differences to existing works. Without using the interpolation operation or the carefully designed two-branch structure , the proposed block architecture is simple and sufficiently general to be plugged into most backbones with minimal modification. Our formulation is mostly similar to that in , which could be viewed as a variant of our method with the spatial granularity SS=1 for all blocks. Instead of performing spatially adaptive inference at the finest pixel level, our granularity SS is optimized under the guidance of our latency prediction model (details are presented in the following Sec. 4.2) to achieve realistic speedup on target computing platforms.

3 Latency prediction model

Hardware modeling. We model a hardware device as multiple processing engines (PEs), and parallel computation can be executed on these PEs. As shown in Figure 4, we model the memory system as a three-level structure : 1) off-chip memory, 2) on-chip global memory, and 3) memory in PE. Such a hardware model enables us to accurately predict the cost on both data movement and computation.

Latency prediction. When simulating the data movement procedure, the efficiency of non-contiguous memory accesses under different granularity SS settings is considered. As for the computation latency, it is important to adopt a proper scheduling strategy to increase the parallelism of computation. Therefore, we search for the optimal scheduling (the configuration of tiling and in-PE parallel computing) of dynamic operations to maximize the utilization of hardware resources. A more detailed description of our latency prediction model is presented in Appendix A.

Empirical validation. We take the first block in ResNet-101 as an example and vary the activation rate rr to evaluate the performance of our prediction model. The comparison between our predictions and the real testing latency on the Nvidia V100 GPU is illustrated in Figure 4, from which we can observe that our predictor can accurately estimate the real latency in a wide range of activation rates.

4 Implementation details

We use general optimization methods like fusing activation functions and batch normalization layers into convolution layers. We also optimize the specific operators in our spatial-wise dynamic convolutional blocks as follows (see also Figure 2 for an overview).

Fusing the gather operation and the dynamic convolution. Traditional approaches first gather the input pixels of the first dynamic convolution in a block (Figure 1 (b)). The gather operation is also a memory-bounded operation. Furthermore, when the size of the convolution kernel exceeds 1×\times1, the area of input patches may overlap, resulting in repeated memory load/store. We fuse the gather operation into the dynamic convolution to reduce the memory access.

Fusing the scatter operation and the add operation. Traditional approaches scatter the output pixels of the last dynamic convolution, and then execute the element-wise addition (Figure 1 (b)). We fuse these two operators to reduce the memory access. The ablation study in Sec. 4.4 validates the effectiveness of the proposed fusing methods.

5 Training

where τ\tau is the Softmax temperature. Following the common practice , we let τ\tau decrease exponentially from 5.0 to 0.1 in training to facilitate the optimization of maskers.

We further propose to leverage the static counterparts of our dynamic networks as “teachers” to guide the optimization procedure. Let y\mathbf{y} and y′\mathbf{y}^{\prime} denote the output logits of a dynamic model (“student”) and its “teacher”, respectively. Our final loss can be written as

Experiments

In this section, we first introduce the experiment settings in Sec. 4.1. Then the latency of different granularity settings are analyzed in Sec. 4.2. The performance of our LASNet on ImageNet is further evaluated in Sec. 4.3, followed by the ablation studies in Sec. 4.4. Visualization results are illustrated in Sec. 4.5, and we finally validate our method on the object detection task (Sec. 4.6). The results on the instance segmentation task are presented in . For simplicity, we add “LAS-” as a prefix before model names to denote our LASNet, e.g., LAS-ResNet-50.

Latency prediction. Various types of hardware platforms are tested, including a server GPU (Tesla V100), a desktop GPU (GTX1080) and edge devices (e.g., Nvidia Nano and Jetson TX2). The major properties considered by our latency prediction model include the number of processing engines (#PE), the floating-point computation in a processing engine (#FP32), the frequency and the bandwidth. It can be observed that the server GPUs generally have a larger #PE than the IoT devices. The batch size is set as 1 for all dynamic models and computing platforms.

Image classification. The image classification experiments are conducted on the ImageNet dataset. Following , we initialize the backbone parameter from a pre-trained checkpointWe use the torchvision pre-trained models at https://pytorch.org/vision/stable/models.html., and finetune the whole network for 100 epochs with the loss function in Eq. (2). We fix α=10,β=0.5\alpha=10,\beta=0.5 and T=4.0T=4.0 for all dynamic models. More details are provided in Appendix B.

2 Latency prediction results

In this subsection, we present the latency prediction results of the spatial-wise dynamic convolutional blocks in two different models: LAS-ResNet-101 (on V100) and LAS-RegNetY-800MF (on TX2). All the blocks have the bottleneck structure with different channel numbers and convolution groups, and the RegNetY is equipped with Squeeze-and-Excitation (SE) modules.

3 ImageNet classification results

We now empirically evaluate our proposed LASNet on the ImageNet dataset. The network performance is measured in terms of the trade-off between classification accuracy and inference efficiency. Both theoretical (i.e. FLOPs) and practical efficiency (i.e. latency) are tested in our experiments.

We first establish our LASNet based on the standard ResNets . Specifically, we build LAS-ResNet-50 and LAS-ResNet-101 by plugging our maskers in the two common ResNet structures.

3.2 Lightweight baseline comparison: RegNets

We further evaluate our LASNet in lightweight CNN architectures, i.e. RegNets-Y . Two different sized models are tested: RegNetY-400MF and RegNetY-800MF. Compared baselines include other types of efficient models, e.g., MobileNets-v2 , ShuffletNets-v2 and CondenseNets .

The results are presented in Figure 7 (b). The x-axis for the three sub-figures are the FLOPs, the latency on TX2, and the latency on the Nvidia nano GPU, respectively. We can observe that our method outperforms various types of static models in terms of the trade-off between accuracy and efficiency. More results on image classification are provided in Appendix C.2.

4 Ablation studies

We conduct ablation studies to validate the effectiveness of our coarse-grained spatially adaptive inference (Sec. 3.2) and operator fusion operations (Sec. 3.4).

Operator fusion. We investigate the effect of our operator fusion introduced in Sec. 3.4. One convolutional block in the first stage of a LAS-ResNet-101 (SS=4, rr=0.6) is tested. The results in Table 1 validate that every step of operator fusion benefits the practical latency of a block, as the overhead on memory access is effectively reduced. Especially, the fusion of the masker operation and the first convolution is crucial to reducing the latency.

5 Visualization

6 COCO Object detection

7 COCO instance segmentation

Conclusion

In this paper, we propose to build latency-aware spatial-wise dynamic networks (LASNet) under the guidance of a latency prediction model. By simultaneously considering the algorithm, the scheduling strategy and the hardware properties, we can efficiently estimate the practical latency of spatial-wise dynamic operators on arbitrary computing platforms. Based on the empirical analysis on the relationship between the latency and the granularity of spatially adaptive inference, we optimize both the algorithm and the scheduling strategies to achieve realistic speedup on many multi-core processors, e.g., the Tesla V100 GPU and the Jetson TX2 GPU. Experiments on image classification, object detection and instance segmentation tasks validate that the proposed method significantly improves the practical efficiency of deep CNNs, and outperforms various competing approaches.

Acknowledgement

This work is supported in part by the National Key R&D Program of China under Grant 2020AAA0105200, the National Natural Science Foundation of China under Grants 62022048 and the Tsinghua University-China Mobile Communications Group Co.,Ltd. Joint Institute..

References

Appendix

Appendix A Latency prediction model.

As the dynamic operators in our method have not been supported by current deep learning libraries, we propose a latency prediction model to efficiently estimate the real latency of these operators on hardware device. The inputs of the latency prediction model include: 1) the structural configuration of a convolutional block, 2) its activation rate rr which decides the computation amount, 3) the spatial granularity SS, and 4) the hardware properties mentioned in Table 4. The latency of a dynamic block is predicted as follows.

Input/output shape definition. The first step of predicting the latency of an operation is to calculate the shape of input and output. Taking the gather-conv2 operation as an example, the input of this operation is the activation with the shape of Cin×H×WC_{in}\times H\times W, where CinC_{in} is the number of input channels, and HH and WW are the resolution of the feature map. The shape of the output tensor is P×Cout×S×SP\times C_{out}\times S\times S, where PP is the number of output patches, CoutC_{out} is the number of output channels and SS is the spatial granularity. Note that PP is obtained based on the output of our maskers.

Operation-to-hardware mapping. Next, we map the operations to hardware. As is mentioned in the paper, we model a hardware device as multiple processing engines (PEs). We assign the computation of each element in the output feature map to a PE. Specifically, we consecutively split the output feature map into multiple tiles. The shape of each tile is TP×TC×TS1×TS2T_{P}\times T_{C}\times T_{S1}\times T_{S2}. These split tiles are assigned to multiple PEs. The computation of the elements in each tile is executed in a PE. We can configure different shapes of tiles. In order to determine the optimal shape of the tile, we make a search space of different tile shapes. The tile shape has 4 dimensions. The candidates of each dimension are power-of-2 and do not exceed the corresponding dimension of the feature map.

The latency of data movement is affected by the granularity SS: when the granularity SS is small, the same input data has a higher probability of being sent to multiple PEs to compute different output patches, which significantly increases the number of on-chip memory movement. And due to the small amount of data transmitted each time and the data is randomly distributed, the efficiency of data movement will be low. This accounts for our experiment results in the paper that a larger SS will effectively improve the practical efficiency.

To summarize, our latency prediction model can predict the real latency of dynamic operators by considering both the data movement cost and the computation cost. Guided by the latency prediction model, we propose our LASNets with coarse-grained spatially adaptive inference (S>1S>1). It is validated in our paper that LASNets achieve better efficiency than previous approaches (S=1S=1), as it effectively reduces the data movement latency, which is rarely considered by other researchers.

Appendix B Detailed experimental settings

In this section, we present the detailed experiment settings which are not provided in the main paper due to the page limit.

Hardware properties considered by our latency prediction model include the number of processing engines (#PE), the floating-point computation in a processing engine (#FP32), the frequency and the bandwidth. We test four types of hardware devices, and their properties are listed in Table 4.

It could be found that the server GPU V100 is the most powerful hardware device, especially with the most number of processing engines (#PE). Therefore, spatially adaptive inference could easily fall into a memory-bounded operation on V100 due to its high parallelism. Our experiment results in Figure 7 (a) and Figure 8 in the paper can reflect this phenomenon: the more flexibility the computation is, the harder to improve the practical efficiency.

In contrast, on the less powerful computing devices such as the IoT devices, the real acceleration is close to the theoretical effect (compare Figure 7 (a) left with Figure 7 (a) middle).

1) Fusing the masker and the first convolution. We mentioned in Sec. 3.4 of the paper that the masker operation is fused with the first 1×\times1 convolution in a block to reduce the cost on memory access. This is feasible because the two operators share the same input feature, and their convolutional kernel sizes are both 1×\times1.

Afterwards, we fuse the masker with the first convolution layer by performing once convolution whose output channel number is C+1C+1, where CC is the original output width of the first convolution. The output of this step is split into a feature map (for further computation) and a mask (for obtaining the index for gathering). Such operator fusion avoids the repeated reading the input feature, and helps reduce the inference latency (see Table 1 in the paper).

2) Fusing the gather operation and the dynamic convolution. To facilitate the scheduling on hardware devices with multiple PEs, the masker generates the indices of activated patches instead of sparse mask at inference time. In this way, it is easy to evenly distribute the computation of output patches to different PEs, thus avoiding unbalanced computation of PEs. Each element in the indices represents the index of an activated patch. PE fetches the input data from the corresponding positions on the feature map according to the index. The output patches could be densely stored in memory. Such operator fusion benefits the contiguous memory access and parallel computation on multiple PEs.

3) Fusing the scatter operation and the add operation.

Similar to the previous operation, each PE fetches a tile of data from the residual feature map according to the index, adds them with the corresponding feature map from previous dynamic convolution, and then stores the results to the corresponding position on the residual feature map according to the index. This optimization can significantly reduce the costs on memory access.

Speed test. We test the latency on real hardware devices to evaluate the accuracy of our latency prediction model. On GPUs, we use TVM and CUDA (version 11.6) for code generation and compilation respectively. The results in Fig. 4 of the paper validate the effectiveness of our model.

B.2 ImageNet classification

As mentioned in the paper, we use pre-trained CNN models in the official torchvision website to initialize our backbone parameters, and finetune the overall models for 100 epochs. The initial learning rate is set as 0.01×\timesbatch size/128, and decays with a cosine shape. The training batch size is determined on the model size and the GPU memory. For example, we train our LAS-ResNet-101 on 8 RTX 3090 GPUs with the batch size of 512, and the batch size for LAS-ResNet-50 is doubled. We use the same weight decay and the standard data augmentation as in the RegNet paper . For our own hyper-parameter τ\tau in Eq. (1) of the paper, this Gumbel temperature τ\tau exponentially decreases from 5 to 0.1 in the training procedure. For the training hyper-parameter in Eq. (2), we simply fix α=10,β=0.5\alpha=10,\beta=0.5 and T=4.0T=4.0 for all dynamic models. We conduct a very simple grid search with a RegNet for β∈{0.3,0.5}\beta\in\{0.3,0.5\} and T∈{1.0,4.0}T\in\{1.0,4.0\} to determine their values.

B.3 COCO object detection & instance segmentation

We use the standard setting suggested in , except that we decrease the learning rate for our pre-trained backbone network. We simply set a learning rate multiplier 0.5 for Faster R-CNN , 0.2 for RetinaNet and 0.5 for Mask R-CNN . As for the additional loss items, the hyper-parameters are kept the same as training our classification models, except that the temperature is fixed as 0.1 in the 12 training epochs.

Appendix C More experimental results

In this section, we report more experimental results which are not presented in the main paper.

C.2 ImageNet classification

C.3 More visualization results

Appendix D Limitations

The current limitations of our work include the following aspects:

1) the latency-ware co-designing framework is only constructed for spatial-wise dynamic networks. Support for more types of dynamic models (e.g. channel skipping) will be explored in the future;

2) To achieve faster inference, the spatial masks are defined the same for all input/output channels, which might limit the flexibility of adaptive inference. Future work may explore more flexible forms of dynamic computation;

3) The combination with other acceleration techniques such as Winograd, and the implementation on more CNN backbones may be worth studying in the future.

Social impact. Our work can help reduce the inference cost of deep CNNs, but the training of our models might potentially increase the carbon emissions.