Lite-HRNet: A Lightweight High-Resolution Network

Changqian Yu, Bin Xiao, Changxin Gao, Lu Yuan, Lei Zhang, Nong Sang, Jingdong Wang

Introduction

Human pose estimation requires high-resolution representation to achieve high performance. Motivated by the increasing demand for model efficiency, this paper studies the problem of developing efficient high-resolution models under computation-limited resources.

Existing efficient networks are mainly designed from two perspectives. One is to borrow the design from classification networks, such as MobileNet and ShuffleNet , to reduce the redundancy in matrix-vector multiplication, where convolution operations dominate the cost. The other is to mediate the spatial information loss with various tricks, such as encoder-decoder architectures , and multi-branch architectures .

We first study a naive lightweight network by simply combining the shuffle block in ShuffleNet and the high-resolution design pattern in HRNet . HRNet has shown a stronger capability among large models in position-sensitive problems, e.g., semantic segmentation, human pose estimation, and object detection. It remains unclear whether high resolution helps for small models. We empirically show that the direct combination outperforms ShuffleNet, MobileNet, and Small HRNetSmall HRNet is available at https://github.com/HRNet/HRNet-Semantic-Segmentation. It simply reduces the depth and the width of the original HRNet..

To further achieve higher efficiency, we introduce an efficient unit, named conditional channel weighting, performing information exchange across channels, to replace the costly pointwise (1×11\times 1) convolution in a shuffle block. The channel weighting scheme is very efficient: the complexity is linear w.r.t the number of channels and lower than the quadratic time complexity for the pointwise convolution. For example, with the multi-resolution features of 64×64×4064\times 64\times 40 and 32×32×8032\times 32\times 80, the conditional channel weighting unit can reduce the shuffle block’s whole computation complexity by 80%80\%.

Unlike the regular convolutional kernel weights learned as model parameters, the proposed scheme weights are conditioned on the input maps and computed across channels through a lightweight unit. Thus, they contain the information in all the channel maps and serve as a bridge to exchange information through channel weighting. Furthermore, we compute the weights from parallel multi-resolution channel maps that are readily available HRNet so that the weights contain richer information and are strengthened. We call the resulting network, Lite-HRNet.

The experimental results show that Lite-HRNet outperforms the simple combination of shuffle blocks and HRNet (which we call naive Lite-HRNet). We believe that the superiority is because the computational complexity reduction is more significant than the loss of information exchange in the proposed conditional channel weighting scheme.

We simply apply the shuffle blocks to HRNet, leading a lightweight network naive Lite-HRNet. We empirically show superior performance over MobileNet, ShuffleNet, and Small HRNet.

We present an improved efficient network, Lite-HRNet. The key point is that we introduce an efficient conditional channel weighting unit to replace the costly 1×11\times 1 convolution in shuffle blocks, and the weights are computed across channels and resolutions.

Lite-HRNet is the state-of-the-art in terms of complexity and accuracy trade-off on COCO and MPII human pose estimation and easily generalized to semantic segmentation task.

Related Work

Efficient blocks for classification. Separable convolutions and group convolutions have been increasingly popular in lightweight networks, such as MobileNet , IGCV3 , and ShuffleNet . Xception and MobileNetV1 disentangle one normal convolution into depthwise convolution and pointwise convolution. MobileNetV2 and IGCV3 further combine linear bottlenecks that are about low-rank kernels. MixNet applies mixed kernels on the depthwise convolutions. EfficientHRNet introduces the mobile convolutions into HigherHRNet .

The information across channels are blocked in group convolutions and depthwise convolutions. The pointwise convolutions are heavily used to address it but are very costly in lightweight network design. To reduce the complexity, grouping 1×11\times 1 convolutions with channel shuffling or interleaving are used to keep information exchange across channels. Our proposed solution is a lightweight manner performing information exchange across channels to replace costly 1×11\times 1 convolutions.

Mediating spatial information loss. The computation complexity is positively related to spatial resolution. Reducing the spatial resolution with mediating spatial information loss is another way to improve efficiency. Encoder-decoder architecture is used to recover the spatial resolution, such as ENet and SegNet . ICNet applies different computations to different resolution inputs to reduce the whole complexity. BiSeNet decouples the detail information and context information with different lightweight sub-networks. Our solution follows the high-resolution pattern in HRNet to maintain the high-resolution representation through the whole process.

Convolutional weight generation and mixing. Dynamic filter networks dynamically generates the convolution filters conditioned on the input. Meta-Network adopts a meta-learner to generate weights to learn cross-task knowledge. CondINS and SOLOV2 apply this design to the instance segmentation task, generating the parameters of the mask sub-network for each instance. CondConv and Dynamic Convolution learn a series of weights to mix the corresponding convolution kernels for each sample, increasing the model capacity.

Attention mechanism can be regarded as a kind of conditional weight generation. SENet uses global information to learn the weights to excite or suppress the channel maps. GENet expands on this by gathering local information to exploit the contextual dependencies. CBAM exploits the channel and spatial attention to refine the features.

The proposed conditional channel weighting scheme can be, in some sense, regarded as a conditional channel-wise 1×11\times 1 convolution. Besides its cheap computation, we exploit an extra effect and use the conditional weights as the bridge to exchange information across channels.

Conditional architecture. Different from normal networks, conditional architecture can achieve dynamic width, depth, or kernels. SkipNet uses a gated network to skip some convolutional blocks to reduce complexity selectively. Spatial Transform Networks learn to warp the feature map conditioned on the input. Deformable Convolution learns the offsets for the convolution kernels conditioned on each spatial location.

Approach

Shuffle blocks. The shuffle block in ShuffleNet V2 first splits the channels into two partitions. One partition passes through a sequence of 1×11\times 1 convolution, 3×33\times 3 depthwise convolution, and 1×11\times 1 convolution, and the output is concatenated with the other partition. Finally, the concatenated channels are shuffled, as illustrated in Figure 1 (a).

HRNet. The HRNet starts from a high-resolution convolution stem as the first stage, gradually adding high-to-low resolution streams one by one as new stages. The multi-resolution streams are connected in parallel. The main body consists of a sequence of stages. In each stage, the information across resolutions is exchanged repeatedly. We follow the Small HRNet designhttps://github.com/HRNet/HRNet-Semantic-Segmentation and use fewer layers and smaller width to form our network. The stem of Small HRNet consists of two 3×33\times 3 convolutions with stride 2. Each stage in the main body contains a sequence of residual blocks and one multi-resolution fusion. Figure 2 illustrates the structure of Small HRNet.

Simple combination. We adopt the shuffle block to replace the second 3×33\times 3 convolution in the stem of Small HRNet, and replace all the normal residual blocks (formed with two 3×33\times 3 convolutions). The normal convolutions in the multi-resolution fusion are replaced by the separable convolutions , resulting in a naive Lite-HRNet.

2 Lite-HRNet

1×11\times 1 convolution is costly. The 1×11\times 1 convolution performs a matrix-vector multiplication at each position:

where X\mathsf{X} and Y\mathsf{Y} are input and output maps, and W\mathbf{W} is the 1×11\times 1 convolutional kernel. It serves a critical role of exchanging information across channels as the shuffle operation and the depthwise convolution have no effect on information exchange across channels.

The 1×11\times 1 convolution is of quadratic time complexity (Θ(C2)\Theta(C^{2})) with respect to the number (CC) of channels. The 3×33\times 3 depthwise convolution is of linear time complexity (Θ(9C)\Theta(9C)In terms of time complexity, the constant 99 should be ignored. We keep it for analysis convenience.). In the shuffle block, the complexity of two 1×11\times 1 convolutions is much higher than that of the depthwise convolution: Θ(2C2)>Θ(9C)\Theta(2C^{2})>\Theta(9C), for the usual case C>5C>5. Table 2 shows an example of the complexity comparison between 1×11\times 1 convolutions and depthwise convolutions.

Conditional channel weighting. We propose to use the element-wise weighting operation to replace the 1×11\times 1 convolution in naive Lite-HRNet, which has ss branches in the ssth stage. The element-wise weighting operation for the ssth resolution branch is written as,

where Ws\mathsf{W}_{s} is a weight map, a 33-d tensor of size Ws×Hs×CsW_{s}\times H_{s}\times C_{s}, and ⊙\odot is the element-wise multiplication operator.

The complexity is linear with respect to the channel number Θ(C)\Theta(C), and much lower than 1×11\times 1 convolution in the shuffle block.

We compute the weights by using the channels for a single resolution and the channels across all the resolutions, as shown in Figure 1 (b), and show that the weights play a role of exchanging information across channels and resolutions.

Cross-resolution weight computation. Considering the ss-th stage, there are ss parallel resolutions, and ss weight maps W1,W2,…,Ws\mathsf{W}_{1},\mathsf{W}_{2},\dots,\mathsf{W}_{s}, each for the corresponding resolution. We compute the ss weight maps from all the channels across resolutions using a lightweight function Hs(⋅)\mathcal{H}_{s}(\cdot),

where {X1,…,Xs}\{\mathsf{X}_{1},\dots,\mathsf{X}_{s}\} are the input maps for the ss resolutions. X1\mathsf{X}_{1} corresponds to the highest resolution, and Xs\mathsf{X}_{s} corresponds to the ss-th highest resolution.

We implement the lightweight function Hs(⋅)\mathcal{H}_{s}(\cdot) as following. We perform adaptive average pooling (AAP⁡\operatorname{AAP}) on {X1,X2,…,Xs−1}\{\mathsf{X}_{1},\mathsf{X}_{2},\dots,\mathsf{X}_{s-1}\}: X1′=AAP⁡(X1)\mathsf{X}_{1}^{\prime}=\operatorname{AAP}(\mathsf{X}_{1}), X2′=AAP⁡(X2)\mathsf{X}_{2}^{\prime}=\operatorname{AAP}(\mathsf{X}_{2}), …\dots, Xs−1′=AAP⁡(Xs−1)\mathsf{X}_{s-1}^{\prime}=\operatorname{AAP}(\mathsf{X}_{s-1}), in which the AAP pools any input size to a given output size Ws×HsW_{s}\times H_{s}. Then we concatenate {X1′,X2′,…,Xs−1′}\{\mathsf{X}_{1}^{\prime},\mathsf{X}_{2}^{\prime},\dots,\mathsf{X}_{s-1}^{\prime}\} and Xs\mathsf{X}_{s} together, followed by a sequence of 1×11\times 1 convolution, ReLU, 1×11\times 1 convolution, and sigmoid, generating weight maps consisting of ss partitions, W1′,W2′,…,Ws\mathsf{W}^{\prime}_{1},\mathsf{W}^{\prime}_{2},\dots,\mathsf{W}_{s} (each for one resolution):

Here, the weights at each position for each resolution depend on the channel feature at the same position from the average-pooled multi-resolution channel maps. This is why we call the scheme as cross-resolution weight computation. The s−1s-1 weight maps, W1′,W2′,…,Ws−1′\mathsf{W}_{1}^{\prime},\mathsf{W}_{2}^{\prime},\dots,\mathsf{W}^{\prime}_{s-1}, are upsampled to the corresponding resolutions, outputting W1,W2,…,Ws−1\mathsf{W}_{1},\mathsf{W}_{2},\dots,\mathsf{W}_{s-1}, for the subsequent element-wise channel weighting.

We show that the weight maps serves as a bridge for information exchange across channels and resolutions. Each element of the weight vector wsi\mathbf{w}_{si} at the position ii (from the weight map Ws\mathsf{W}_{s}) receives the information from all the input channels of all the ss resolutions at the same pooling region, which is easily verified from the operations in Equation 4. Through such a weight vector, each of the output channels at this position,

receives the information from all the input channels at the same position across all the resolutions. In other words, the channel weighting scheme plays the role as well as the 1×11\times 1 convolution in terms of exchanging information.

On the other hand, the function Hs(⋅)\mathcal{H}_{s}(\cdot) is applied on the small resolution, and thus the computation complexity is very light. Table 2 illustrates that the whole unit has much lower complexity than 1×11\times 1 convolution.

Spatial weight computation. For each resolution, we also compute the spatial weights which are homogeneous to spatial positions: the weight vector wsi\mathbf{w}_{si} at all positions are the same. The weights depend on all the pixels of the input channels in a single resolution:

Here, the function Fs(⋅)\mathcal{F}_{s}(\cdot) is implemented as: Xs→GAP⁡→FC⁡→ReLU⁡→FC⁡→sigmoid⁡→ws\mathsf{X}_{s}\rightarrow\operatorname{GAP}\rightarrow\operatorname{FC}\rightarrow\operatorname{ReLU}\rightarrow\operatorname{FC}\rightarrow\operatorname{sigmoid}\rightarrow\mathbf{w}_{s}. The global average pooling (GAP⁡\operatorname{GAP}) operator serves as a role of gathering the spatial information from all the positions.

By weighting the channels with the spatial weights, ysi=ws⊙xsi\mathbf{y}_{si}=\mathbf{w}_{s}\odot\mathbf{x}_{si}, each element in the output channels receives the contribution from all the positions of all the input channels. We compare the complexity between 1×11\times 1 convolutions and conditional channel weighting unit in Table 2.

Instantiation. The Lite-HRNet consists of a high-resolution stem and the main body to maintain the high-resolution representation. The stem has one 3×33\times 3 convolution with stride 2 and a shuffle block, as the first stage. The main body has a sequence of modularized modules. Each module consists of two conditional channel weighting blocks and one multi-resolution fusion. Each resolution branch’s channel dimensions are CC, 2C2C, 4C4C, and 8C8C, respectively. Table 1 describes the detailed structures.

Connection. The conditional channel weighting scheme shares the same philosophy to the conditional convolutions , dynamic filters , and squeeze-excite-network . Those works learn the convolution kernels or the mixture weights by sub-network conditioned on the input features for increasing the model capacity. Our method instead exploits an extra effect and uses the weights learned from all the channels as a bridge to exchange information across channels and resolutions. It can replace costly 1×11\times 1 convolutions in lightweight networks. Besides, we introduce multi-resolution information to boost weight learning.

Experiments

We evaluate our approach on two human pose estimation datasets, COCO and MPII . Following the state-of-the-art top-down framework, our approach estimates KK heatmaps to indicate the keypoint location confidence. We perform a comprehensive ablation on COCO and report the comparisons with other methods on both datasets.

Datasets. COCO has over 200K200K images and 250K250K person instances with 17 keypoints. Our models are trained on train2017 dataset (includes 57K57K images and 150K150K person instances) and validated on val2017 (includes 5K5K images) and test-dev2017 (includes 20K20K images).

The MPII Human Pose dataset contains around 25K25K images with full-body pose annotations taken from real-world activities. There are over 40K40K person instances, split 12K12K instances for testing, and others for training.

Training. The network is trained on 8 NVIDIA V100 GPUs with mini-batch size 32 per GPU. We adopt Adam optimizer with an initial learning rate of 2e−32e^{-3}.

The human detection boxes are expanded to have a fixed aspect ratio of 4: 3, and then crop the box from the images. The image size is resized to 256×192256\times 192 or 384×288384\times 288 for COCO, and 256×256256\times 256 for MPII. Each image will go through a series of data augmentation operations, containing random rotation ([<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><moseparator="true">,</mo></mrow><annotationencoding="application/x−tex">,</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3em;vertical−align:−0.1944em;"></span><spanclass="mpunct">,</span></span></span></span></span>][<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">,</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3em;vertical-align:-0.1944em;"></span><span class="mpunct">,</span></span></span></span></span>]), random scale ([0.75,1.25][0.75,1.25]), and random flipping for both datasets and additional half body data augmentation for COCO.

Testing. For COCO, following , we adopt the two-stage top-down paradigm (detect the person instance via a person detector and predict keypoints) with the person detectors provided by SimpleBaseline . For MPII, we adopt the standard testing strategy to use the provided person boxes. We estimate the heat maps via a post-gaussian filter and average the original and flipped images’ predicted heat maps. A quarter offset in the direction from the highest response to the second-highest response is applied to obtain each keypoint location.

Evaluation. We adopt the OKS-based mAP metric on COCO, where OKS⁡\operatorname{OKS} (Object Keypoint Similarity) defines the similarity between different human poses. We report standard average precision and recall scores: AP⁡\operatorname{AP} (the mean of AP⁡\operatorname{AP} scores at 10 positions, OKS⁡=0.50,0.55,…,0.90,0.95\operatorname{OKS}=0.50,0.55,\dots,0.90,0.95), AP⁡50\operatorname{AP}^{50} (AP⁡\operatorname{AP} at OKS⁡=0.50\operatorname{OKS}=0.50), AP⁡75\operatorname{AP}^{75}, AR⁡\operatorname{AR} and AR⁡50\operatorname{AR}^{50}. For MPII, we use the standard metric PCKH⁡@0.5\operatorname{PCKH}@0.5 (head-normalized probability of correct keypoint) to evaluate the performance.

2 Results

COCO val. The results of our method and other state-of-the-art methods are reported in Table 3. Our Lite-HRNet-30, trained from scratch with the 256×192256\times 192 input size, achieves 67.267.2 AP score, outperforming other light-weight methods. Compared to MobileNetV2, Lite-HRNet improves AP by 2.62.6 points with only 20%20\% GFLOPs and parameters. Compared to ShuffleNetV2, our Lite-HRNet-18 and Lite-HRNet-30 achieve 4.94.9 and 7.37.3 points gain, respectively. The complexity of our network is much lower than ShuffleNetV2. Compared to Small HRNet-W16, Lite-HRNet improves over 1010 AP points. Compared to large networks, e.g., Hourglass and CPN, our networks achieve comparable AP score with far lower complexity.

With the input size 384×288384\times 288, our Lite-HRNet-18 and Lite-HRNet-30 achieve 67.667.6 and 70.470.4 AP, respectively. Due to the efficient conditional channel weighting, Lite-HRNet achieves a better balance between accuracy and computational complexity, as shown in Figure 4 (a). Figure 3 shows the visual results on COCO from Lite-HRNet-30.

COCO test-dev. Table 4 reports the comparison results of our networks and other state-of-the-art methods. Our Lite-HRNet-30 achieves 69.769.7 AP score. It is significantly better than the small networks, and is more efficient in terms of GFLOPs and parameters. Compared to the large networks, our Lite-HRNet-30 outperforms Mask-RCNN , G-RMI , and Integral Pose Regression . Although there is a performance gap with some large networks, our networks have far lower GFLOPs and parameters.

MPII val. Table 5 reports the results of our networks and other lightweight networks. Our Lite-HRNet-18 achieves better accuracy with much lower GFLOPs than MobileNetV2, MobileNetV3, ShuffleNetV2, Small HRNet-W16. With increasing the model size, as Lite-HRNet-30, the improvement gap becomes larger. Our Lite-HRNet-30 achieves 87.087.0 PCKh⁡@0.5\operatorname{PCKh}@0.5, improving MobileNetV2, MobileNetV3, ShuffleNetV2 and Small HRNet-W16 by 1.61.6, 2.72.7, 4.24.2, and 6.86.8 points, respectively. Figure 4 (b) shows the comparison of accuracy and complexity.

3 Ablation Study

We perform ablations on two datasets: COCO and MPII, and report the results on the validation sets. The input size is 256×192256\times 192 for COCO, and 256×256256\times 256 for MPII.

Naive Lite-HRNet vs. Small HRNet. We empirically study that the shuffle blocks combined into HRNet improve performance. Figure 4 shows the comparison to Small HRNet-W1616Available from https://github.com/HRNet/HRNet-Semantic-Segmentation). We can see that naive Lite-HRNet achieves higher AP scores with lower computation complexity. On COCO val, naive Lite-HRNet improves AP over the Small HRNet-W16 by 7.3 points, and the GFLOPs and parameters are less than half. When increasing to similar parameters as wider naive Lite-HRNet, the improvement is enlarged to 10.5 points, as shown in Figure 4 (a). On MPII val, naive Lite-HRNet outperforms the Small HRNet-W16 by 5.1 points, while the wider network outperforms 6.6 points, as illustrated in Figure 4 (b).

Conditional channel weighting vs. 1×11\times 1 convolution. We compare the performance between 1×11\times 1 convolution (wider naive Lite-HRNet) and conditional channel weighting (Lite-HRNet). We simply remove one or two 1×11\times 1 convolutions in the shuffle blocks in wider naive Lite-HRNet.

Table 6 shows the studies on the COCO val and MPII val sets. 1×11\times 1 convolutions can exchange the information across channels, important to representation learning. On COCO val, dropping two 1×11\times 1 convolutions leads to 4.44.4 AP points decrease for wider naive Lite-HRNet, and also reduces almost 40%40\% FLOPs.

Our conditional channel weighting improves by 3.53.5 AP points over dropping two 1×11\times 1 convolutions with only increasing 1616M FLOPs. The AP score is comparable with the wider naive Lite-HRNet by using only 65%65\% FLOPs. Increasing the depth of Lite-HRNet leads to 1.51.5 AP improvements with similar FLOPs as wider naive Lite-HRNet and slightly larger #parameters than wider naive Lite-HRNet. The observations on MPII val are consistent (see Table 6). The AP improvement is because that our lightweight weighting operations can make the network capacity improved, by exploring the multi-resolution information using cross-resolution channel weighting and deepening the network, if taking similar FLOPs with naive version.

Spatial and multi-resolution weights. We empirically study how spatial weights and multi-resolution weights influence the performance, as shown in Table 7.

On COCO val, the spatial weights achieve 1.3 AP increase, and the multi-resolution weights obtain 1.7 point gain. The FLOPs of both operations are cheap. With both spatial and cross-resolution weights, our network improves by 3.53.5 points. Table 7 reports the consistent improvements on MPII val. These studies validate the efficiency and effectiveness of the spatial and cross-resolution weights.

We conduct the experiments by changing the arrangement order of the spatial weighting and cross-resolution weighting, which achieves similar performance. The experiments with only two spatial weights or two cross-resolution weights, lead to an almost 0.30.3 drop.

4 Application to Semantic Segmentation

Dataset. Cityscapes includes 30 classes and 19 of them are used for semantic segmentation task. The dataset contains 2,975, 500, and 1,525 finely-annotated images for training, validation, and test sets, respectively. In our experiments, we only use the fine annotated images.

Training. Our models are trained from scratch with the SGD algorithm . The initial rate is set to 1e−21e^{-2} with a “poly” learning rate strategy with a multiplier of (1−itermax_iters)0.9(1-\frac{iter}{max\_iters})^{0.9} each iteration. The total iterations are 160K with 16 batch size, and the weight decay is 5e−45e^{-4}. We randomly horizontally flip, scale ([0.5,2][0.5,2]), and crop the input images to a fixed size (512×1024512\times 1024) for training.

Results. We do not adopt testing tricks, e.g., sliding-window and multi-scale evaluation, beneficial to performance improvement but time-consuming. Table 8 shows that Lite-HRNet-18 achieves 72.8%72.8\% mIoU with only 1.951.95 GFLOPs and Lite-HRNet-30 achieves 75.3%75.3\% mIoU with 3.023.02 GFLOPs, outperforming the hand-crafted methods and NAS-based methods , and comparable with SwiftNetRN-18 that is far computationally intensive (104 GFLOPs).

Acknowledgements

This work is supported by the National Natural Science Foundation of China (No. 61433007 and 61876210).

References