High-Resolution Representations for Labeling Pixels and Regions

Ke Sun, Yang Zhao, Borui Jiang, Tianheng Cheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, Jingdong Wang

Introduction

Deeply-learned representations have been demonstrated to be strong and achieved state-of-the-art results in many vision tasks. There are two main kinds of representations: low-resolution representations that are mainly for image classification, and high-resolution representations that are essential for many other vision problems, e.g., semantic segmentation, object detection, human pose estimation, etc. The latter one, the interest of this paper, remains unsolved and is attracting a lot of attention.

There are two main lines for computing high-resolution representations. One is to recover high-resolution representations from low-resolution representations outputted by a network (e.g., ResNet) and optionally intermediate medium-resolution representations, e.g., Hourglass , SegNet , DeconvNet , U-Net , and encoder-decoder . The other one is to maintain high-resolution representations through high-resolution convolutions and strengthen the representations with parallel low-resolution convolutions . In addition, dilated convolutions are used to replace some strided convolutions and associated regular convolutions in classification networks to compute medium-resolution representations .

We go along the research line of maintaining high-resolution representations and further study the high-resolution network (HRNet), which is initially developed for human pose estimation , for a broad range of vision tasks. An HRNet maintains high-resolution representations by connecting high-to-low resolution convolutions in parallel and repeatedly conducting multi-scale fusions across parallel convolutions. The resulting high-resolution representations are not only strong but also spatially precise.

We make a simple modification by exploring the representations from all the high-to-low resolution parallel convolutions other than only the high-resolution representations in the original HRNet . This modification adds a small overhead and leads to stronger high-resolution representations. The resulting network is named as HRNetV22. We empirically show the superiority to the original HRNet.

We apply our proposed network to semantic segmentation/facial landmark detection through estimating segmentation maps/facial landmark heatmaps from the output high-resolution representations. In semantic segmentation, the proposed approach achieves state-of-the-art results on PASCAL Context, Cityscapes, and LIP with similar model sizes and lower computation complexity. In facial landmark detection, our approach achieves overall best results on four standard datasets: AFLW, COFW, 300300W, and WFLW.

In addition, we construct a multi-level representation from the high-resolution representation, and apply it to the Faster R-CNN object detection framework and its extended frameworks, Mask R-CNN and Cascade R-CNN . The results show that our method gets great detection performance improvement and in particular dramatic improvement for small objects. With single-scale training and testing, the proposed approach achieves better COCO object detection results than existing single-model methods.

Related Work

Strong high-resolution representations play an essential role in pixel and region labeling problems, e.g., semantic segmentation, human pose estimation, facial landmark detection, and object detection. We review representation learning techniques developed mainly in the semantic segmentation, facial landmark detection and object detection areasThe techniques developed for human pose estimation are reviewed in ., from low-resolution representation learning, high-resolution representation recovering, to high-resolution representation maintaining.

Learning low-resolution representations. The fully-convolutional network (FCN) approaches compute low-resolution representations by removing the fully-connected layers in a classification network, and estimate from their coarse segmentation confidence maps. The estimated segmentation maps are improved by combining the fine segmentation score maps estimated from intermediate low-level medium-resolution representations , or iterating the processes . Similar techniques have also been applied to edge detection, e.g., holistic edge detection .

The fully convolutional network is extended, by replacing a few (typically two) strided convolutions and the associated convolutions with dilated convolutions, to the dilation version, leading to medium-resolution representations . The representations are further augmented to multi-scale contextual representations through feature pyramids for segmenting objects at multiple scales.

Recovering high-resolution representations. An upsample subnetwork, like a decoder, is adopted to gradually recover the high-resolution representations from the low-resolution representations outputted by the downsample process. The upsample subnetwork could be a symmetric version of the downsample subnetwork, with skipping connection over some mirrored layers to transform the pooling indices, e.g., SegNet and DeconvNet , or copying the feature maps, e.g., U-Net and Hourglass , encoder-decoder , FPN , and so on. The full-resolution residual network introduces an extra full-resolution stream that carries information at the full image resolution, to replace the skip connections, and each unit in the downsample and upsample subnetworks receives information from and sends information to the full-resolution stream.

The asymmetric upsample process is also widely studied. RefineNet improves the combination of upsampled representations and the representations of the same resolution copied from the downsample process. Other works include: light upsample process ; light downsample and heavy upsample processes , recombinator networks ; improving skip connections with more or complicated convolutional units , as well as sending information from low-resolution skip connections to high-resolution skip connections or exchanging information between them ; studying the details the upsample process ; combining multi-scale pyramid representations ; stacking multiple DeconvNets/U-Nets/Hourglass with dense connections .

Maintaining high-resolution representations. High-resolution representations are maintained through the whole process, typically by a network that is formed by connecting multi-resolution (from high-resolution to low-resolution) parallel convolutions with repeated information exchange across parallel convolutions. Representative works include GridNet , convolutional neural fabrics , interlinked CNNs , and the recently-developed high-resolution networks (HRNet) that is our interest.

The two early works, convolutional neural fabrics and interlinked CNNs , lack careful design on when to start low-resolution parallel streams and how and when to exchange information across parallel streams, and do not use batch normalization and residual connections, thus not showing satisfactory performance.

GridNet is like a combination of multiple U-Nets and includes two symmetric information exchange stages: the first stage only passes information from high-resolution to low-resolution, and the second stage only passes information from low-resolution to high-resolution. This limits its segmentation quality.

Learning High-Resolution Representations

The high-resolution network , which we named HRNetV11 for convenience, maintains high-resolution representations by connecting high-to-low resolution convolutions in parallel, where there are repeated multi-scale fusions across parallel convolutions.

Architecture. The architecture is illustrated in Figure 1. There are four stages, and the 22nd, 33rd and 44th stages are formed by repeating modularized multi-resolution blocks. A multi-resolution block consists of a multi-resolution group convolution and a multi-resolution convolution which is illustrated in Figure 2 (a) and (b). The multi-resolution group convolution is a simple extension of the group convolution, which divides the input channels into several subsets of channels and performs a regular convolution over each subset over different spatial resolutions separately.

The multi-resolution convolution is depicted in Figure 2 (b). It resembles the multi-branch full-connection manner of the regular convolution, illustrated in in Figure 2 (c). A regular convolution can be divided as multiple small convolutions as explained in . The input channels are divided into several subsets, and the output channels are also divided into several subsets. The input and output subsets are connected in a fully-connected fashion, and each connection is a regular convolution. Each subset of output channels is a summation of the outputs of the convolutions over each subset of input channels.

The differences lie in two-fold. (i) In a multi-resolution convolution each subset of channels is over a different resolution. (ii) The connection between input channels and output channels needs to handle The resolution decrease is implemented in by using several 22-strided 3×33\times 3 convolutions. The resolution increase is simply implemented in by bilinear (nearest neighbor) upsampling.

Modification. In the original approach HRNetV11, only the representation (feature maps) from the high-resolution convolutions in are outputted, which is illustrated in Figure 3 (a). This means that only a subset of output channels from the high-resolution convolutions is exploited and other subsets from low-resolution convolutions are lost.

We make a simple yet effective modification by exploiting other subsets of channels outputted from low-resolution convolutions. The benefit is that the capacity of the multi-resolution convolution is fully explored. This modification only adds a small parameter and computation overhead.

We rescale the low-resolution representations through bilinear upsampling to the high resolution, and concatenate the subsets of representations, illustrated in Figure 3 (b), resulting in the high-resolution representation, which we adopt for estimating segmentation maps/facial landmark heatmaps. In application to object detection, we construct a multi-level representation by downsampling the high-resolution representation with average pooling to multiple levels, which is depicted in Figure 3 (c). We name the two modifications as HRNetV22 and HRNetV22p, respectively, and empirically compare them in Section 4.4.

Instantiation We instantiate the network using a similar manner as HRNetV11 https://github.com/leoxiaobin/deep-high-resolution-net.pytorch. The network starts from a stem that consists of two strided 3×33\times 3 convolutions decreasing the resolution to 1/41/4. The 11st stage contains 44 residual units where each unit is formed by a bottleneck with the width 6464, and is followed by one 3×33\times 3 convolution reducing the width of feature maps to CC. The 22nd, 33rd, 44th stages contain 11, 44, 33 multi-resolution blocks, respectively. The widths (number of channels) of the convolutions of the four resolutions are CC, 2C2C, 4C4C, and 8C8C, respectively. Each branch in the multi-resolution group convolution contains 44 residual units and each unit contains two 3×33\times 3 convolutions in each resolution.

In applications to semantic segmentation and facial landmark detection, we mix the output representations (Figure 3 (b)), from all the four resolutions through a 1×11\times 1 convolution, and produce a 15C15C-dimensional representation. Then, we pass the mixed representation at each position to a linear classifier/regressor with the softmax/MSE loss to predict the segmentation maps/facial landmark heatmaps. For semantic segmentation, the segmentation maps are upsampled (44 times) to the input size by bilinear upsampling for both training and testing. In application to object detection, we reduce the dimension of the high-resolution representation to 256256, similar to FPN , through a 1×11\times 1 convolution before forming the feature pyramid in Figure 3 (c).

Experiments

Semantic segmentation is a problem of assigning a class label to each pixel. We report the results over two scene parsing datasets, PASCAL Context and Cityscapes , and a human parsing dataset, LIP . The mean of class-wise intersection over union (mIoU) is adopted as the evaluation metric.

Cityscapes. The Cityscapes dataset contains 5,0005,000 high quality pixel-level finely annotated scene images. The finely-annotated images are divided into 2,975/500/1,5252,975/500/1,525 images for training, validation and testing. There are 3030 classes, and 1919 classes among them are used for evaluation. In addition to the mean of class-wise intersection over union (mIoU), we report other three scores on the test set: IoU category (cat.), iIoU class (cla.) and iIoU category (cat.).

We follow the same training protocol . The data are augmented by random cropping (from 1024×20481024\times 2048 to 512×1024512\times 1024), random scaling in the range of [0.5,2][0.5,2], and random horizontal flipping. We use the SGD optimizer with the base learning rate of 0.010.01, the momentum of 0.90.9 and the weight decay of 0.00050.0005. The poly learning rate policy with the power of 0.90.9 is used for dropping the learning rate. All the models are trained for 120K120K iterations with the batch size of 1212 on 44 GPUs and syncBN.

Table 1 provides the comparison with several representative methods on the Cityscapes validation set in terms of parameter and computation complexity and mIoU class. (i) HRNetV22-W4040 (4040 indicates the width of the high-resolution convolution), with similar model size to DeepLabv33+ and much lower computation complexity, gets better performance: 4.74.7 points gain over UNet++, 1.71.7 points gain over DeepLabv3 and about 0.50.5 points gain over PSPNet, DeepLabv3+. (ii) HRNetV22-W4848, with similar model size to PSPNet and much lower computation complexity, achieves much significant improvement: 5.65.6 points gain over UNet++, 2.62.6 points gain over DeepLabv3 and about 1.41.4 points gain over PSPNet, DeepLabv3+. In the following comparisons, we adopt HRNetV22-W4848 that is pretrained on ImageNet The description about ImageNet pretraining is given in the Appendix. and has similar model size as most Dilated-ResNet-101101 based methods.

Table 2 provides the comparison of our method with state-of-the-art methods on the Cityscapes test set. All the results are with six scales and flipping. Two cases w/o using coarse data are evaluated: One is about the model learned on the train set, and the other is about the model learned on the train+valid set. In both cases, HRNetV22-W4848 achieves the best performance and outperforms the previous state-of-the-art by 11 point.

PASCAL context. The PASCAL context dataset includes 4,9984,998 scene images for training and 5,1055,105 images for testing with 5959 semantic labels and 11 background label.

The data augmentation and learning rate policy are the same as Cityscapes. Following the widely-used training strategy , we resize the images to 480×480480\times 480 and set the initial learning rate to 0.0040.004 and weight decay to 0.00010.0001. The batch size is 1616 and the number of iterations is 60K60K.

We follow the standard testing procedure . The image is resized to 480×480480\times 480 and then fed into our network. The resulting 480×480480\times 480 label maps are then resized to the original image size. We evaluate the performance of our approach and other approaches using six scales and flipping.

Table 3 provides the comparison of our method with state-of-the-art methods. There are two kinds of evaluation schemes: mIoU over 5959 classes and 6060 classes (5959 classes + background). In both cases, HRNetV22-W4848 performs superior to previous state-of-the-arts.

LIP. The LIP dataset contains 50,46250,462 elaborately annotated human images, which are divided into 30,46230,462 training images, and 10,00010,000 validation images. The methods are evaluated on 2020 categories (1919 human part labels and 11 background label). Following the standard training and testing settings , the images are resized to 473×473473\times 473 and the performance is evaluated on the average of the segmentation maps of the original and flipped images.

The data augmentation and learning rate policy are the same as Cityscapes. The training strategy follows the recent setting . We set the initial learning rate to 0.0070.007 and the momentum to 0.90.9 and the weight decay to 0.00050.0005. The batch size is 4040 and the number of iterations is 110110K.

Table 4 provides the comparison of our method with state-of-the-art methods. The overall performance of HRNetV22-W4848 performs the best with fewer parameters and lighter computation cost. We also would like to mention that our networks do not use extra information such as pose or edge.

2 COCO Object Detection

We apply our multi-level representations (HRNetV22p)Same as FPN , we also use 55 levels., shown in Figure 3 (c), in the Faster R-CNN and Mask R-CNN frameworks. We perform the evaluation on the MS-COCO 20172017 detection dataset, which contains ∼118\sim 118k images for training, 55k for validation (val) and ∼20\sim 20k testing without provided annotations (test-dev). The standard COCO-style evaluation is adopted.

We train the models for both our HRNetV22p and the ResNet on the public mmdetection platform with the provided training setup, except that we use the learning rate schedule suggested in for 2×2\times. The data is augmented by standard horizontal flipping. The input images are resized such that the shorter edge is 800 pixels . Inference is performed on a single image scale.

Table 5 summarizes #parameters and GFLOPs. Table 6 and Table 7 report the detection results on COCO val. There are several observations. (i) The model size and computation complexity of HRNetV22p-W1818 (HRNetV22p-W3232) are smaller than ResNet-5050-FPN (ResNet-101101-FPN). (ii) With 1×1\times, HRNetV2p-W3232 performs better than ResNet-101101-FPN. HRNetV2p-W1818 performs worse than ResNet-5050-FPN, which might come from insufficient optimization iterations. (iii) With 2×2\times, HRNetV2p-W1818 and HRNetV2p-W3232 perform better than ResNet-5050-FPN and ResNet-101101-FPN, respectively.

Table 8 reports the comparison of our network to state-of-the-art single-model object detectors on COCO test-dev without using multi-scale training and multi-scale testing that are done in . In the Faster R-CNN framework, our networks perform better than ResNets with similar parameter and computation complexity: HRNetV22p-W3232 vs. ResNet-101101-FPN, HRNetV22p-W4040 vs. ResNet-152152-FPN, HRNetV22p-W4848 vs. X-101101-64×464\times 4d-FPN. In the Cascade R-CNN framework, our HRNetV22p-W3232 performs better.

3 Facial Landmark Detection

Facial landmark detection a.k.a. face alignment is a problem of detecting the keypoints from a face image. We perform the evaluation over four standard datasets: WFLW , AFLW , COFW , and 300300W . We mainly use the normalized mean error (NME) for evaluation. We use the inter-ocular distance as normalization for WFLW, COFW, and 300300W, and the face bounding box as normalization for AFLW. We also report area-under-the-curve scores (AUC) and failure rates.

We follow the standard scheme for training. All the faces are cropped by the provided boxes according to the center location and resized to 256×256256\times 256. We augment the data by ±30\pm 30 degrees in-plane rotation, 0.75−1.250.75-1.25 scaling, and randomly flipping. The base learning rate is 0.00010.0001 and is dropped to 0.000010.00001 and 0.0000010.000001 at the 3030th and 5050th epochs. The models are trained for 6060 epochs with the batch size of 1616 on one GPU. Different from semantic segmentation, the heatmaps are not upsampled from 1/41/4 to the input size, and the loss function is optimized over the 1/41/4 maps.

At testing, each keypoint location is predicted by transforming the highest heatvalue location from 1/41/4 to the original image space and adjusting it with a quarter offset in the direction from the highest response to the second highest response .

We adopt HRNetV22-W1818 for face landmark detection whose parameter and computation cost are similar to or smaller than models with widely-used backbones: ResNet-5050 and Hourglass . HRNetV22-W1818: #parameters =9.3=9.3M, GFLOPs =4.3=4.3G; ResNet-5050: #parameters =25.0=25.0M, GFLOPs =3.8=3.8G; Hourglass: #parameters =25.1=25.1M, GFLOPs =19.1=19.1G. The numbers are obtained on the input size 256×256256\times 256. It should be noted that the facial landmark detection methods adopting ResNet-5050 and Hourglass as backbones introduce extra parameter and computation overhead.

WFLW. The WFLW dataset is a recently-built dataset based on the WIDER Face . There are 7,5007,500 training and 2,5002,500 testing images with 9898 manual annotated landmarks. We report the results on the test set and several subsets: large pose (326326 images), expression (314314 images), illumination (698698 images), make-up (206206 images), occlusion (736736 images) and blur (773773 images).

Table 9 provides the comparison of our method with state-of-the-art methods. Our approach is significantly better than other methods on the test set and all the subsets, including LAB that exploits extra boundary information and PDB that uses stronger data augmentation .

AFLW. The AFLW dataset is a widely used benchmark dataset, where each image has 1919 facial landmarks. Following , we train our models on 20,00020,000 training images, and report the results on the AFLW-Full set (4,3864,386 testing images) and the AFLW-Frontal set (13141314 testing images selected from 43864386 testing images).

Table 10 provides the comparison of our method with state-of-the-art methods. Our approach achieves the best performance among methods without extra information and stronger data augmentation and even outperforms DCFE with extra 33D information. Our approach performs slightly worse than LAB that uses extra boundary information and PDB that uses stronger data augmentation.

COFW. The COFW dataset consists of 1,3451,345 training and 507507 testing faces with occlusions, where each image has 2929 facial landmarks.

Table 11 provides the comparison of our method with state-of-the-art methods. HRNetV22 outperforms other methods by a large margin. In particular, it achieves the better performance than LAB with extra boundary information and PDB with stronger data augmentation.

300300W. The dataset is a combination of HELEN , LFPW , AFW , XM2VTS and IBUG datasets, where each face has 6868 landmarks. Following , we use the 3,1483,148 training images, which contains the training subsets of HELEN and LFPW and the full set of AFW. We evaluate the performance using two protocols, full set and test set. The full set contains 689689 images and is further divided into a common subset (554554 images) from HELEN and LFPW, and a challenging subset (135135 images) from IBUG. The official test set, used for competition, contains 600600 images (300300 indoor and 300300 outdoor images).

Table 12 provides the results on the full set, and its two subsets: common and challenging. Table 13 provides the results on the test set. In comparison to Chen et al. that uses Hourglass with large parameter and computation complexity as the backbone, our scores are better except the AUC0.08 scores. Our HRNetV22 gets the overall best performance among methods without extra information and stronger data augmentation, and is even better than LAB with extra boundary information and DCFE that explores extra 33D information.

4 Empirical Analysis

We compare the modified networks, HRNetV22 and HRNetV22p, to the original network (shortened as HRNetV11) on semantic segmentation and COCO object detection. The segmentation and object detection results, given in Figure 4 (a) and Figure 4 (b), imply that HRNetV22 outperforms HRNetV11 significantly, except that the gain is minor in the large model case in segmentation for Cityscapes. We also test a variant (denoted by HRNetV11h), which is built by appending a 1×11\times 1 convolution to increase the dimension of the output high-resolution representation. The results in Figure 4 (a) and Figure 4 (b) show that the variant achieves slight improvement to HRNetV11, implying that aggregating the representations from low-resolution parallel convolutions in our HRNetV22 is essential for increasing the capability.

Conclusions

In this paper, we empirically study the high-resolution representation network in a broad range of vision applications with introducing a simple modification. Experimental results demonstrate the effectiveness of strong high-resolution representations and multi-level representations learned by the modified networks on semantic segmentation, facial landmark detection as well as object detection. The project page is https://jingdongwang2017.github.io/Projects/HRNet/.

Appendix: Network Pretraining

We pretrain our network, which is augmented by a classification head shown in Figure 5, on ImageNet . The classification head is described as below. First, the four-resolution feature maps are fed into a bottleneck and the output channels are increased from CC, 2C2C, 4C4C, and 8C8C to 128128, 256256, 512512, and 10241024, respectively. Then, we downsample the high-resolution representation by a 22-strided 3×33\times 3 convolution outputting 256256 channels and add it to the representation of the second-high-resolution. This process is repeated two times to get 10241024 feature channels over the small resolution. Last, we transform the 10241024 channels to 20482048 channels through a 1×11\times 1 convolution, followed by a global average pooling operation. The output 20482048-dimensional representation is fed into the classifier.

We adopt the same data augmentation scheme for training images as in , and train our models for 100100 epochs with a batch size of 256256. The initial learning rate is set to 0.10.1 and is reduced by 1010 times at epoch 3030, 6060 and 9090. We use SGD with a weight decay of 0.00010.0001 and a Nesterov momentum of 0.90.9. We adopt standard single-crop testing, so that 224×224224\times 224 pixels are cropped from each image. The top-11 and top-55 error are reported on the validation set.

Table 14 shows our ImageNet classification results. As a comparison, we also report the results of ResNets. We consider two types of residual units: One is formed by a bottleneck, and the other is formed by two 3×33\times 3 convolutions. We follow the PyTorch implementation of ResNets and replace the 7×77\times 7 convolution in the input stem with two 22-strided 3×33\times 3 convolutions decreasing the resolution to 1/41/4 as in our networks. When the residual units are formed by two 3×33\times 3 convolutions, an extra bottleneck is used to increase the dimension of output feature maps from 512512 to 20482048. One can see that under similar #parameters and GFLOPs, our results are comparable to and slightly better than ResNets.

In addition, we look at the results of two alternative schemes: (i) the feature maps on each resolution go through a global pooling separately and then are concatenated together to output a 15C15C-dimensional representation vector, named HRNet-Wxx-Ci; (ii) the feature maps on each resolution are fed into several 22-strided residual units (bottleneck, each dimension is increased to the double) to increase the dimension to 512512, and concatenate and average-pool them together to reach a 20482048-dimensional representation vector, named HRNet-Wxx-Cii, which is used in . Table 15 shows such an ablation study. One can see that the proposed manner is superior to the two alternatives.

References