Deep High-Resolution Representation Learning for Human Pose Estimation
Ke Sun, Bin Xiao, Dong Liu, Jingdong Wang
Introduction
D human pose estimation has been a fundamental yet challenging problem in computer vision. The goal is to localize human anatomical keypoints (e.g., elbow, wrist, etc.) or parts. It has many applications, including human action recognition, human-computer interaction, animation, etc. This paper is interested in single-person pose estimation, which is the basis of other related problems, such as multi-person pose estimation , video pose estimation and tracking , etc.
The recent developments show that deep convolutional neural networks have achieved the state-of-the-art performance. Most existing methods pass the input through a network, typically consisting of high-to-low resolution subnetworks that are connected in series, and then raise the resolution. For instance, Hourglass recovers the high resolution through a symmetric low-to-high process. SimpleBaseline adopts a few transposed convolution layers for generating high-resolution representations. In addition, dilated convolutions are also used to blow up the later layers of a high-to-low resolution network (e.g., VGGNet or ResNet) .
We present a novel architecture, namely High-Resolution Net (HRNet), which is able to maintain high-resolution representations through the whole process. We start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the multi-resolution subnetworks in parallel. We conduct repeated multi-scale fusions by exchanging the information across the parallel multi-resolution subnetworks over and over through the whole process. We estimate the keypoints over the high-resolution representations output by our network. The resulting network is illustrated in Figure 1.
Our network has two benefits in comparison to existing widely-used networks for pose estimation. (\romannum1) Our approach connects high-to-low resolution subnetworks in parallel rather than in series as done in most existing solutions. Thus, our approach is able to maintain the high resolution instead of recovering the resolution through a low-to-high process, and accordingly the predicted heatmap is potentially spatially more precise. (\romannum2) Most existing fusion schemes aggregate low-level and high-level representations. Instead, we perform repeated multi-scale fusions to boost the high-resolution representations with the help of the low-resolution representations of the same depth and similar level, and vice versa, resulting in that high-resolution representations are also rich for pose estimation. Consequently, our predicted heatmap is potentially more accurate.
We empirically demonstrate the superior keypoint detection performance over two benchmark datasets: the COCO keypoint detection dataset and the MPII Human Pose dataset . In addition, we show the superiority of our network in video pose tracking on the PoseTrack dataset .
Related Work
Most traditional solutions to single-person pose estimation adopt the probabilistic graphical model or the pictorial structure model , which is recently improved by exploiting deep learning for better modeling the unary and pair-wise energies or imitating the iterative inference process . Nowadays, deep convolutional neural network provides dominant solutions . There are two mainstream methods: regressing the position of keypoints , and estimating keypoint heatmaps followed by choosing the locations with the highest heat values as the keypoints.
Most convolutional neural networks for keypoint heatmap estimation consist of a stem subnetwork similar to the classification network, which decreases the resolution, a main body producing the representations with the same resolution as its input, followed by a regressor estimating the heatmaps where the keypoint positions are estimated and then transformed in the full resolution. The main body mainly adopts the high-to-low and low-to-high framework, possibly augmented with multi-scale fusion and intermediate (deep) supervision.
High-to-low and low-to-high. The high-to-low process aims to generate low-resolution and high-level representations, and the low-to-high process aims to produce high-resolution representations . Both the two processes are possibly repeated several times for boosting the performance .
Representative network design patterns include: (\romannum1) Symmetric high-to-low and low-to-high processes. Hourglass and its follow-ups design the low-to-high process as a mirror of the high-to-low process. (\romannum2) Heavy high-to-low and light low-to-high. The high-to-low process is based on the ImageNet classification network, e.g., ResNet adopted in , and the low-to-high process is simply a few bilinear-upsampling or transpose convolution layers. (\romannum3) Combination with dilated convolutions. In , dilated convolutions are adopted in the last two stages in the ResNet or VGGNet to eliminate the spatial resolution loss, which is followed by a light low-to-high process to further increase the resolution, avoiding expensive computation cost for only using dilated convolutions . Figure 2 depicts four representative pose estimation networks.
Multi-scale fusion. The straightforward way is to feed multi-resolution images separately into multiple networks and aggregate the output response maps . Hourglass and its extensions combine low-level features in the high-to-low process into the same-resolution high-level features in the low-to-high process progressively through skip connections. In cascaded pyramid network , a globalnet combines low-to-high level features in the high-to-low process progressively into the low-to-high process, and then a refinenet combines the low-to-high level features that are processed through convolutions. Our approach repeats multi-scale fusion, which is partially inspired by deep fusion and its extensions .
Intermediate supervision. Intermediate supervision or deep supervision, early developed for image classification , is also adopted for helping deep networks training and improving the heatmap estimation quality, e.g., . The hourglass approach and the convolutional pose machine approach process the intermediate heatmaps as the input or a part of the input of the remaining subnetwork.
Our approach. Our network connects high-to-low subnetworks in parallel. It maintains high-resolution representations through the whole process for spatially precise heatmap estimation. It generates reliable high-resolution representations through repeatedly fusing the representations produced by the high-to-low subnetworks. Our approach is different from most existing works, which need a separate low-to-high upsampling process and aggregate low-level and high-level representations. Our approach, without using intermediate heatmap supervision, is superior in keypoint detection accuracy and efficient in computation complexity and parameters.
There are related multi-scale networks for classification and segmentation . Our work is partially inspired by some of them , and there are clear differences making them not applicable to our problem. Convolutional neural fabrics and interlinked CNN fail to produce high-quality segmentation results because of a lack of proper design on each subnetwork (depth, batch normalization) and multi-scale fusion. The grid network , a combination of many weight-shared U-Nets, consists of two separate fusion processes across multi-resolution representations: on the first stage, information is only sent from high resolution to low resolution; on the second stage, information is only sent from low resolution to high resolution, and thus less competitive. Multi-scale densenets does not target and cannot generate reliable high-resolution representations.
Approach
Human pose estimation, a.k.a. keypoint detection, aims to detect the locations of keypoints or parts (e.g., elbow, wrist, etc) from an image of size . The state-of-the-art methods transform this problem to estimating heatmaps of size , , where each heatmap indicates the location confidence of the th keypoint.
We follow the widely-adopted pipeline to predict human keypoints using a convolutional network, which is composed of a stem consisting of two strided convolutions decreasing the resolution, a main body outputting the feature maps with the same resolution as its input feature maps, and a regressor estimating the heatmaps where the keypoint positions are chosen and transformed to the full resolution. We focus on the design of the main body and introduce our High-Resolution Net (HRNet) that is depicted in Figure 1.
Sequential multi-resolution subnetworks. Existing networks for pose estimation are built by connecting high-to-low resolution subnetworks in series, where each subnetwork, forming a stage, is composed of a sequence of convolutions and there is a down-sample layer across adjacent subnetworks to halve the resolution.
Let be the subnetwork in the th stage and be the resolution index (Its resolution is of the resolution of the first subnetwork). The high-to-low network with (e.g., ) stages can be denoted as:
Parallel multi-resolution subnetworks. We start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one, forming new stages, and connect the multi-resolution subnetworks in parallel. As a result, the resolutions for the parallel subnetworks of a later stage consists of the resolutions from the previous stage, and an extra lower one.
An example network structure, containing parallel subnetworks, is given as follows,
Repeated multi-scale fusion. We introduce exchange units across parallel subnetworks such that each subnetwork repeatedly receives the information from other parallel subnetworks. Here is an example showing the scheme of exchanging information. We divided the third stage into several (e.g., ) exchange blocks, and each block is composed of parallel convolution units with an exchange unit across the parallel units, which is given as follows,
where represents the convolution unit in the th resolution of the th block in the th stage, and is the corresponding exchange unit.
The function consists of upsampling or downsampling from resolution to resolution . We adopt strided convolutions for downsampling. For instance, one strided convolution with the stride for downsampling, and two consecutive strided convolutions with the stride for downsampling. For upsampling, we adopt the simple nearest neighbor sampling following a convolution for aligning the number of channels. If , is just an identify connection: .
Heatmap estimation. We regress the heatmaps simply from the high-resolution representations output by the last exchange unit, which empirically works well. The loss function, defined as the mean squared error, is applied for comparing the predicted heatmaps and the groundtruth heatmaps. The groundtruth heatmpas are generated by applying D Gaussian with standard deviation of pixel centered on the grouptruth location of each keypoint.
Network instantiation. We instantiate the network for keypoint heatmap estimation by following the design rule of ResNet to distribute the depth to each stage and the number of channels to each resolution.
The main body, i.e., our HRNet, contains four stages with four parallel subnetworks, whose the resolution is gradually decreased to a half and accordingly the width (the number of channels) is increased to the double. The first stage contains residual units where each unit, the same to the ResNet-, is formed by a bottleneck with the width , and is followed by one convolution reducing the width of feature maps to . The nd, rd, th stages contain , , exchange blocks, respectively. One exchange block contains residual units where each unit contains two convolutions in each resolution and an exchange unit across resolutions. In summary, there are totally exchange units, i.e., multi-scale fusions are conducted.
In our experiments, we study one small net and one big net: HRNet-W and HRNet-W, where and represent the widths () of the high-resolution subnetworks in last three stages, respectively. The widths of other three parallel subnetworks are for HRNet-W, and for HRNet-W.
Experiments
Dataset. The COCO dataset contains over images and person instances labeled with keypoints. We train our model on COCO train dataset, including images and person instances. We evaluate our approach on the val set and test-dev set, containing images and images, respectively.
Evaluation metric. The standard evaluation metric is based on Object Keypoint Similarity (OKS): Here is the Euclidean distance between the detected keypoint and the corresponding ground truth, is the visibility flag of the ground truth, is the object scale, and is a per-keypoint constant that controls falloff. We report standard average precision and recall scoreshttp://cocodataset.org/#keypoints-eval: ( at ) , (the mean of scores at positions, ; for medium objects, for large objects, and at .
Training. We extend the human detection box in height or width to a fixed aspect ratio: , and then crop the box from the image, which is resized to a fixed size, or . The data augmentation includes random rotation (), random scale (), and flipping. Following , half body data augmentation is also involved.
Testing. The two-stage top-down paradigm similar as is used: detect the person instance using a person detector, and then predict detection keypoints.
We use the same person detectors provided by SimpleBaselinehttps://github.com/Microsoft/human-pose-estimation.pytorch for both validation set and test-dev set. Following the common practice , we compute the heatmap by averaging the headmaps of the original and flipped images. Each keypoint location is predicted by adjusting the highest heatvalue location with a quarter offset in the direction from the highest response to the second highest response.
Results on the validation set. We report the results of our method and other state-of–the-art methods in Table 1. Our small network - HRNet-W, trained from scratch with the input size , achieves an AP score, outperforming other methods with the same input size. (\romannum1) Compared to Hourglass , our small network improves AP by points, and the GFLOPs of our network is much lower and less than half, while the number of parameters are similar and ours is slightly larger. (\romannum2) Compared to CPN w/o and w/ OHKM, our network, with slightly larger model size and slightly higher complexity, achieves and points gain, respectively. (\romannum3) Compared to the previous best-performed SimpleBaseline , our small net HRNet-W obtains significant improvements: points gain for the backbone ResNet- with a similar model size and GFLOPs, and points gain for the backbone ResNet- whose model size (#Params) and GLOPs are twice as many as ours.
Our nets can benefit from (\romannum1) training from the model pretrained for the ImageNet classification problem: The gain is points for HRNet-W; (\romannum2) increasing the capacity by increasing the width: Our big net HRNet-W gets and improvements for the input sizes and , respectively.
Considering the input size , our HRNet-W and HRNet-W, get the and AP, which have and improvements compared to the input size . In comparison to the SimpleBaseline that uses ResNet- as the backbone, our HRNet-W and HRNet-W attain and points gain in terms of AP at and computational cost, respectively.
Results on the test-dev set. Table 2 reports the pose estimation performances of our approach and the existing state-of-the-art approaches. Our approach is significantly better than bottom-up approaches. On the other hand, our small network, HRNet-W, achieves an AP of . It outperforms all the other top-down approaches, and is more efficient in terms of model size (#Params) and computation complexity (GFLOPs). Our big model, HRNet-W, achieves the highest AP. Compared to the SimpleBaseline with the same input size, our small and big networks receive and improvements, respectively. With additional data from AI Challenger for training, our single big network can obtain an AP of .
2 MPII Human Pose Estimation
Dataset. The MPII Human Pose dataset consists of images taken from a wide-range of real-world activities with full-body pose annotations. There are around images with subjects, where there are subjects for testing and the remaining subjects for the training set. The data augmentation and the training strategy are the same to MS COCO, except that the input size is cropped to for fair comparison with other methods.
Testing. The testing procedure is almost the same to that in COCO except that we adopt the standard testing strategy to use the provided person boxes instead of detected person boxes. Following , a six-scale pyramid testing procedure is performed.
Evaluation metric. The standard metric , the PCKh (head-normalized probability of correct keypoint) score, is used. A joint is correct if it falls within pixels of the groundtruth position, where is a constant and is the head size that corresponds to of the diagonal length of the ground-truth head bounding box. The PCKh () score is reported.
Results on the test set. Tables 3 and 4 show the PCKh results, the model size and the GFLOPs of the top-performed methods. We reimplement the SimpleBaseline by using ResNet- as the backbone with the input size . Our HRNet-W achieves a PKCh@ score, and outperforms the stacked hourglass approach and its extensions . Our result is the same as the best one among the previously-published results on the leaderboard of Nov. th, http://human-pose.mpi-inf.mpg.de/#results. We would like to point out that the approach , complementary to our approach, exploits the compositional model to learn the configuration of human bodies and adopts multi-level intermediate supervision, from which our approach can also benefit. We also tested our big network - HRNet-W and obtained the same result . The reason might be that the performance in this datatset tends to be saturate.
3 Application to Pose Tracking
Dataset. PoseTrack is a large-scale benchmark for human pose estimation and articulated tracking in video. The dataset, based on the raw videos provided by the popular MPII Human Pose dataset, contains video sequences with frames. The video sequences are split into , , videos for training, validation, and testing, respectively. The length of the training videos ranges between frames, and frames from the center of the video are densely annotated. The number of frames in the validation/testing videos ranges between frames. The frames around the keyframe from the MPII Pose dataset are densely annotated, and afterwards every fourth frame is annotated. In total, this constitutes roughly labeled frames and pose annotations.
Evaluation metric. We evaluate the results from two aspects: frame-wise multi-person pose estimation, and multi-person pose tracking. Pose estimation is evaluated by the mean Average Precision (mAP) as done in . Multi-person pose tracking is evaluated by the multi-object tracking accuracy (MOTA) . Details are given in .
Testing. We follow to track poses across frames. It consists of three steps: person box detection and propagation, human pose estimation, and pose association cross nearby frames. We use the same person box detector as used in SimpleBaseline , and propagate the detected box into nearby frames by propagating the predicted keypoints according to the optical flows computed by FlowNet 2.0 https://github.com/NVIDIA/flownet2-pytorch, followed by non-maximum suppression for box removing. The pose association scheme is based on the object keypoint similarity between the keypoints in one frame and the keypoints propagated from the nearby frame according to the optical flows. The greedy matching algorithm is then used to compute the correspondence between keypoints in nearby frames. More details are given in .
Results on the PoseTrack test set. Table 5 reports the results. Our big network - HRNet-W achieves the superior result, a mAP score and a MOTA score. Compared with the second best approach, the FlowTrack in SimpleBaseline , that uses ResNet- as the backbone, our approach gets and points gain in terms of mAP and MOTA, respectively. The superiority over the FlowTrack is consistent to that on the COCO keypoint detection and MPII human pose estimation datasets. This further implies the effectiveness of our pose estimation network.
4 Ablation Study
We study the effect of each component in our approach on the COCO keypoint detection dataset. All results are obtained over the input size of except the study about the effect of the input size.
Repeated multi-scale fusion. We empirically analyze the effect of the repeated multi-scale fusion. We study three variants of our network. (a) W/o intermediate exchange units ( fusion): There is no exchange between multi-resolution subnetworks except the last exchange unit. (b) W/ across-stage exchange units only ( fusions): There is no exchange between parallel subnetworks within each stage. (c) W/ both across-stage and within-stage exchange units (totally fusion): This is our proposed method. All the networks are trained from scratch. The results on the COCO validation set given in Table 6 show that the multi-scale fusion is helpful and more fusions lead to better performance.
Resolution maintenance. We study the performance of a variant of the HRNet: all the four high-to-low resolution subnetworks are added at the beginning and the depth are the same; the fusion schemes are the same to ours. Both our HRNet-W and the variant (with similar #Params and GFLOPs) are trained from scratch and tested on the COCO validation set. The variant achieves an AP of , which is lower than the AP of our small net, HRNet-W. We believe that the reason is that the low-level features extracted from the early stages over the low-resolution subnetworks are less helpful. In addition, the simple high-resolution network of similar parameter and computation complexities without low-resolution parallel subnetworks shows much lower performance .
Representation resolution. We study how the representation resolution affects the pose estimation performance from two aspects: check the quality of the heatmap estimated from the feature maps of each resolution from high to low, and study how the input size affects the quality.
We train our small and big networks initialized by the model pretrained for the ImageNet classification. Our network outputs four response maps from high-to-low solutions. The quality of heatmap prediction over the lowest-resolution response map is too low and the AP score is below points. The AP scores over the other three maps are reported in Figure 5. The comparison implies that the resolution does impact the keypoint prediction quality.
Figure 6 shows how the input image size affects the performance in comparison with SimpleBaseline (ResNet-50) . We can find that the improvement for the smaller input size is more significant than the larger input size, e.g., the improvement is points for and points for . The reason is that we maintain the high resolution through the whole process. This implies that our approach is more advantageous in the real applications where the computation cost is also an important factor. On the other hand, our approach with the input size outperforms the SimpleBaseline with the large input size of .
Conclusion and Future Works
In this paper, we present a high-resolution network for human pose estimation, yielding accurate and spatially-precise keypoint heatmaps. The success stems from two aspects: (\romannum1) maintain the high resolution through the whole process without the need of recovering the high resolution; and (\romannum2) fuse multi-resolution representations repeatedly, rendering reliable high-resolution representations.
The future works include the applications to other dense prediction tasks, e.g., semantic segmentation, object detection, face alignment, image translation, as well as the investigation on aggregating multi-resolution representations in a less light way. All them are available at https://jingdongwang2017.github.io/Projects/HRNet/index.html.
Appendix
We provide the results on the MPII validation set . Our models are trained on a subset of MPII training set and evaluate on a heldout validation set of 2975 images. The training procedure is the same to that for training on the whole MPII training set. The heatmap is computed as the average of the heatmaps of the original and flipped images for testing. Following , we also perform six-scale pyramid testing procedure (multi-scale testing). The results are shown in Table 7.
More Results on the PoseTrack Dataset
We provide the results for all the keypoints on the PoseTrack dataset . Table 8 shows the multi-person pose estimation performance on the PoseTrack dataset. Our HRNet-W achieves 77.3 and 74.9 points mAP on the validation and test setss, and outperforms previous state-of-the-art method by 0.6 points and 0.3 points respectively. We provide more detailed results of multi-person pose tracking performance on the PoseTrack2017 test set as a supplement of the results reported in the paper, shown in Table 9.
Results on the ImageNet Validation Set
We apply our networks to image classification task. The models are trained and evaluated on the ImageNet 2013 classification dataset . We train our models for 100 epochs with a batch size of 256. The initial learning rate is set to 0.1 and is reduced by 10 times at epoch 30, 60 and 90. Our models can achieve comparable performance as those networks specifically designed for image classification, such as ResNet . Our HRNet-W has a single-model top-5 validation error of 6.5% and has a single-model top-1 validation error of 22.7% with the single-crop testing. Our HRNet-W gets better performance: 6.1% top-5 errors and 22.1% top-1 error. We use the models trained on the ImageNet dataset to initialize the parameters of our pose estimation networks.
Acknowledgements. The authors thank Dianqi Li and Lei Zhang for helpful discussions.