Bottom-Up Human Pose Estimation Via Disentangled Keypoint Regression
Zigang Geng, Ke Sun, Bin Xiao, Zhaoxiang Zhang, Jingdong Wang
Introduction
Human pose estimation is a problem of predicting the keypoint positions of each person from an image, i.e., localize the keypoints as well as identify the keypoints belonging to the same person. There are broad applications, including action recognition, human-computer interaction, smart photo editing, pedestrian tracking, etc.
There are two main paradigms: top-down and bottom-up. The top-down paradigm first detects the person and then performs single-person pose estimation for each detected person. The bottom-up paradigm either directly regresses the keypoint positions belonging to the same person, or detects and groups the keypoints, such as affinity linking , associative embedding , HGG and HigherHRNet . The top-down paradigm is more accurate but more costly due to an extra person detection process, and the bottom-up paradigm, the interest of this paper, is more efficient.
The recently-developed pixel-wise keypoint regression approach, CenterNet , estimates the keypoint positions together for each pixel from the representation at the pixel. Direct regression to keypoint positions in CenterNet performs reasonably. But the regressed keypoints are spatially not accurate and the performance is worse than the keypoint detection and grouping scheme. Figure 1 (left) shows two examples in which the salient areas for keypoint regression spread broadly and the regression quality is not satisfactory.
We argue that regressing the keypoint positions accurately needs to learn representations that focus on the keypoint regions. Starting from this regression by focusing concept, we present a simple yet effective approach, named disentangled keypoint regression (DEKR). We adopt adaptive convolutions, through pixel-wise spatial transformer (a pixel-wise extension of spatial transformer network ), to activate the pixels lying in the keypoint regions, and then learn the representations from these activated pixels, so that the learned representations can focus on the keypoint regions.
We further decouple the representation learning for one keypoint from other keypoints. We adopt a separate regression scheme through a multi-branch structure: each branch learns a representation for one keypoint with adaptive convolutions dedicated for the keypoint and regresses the position for the corresponding keypoint. Figure 1 (right) illustrates that our approach is able to learn highly concentrative representations, each of which focuses on the corresponding keypoint region.
Experimental results demonstrate that the proposed DEKR approach improves the localization quality of the regressed keypoint positions. Our approach, that performs direct keypoint regression without matching the regression results to the closest keypoints detected from the keypoint heatmaps, outperforms keypoint detection and grouping methods and achieves superior performance over previous state-of-the-art bottom-up pose estimation methods on two benchmark datasets, COCO and CrowdPose.
Our contributions to bottom-up human pose estimation are summarized as follows.
We argue that the representations for regressing the positions of the keypoints accurately need to focus on the keypoint regions.
The proposed DEKR approach is able to learn disentangled representations through two simple schemes, adaptive convolutions and multi-branch structure, so that each representation focuses on one keypoint region and the prediction of the corresponding keypoint position from such representation is accurate.
The proposed direct regression approach outperforms keypoint detection and grouping schemes and achieves new state-of-the-art bottom-up pose estimation results on the benchmark datasets, COCO and CrowdPose.
Related Work
The convolutional neural network (CNN) solutions to human pose estimation have shown superior performance over the conventional methods, such as the probabilistic graphical model or the pictorial structure model . Early CNN approaches directly predict the keypoint positions for single-person pose estimation, which is later surpassed by the heatmap estimation based methods . The geometric constraints and structured relations among body keypoints are studied for performance improvement .
Top-down paradigm. The top-down methods perform single-person pose estimation by firstly detecting each person from the image. Representative works include: HRNet , PoseNet , RMPE , convolutional pose machine , Hourglass , Mask R-CNN , CFN , Integral pose regression , CPN , simple baseline , CSM-SCARB , Graph-PCNN , RSN , and so on. These methods exploit the advances in person detection as well as extra person bounding-box labeling information. The top-down paradigm, though achieving satisfactory performance, takes extra cost in person box detection.
Other developments include improving the keypoint localization from the heatmap , refining pose estimation , better data augmentation , developing a multi-task learning architecture combining detection, segmentation and pose estimation , and handling the occlusion issue .
Bottom-up paradigm. Most existing bottom-up methods mainly focus on how to associate the detected keypoints that belong to the same person together. The pioneering work, DeepCut , DeeperCut , and L-JPA formulate the keypoint association problem as an integer linear program, which however takes longer processing time (e.g., the order of hours).
Various grouping techniques are developed, such as part-affinity fields in OpenPose and its extension in PifPaf , associative embedding , greedy decoding with hough voting in PersonLab , and graph clustering in HGG .
Several recent works densely regress a set of pose candidates, where each candidate consists of the keypoint positions that might be from the same person. Unfortunately, the regression quality is not high, and the localization quality is weak. A post-processing scheme, matching the regressed keypoint positions to the closest keypoints (which is spatially more accurate) detected from the keypoint heatmaps, is usually adopted to improve the regression results.
Our approach aims to improve the direct regression results, by exploring our regression by focusing idea. We learn disentangled representations, each of which is dedicated for one keypoint and learns from the adaptively activated pixels, so that each representation focuses on the corresponding keypoint area. As a result, the position prediction for one keypoint from the corresponding disentangled representation is spatially accurate. Our approach is superior to and differs from that uses the mixture density network for handling uncertainty to improve direct regression results.
Disentangled representation learning. Disentangled representations have widely been studied in computer vision , e.g., disentangling the representations into content and pose , disentangling motion from content , disentangling pose and appearance .
Our proposed disentangled regression in some sense can be regarded as disentangled representation learning: learn the representation for each keypoint separately from the corresponding keypoint region. The idea of representation disentanglement for pose estimation is also explored in the top-down approach, part-based branching network (PBN) , which learns high-quality heatmaps by disentangling representations into each part group. They are clearly different: our approach learns representations focusing on each keypoint region for position regression, and PBN de-correlates the appearance representations among different part groups.
Approach
Given an image , multi-person pose estimation aims to predict the human poses, where each pose consists of keypoints, such as shoulder, elbow, and so on. Figure 2 illustrates the multi-person pose estimation problem.
The pixel-wise keypoint regression framework estimates a candidate pose at each pixel (called center pixel), by predicting an -dimensional offset vector from the center pixel for the keypoints. The offset maps , containing the offset vectors at all the pixels, are estimated through a keypoint regression head,
where is the feature computed from a backbone, HRNet in this paper, and is the keypoint position regression head predicting the offset maps .
The structure of the proposed disentangled keypoint regression (DEKR) head is illustrated in Figure 3. DEKR adopts the multi-branch parallel adaptive convolutions to learn disentangled representations for the regression of the keypoints, so that each representation focuses on the corresponding keypoint region.
Adaptive activation. One normal convolution (e.g., convolution) only sees the pixels nearby the center pixel . A sequence of several normal convolutions may see the pixels farther from the center pixel that might lie in the keypoint region, but might not focus on and highly activate these pixels.
We adopt the adaptive convolutions, to learn representations focusing on the keypoint region. The adaptive convolution is a modification of a normal convolution (e.g., convolution):
Here, is the center (D) position, and is the offset, corresponds to the th activated pixel. are the kernel weights.
Separate regression. The offset regressor in is a single branch, and estimates all the offsets together from a single feature for each position. We propose to use a -branch structure, where each branch performs the adaptive convolutions and then regresses the offset for the corresponding keypoint.
We divide the feature maps output from the backbone into feature maps, , and estimate the offset map for each keypoint from the corresponding feature map:
where is the th regressor on the th branch, and is the offset map for the th keypoint. The regressors, have the same structures, and their parameters are learned independently.
Each branch in separate regression is able to learn its own adaptive convolutions, and accordingly focuses on activating the pixels in the corresponding keypoint region (see Figure 4 (b - e)). In the single-branch case, the pixels around all the keypoints are activated, and the activation is not focused (see Figure 4 (a)).
The multi-branch structure explicitly decouples the representation learning for one keypoint from other keypoints, and thus improves the regression quality. In contrast, the single-branch structure has to decouple the feature learning implicitly which increases the optimization difficulty. Our results in Figure 5 show that the multi-branch structure reduces the regression loss.
2 Loss Function
Regression loss. We use the normalized smooth loss to form the pixel-wise keypoint regression loss:
Here, is the size of the corresponding person instance, and and are the height and the width of the instance box. is the set of the positions that have groundtruth poses. (), a column of the offset maps (), is the -dimensional estimated (groundtruth) offset vector for the position .
Keypoint and center heatmap estimation loss. We also estimate keypoint heatmaps each corresponding to a keypoint type and the center heatmap indicating the confidence that each pixel is the center of some person, using a separate heatmap estimation branch,
The heatmaps are used for scoring and ranking the regressed poses. The heatmap estimation loss function is formulated as the weighted distances between the predicted heat values and the groundtruth heat values:
Here, is the entry-wise -norm. is the element-wise product operation. has masks, and the size is . The th mask, , is formed so that the mask weight of the positions not lying in the th keypoint region is , and others are . The same is done for the mask for the center heatmap. and are the target keypoint and center heatmaps.
Whole loss. The whole loss function is the sum of the heatmap loss and the regression loss:
where is a trade-off weight, and set as in our experiments.
3 Inference
A testing image is fed into the network, outputting the regressed pose at each position, and the keypoint and center heatmaps. We first perform the center NMS process on the center heatmap to remove non-locally maximum positions and the positions whose center heat value is not higher than . Then we perform the pose NMS process over the regressed poses at the positions remaining after center NMS, to remove some overlapped regressed poses, and maintain at most candidates. The score used in pose NMS is the average of the heat values at the regressed keypoints, which is helpful to keep candidate poses with highly accurately localized keypoints.
We rank the remaining candidate poses using the score that is estimated by jointly considering their corresponding center heat values, keypoint heat values and their shape scores. The shape feature includes the distance and the relative offset between a pair of neighboring keypoints A neighboring pair corresponds to a stick in the COCO dataset, and there are sticks (denoted by ) in the COCO dataset.: and , and keypoint heat values indicating the visibility of each keypoint. We input the three kinds of features to a scoring net, consisting of two fully-connected layers (each followed by a ReLU layer), and a linear prediction layer, which aims to learn the OKS score for the corresponding predicted pose with the real OKS as the target on the training set.
4 Discussions
Separate regression, group convolution and complexities. In the multi-branch structure, we divide the channel maps into non-overlapping partitions, and feed each partition into each branch, for learning disentangled representations. This process resembles group convolution. The difference lies in: group convolution usually increases the capacity of the whole representation through reducing the redundancy and increasing the width within the computation and parameter budget, while our approach does not change the width and aims to learn rich representations focusing on each keypoint.
Let’s look at an example with HRNet-W as the backbone. The standard process in HRNet-W concatenates the channels obtained from resolutions and feeds the concatenated channels to a convolution, outputting channels. When applied to our disentangled regressors, we modify the convolution to output () channels, so that each partition has channels. This modification does not increase the width. The parameter complexity and the computation complexity for the regression head are reduced, and in particular the overall computation complexity is significantly reduced. The detailed numbers are given in Table 1.
Separate group regression. It is noticed that the salient regions and the activated pixels for some keypoints might have some overlapping. For example, the salient regions of the five keypoints in the head are overlapped, and three keypoints in the arms have similar characteristics. We investigate the performance if grouping some keypoints into a single branch instead of letting each branch handle one single keypoint. We consider two grouping schemes. (1) Five keypoints in the head use a single branch, and there are totally branches. (2) Five keypoints in the head use a single branch, the three keypoints in left arm (right arm, left leg, right leg) use a single branch. There are totally branches. Empirical results show that separate group regression performs worse than separate regression, e.g., the AP score for five branches decreases by on COCO validation with the backbone HRNet-W.
Experiments
Dataset. We evaluate the performance on the COCO keypoint detection task . The train set includes images and person instances annotated with keypoints, the val set contains images, and the test-dev set consists of images. We train the models on the train set and report the results on the val and test-dev sets.
Training set construction. The training sets consist of keypoint and center heatmaps, and offset maps.
Groundtruth keypoint and center heatmaps: The groundtruth keypoint heatmaps for each image contains maps, and each map corresponds to one keypoint type. We build them as done in : assigning a heat value using the Gaussian function centered at a point around each groundtruth keypoint. The center heatmap is similarly constructed and described in the following.
Groundtruth offset maps: The groundtruth offset maps for each image are constructed from all the poses . We use the th pose as an example and others are the same. We compute the center position and the offsets as the groundtruth offsets for the pixel corresponding to the center position. We use an expansion scheme to augment the center point to the center region: , which are the central positions around the center point with the radius , and accordingly update the offsets. The positions not lying in the region have no offset values.
Each central position has a confidence value indicating how confident it is the center and computed using the way forming the groundtruth center heatmap In case that one position belongs to two or more central regions, we choose only one central region whose center is the closest to that position.. The positions not lying in the region have zero heat value.
Evaluation metric. We follow the standard evaluation metrichttp://cocodataset.org/#keypoints-eval and use OKS-based metrics for COCO pose estimation. We report average precision and average recall scores with different thresholds and different object sizes: , , , , , , , and .
Training. The data augmentation follows and includes random rotation (), random scale () and random translation ($512\times 51232640\times 64048$ with random flipping as training samples.
Testing. We resize the short side of the images to and keep the aspect ratio between height and width, and compute the heatmap and pose positions by averaging the heatmaps and pixel-wise keypoint regressions of the original and flipped images. Following , we adopt three scales and in multi-scale testing. We average the three heatmaps over three scales and collect the regressed results from the three scales as the candidates.
2 Results
COCO Validation. Table 2 shows the comparisons of our method and other state-of-the-art methods. Table 3 presents the parameter and computation complexities for our approach and the representative top competitors, such as AE-Hourglass , PersonLab , and HrHRNet .
Our approach, using HRNet-W as the backbone, achieves AP score. Compared to the methods with similar GFLOPs, CenterNet-DLA and PersonLab (with the input size ), our approach achieves over improvement. In comparison to CenterNet-HG whose model size is far larger than HRNet-W, our gain is . Our baseline result (Table 5) is lower than CenterNet-HG that adopts post-processing to match the predictions to the closest keypoints identified from keypoint heatmaps. This implies that our gain comes from our methodology.
Our approach benefits from large input size and large model size. Our approach, with HRNet-W as the backbone and the input size , obtains the best performance and gain over HRNet-W. Compared with state-of-the-art methods, our approach gets gain over CenterNet-HG, gain over PersonLab (the input size ), gain over PifPaf whose GFLOPs are more than twice as many as ours, and gain over HrHRNet-W that uses higher resolution representations.
Following , we report the results with multi-scale testing. This brings about gain for HRNet-W, gain for HRNet-W.
COCO test-dev. The results of our approach and other state-of-the-art methods on the test-dev dataset are presented in Table 4. Our approach with HRNet-W as the backbone achieves AP score, and significantly outperforms the methods with the similar model size. Our approach with HRNet-W as the backbone gets the best performance , leading to gain over PersonLab, gain over PifPaf , and gain over HrHRNet .
With multi-scale testing, our approach with HRNet-W achieves , even better than PersonLab with a larger model size. Our approach with HRNet-W achieves AP score, much better than associative embedding , gain over PersonLab, and gain over HrHRNet .
3 Empirical Analysis
Ablation study. We study the effects of the two components: adaptive activation (AA) and separate regression (SR). We use the backbone HRNet-W as an example. The observations are consistent for HRNet-W.
The ablation study results are presented in Table 5. We can observe: (1) adaptive activation (AA) achieves the gain over the regression baseline (). (2) separate regression (SR) further improves the AP score by . (3) separate regression w/o adaptive activation gets AP gain. The whole gain is .
We further analyze how each component contributes to the performance improvement by using the coco-analyze tool . Four error types are studied: (i) Jitter error: small localization error; (ii) Miss error: large localization error; (iii) Inversion error: confusion between keypoints within an instance. (iv) Swap error: confusion between keypoints of different instances. The detailed definitions are in .
Table 5 shows the errors of the four types for four schemes. The two components, adaptive activation (AA) and separate regression (SR), mainly influence the two localization errors, Jitter and Miss. Adaptive activation reduces the Jitter error and the Miss error by and , respectively. Separate regression further reduces the two errors by and . The other two errors are almost not changed. This indicates that the proposed two components indeed improve the localization quality.
Comparison with grouping detected keypoints. It is reported in HigherHRNet that associative embedding with HRNet-W achieves an AP score on COCO validation. The regression baseline using the same backbone HRNet-W gets an lower AP score (Table 5). The proposed two components lead to an AP score , higher than associative embedding HRNet-W.
Matching regression to the closest keypoint detection. The CenterNet approach performs a post-processing step to refine the regressed keypoint positions by absorbing the regressed keypoint to the closest keypoint among the keypoints identified from the keypoint heatmaps.
We tried this absorbing scheme. The results are presented in Table 6. We can see that the absorbing scheme does not improve the performance in the single-scale testing case. The reason might be that the keypoint localization quality of our approach is very close to that of keypoint identification from the heatmap. In the multi-scale testing case, the absorbing scheme improves the results. The reason is that keypoint position regression is conducted separately for each scale and the absorbing scheme makes the regression results benefit from the heatmap improved from multiple scales. Our current focus is not on multi-scale testing whose practical value is not as high as single-scale testing. We leave finding a better multi-scale testing scheme as our future work.
4 CrowdPose
Dataset. We evaluate our approach on the CrowdPose dataset that is more challenging and includes many crowded scenes. The train set contains images, the val set includes images and the test set consists of images. We train our models on the CrowdPose train and val sets and report the results on the test set as done in .
Evaluation metric. The standard average precision based on OKS which is the same as COCO is adopted as the evaluation metrics. The CrowdPose dataset is split into three crowding levels: easy, medium, hard. We report the following metrics: , , , as well as , and for easy, medium and hard images.
Test set results. The results of our approach and other state-of-the-art methods on the test set are showed in Table 7. Our approach with HRNet-W as the backbone achieves AP and is better than HrHRNet-W () that is a keypoint detection and grouping approach with a backbone that is designed for improving heatmaps. With multi-scale testing, our approach with HRNet-W achieves 68.0 AP score and by a further matching process (see Table 6) the performance is improved, leading to 0.7 gain over HrHRNet-W .
Conclusions
The proposed direct regression approach DEKR improves the keypoint localization quality and achieves state-of-the-art bottom-up pose estimation results. The success stems from that we disentangle the representations for regressing different keypoints so that each representation focuses on the corresponding keypoint region. We believe that the idea of regression by focusing and disentangled keypoint regression can benefit some other methods, such as CornetNet and CenterNet for object detection.