Learning to Estimate 3D Human Pose and Shape from a Single Color Image

Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou, Kostas Daniilidis

Introduction

Estimating the full body 3D pose and shape of humans from images has been a challenging goal of computer vision going all the way back to the work of Hogg . The inherent ambiguity of the problem has forced the researchers to use monocular image sequences for inference , employ multiple camera views , or even explore alternative sensors, like Kinect or IMUs . In these settings, the body shape reconstruction results are remarkable. However, estimating 3D pose and shape from single color images remains the ultimate goal for 3D human analysis.

Considering the particularly challenging nature of such a problem, the literature remains undeniably sparse. Most approaches rely on iterative optimization, attempting to estimate a full body 3D shape that is consistent with 2D image observations, like silhouettes, edges, shading, or 2D keypoints . Despite the significant runtime required to solve the complicated optimization problem, the common failures because of local minima, and the error-prone reliance on ambiguous 2D cues, optimization-based solutions remain the leading paradigm for this problem . Even the emergence of deep learning has not changed significantly the landscape. ConvNets did not seem as a viable candidate for this problem because they require a huge amount of training data and they are infamous for their low resolution 3D predictions . The goal of our work is to demonstrate that ConvNets can indeed offer an attractive solution for this problem, by proposing an efficient and effective direct prediction approach, which is competitive and even outperforms iterative optimization methods.

To make this feasible, a critical design choice for our approach is the incorporation of a parametric statistical body shape model (SMPL ) within our end-to-end framework, presented in Figure 1. The advantage of such a representation is that we can generate high quality 3D meshes in the form of 6890 vertices while estimating only a small number of parameters, i.e., 72 for pose and 10 for shape. This low-dimensional parameterization makes the model friendly for direct network prediction. In fact, this prediction is feasible and accurate by using only 2D keypoints and silhouettes as input. This allows us to relax the limiting assumption that natural images with 3D shape ground truth are available for training. In contrast, we can leverage the available 2D image annotations (e.g., ) to train for image-to-2D inference, while using instances of the parametric model to train for 2D-to-3D shape inference. Simultaneously, another major advantage of employing this parametric model is that its structure allows us to generate the estimated 3D mesh at training time and optimize directly for the surface, by using a 3D per-vertex loss. This loss has better correlation with the vertex-to-vertex 3D error that is typically used for evaluation and improves training compared to naive parameter regression. Finally, we propose to employ a differentiable renderer to project the generated 3D mesh back to the 2D image. This enables end-to-end finetuning of the network by optimizing for the consistency of the projection with annotated 2D observations, i.e., 2D keypoints and masks. The complete framework offers a modular direct prediction solution to the problem of 3D human pose and shape estimation from a single color image and outperforms previous approaches on the relevant benchmarks.

Our main contributions can be summarized as follows:

an end-to-end framework for 3D human pose and shape estimation from a single color image.

incorporation of a parametric statistical shape model, SMPL, within the end-to-end framework, enabling:

prediction of the SMPL model parameters from ConvNet-estimated 2D keypoints and masks to avoid training on synthetic image examples.

generation of the 3D body mesh at training time and supervision based on the 3D shape consistency.

use of a differentiable renderer for 3D mesh projection and refinement of the network with supervision based on the consistency with 2D annotations.

superior performance compared to previous approaches for 3D human pose and shape estimation at significantly faster running time.

Related work

3D human pose estimation: In order to estimate a convincing 3D reconstruction of the human body, it is crucial to get an accurate prediction of the 3D pose of the person. Many recent works follow the end-to-end paradigm , using images as input to predict 3D joint locations , regress 3D heatmaps , or classify the image in a particular pose class . Unfortunately, an important constraint is that most of these ConvNets require images with 3D pose ground truth for training, limiting the available training data sources. Other approaches commit to the 2D pose estimates provided by state-of-the-art ConvNets and focus on the 3D pose reconstruction , recover 3D pose exemplars , or produce multiple 3D pose candidates consistent with the 2D pose . Notably, Martinez et al. demonstrate state-of-the-art results using a simple multi-layer perceptron which regresses the 3D joint locations from 2D pose input. Our goal is significantly different from the aforementioned works, since instead of a rough stickman-like figure, we estimate the whole surface geometry of the human body.

Human shape estimation: Concurrently with advances in 3D human pose, a different set of works addressed the problem of human shape estimation. In this case, given a single image, most methods attempt to estimate the parameters of a statistical body shape model like SCAPE or SMPL . The input is usually silhouettes, while regression forests and ConvNets have been proposed for the prediction. Knowledge of human shape is useful for biometric applications, however we argue that for 3D perception the potential and the challenges are significantly greater when pose and shape are inferred jointly.

Joint 3D human pose and shape estimation: Despite individual advances in pose and shape prediction, their joint estimation makes the task significantly harder. This has consistently fostered research in non single image scenarios, for more robust results. Xu et al. propose a pipeline for full performance capture from monocular video assuming knowledge of the shape mesh for the observed subject. Alldieck et al. estimate pose and shape jointly from monocular video relying on optical flow cues. Rhodin et al. and Huang et al. use images from multiple calibrated cameras and rely on keypoint detections, silhouettes and temporal consistency to recover a reconstruction of the body. An alternative setting is proposed by Weiss et al. making use of the depth modality of the Kinect sensor to tackle the same problem. In the same spirit of exploring different sensors, von Marcard et al. use a sparse set of IMUs on the subject to recover pose and shape jointly.

3D human pose and shape from a single color image: In the most challenging case of using only a single color image as input, the work of Sigal et al. is among the first to estimate high quality 3D shape estimates, by fitting the parametric model SCAPE to ground truth image silhouettes. Guan et al. use silhouettes, edges and shading as cues during the fitting process, but still require initialization through a user specified 2D skeleton. A fully automatic approach was proposed very recently by Bogo et al. . They use 2D keypoint detections from a 2D pose ConvNet and fit the parametric model SMPL to these 2D locations. Their 3D pose results are very accurate, but shape remains highly underconstrained. To improve upon this, Lassner et al. extends the fitting using silhouettes provided by a segmentation ConvNet. The common theme of these works is that they pose an optimization problem and attempt to fit a body model to a set of 2D observations. The drawback though is that solving this iterative optimization problem is very slow, it can easily fail because of local minima, and it relies a lot on error-prone 2D observations.

Alternatively, direct prediction approaches estimate 3D pose and shape in a discriminative way, without explicitly optimizing a specific objective during inference. Relevant to this paradigm is the work of Lassner et al. , where a ConvNet detects 91 landmarks of the human body and then a random forest estimates the 3D body and shape from these detections. However, to train for these landmarks, they still require alignment of body shapes with images. In contrast, we demonstrate that only a much smaller set of annotations are critical for the reconstruction, i.e., 2D joints and masks, which can be provided by human annotators and are abundant for in-the-wild images , while we also incorporate everything within a unified end-to-end framework. Concurrently, Tan et al. use an encoder-decoder ConvNet, where the decoder is trained to predict the silhouette corresponding to SMPL parameters. We differ to them by identifying that from these parameters we can analytically generate the body mesh and project it to the image in a differentiable way (as in for face models), avoiding half a million of extra learnable weights. Instead, we focus our computational and learning effort in the image to 3D shape part of the framework. Our work is also related to the concurrent work of Tung et al. , however our framework can be trained from scratch instead of relying on synthetic image data for pretraining, and we demonstrate state-of-the-art results for model-based 3D pose and shape prediction.

Human body shape models

Statistical body shape models, like SCAPE or SMPL , are powerful tools, which provide significant opportunities for an end-to-end framework. One of the important advantages is their low-dimensional parameter space, which is very suitable for direct network prediction. With this parameter representation, we can keep the output prediction space small, compared to voxelized or point cloud representations. Simultaneously, the low dimensional prediction does not sacrifice the quality of the output, since we can still generate high quality 3D meshes from the estimated parameters. Furthermore, from a learning perspective, we bypass the problem of learning the statistics of the human body, and devote the network capacity at the inference of the model parameters from image evidence. In contrast, approaches without the aid of a model put additional burden on the learning side, which often leads to embarrassing prediction errors (e.g., failing to reconstruct limbs under occlusion, missing body details, etc). Moreover, most models offer a convenient disentanglement of pose and shape which is useful to independently focus on the factors that affect each one of the two. Last but certainly not least for end-to-end approaches, the function which generates the 3D mesh from parameter inputs is differentiable, making the models compatible with current end-to-end pipelines.

Technical approach

The conventional ConvNet-based approach for our task would be to acquire a large amount of color images with 3D shape ground truth and train the network with these input-output pairs. However, except for small-scale datasets or synthetically generated image examples this type of data is typically unavailable. Therefore, to deal with this task, we need to rethink the typical pipeline. Our main goal is to leverage all the resources we have available and use our insights for the problem to build an effective framework. As a first step, from findings of prior work, we identify that 3D pose can be estimated reliably from 2D pose estimates , while the shape can be inferred from silhouette measurements . This observation conveniently decomposes the problem in a) estimation of keypoints and masks from color images and, b) prediction of 3D pose and shape from the 2D evidence. The advantage of this practice is that the framework can be trained without requiring images with 3D shape ground truth.

The first step of our framework focuses on 2D keypoint and silhouette estimation. This part is motivated by the availability of large-scale benchmarks with 2D joints and mask annotations. Considering the volume and the variability of this data, we leverage it to train a ConvNet for 2D pose and silhouette prediction, that is particularly reliable under various imaging conditions and poses.

In the past, two individual ConvNets have been used to provide 2D keypoints and masks . In contrast, for a more elegant solution, we train a single ConvNet, which we denote as Human2D, that generates two outputs, one for keypoints and one for silhouettes. Human2D follows the Stacked Hourglass design , using two hourglasses, which was found to be a good trade-off between accuracy and running time. The keypoint output is in the form of heatmaps , where an MSE loss, Lhm\mathcal{L}_{hm}, between the ground truth and the predicted heatmaps is used for supervision. The silhouette output has two channels (body and background) and is supervised using a pixelwise binary cross entropy loss, Lsil\mathcal{L}_{sil}. For training, we combine the two losses: Lhg=λLhm+Lsil\mathcal{L}_{hg}=\lambda\mathcal{L}_{hm}+\mathcal{L}_{sil}, where λ=100\lambda=100. This ConvNet falls under the multi-task learning paradigm . Through sharing, the two tasks might benefit each other, but multi-task learning can also pose certain challenges (e.g., appropriate weighting of the losses), as Kokkinos identifies .

2 3D pose and shape prediction

The second step is significantly more challenging, requiring estimation of the full body 3D pose and shape from 2D keypoints and silhouettes. Silhouettes and/or keypoints have been used extensively for 3D model fitting through iterative optimization . Here, we demonstrate that this mapping can also be learned from data while it is possible to get a reliable prediction in a single estimation step.

For this mapping, we train two network components: (a) the PosePrior, which uses 2D keypoint locations as input together with the confidence of the detections (realised by the maximum value of each heatmap) and estimates the pose coefficients θ\bm{\theta}, and (b) the ShapePrior, which uses the silhouette as input and estimates the shape coefficients β\bm{\beta}. In general, the silhouette can be helpful for 3D pose inference and vice versa . However, empirically we discovered this disentanglement to provide more stable and accurate 3D predictions, while it also leads to a more modular pipeline (e.g. updating only the PosePrior, without retraining the whole network). Regarding the architecture, the PosePrior uses two bilinear units , where the input is the 2D keypoint locations and the maximum responses from each heatmap, and the output is the 72 SMPL pose parameters θ\bm{\theta}. The ShapePrior uses a simple architecture with five 3×33\times 3 convolutional layers, each one followed by max-pooling, and an additional bilinear unit at the end with 10 outputs, corresponding to the SMPL shape parameters β\bm{\beta}.

The form of the input (2D keypoints and masks) and the output (shape and pose parameters) allows us to produce large amount of training data by generating instances of the SMPL model with different 3D pose and shape (Figure 2). In fact, we can leverage MoCap data (e.g., ) to sample 3D poses, and body scans (e.g., ) to sample body shapes. For the input, we only need to project the 3D model to the image plane (possibly from different viewpoints), and compute silhouettes and 2D keypoint locations to generate input-output pairs for training. This data generation is feasible, exactly because we used an intermediate silhouette and keypoints representation. In contrast, attempting to learn a mapping directly from color images would require generation of synthetic image examples , which typically do not reach the variability of in-the-wild images.

In the previous paragraphs, we deliberately avoided discussing the supervision of the Priors networks. Past works have examined supervision schemes using a typical L2\mathcal{L}_{2} loss between the predicted and ground truth parameters. One shortcoming of this naive parameter regression approach, is that different parameters might have effects of different scale on the final reconstruction (e.g., the global body rotation is much more crucial than the local rotation of the hand with respect to the wrist). To avoid hand-selecting or tuning the supervision for each parameter, we aim for a more global solution. Our approach entails the generation of the full body mesh at training time, where we optimize explicitly for the predicted surface by applying a 3D per-vertex loss. Since the function M(β,θ;Φ)\mathcal{M}(\bm{\beta},\bm{\theta};\Phi) is differentiable, we can backpropagate through it and handle this mesh generator as a typical layer of our network, without any learnable parameters. Given the predicted mesh vertices P^i\hat{P}_{i} and the corresponding groundturth vertices PiP_{i}, we can supervise the network with a 3D per-vertex loss:

which considers all the vertices equally and has better correlation with the 3D per-vertex error which is usually employed for evaluation. Alternatively, if the focus is mainly on 3D pose, we can also supervise the network considering only the MM relevant 3D joints JiJ_{i}, which are trivially exposed by the model as a sparse linear combination of the mesh vertices. In this case, denoting with J^i\hat{J}_{i} the estimated joints, the corresponding loss can be expressed as:

Empirically, we found that the best training strategy is to initially get a reasonable initialization for the network parameters using an L2\mathcal{L}_{2} parameter loss, and then activate also the vertex loss LM\mathcal{L_{\mathcal{M}}} (or the joints loss LJ\mathcal{L_{\mathcal{J}}} if the focus is on pose only), to train a better model.

3 Differentiable renderer

Our previous analysis relaxed the assumption that images with 3D shape ground truth are available for training and relied on geometric 3D data (MoCap and body scans). In some cases though, even this type of data might be unavailable. For example, LSP has gymnastics or parkour poses which are not represented in typical MoCap. Luckily, our generated 3D mesh has potential to leverage these 2D annotations for training purposes.

where μ=10\mu=10. The goal of this type of supervision is twofold: (a) it can be employed for end-to-end refinement of the network, using only images with 2D keypoints and/or masks for training, and (b) it can be useful to mildly adapt a generic pose or shape prior to a new setting (e.g., new dataset), where only 2D annotations are available.

Empirical evaluation

This section focuses on the empirical evaluation of the proposed approach. First, we present the benchmarks that we employed for quantitative and qualitative evaluation. Then, we provide some essential implementation details of the approach. Finally, quantitative and qualitative results are presented on the selected datasets.

For the empirical evaluation, we employed two recent benchmarks that provide color images with 3D body shape ground truth, the UP-3D dataset and the SURREAL dataset . Additionally, we used the Human3.6M dataset for further evaluation of the 3D pose accuracy.

UP-3D: It is a recent dataset that collects color images from 2D human pose benchmarks, like LSP and MPII and uses an extended version of SMPLify to provide 3D human shape candidates. The candidates were evaluated by human annotators to select only the images with good 3D shape fits. It comprises 8515 images, where 7818 are used for training and 1389 for testing. We report results on this test set, while we also consider subsets, based on the original dataset (LSP, MPII, or FashionPose) of the UP-3D images. Finally, we examine a reduced test set of 139 images, selected by Tan et al. aiming to limit the range for the global rotation. We report results using the mean per-vertex error, between predicted and ground truth shape.

SURREAL: It is a recent dataset which provides synthetic image examples with 3D shape ground truth. The dataset draws poses from MoCap and body shapes from body scans to generate valid SMPL instances for each image. The synthetic images are not very realistic, but the accurate ground truth, makes it a useful benchmark for evaluation. We report results on the Human3.6M part of the dataset, considering all test videos and keeping every fifth frame of each video to avoid excessive redundancy in the data. Results are reported using the mean per-vertex error.

Human3.6M: It is a large-scale indoor dataset that contains multiple subjects performing typical actions like “Eating” and “Walking”. We follow the protocol of Bogo et al. using all videos of subjects S9 and S11 from ‘cam3’ for evaluation. The original videos are downsampled from 50fps to 10fps to remove redundancy as is done in . Results are reported using the reconstruction error.

2 Implementation details

The Human2D network is trained on MPII , LSP and LSP-extended data, using the silhouettes from Lassner et al. . We use a batch size of 4, learning rate set to 3e-4, and rmsprop for the optimization. Augmentation for rotation (±30∘\pm 30^{\circ}), scale (0.75-1.25) and flipping (left-right) is used. The training lasts for 1.2M iterations.

For the Priors networks, we train with a batch size of 256, learning rate set to 3e-4, and using rmsprop for the optimization. Initially, the networks are trained for 40k iterations using an L2\mathcal{L}_{2} parameter loss, and then for 60k more iterations using also LM\mathcal{L}_{\mathcal{M}} (or LJ\mathcal{L}_{\mathcal{J}} if we focus on pose only) weighted equally with the parameter loss.

The end-to-end refinement with the reprojection loss lasts for 2k iterations with a batch size of 4, learning rate set to 8e-5, and using rmsprop for the optimization. To improve training robustness, the end-to-end updates are alternated with individual updates of the Human2D and the Priors networks (as described in the previous two paragraphs). This helps the individual components to maintain their original purpose, while we are also leveraging the strength of end-to-end training to integrate them together.

3 Component evaluation

In this section, we evaluate the components of our approach, using the UP-3D dataset. We train two different versions of our system, where for Priors we leverage data either from UP-3D (provided by Lassner et al. ), or from CMU MoCap (provided by Varol et al. ). The Human2D network remains the same in both cases.

Our experiment focuses on the type of supervision. Naively training the Priors networks using an L2\mathcal{L}_{2} loss for the θ\bm{\theta} and β\bm{\beta} parameters , keeps the prediction error high as can be seen in Table 1 (line 1). Alternatively, we can transform the θ\bm{\theta} parameters from axis-angle representation to rotation matrix using the Rodrigues’ rotation formula , and apply an L2\mathcal{L}_{2} loss on this representation instead (line 2). This leads to more stable training and better performance, as has also been observed by Lassner et al. . However, generating the body mesh and further training of the network using our proposed per-vertex supervision (line 3) is even more appropriate and elevates our framework to state-of-the-art performance (see Section 5.4). Finally, the additional end-to-end finetuning with 2D annotations and the reprojection error (line 4) offers a mild refinement to the network. In the UP-3D case, the benefit is small, since the Priors have already observed very similar examples with full 3D ground truth, so 2D annotations become redundant. However, when training the Priors with CMU data, the domain shift, from CMU poses to UP-3D poses is significant, so these 2D annotations offers a clear performance benefit. This is an interesting empirical result demonstrating that training with reprojection losses can be useful not only for end-to-end refinement, but it can also assist the network with novel information recovered from 2D annotations. Some qualitative results from UP-3D using our best model are presented in Figure 3.

4 Comparison with state-of-the-art

UP-3D: We compare with two state-of-the-art direct prediction approaches by Lassner et al. and Tan et al. . We do not include the SMPLify method since a version of this algorithm was used to generate the ground truth for this dataset, so we observed that many estimated reconstructions had only minimal differences from the ground truth. For we use the publicly available code to generate predictions. The complete results are presented in Table 2. Our approach outperforms the other two baselines by significant margins. It is interesting to note that a version of , which uses over 100k images (most of them synthetic) with ground truth pose and shape parameters to directly supervise the network (line ‘Direct’) is outperformed by our approach which does not have access to this data. Finally, in Figure 3, we provide a qualitative comparison with our closest competitor, the direct prediction approach of .

SURREAL: We compare with two state-of-the-art approaches, one based on iterative optimization, SMPLify , and one based on direct prediction . We use the publicly available code for both approaches to generate predictions. For our approach, we train the PosePrior using CMU data which we found to be more general than UP-3D. Also, we train two ShapePriors, for female and male subjects respectively, since the gender is known for this dataset. We emphasize that the testing was conducted on the Human3.6M part of the dataset to avoid any overlap with the training of the different methods (in terms of images or priors). The complete results are presented in Table 3. Since Lassner et al. provide only a non gender-specific model for shape, we also report results considering only the pose estimates, and assuming known shape parameters. Our approach outperforms the other two baselines. For this dataset we observed that because of the challenging color images (low illumination, out-of-context backgrounds, etc), the 2D detections where more noisy than usual, providing some hard failures for the iterative optimization approach . In contrast, our approach was more resistant to these noisy cases recovering a coherent 3D shape in most cases.

Human3.6M: Finally, for Human3.6M we evaluate only the estimated 3D pose, since there is no body shape ground truth available. Our network is the same as before (Priors trained on CMU), although, we use the 3D joints error for supervision (equation 2), since the focus is on pose. Among others, we compare with the SMPLify method and the direct prediction approach of Lassner et al. . Similarly to the other approaches we compare with, we do not use any data from this dataset for training. The detailed results are presented in Table 4. Our approach again outperforms the other baselines. Some works have reported better results results on Human3.6M (e.g., ), but they do so only by leveraging the training data of this dataset for training.

5 Boosting SMPLify

In the previous section, we validated that our direct prediction approach can achieve state-of-the-art results with a single prediction step. However, we aspire our method to have greater applicability, by being complementary to iterative optimization solutions. In fact, here we demonstrate that our direct predictions can be a useful initialization and provide a reliable anchor for the SMPLify approach .

To keep it simple, we make only minor modifications to the SMPLify optimization. First, we use our predicted pose as an initialization, instead of the typical mean pose. Additionally, we avoid the hierarchical four-step optimization, and we limit the whole procedure in a single step. The reason for the multi-stage optimization is to explore the pose space and get a roughly correct pose estimate. However, using our predicted pose as initialization makes this search unnecessary, so we require only the last step of the previously complex optimization scheme. Finally, we add one more data term to the optimization: Eanchor(θ)=∑iρ(θi−θiinit)E_{anchor}(\bm{\theta})=\sum_{i}\rho(\theta_{i}-\theta^{init}_{i}), to avoid deviations from our predicted, anchor pose. Similarly to , we use the Geman-McClure penalty function, ρ\rho , for the optimization. This anchoring, does not typically have effect on the quality of the output, but it can accelerate the convergence. We can also use the shape parameters as anchor, but we observed that pose had greater effect than shape on the optimization.

For our evaluation, we use the public implementation of SMPLify and we run the original code, as well as our anchored version, on the LSP test set. The anchored version is three times faster on average than vanilla SMPLify. More importantly, this speedup comes also with a quantitative performance benefit. In Table 5 we present the segmentation accuracy of different SMPLify versions, by projecting the 3D shape estimate on the image. To demonstrate that the performance benefit of our anchored version is non-trivial, we report the results for running SMPLify on the ground truth 2D joints and silhouettes. Improved fits from the anchored version are presented in Figure 5. These results validate the additional benefit of our direct prediction approach, since it can also enhance current pipelines that rely on iterative optimization.

6 Running time

Our approach requires a single forward pass from the ConvNet to estimate the full body 3D human pose and shape. This translates to only 50ms on a Titan X GPU. In comparison, SMPLify report roughly 1 minute for the optimization, while the publicly available (unoptimized) code runs on 3 minutes per image on average. When the number of landmarks increases to 91, Lassner et al. report that the SMPLify optimization can get two times slower. This makes our direct prediction approach more than three orders of magnitude faster than the state-of-the-art iterative optimization approaches. Regarding other direct prediction approaches, Lassner et al. reports runtime of 378ms, but we demonstrate significantly better performance with our end-to-end framework.

Summary

The goal of this paper was to present a viable ConvNet-based approach to predict 3D human pose and shape from a single color image. A central part of our solution was the incorporation of a body shape model, SMPL, in the end-to-end framework. Through this inclusion we enabled: a) prediction of the parameters from 2D keypoints and silhouettes, b) generation of the full body 3D mesh at training time using supervision for the surface with a per-vertex loss, and c) integration of a differentiable renderer for further end-to-end refinement using 2D annotations. Our approach achieved state-of-the-art results on relevant benchmarks, outperforming previous direct prediction and optimization-based solutions for 3D pose and shape prediction. Finally, considering the efficiency of our approach, we demonstrated its potential to accelerate and improve typical iterative optimization pipelines.

Project Page: https://www.seas.upenn.edu/~pavlakos/projects/humanshape

Acknowledgements: We gratefully appreciate support through the following grants: NSF-IIP-1439681 (I/UCRC), ARL RCTA W911NF-10-2-0016, ONR N00014-17-1-2093, DARPA FLA program and NSF/IUCRC.

References