End-to-end Recovery of Human Shape and Pose
Angjoo Kanazawa, Michael J. Black, David W. Jacobs, Jitendra Malik
Introduction
We present an end-to-end framework for recovering a full 3D mesh of a human body from a single RGB image. We use the generative human body model, SMPL , which parameterizes the mesh by 3D joint angles and a low-dimensional linear shape space. As illustrated in Figure 1, estimating a 3D mesh opens the door to a wide range of applications such as foreground and part segmentation, which is beyond what is practical with a simple skeleton. The output mesh can be immediately used by animators, modified, measured, manipulated and retargeted. Our output is also holistic – we always infer the full 3D body even in cases of occlusion and truncation.
Note that there is a great deal of work on the 3D analysis of humans from a single image. Most approaches, however, focus on recovering 3D joint locations. We argue that these joints alone are not the full story. Joints are sparse, whereas the human body is defined by a surface in 3D space.
Additionally, joint locations alone do not constrain the full DoF at each joint. This means that it is non-trivial to estimate the full pose of the body from only the 3D joint locations. In contrast, we output the relative 3D rotation matrices for each joint in the kinematic tree, capturing information about 3D head and limb orientation. Predicting rotations also ensures that limbs are symmetric and of valid length. Our model implicitly learns the joint angle limits from datasets of 3D body models.
Existing methods for recovering 3D human mesh today focus on a multi-stage approach . First they estimate 2D joint locations and, from these, estimate the 3D model parameters. Such a stepwise approach is typically not optimal and here we propose an end-to-end solution to learn a mapping from image pixels directly to model parameters.
There are several challenges, however, in training such a model in an end-to-end manner. First is the lack of large-scale ground truth 3D annotation for in-the-wild images. Existing datasets with accurate 3D annotations are captured in constrained environments. Models trained on these datasets do not generalize well to the richness of images in the real world. Another challenge is in the inherent ambiguities in single-view 2D-to-3D mapping. Most well known is the problem of depth ambiguity where multiple 3D body configurations explain the same 2D projections . Many of these configurations may not be anthropometrically reasonable, such as impossible joint angles or extremely skinny bodies. In addition, estimating the camera explicitly introduces an additional scale ambiguity between the size of the person and the camera distance.
In this paper we propose a novel approach to mesh reconstruction that addresses both of these challenges. A key insight is that there are large-scale 2D keypoint annotations of in-the-wild images and a separate large-scale dataset of 3D meshes of people with various poses and shapes. Our key contribution is to take advantage of these unpaired 2D keypoint annotations and 3D scans in a conditional generative adversarial manner. The idea is that, given an image, the network has to infer the 3D mesh parameters and the camera such that the 3D keypoints match the annotated 2D keypoints after projection. To deal with ambiguities, these parameters are sent to a discriminator network, whose task is to determine if the 3D parameters correspond to bodies of real humans or not. Hence the network is encouraged to output parameters on the human manifold and the discriminator acts as weak supervision. The network implicitly learns the angle limits for each joint and is discouraged from making people with unusual body shapes.
An additional challenge in predicting body model parameters is that regressing to rotation matrices is challenging. Most approaches formulate rotation estimation as a classification problem by dividing the angles into bins . However differentiating angle probabilities with respect to the reprojection loss is non-trivial and discretization sacrifices precision. Instead we propose to directly regress these values in an iterative manner with feedback. Our framework is illustrated in Figure 2.
Our approach is similar to 3D interpreter networks in the use of reprojection loss and the more recent adversarial inverse graphics networks for the use of the adversarial prior. We go beyond the existing techniques in multiple ways:
We infer 3D mesh parameters directly from image features, while previous approaches infer them from 2D keypoints. This avoids the need for two stage training and also avoids throwing away valuable information in the image such as context.
Going beyond skeletons, we output meshes, which are more complex and more appropriate for many applications. Again, no additional inference step is needed.
Our framework is trained in an end-to-end manner. We out-perform previous approaches that output 3D meshes in terms of 3D joint error and run time.
We show results with and without paired 2D-to-3D data. Even without using any paired 2D-to-3D supervision, our approach produces reasonable 3D reconstructions. This is most exciting because it opens up possibilities for learning 3D from large amounts of 2D data.
Since there are no datasets for evaluating 3D mesh reconstructions of humans from in-the-wild images, we are bound to evaluate our approach on the standard 3D joint location estimation task. Our approach out performs previous methods that estimate SMPL parameters from 2D joints and is competitive with approaches that only output 3D skeletons. We also evaluate our approach on an auxiliary task of human part segmentation. We qualitatively evaluate our approach on challenging images in-the-wild and show results sampled at different error percentiles. Our model and code is available for research purposes at https://akanazawa.github.io/hmr/.
Related Work
Many papers formulate human pose estimation as the problem of locating the major 3D joints of the body from an image, a video sequence, either single-view or multi-view. We argue that this notion of “pose” is overly simplistic but it is the major paradigm in the field. The approaches are split into two categories: two-stage and direct estimation.
Two stage methods first predict 2D joint locations using 2D pose detectors or ground truth 2D pose and then predict 3D joint locations from the 2D joints either by regression or model fitting, where a common approach exploits a learned dictionary of 3D skeletons . In order to constrain the inherent ambiguity in 2D-to-3D estimation, these methods use various priors . Most methods make some assumption about the limb-length or proportions . Akhter and Black learn a novel pose prior that captures pose-dependent joint angle limits. Two stage-methods have the benefit of being more robust to domain shift, but rely too much on 2D joint detections and may throw away image information in estimating 3D pose.
Video datasets with ground truth motion capture like HumanEva and Human3.6M define the problem in terms of 3D joint locations. They provide training data that lets the 3D joint estimation problem be formulated as a standard supervised learning problem. Thus, many recent methods estimate 3D joints directly from images in a deep learning framework . Dominant approaches are fully-convolutional, except for the very recent method of Xiao et al. that regresses bones and obtains excellent results on the 3D pose benchmarks. Many methods do not solve for the camera, but estimate the depth relative to root and use a predefined global scale based the average length of bones . Recently Rogez et al. combine human detection with 3D pose prediction. The main issue with these direct estimation methods is that images with accurate ground truth 3D annotations are captured in controlled MoCap environments. Models trained only on these images do not generalize well to the real world.
Weakly-supervised 3D: Recent work tackles this problem of the domain gap between MoCap and in-the-wild images in an end-to-end framework. Rogez and Schmid artificially endow 3D annotations to images with 2D pose annotation using MoCap data. Several methods train on both in-the-wild and MoCap datasets jointly. Still others use pre-trained 2D pose networks and also use 2D pose prediction as an auxiliary task. When 3D annotation is not available, Zhou et al. gain weak supervision from a geometric constraint that encourages relative bone lengths to stay constant. In this work, we output 3D joint angles and 3D shape, which subsumes these constraints that the limbs should be symmetric. We employ a much stronger form of weak supervision by training an adversarial prior.
Methods that output more than 3D joints: There are multiple methods that fit a parametric body model to manually extracted silhouettes and a few manually provided correspondences . More recent works attempt to automate this effort. Bogo et al. propose SMPLify, an optimization-based method to recover SMPL parameters from 14 detected 2D joints that leverages multiple priors. However, due to the optimization steps the approach is not real-time, requiring 20-60 seconds per image. They also make a priori assumptions about the joint angle limits. Lassner et al. take curated results from SMPLify to train 91 keypoint detectors corresponding to traditional body joints and points on the surface. They then optimize the SMPL model parameters to fit the keypoints similarly to . They also propose a random forest regression approach to directly regress SMPL parameters, which reduces run-time at the cost of accuracy. Our approach out-performs both methods, directly infers SMPL parameters from images instead of detected 2D keypoints, and runs in real time.
VNect fits a rigged skeleton model over time to estimated 2D and 3D joint locations. While they can recover 3D rotations of each joint after optimization, we directly output rotations from images as well as the surface vertices. Similarly Zhou et al. directly regress joint rotations of a fixed kinematic tree. We output shape as well as the camera scale and out-perform their approach in 3D pose estimation.
There are other related methods that predict SMPL-related outputs: Varol et al. use a synthetic dataset of rendered SMPL bodies to learn a fully convolutional model for depth and body part segmentation. DenseReg similarly outputs a dense correspondence map for human bodies. Both are 2.5D projections of the underlying 3D body. In this work, we recover all SMPL parameters and the camera, from which all of these outputs can be obtained.
Kulkarni et al. use a generative model of body shape and pose with a probabilistic programming framework to estimate body pose from single image. They deal with visually simple images and do not evaluate 3D pose accuracy. More recently Tan et al. infer SMPL parameters by first learning a silhouette decoder of SMPL parameters using synthetic data, and then learning an image encoder with the decoder fixed to minimize the silhouette reprojection loss. However, the reliance on silhouettes limits their approach to frontal images and images of humans without any occlusion. Concurrently Tung et al. predict SMPL parameters from an image and a set of 2D joint heatmaps. The model is pretrained on a synthetic dataset and fine-tuned at test time over two consecutive video frames to minimize the reprojection loss of keypoints, silhouettes and optical flow. Our approach can be trained without any paired supervision, does not require 2D joint heatmaps as an input and we test on images without fine-tuning. Additionally, we also demonstrate our approach on images of humans in-the-wild with clutter and occlusion.
Model
We propose to reconstruct a full 3D mesh of a human body directly from a single RGB image centered on a human in a feedforward manner. During training we assume that all images are annotated with ground truth 2D joints. We also consider the case in which some have 3D annotations as well. Additionally we assume that there is a pool of 3D meshes of human bodies of varying shape and pose. Since these meshes do not necessarily have a corresponding image, we refer to this data as unpaired .
Figure 2 shows the overview of the proposed network architecture, which can be trained end-to-end. Convolutional features of the image are sent to the iterative 3D regression module whose objective is to infer the 3D human body and the camera such that its 3D joints project onto the annotated 2D joints. The inferred parameters are also sent to an adversarial discriminator network whose task is to determine if the 3D parameters are real meshes from the unpaired data. This encourages the network to output 3D human bodies that lie on the manifold of human bodies and acts as a weak-supervision for in-the-wild images without ground truth 3D annotations. Due to the rich representation of the 3D mesh model, this data-driven prior can capture joint angle limits, anthropometric constraints (e.g. height, weight, bone ratios), and subsumes the geometric priors used by models that only predict 3D joint locations . When ground truth 3D information is available, we may use it as an intermediate loss. In all, our overall objective is
where is an orthographic projection.
2 Iterative 3D Regression with Feedback
The goal of the 3D regression module is to output given an image encoding such that the joint reprojection error
However, directly regressing in one go is a challenging task, particularly because includes rotation parameters. In this work, we take inspiration from previous works and regress in an iterative error feedback (IEF) loop, where progressive changes are made recurrently to the current estimate. Specifically, the 3D regression module takes the image features and the current parameters as an input and outputs the residual . The parameter is updated by adding this residual to the current estimate . The initial estimate is set as the mean . In the estimates are rendered to an image space to concatenate with the image input. In this work, we keep everything in the latent space and simply concatenate the features as the input to the regressor. We find that this works well and is suitable when differentiable rendering of the parameters is non-trivial.
Additional direct 3D supervision may be employed when paired ground truth 3D data is available. The most common form of 3D annotation is the 3D joints. Supervision in terms of SMPL parameters [\mbox{\boldmath\beta},\mbox{\boldmath\theta}] may be obtained through MoSh when raw 3D MoCap marker data is available. Below are the definitions of the 3D losses. We show results with and without using any direct supervision .
Both use a “bounded” correction target to supervise the regression output at each iteration. However this assumes that the ground truth estimate is always known, which is not the case in our setup where many images do not have ground truth 3D annotations. As noted by these approaches, supervising each iteration with the final objective forces the regressor to overshoot and get stuck in local minima. Thus we only apply and on the final estimate , but apply the adversarial loss on the estimate at every iteration forcing the network to take corrective steps that are on the manifold of 3D human bodies.
3 Factorized Adversarial Prior
The reprojection loss encourages the network to produce a 3D body that explains the 2D joint locations, however anthropometrically implausible 3D bodies or bodies with gross self-intersections may still minimize the reprojection loss. To regularize this, we use a discriminator network that is trained to tell whether SMPL parameters correspond to a real body or not. We refer to this as an adversarial prior as in since the discriminator acts as a data-driven prior that guides the 3D inference.
A further benefit of employing a rich, explicit 3D representation like SMPL is that we precisely know the meaning of the latent space. In particular SMPL has a factorized form that we can take advantage of to make the adversary more data efficient and stable to train. More concretely, we mirror the shape and pose decomposition of SMPL and train a discriminator for shape and pose independently. The pose is based on a kinematic tree, so we further decompose the pose discriminators and train one for each joint rotation. This amounts to learning the angle limits for each joint. In order to capture the joint distribution of the entire kinematic tree, we also learn a discriminator that takes in all the rotations. Since the input to each discriminator is very low dimensional (10-D for , 9-D for each joint and -D for all joints), they can each be small networks, making them rather stable to train. All pose discriminators share a common feature space of rotation matrices and only the final classifiers are learned separately.
Unlike previous approaches that make a priori assumptions about the joint limits , we do not predefine the degrees of freedom of the kinematic skeleton model. Instead this is learned in a data-driven manner through this factorized adversarial prior. Without the factorization, the network does not learn to properly regularize the pose and shape, producing visually displeasing results. The importance of the adversarial prior is paramount when no paired 3D supervision is available. Without the adversarial prior the network produces totally unconstrained human bodies as we show in section 4.3.
While mode collapse is a common issue in GANs we do not really suffer from this because the network not only has to fool the discriminator but also has to minimize the reprojection error. The images contain all the modes and the network is forced to match them all. The factorization may further help to avoid mode collapse since it allows generalization to unseen body shape and poses combinations.
In all we train discriminators. Each discriminator outputs values between , representing the probability that came from the data. In practice we use the least square formulation for its stability. Let represent the encoder including the image encoder and the 3D module. Then the adversarial loss function for the encoder is
and the objective for each discriminator is
We optimize and all s jointly.
4 Implementation Details
Datasets: The in-the-wild image datasets annotated with 2D keypoints that we use are LSP, LSP-extended MPII and MS COCO . We filter images that are too small or have less than 6 visible keypoints and obtain training sets of sizes , , and images respectively. We use the standard train/test split of these datasets. All test results are obtained using the ground truth bounding box.
For the 3D datasets we use Human3.6M and MPI-INF-3DHP . We leave aside sequences from training Subject 8 of MPI-INF-3DHP as the validation set to tune hyper-parameters, and use the full training set for the final experiments. Both datasets are captured in a controlled environment and provide 150k training images with 3D joint annotations. For Human3.6M, we also obtain ground truth SMPL parameters for the training images using MoSh from the raw 3D MoCap markers. The unpaired data used to train the adversarial prior comes from MoShing three MoCap datasets: CMU , Human3.6M training set and the PosePrior dataset , which contains an extensive variety of extreme poses. These consist of 390k, 150k and 180k samples respectively.
All images are scaled to preserving the aspect ratio such that the diagonal of the tight bounding box is roughly 150px (see ). The images are randomly scaled, translated, and flipped. Mini-batch size is 64. When paired 3D supervision is employed each mini-batch is balanced such that it consists of half 2D and half 3D samples. All experiments use all datasets with paired 3D loss unless otherwise specified.
The definition of the joints in SMPL do not align perfectly with the common joint definitions used by these datasets. We follow and use a regressor to obtain the 14 joints of Human3.6M from the reconstructed mesh. In addition, we also incorporate the 5 face keypoints from the MS COCO dataset . New keypoints can easily be incorporated with the mesh representation by specifying the corresponding vertex IDsThe vertex ids in 0-indexing are nose: 333, left eye: 2801, right eye: 6261, left ear: 584, right ear: 4072.. In total the reprojection error is computed over keypoints.
Experimental Results
Although we recover much more than 3D skeletons, evaluating the result is difficult since no ground truth mesh 3D annotations exist for current datasets. Consequently we evaluate quantitatively on the standard 3D joint estimation task. We also evaluate an auxiliary task of body part segmentation. In Figure 1 we show qualitative results on challenging images from MS COCO with occlusion, clutter, truncation, and complex poses. Note how our model recovers head and limb orientations. In Figure 3 we show results on the test set of Human3.6M, MPI-INF-3DHP, LSP and MS COCO at various error percentiles. Our approach recovers reasonable reconstructions even at 95th percentile error. Please see the project websitehttps://akanazawa.github.io/hmr/ for more results. In all figures, results on the model trained with and without paired 2D-to-3D supervision are rendered in light blue and light pink colors respectively.
We evaluate 3D joint error on Human3.6M, a standard 3D pose benchmark captured in a lab environment. We also compare with the more recent MPI-INF-3DHP , a dataset covering more poses and actor appearances than Human3.6M. While the dataset is more diverse, it is still far from the complexity and richness of in-the-wild images.
We report using several error metrics that are used for evaluating 3D joint error. Most common evaluations report the mean per joint position error (MPJPE) and Reconstruction error, which is MPJPE after rigid alignment of the prediction with ground truth via Procrustes Analysis ). Reconstruction error removes global misalignments and evaluates the quality of the reconstructed 3D skeleton.
Human3.6M We evaluate on two common protocols. The first, denoted P1, is trained on 5 subjects (S1, S5, S6, S7, S8) and tested on 2 (S9, S11). Following previous work , we downsample all videos from 50fps to 10fps to reduce redundancy. The second protocol , P2, uses the same train/test set, but is tested only on the frontal camera (camera 3) and reports reconstruction error.
We compare results for our method (HMR) on P2 in Table 1 with two recent approaches that also output SMPL parameters from a single image. Both approaches require 2D keypoint detection as input and we out-perform both by a large margin. We show results on P1 in Table 2. Here we also out-perform the recent approach of Zhou et al. , which also outputs 3D joint angles in a kinematic tree instead of joint positions. Note that they specify the DoF of each joint by hand, while we learn this from data. They also assume a fixed bone length while we solve for shape. HMR is competitive with recent state-of-the-art methods that only predict the 3D joint locations.
We note that MPJPE does not appear to correlate well with the visual quality of the results. We find that many results with high MPJPE appear quite reasonable as shown in Figure 3, which shows results at various error percentiles.
MPI-INF-3DHP The test set of MPI-INF-3DHP consists of 2929 valid frames from 6 subjects performing 7 actions. This dataset is collected indoors and outdoors with a multi-camera marker-less MoCap system. Because of this, the ground truth 3D annotations have some noise. In addition to MPJPE, we report the Percentage of Correct Keypoints (PCK) thresholded at 150mm and the Area Under the Curve (AUC) over a range of PCK thresholds .
The results are shown in Table 3. All methods use the perspective correction of . We also report metrics after rigid alignment for HMR and VNect using the publicly available code . We report VNect results without post-processing optimization over time. Again, we are competitive with approaches that are trained to output 3D joints and we improve upon VNect after rigid alignment.
2 Human Body Segmentation
We also evaluate our approach on the auxiliary task of human body segmentation on the 1000 test images of LSP labeled by . The images have labels for six body part segments and the background. Note that LSP contains complex poses of people playing sports and no ground truth 3D labels are available for training. We do not use the segmentation label during training either.
We report the segmentation accuracy and average F1 score over all parts including the background as done in . We also report results on foreground-background segmentation. Note that the part definition segmentation of the SMPL mesh is not exactly the same as that of annotation; this limits the best possible accuracy to be less than 100%.
Results are shown in Table 4. Our results are comparable to the SMPLify oracle , which uses ground truth segmentation and keypoints as the optimization target. It also out-performs the Decision Forests of . Note that HMR is also real-time given a bounding box.
3 Without Paired 3D Supervision
So far we have used paired 2D-to-3D supervision, i.e. whenever available. Here we evaluate a model trained without any paired 3D supervision. We refer to this setting as HMR unpaired and report numerical results in all the tables. All methods that report results on the 3D joint estimation task rely on direct 3D supervision and cannot train without it. Even methods that are based on a reprojection loss require paired 2D-to-3D training data.
The results are surprisingly competitive given this challenging setting. Note that the adversarial prior is essential for training without paired 2D-to-3D data. Figure 5 shows that a model trained with neither the paired 3D supervision nor the adversarial loss produces monsters with extreme shape and poses. It remains open whether increasing the amount of 2D data will significantly increase 3D accuracy.
Conclusion
In this paper we present an end-to-end framework for recovering a full 3D mesh model of a human body from a single RGB image. We parameterize the mesh in terms of 3D joint angles and a low dimensional linear shape space, which has a variety of practical applications. In this past few years there has been rapid progress in single-view 3D pose prediction on images captured in a controlled environment. Although the performance on these benchmarks is starting to saturate, there has not been much progress on 3D human reconstruction from images in-the-wild. Our results without using any paired 3D data are promising since they suggest that we can keep on improving our model using more images with 2D labels, which are relatively easy to acquire, instead of ground truth 3D, which is considerably more challenging to acquire in a natural setting.
Acknowledgements. We thank N. Mahmood for the SMPL model fits to mocap data and the mesh retargeting for character animation, D. Mehta for his assistance on MPI-INF-3DHP, and S. Tulsiani, A. Kar, S. Gupta, D. Fouhey and Z. Liu for helpful discussions. This research was supported by BAIR sponsors and NSF Award IIS-1526234.