VIBE: Video Inference for Human Body Pose and Shape Estimation
Muhammed Kocabas, Nikos Athanasiou, Michael J. Black
Introduction
Tremendous progress has been made on estimating 3D human pose and shape from a single image . While this is useful for many applications, it is the motion of the body in the world that tells us about human behavior. As noted by Johansson even a few moving point lights on a human body in motion informs us about behavior. Here we address how to exploit temporal information to more accurately estimate the 3D motion of the body from monocular video. While this problem has received over 30 years of study, we may ask why reliable methods are still not readily available. Our insight is that the previous temporal models of human motion have not captured the complexity and variability of real human motions due to insufficient training data. We address this problem here with a new temporal neural network and training approach, and show that it significantly improves 3D human pose estimation from monocular video.
Existing methods for video pose and shape estimation often fail to produce accurate predictions as illustrated in Fig. 1 (top). A major reason behind this is the lack of in-the-wild ground-truth 3D annotations, which are non-trivial to obtain even for single images. Previous work combines indoor 3D datasets with videos having 2D ground-truth or pseudo-ground-truth keypoint annotations. However, this has several limitations: (1) indoor 3D datasets are limited in the number of subjects, range of motions, and image complexity; (2) the amount of video labeled with ground-truth 2D pose is still insufficient to train deep networks; and (3) pseudo-ground-truth 2D labels are not reliable for modeling 3D human motion.
To address this, we take inspiration from Kanazawa et al. who train a single-image pose estimator using only 2D keypoints and an unpaired dataset of static 3D human shapes and poses using an adversarial training approach. For video sequences, there already exist in-the-wild videos with 2D keypoint annotations. The question is then how to obtain realistic 3D human motions in sufficient quality for adversarial training. For that, we leverage the large-scale 3D motion-capture dataset called AMASS , which is sufficiently rich to learn a model of how people move. Our approach learns to estimate sequences of 3D body shapes poses from in-the-wild videos such that a discriminator cannot tell the difference between the estimated motions and motions in the AMASS dataset. As in , we also use 3D keypoints when available.
The output of our method is a sequence of pose and shape parameters in the SMPL body model format , which is consistent with AMASS and the recent literature. Our method learns about the richness of how people appear in images and is grounded by AMASS to produce valid human motions. Specifically, we leverage two sources of unpaired information by training a sequence-based generative adversarial network (GAN) . Here, given the video of a person, we train a temporal model to predict the parameters of the SMPL body model for each frame while a motion discriminator tries to distinguish between real and regressed sequences. By doing so, the regressor is encouraged to output poses that represent plausible motions through minimizing an adversarial training loss while the discriminator acts as weak supervision. The motion discriminator implicitly learns to account for the statics, physics and kinematics of the human body in motion using the ground-truth motion-capture (mocap) data. We call our method VIBE, which stands for “Video Inference for Body Pose and Shape Estimation.”
During training, VIBE takes in-the-wild images as input and predicts SMPL body model parameters using a convolutional neural network (CNN) pretrained for single-image body pose and shape estimation followed by a temporal encoder and body parameter regressor used in . Then, a motion discriminator takes predicted poses along with the poses sampled from the AMASS dataset and outputs a real/fake label for each sequence. We implement both the temporal encoder and motion discriminator using Gated Recurrent Units (GRUs) to capture the sequential nature of human motion. The motion discriminator employs a learned attention mechanism to amplify the contribution of distinctive frames. The whole model is supervised by an adversarial loss along with regression losses to minimize the error between predicted and ground-truth keypoints, pose, and shape parameters.
At test time, given a video, we use the pretrained CNN and our temporal module to predict pose and shape parameters for each frame. The method works for video sequences of arbitrary length. We perform extensive experiments on multiple datasets and outperform all state-of-the-art methods; see Fig. 1 (bottom) for an example of VIBE’s output. Importantly, we show that our video-based method always outperforms single-frame methods by a significant margin on the challenging 3D pose estimation benchmarks 3DPW and MPI-INF-3DHP . This clearly demonstrates the benefit of using video in 3D pose estimation.
In summary, the key contributions in this paper are: First, we leverage the AMASS dataset of motions for adversarial training of VIBE. This encourages the regressor to produce realistic and accurate motions. Second, we employ an attention mechanism in the motion discriminator to weight the contribution of different frames and show that this improves our results over baselines. Third, we quantitatively compare different temporal architectures for 3D human motion estimation. Fourth, we achieve state-of-the-art results on major 3D pose estimation benchmarks. Code and pretrained models are available for research purposes at https://github.com/mkocabas/VIBE.
Related Work
3D pose and shape from a single image. Parametric 3D human body models are widely used as the output target for human pose estimation because they capture the statistics of human shape and provide a 3D mesh that can be used for many tasks. Early work explores “bottom up” regression approaches, “top down” optimization approaches, and multi-camera settings using keypoints and silhouettes as input . These approaches are brittle, require manual intervention, or do not generalize well to images in the wild. Bogo et al. propose SMPLify, one of the first end-to-end approaches, which fits the SMPL model to the output of a CNN keypoint detector . Lassner et al. use silhouettes along with keypoints during fitting. Recently, deep neural networks are trained to directly regress the parameters of the SMPL body model from pixels . Due to the lack of in-the-wild 3D ground-truth labels, these methods use weak supervision signals obtained from a 2D keypoint re-projection loss , use body/part segmentation as an intermediate representation , or employ a human in the loop . Kolotouros et al. combine regression-based and optimization-based methods in a collaborative fashion by using SMPLify in the training loop. At each step of the training, the deep network initializes the SMPLify optimization method that fits the body model to 2D joints, producing an improved fit that is used to supervise the network. Alternatively, several non-parametric body mesh reconstruction methods has been proposed. Varol et al. use voxels as the output body representation. Kolotouros et al. directly regress vertex locations of a template body mesh using graph convolutional networks . Saito et al. predict body shapes using pixel-aligned implicit functions followed by a mesh reconstruction step. Despite capturing the human body from single images, when applied to video, these methods yield jittery, unstable results.
3D pose and shape from video. The capture of human motion from video has a long history. In early work, Hogg et al. fit a simplified human body model to images features of a walking person. Early approaches also exploit methods like PCA and GPLVMs to learn motion priors from mocap data but these approaches were limited to simple motions. Many of the recent deep learning methods that estimate human pose from video focus on joint locations only. Several methods use a two-stage approach to “lift” off-the-shelf 2D keypoints into 3D joint locations. In contrast, Mehta et al. employ end-to-end methods to directly regress 3D joint locations. Despite impressive performance on indoor datasets like Human3.6M , they do not perform well on in-the-wild datasets like 3DPW and MPI-INF-3DHP . Several recent methods recover SMPL pose and shape parameters from video by extending SMPLify over time to compute a consistent body shape and smooth motions . Particularly, Arnab et al. show that Internet videos annotated with their version of SMPLify help to improve HMR when used for fine tuning. Kanazawa et al. learn human motion kinematics by predicting past and future framesNote that they refer to kinematics over time as dynamics.. They also show that Internet videos annotated using a 2D keypoint detector can mitigate the need for the in-the-wild 3D pose labels. Sun et al. propose to use a transformer-based temporal model to improve the performance further. They propose an unsupervised adversarial training strategy that learns to order shuffled frames.
GANs for sequence modeling. Generative adversarial networks GANs have had a significant impact on image modeling and synthesis. Recent works have incorporated GANs into recurrent architectures to model sequence-to-sequence tasks like machine translation . Research in motion modelling has shown that combining sequential architectures and adversarial training can be used to predict future motion sequences based on previous ones or to generate human motion sequences . In contrast, we focus on adversarially refining predicted poses conditioned on the sequential input data. Following that direction, we employ a motion discriminator that encodes pose and shape parameters in a latent space using a recurrent architecture and an adversarial objective taking advantage of 3D mocap data .
Approach
The overall framework of VIBE is summarized in Fig. 2. Given an input video of length , of a single person, we extract the features of each frame using a pretrained CNN. We train a temporal encoder composed of bidirectional Gated Recurrent Units (GRU) that outputs latent variables containing information incorporated from past and future frames. Then, these features are used to regress the parameters of the SMPL body model at each time instance.
Given a video sequence, VIBE computes where are the pose parameters at time step and is the single body shape prediction for the sequence. Specifically, for each frame we predict the body shape parameters. Then, we apply average pooling to get a single shape () across the whole input sequence. We refer to the model described so far as the temporal generator . Then, output, , from and samples from AMASS, , are given to a motion discriminator, , in order to differentiate fake and real examples.
Overall, the loss of the proposed temporal encoder is composed of 2D (), 3D (), pose () and shape () losses when they are available. This is combined with an adversarial loss. Specifically the total loss of the is:
where is the adversarial loss explained below.
2 Motion Discriminator
The body discriminator and the reprojection loss used in enforce the generator to produce feasible real world poses that are aligned with 2D joint locations. However, single-image constraints are not sufficient to account for sequences of poses. Multiple inaccurate poses may be recognized as valid when the temporal continuity of movement is ignored. To mitigate this, we employ a motion discriminator, , to tell whether the generated sequence of poses corresponds to a realistic sequence or not. The output, , of the generator is given as input to a multi-layer GRU model depicted in Fig. 3, which estimates a latent code at each time step where . In order to aggregate hidden states we use self attention elaborated below. Finally, a linear layer predicts a value representing the probability that belongs to the manifold of plausible human motions. The adversarial loss term that is backpropagated to is:
and the objective for is:
where is a real motion sequence from the AMASS dataset, while is a generated motion sequence. Since is trained on ground-truth poses, it also learns plausible body pose configurations, hence alleviating the need for a separate single-frame discriminator .
Self-Attention Mechanism.
Recurrent networks update their hidden states as they process input sequentially. As a result, the final hidden state holds a summary of the information in the sequence. We use a self-attention mechanism to amplify the contribution of the most important frames in the final representation instead of using either the final hidden state or a hard-choice pooling of the hidden state feature space of the whole sequence. By employing an attention mechanism, the representation of the input sequence is a learned convex combination of the hidden states. The weights are learned by a linear MLP layer , and are then normalized using softmax to form a probability distribution. Formally:
We compare our dynamic feature weighting with a static pooling schema. Specifically, the features , representing the hidden state at each frame, are averaged and max pooled. Then, those two representations and are concatenated to constitute the final static vector, , used for the fake/real decision.
3 Training Procedure
Experiments
We first describe the datasets used for training and evaluation. Next, we compare our results with previous frame-based and video-based state-of-the-art approaches. We also conduct ablation experiments to show the effect of our contributions. Finally, we present qualitative results in Fig. 4.
Following previous work , we use batches of mixed 2D and 3D datasets. PennAction and PoseTrack are the only ground-truth 2D video datasets we use, while InstaVariety and Kinetics-400 are pseudo ground-truth datasets annotated using a 2D keypoint detector . For 3D annotations, we employ 3D joint labels from MPI-INF-3DHP and Human3.6M . When used, 3DPW and Human3.6M provide SMPL parameters that we use to calculate . AMASS is used for adversarial training to obtain real samples of 3D human motion. We also use the 3DPW training set to perform ablation experiments; this demonstrate the strength of our model on in-the-wild data.
Evaluation.
For evaluation, we use 3DPW , MPI-INF-3DHP , and Human3.6M . We report results with and without the 3DPW training to enable direct comparison with previous work that does not use 3DPW for training. We report Procrustes-aligned mean per joint position error (PA-MPJPE), mean per joint position error (MPJPE), Percentage of Correct Keypoints (PCK) and Per Vertex Error (PVE). We compare VIBE with state-of-the-art single-image and temporal methods. For 3DPW, we report acceleration error (), calculated as the difference in acceleration between the ground-truth and predicted 3D joints.
1 Comparison to state-of-the-art results
Table 1 compares VIBE with previous state-of-the-art frame-based and temporal methods. VIBE (direct comp.) corresponds to our model trained using the same datasets as Temporal-HMR , while VIBE also uses the 3DPW training set. As standard practice, previous methods do not use 3DPW, however we want to demonstrate that using 3DPW for training improves in-the-wild performance of our model. Our models in Table 1 use pretrained HMR from SPIN as a feature extractor. We observe that our method improves the results of SPIN, which is the previous state-of-the-art. Furthermore, VIBE outperforms all previous frame-based and temporal methods on the challenging in-the-wild 3DPW and MPI-INF-3DHP datasets by a significant amount, while achieving results on-par with SPIN on Human3.6M. Note that, Human3.6M is an indoor dataset with a limited number of subjects and minimal background variation, while 3DPW and MPI-INF-3DHP contain challenging in-the-wild videos.
We observe significant improvements in the MPJPE and PVE metrics since our model encourages temporal pose and shape consistency. These results validate our hypothesis that the exploitation of human motion is important for improving pose and shape estimation from video. In addition to the reconstruction metrics, e.g. MPJPE, PA-MPJPE, we also report acceleration error (Table 1). While we achieve smoother results compared with the baseline frame-based methods , Temporal-HMR yields even smoother predictions. However, we note that Temporal-HMR applies aggressive smoothing that results in poor accuracy on videos with fast motion or extreme poses. There is a trade-off between accuracy and smoothness. We demonstrate this finding in a qualitative comparison between VIBE and Temporal-HMR in Fig. 5. This figure depicts how Temporal-HMR over-smooths the pose predictions while sacrificing accuracy. Visualizations from alternative viewpoints in Fig. 4 show that our model is able to recover the correct global body rotation, which is a significant problem for previous methods. This is further quantitatively demonstrated by the improvements in the MPJPE and PVE errors. For video results see the GitHub page.
2 Ablation Experiments
Table 2 shows the performance of models with and without the motion discriminator, . First, we use the original HMR model proposed by as a feature extractor. Once we add our generator, , we obtain slightly worse but smoother results than the frame-based model due to lack of sufficient video training data. This effect has also been observed in the Temporal-HMR method . Using helps to improve the performance of while yielding smoother predictions.
When we use the pretrained HMR from , we observe a similar boost when using over using only . We also experimented with MPoser as a strong baseline against . MPoser acts as a regularizer in the loss function to ensure valid pose sequence predictions. Even though, MPoser performs better than using only , it is worse than using . One intuitive explanation for this is that, even though AMASS is the largest mocap dataset available, it fails to cover all possible human motions occurring in in-the-wild videos. VAEs, due to over-regularization attributed to the KL divergence term , fail to capture real motions that are poorly represented in AMASS. In contrast, GANs do not suffer from this problem . Note that, when trained on AMASS, MPoser gives 4.5mm PVE on a held out test set, while the frame-based VPoser gives 6.0mm PVE error; thus modeling motion matters. Overall, results shown in Table 1 demonstrate that introducing improves performance in all cases. Although one may think that the motion discriminator might emphasize on motion smoothness over single pose correctness, our experiments with a pose only, motion only, and both modules revealed that the motion discriminator is capable of refining single poses while producing smooth motion.
Dynamic feature aggregation in significantly improves the final results compared to static pooling ( - concat), as demonstrated in Table 3. The self-attention mechanism enables to learn how the frames correlate temporally instead of hard-pooling their features. In most of the cases, the use of self attention yields better results. Even with an MLP hidden size of , adding one more layer outperforms static aggregation. The attention mechanism is able to produce better results because it can learn a better representation of the motion sequence by weighting features from each individual frame. In contrast, average and max pooling the features produces a rough representation of the sequence without considering each frame in detail. Self-attention involves learning a coefficient for each frame to re-weight its contribution in the final vector () producing a more fine-grained output. That validates our intuition that attention is helpful for modeling temporal dependencies in human motion sequences.
Conclusion
While current 3D human pose methods work well, most are not trained to estimate human motion in video. Such motion is critical for understanding human behavior. Here we explore several novel methods to extend static methods to video: (1) we introduce a recurrent architecture that propagates information over time; (2) we introduce discriminative training of motion sequences using the AMASS dataset; (3) we introduce self-attention in the discriminator so that it learns to focus on the important temporal structure of human motion; (4) we also learn a new motion prior (MPoser) from AMASS and show it also helps training but is less powerful than the discriminator. We carefully evaluate our contributions in ablation studies and show how each choice contributes to our state-of-the-art performance on video benchmark datasets. This provides definitive evidence for the value of training from video.
Future work should explore using video for supervising single-frame methods by fine tuning the HMR features, examine whether dense motion cues (optical flow) could help, use motion to disambiguate the multi-person case, and exploit motion to track through occlusion. In addition, we aim to experiment with other attentional encoding techniques such as transformers to better estimate body kinematics.
Acknowledgements: We thank Joachim Tesch for helping with Blender rendering. We thank all Perceiving Systems department members for their feedback and the fruitful discussions. This research was partially supported by the Max Planck ETH Center for Learning Systems and the Max Planck Graduate Center for Computer and Information Science.
Disclosure: MJB has received research gift funds from Intel, Nvidia, Adobe, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, MPI. MJB has financial interests in Amazon and Meshcapade GmbH.
References
Appendix – Supplmentary Material
The architecture is depicted in Figure 6. After feature extraction using ResNet50, we use a 2-layer GRU network followed by a linear projection layer. The pose and shape parameters are then estimated by a SMPL parameter regressor. We employ a residual connection to assist the network during training. The SMPL parameter regressor is initialized with the pre-trained weights from HMR . We decrease the learning rate if the reconstruction does not improve for more than 5 epochs.
Motion Discriminator.
We employ 2 GRU layers with a hidden size of 1024. For our most accurate results we used a self-attention mechanism with 2 MLP layers, each with 1024 neurons, and a dropout rate of . For the ablation experiments we keep the same parameters changing the number of neurons and the number of MLP layers only.
Loss.
We use different weight coefficients for each term in the loss function. The 2D and 3D keypoint loss coefficients are respectively, while , . We set the motion discriminator adversarial loss term, as . We use 2 GRU layers with a hidden dimension size of .
2 Datasets
Below is a detailed summary of the different datasets we use for training and testing.
MPI-INF-3DHP is a multi-view, mostly indoor, dataset captured using a markerless motion capture system. We use the training set proposed by the authors, which consists of 8 subjects and 16 videos per subject. We evaluate on the official test set.
Human3.6M contains 15 action sequences of several individuals, captured in a controlled environment. There are 1.5 million training images with 3D annotations. We use SMPL parameters computed using MoSh for training. Following previous work, our model is trained on 5 subjects (S1, S5, S6, S7, S8) and tested on the other 2 subjects (S9, S11). We subsampled the dataset to 25 frames per second for training.
3DPW is an in-the-wild 3D dataset that captures SMPL body pose using IMU sensors and hand-held cameras. It contains 60 videos (24 train, 24 test, 12 validation) of several outdoor and indoor activities. We use it both for evaluation and training.
PennAction dataset contains 2326 video sequences of 15 different actions and 2D human keypoint annotations for each sequence. The sequence annotations include the class label, human body joints (both 2D locations and visibility), 2D bounding boxes, and training/testing labels.
InstaVariety is a recently curated dataset using Instagram videos with particular action hashtags. It contains 2D annotations for about 24 hours of video. The 2D annotations were extracted using OpenPose and Detect and Track in the case of multi-person scenes.
PoseTrack is a benchmark for multi-person pose estimation and tracking in videos. It contains 1337 videos, split into 792, 170 and 375 videos for training, validation and testing respectively. In the training split, 30 frames in the center of the video are annotated. For validation and test sets, besides the aforementioned 30 frames, every fourth frame is also annotated. The annotations include 15 body keypoint locations, a unique person ID, and a head and a person bounding box for each person instance in each video. We use PoseTrack during training.
3 Evaluation
Here we describe the evaluation metrics and procedures we used in our experiments. For direct comparison we used the exact same setup as in . Our best results are achieved with a model that includes the 3DPW training dataset in our training loop. Even without 3DPW training data, our method is more accurate than the previous SOTA. For consistency with prior work, we use the Human3.6M training set when evaluating on the Human3.6M test set. In our best performing model on 3DPW, we do not use Human3.6M training data. We observe, however, that strong performance on the Human3.6M does not translate to strong in-the-wild pose estimation.
We use standard evaluation metrics for each respective dataset. First, we report the widely used MPJPE (mean per joint position error), which is calculated as the mean of the Euclidean distances between the ground-truth and the predicted joint positions after centering the pelvis joint on the ground truth location (as is common practice). Also we report PA-MPJPE (Procrustes Aligned MPJPE), which is calculated similarly to MPJPE but after a rigid alignment of the predicted pose to the and ground-truth pose. Furthermore, we calculate Per-Vertex-Error (PVE), which is denoted by the Euclidean distance between the ground truth and predicted mesh vertices that are the output of SMPL layer. We also use the Percentage of Correct Keypoints metric (PCK) . The PCK counts as correct the cases where the Euclidean distance between the actual and predicted joint positions is below a predefined threshold. Finally, we report acceleration error, which was reported in . Acceleration error is the mean difference between ground-truth and predicted 3D acceleration for every joint ().