3D Human Motion Estimation via Motion Compression and Refinement

Zhengyi Luo, S. Alireza Golestaneh, Kris M. Kitani

Introduction

Estimating the 3D pose sequence of a person from a single video requires a computational model that can extract the underlying kinematics of human motion while also preserving motions that are unique to the person being captured. Since people share a similar body structure (e.g., same number of joints) and similar physical constraints (e.g., joint limitations), it is possible to learn a generalized kinematic model that can be matched against the image to infer the general motion of a person. However, since generalized models of motion can also fail to model person-specific motions, it may also be necessary to ‘add back in’ or refine the general motion estimates using image evidence. In this work, we present a two-stage 3D motion estimation method that first extracts coarse kinematic sequences of a person in a video and then refines that sequence to produce a more accurate 3D motion estimate. We show that by decomposing the inference process into (1) a general model of motion and (2) a person-specific model of motion, we are able to obtain more accurate and smooth estimates.

Over the past years, significant progress has been made on improving the accuracy of 3D human pose estimation. Impressive results have been obtained through human mesh recovery from single images and videos . The main metric used to evaluate these methods is the Mean Per Joint Position Error (MPJPE), which measures the performance in terms of the relative joint positions computed for each frame of a video. However, less emphasis has been given to the temporal smoothness of the estimated motion. Optimizing for this metric, the tendency is to generate pose estimates that ‘jitter’ near the true pose. This is expected as the MPJPE only penalizes for spatial errors and is not designed to account for temporal consistency. As the methods for 3D pose estimation have improved in recent years, the ‘jitter’ has become less pronounced, especially when applied to dynamic scenes with vibrant moving backgrounds and camera motion. However, by rendering the estimated 3D pose of state-of-the-art methods on a plain background, the ‘jitter’ can still be observed, resulting in an overall unnatural motion estimation.

The issue of temporal smoothness is a known problem and has been addressed in part by prior work. Large-scale motion datasets such as Archive of Motion Capture as Surface Shapes (AMAAS) ) and adversarial loss have enabled methods to improve both pose accuracy and temporal smoothness . Other methods have been developed to enforce temporal smoothness by letting the model predict frame ordering. However, using prior knowledge only in the loss function, it is hard to find the balance between smoothness and accuracy.

In this work, we argue that striking such balance between smoothness and accuracy can be done through an explicit breakdown of coarse and fine motion. First, we learn a coarse motion model by observing a large dataset of human motion–since human motion is typically smooth (e.g., we usually do not shake as we walk), if one were to fit a model to a large set of human motions, most of the data would lie in a sub-space in which motions are smooth. This implies that if we were to compress human motion data, it should learn a latent subspace in which the motions are inherently smooth and coarse. Using this latent space as the regression target, we can directly infer coarse human motion from the input videos. The problem of using such a human motion subspace, of course, is that rare motions (e.g., sudden motions) are removed from the motion data. To retain such motion, we also argue that producing the final 3D motion estimate can be treated as a refinement step to ”add back” the fine details to the coarse motion sequence.

To validate our arguments, we propose a two-stage 3D human motion estimation method that first estimates a coarse human pose sequence through data compression using a Variational Autoencoder (VAE), which we call the Variational Motion Estimator (VME). Then we take the output of the VME and refine the pose estimate using the image evidence with a pose regressor, which we call the Motion Refinement Regressor (MRR).

In summary, we propose a video-based 3D human motion estimation method that focuses on producing smooth and accurate human motion sequences. Our main contributions are as follows: (1) we propose a two-stage motion estimation method for ensuring temporal smoothness and accurate pose estimates; (2) we describe a procedure to learn a robust Variational Autoencoder that can serve as a latent human motion subspace for estimating coarse 3D human motion from videos; (3) we demonstrate state-of-the-art pose/motion estimation performance on challenging in-the-wild dataset such as 3DPW , reducing acceleration error by 54.3%54.3\% while achieving state-of-the-art MPJPE results.

Related Works

In this section, we will first review the relevant work in human shape and pose recovery from a single image and from videos–human motion recovery can be treated as a subset of human pose estimation, as human motion is a sequence of human poses. Then we will review how existing methods use human motion as a prior and how popular methods map motion sequences to a low-dimensional space.

Here we focus on model-based methods that jointly recover human shape and pose. We choose to use a parametric 3D human body model since it can be easily turned into a 3D human mesh that is usable for downstream tasks such as animation. Directly fitting a parametric 3D human body to image input has gained substantial traction over the years, morphing from methods that require silhouette or human input , to ones that can directly fit model parameters to 2D joint positions , and to ones that can directly estimate shape and pose from images . Due to the lack of ground truth 3D labels, these methods use a weakly supervised approach to fit the 3D human body to 2D joint positions , body part segmentation , or dense pixel correspondence . Although these methods achieve amazing results, their extracted motion tends to be unstable due to the lack of temporal information.

2 Recovering 3D human pose and shape from video

Using temporal information to aid 3D human pose estimation is a natural extension to single frame methods. focus on ”lifting” predicted joint positions from 2D to 3D, and uses LSTM , Temporal Convolution , and fully connected layers to exploit temporal information. , on the other hand, predict 3D joint positions directly from images and use a temporal filter to postprocess the motion sequence. For methods that jointly recover shape and pose. HMMR , Sun et al. and VIBE are the best performing models that exploit temporal information. HMMR proposes to enforce temporal consistency by letting the model predict future and past frames of motion. Sun et al. learns temporal information by predicting the ordering of shuffled frames. VIBE utilizes temporal information by employing a Gated Recurrent Unit (GRU) to convert the input frames of features into a sequence of temporally correlated latent features.

3 Human Pose and Motion Prior

Using prerecorded human motion sequences as a prior has also been explored in various tasks related to human motion. Earlier work like tries to quantify unnaturalness in animated human motion by statistically analyzing existing motion capture (MoCap) sequences, and propose to use learned MoCap motions to aid 3D motion tracking. More recently, use an adversarial discriminator at a per-frame level to ensure that the recovered pose is a valid human pose. uses a pretrained pose VAE’s latent space for a similar purpose. proposes to use the discriminator at a temporal level, discriminating against a whole motion sequence. All above methods use a pose or motion prior in an adversarial way, utilizing the prior knowledge in the loss function.

4 Human Motion Representation

Compressing human motion into a compact latent representation plays an important role in tasks such as human motion generation , human motion generation across different modalities , and trajectory forecasting . Existing methods leverage different generative models such as VAE , generative flow , or generative adversarial networks to achieve a compact motion representation in the latent space.

Approach

As discussed in Sections 1 and 2, existing human motion estimation methods often find it difficult to achieve a balance between temporal smoothness and accuracy. To tackle this, we propose MEVA, Motion Estimation via Variational Autoencoder, a framework that learns the overall coarse motion first then adds back fine detailed motion as a residual. MEVA processes the inputs in three steps: it first extracts correlated temporal features using a spatio-temporal feature extractor (STE), and then captures the overall coarse motion through a variational motion estimator (VME), and finally uses a motion residual regressor (MRR) to add back the fine motion details. Our overall framework can be shown in Fig. 2. In this section, we will first setup the overall problem and then present the details of our framework. Finally, we will discuss the training procedures.

Given an input video VT={It}t=1TV_{T}=\{I_{t}\}_{t=1}^{T}, where ItI_{t} denotes the ttht^{th} frame, our task is to recover coherent human motion sequences MT={θt}t=1TM_{T}=\{\theta_{t}\}_{t=1}^{T} where each θt\theta_{t} represents the human pose for the ttht^{th} frame. To represent the human motion, we utilize the SMPL 3D human mesh model . We choose SMPL parametrization over 3D joint positions or other human models due to its versatility: SMPL parameters can be easily converted to 3D joint positions and human mesh. Specifically, a human body is represented by its shape β\beta and pose θ\theta, denoted by Θ={β,θ}\Theta=\{\beta,\theta\}. Given θ\theta and β\beta, let SS denote the pretrained SMPL function, where S(Θ):β,θ→R6890×3S(\Theta):\beta,\theta\rightarrow R^{6890\times 3} ( 68906890 is the number of vertices of the resulting triangular human mesh). The pose parameter θ∈R24×N\theta\in R^{24\times N} stands for the joint angles for the 23 joints plus the root orientation. NN is the dimension of the chosen rotation representation (N=3N=3 for axis/euler angle, N=4N=4 for quaternions, N=6N=6 for a 6 degrees-of-freedom rotation representation ). The shape parameters β∈R10\beta\in R^{10} represent the linear coefficient for the principal component of the parametric human shape space. Given a set of β\beta and θ\theta, SS can recover the 3D joint positions through a pretrained mesh vertex regressor P:jp3d=P{S(β,θ)}∈RN×3P:jp^{3d}=P\{S(\beta,\theta)\}\in R^{N\times 3}. To project the 3D joint positions back to 2D images, a weak perspective camera π={s,tx,ty}\pi=\{s,t_{x},t_{y}\} needs to be estimated: jp2d=Π(P{S(β,θ)})∈RN×2jp^{2d}=\Pi(P\{S(\beta,\theta)\})\in R^{N\times 2}.

Intuitively, recovering human motion from video frames does not require recovering the human shape: one can directly learn a mapping from input video frames to the estimated human motion if sufficient paired ground truth data exist. However, videos paired with ground truth motion annotation (SMPL sequences) require professional capture equipment such as a motion capture (MoCap) rig, which is still rather rare compared to annotated 2D pose datasets. In the absence of 3D data samples, it is critical for our model to learn motion sequences in a semisupervised fashion (i.e., from videos with 3D or 2D labeled joint positions following the approach utilized in ).

Overall, our motion estimation objective is to learn a function MEVA(V):VT→MT\text{MEVA}(V):V_{T}\rightarrow{M_{T}} where VT={It}t=1TV_{T}=\{I_{t}\}_{t=1}^{T} and MT={θt}t=1TM_{T}=\{\theta_{t}\}_{t=1}^{T}.

2 Spatio-Temporal Feature Extractor (STE)

Human motion is inherently temporal and correlated, and past movement can give cues about future motion. Thus, instead of extracting per-frame visual features using a feed-forward convolutional network independently, we can produce temporally correlated features that lead to better motion sequence modeling. Similar to , we use a GRU-based temporal feature extractor (STE) that encodes the input video frames I1,I2,I3,...ITI_{1},I_{2},I_{3},...I_{T} into a sequence of temporally correlated features f1′,f2′,f3′,...fT′f^{\prime}_{1},f^{\prime}_{2},f^{\prime}_{3},...f^{\prime}_{T}.

3 Variational Motion Estimator (VME)

To learn a human motion subspace that can encapsulate a broad spectrum of human motion, we choose to use a Variational Autoencoder (VAE). VAEs can effectively capture a large number of possible data modes by explicitly mapping each data point to a latent code, and by imposing a Gaussian prior on the learned latent space, similar motions’ latent codes will be near each other. Thus, the VAE’s latent space allows for more overlap between codes and therefore enforces smoothness in the latent space. Having a smooth latent code is essential in improving the generalizability of the model since the space of possible human motion is highly correlated and limited. Formally, following the previous work on VAEs , the objective is to maximize the evidence lower bound of the log-likelihood pλ(x)p_{\lambda}(x) (λ\lambda and ϕ\phi denotes the function parametrization):

where xx is the input and the latent code z∼N(0,I)z\sim\mathcal{N}(\mathbf{0},\mathbf{I}).

In the context of encoding human motion via VAE, the encoder EvaeE_{vae} takes in a sequence of WW frames of human motion represented in terms of SMPL pose parameters: x=MW=[θw1,θw2,θw3,...]∈RW×144x=M_{W}=[\theta_{w1},\theta_{w2},\theta_{w3},...]\in R^{W\times 144} and outputs the latent code zz. A single frame of SMPL pose is represented in joint rotations, resulting by a 24×6=14424\times 6=144 dimensional input (a 6 degrees-of-freedom rotation representation is used for continuity purpose). The decoder DvaeD_{vae} takes in the latent code zz and reconstructs the motion M^W\hat{M}_{W}. Both the encoder EvaeE_{vae} and decoder DvaeD_{vae} are implemented as GRUs, and the detailed architectures are given in Fig.3. Based on the Gaussian parameterization of the VAE, the objective function of Eq.(1) can be written Eq.(2)

where SS is the number of samples for the current batch, SzS_{z} is the dimension of the current latent variable, and β\beta is the weighing parameter. Once the VAE is trained and converged to a desirable reconstruction accuracy, the decoder DvaeD_{vae} is frozen for later use. During inference, given a latent code z∈R1×Szz\in R^{1\times S_{z}}, DvaeD_{vae} can decode it back into a sequence of human motion: MW∈Rw×144M_{W}\in R^{w\times 144}.

3.2 Human Motion Data augmentation

Our learned VAE should be able to generalize to unseen human motion sequences and achieve high reconstruction accuracy to ensure that the learned latent space can indeed serve as a comprehensive human motion subspace. Using an already large-scale human motion dataset AMASS (13k motion samples with varying length), our trained VAE still has poor generalizability on unseen sequences (for details refer to Sec.4.3.1). Thus, we devise an elaborate data augmentation scheme that can augment the existing motion and produce viable yet distinct human motion sequences. While data augmentation has been studied extensively in image processing, to the best of our knowledge, few attempts have been done in augmenting a human motion dataset. When used in trajectory forecasting and human motion generation , the generalizability of the VAE latent space has not been discussed extensively since the models only need to generate new motion sequences and do not emphasize on the ability to encode unseen motion sequences.

Given a TT frame human motion sequence in SMPL parameters MT∈RT×144M_{T}\in R^{T\times 144} with a frame-rate FamassF_{amass}, we employ the following data augmentation scheme:

Speeding up and slowing down: based on FamassF_{amass}, we can uniformly up-sample or downsample the frames and produce novel sequences that are still plausible and natural human motion.

Flipping left and right: The same action, performed using either the left or right hand, will remain a valid human motion. Thus, we can follow the kinematic tree of the SMPL model and mirror the motion across the left and right and generate a new motion sequence.

Random root rotation: We randomly sample a root rotation from a unit sphere to capture different root orientations for possible human motion. Different pose estimators may assume different ground planes and coordinate systems, so SMPL parameters usually come in different root orientations. Sampling random root rotation helps the model cope with different possible coordinate frame choices.

3.3 Learning Smooth Motion from videos

After learning a comprehensive human motion subspace using the VAE, we learn an additional encoder EmotionE_{motion} that can directly extract coarse motion sequences from video features, mapping to the same latent space as EvaeE_{vae}. Given an input sequence of video features fW={fw}w=1Wf_{W}=\{f_{w}\}_{w=1}^{W}, the encoder EmotionE_{motion}’s task is to compress the input features into a latent code zz that best summarizes the current observation as a coarse human motion sequence. We use the pretrained decoder DvaeD_{vae} from the motion VAE to force EmotionE_{motion} to sample from our pretrained motion subspace. Constraining the latent space of the EmotionE_{motion} to a pretrained human motion subspace provides a strong human motion prior that greatly aids the learning process of EmotionE_{motion}. Combining EmotionE_{motion} and DvaeD_{vae}, we form our Variational Motion Estimator (VME).

4 Motion Residual Regressor (MRR)

As noted in the previous section, the learned motion sequences using the VAE’s latent space as the target are inherently smooth and coarse, capturing the overall motion signature of the current video frames through information compression. To capture the details, we utilize a SMPL regressor from that can iteratively refine the estimated poses. The regressor takes in an initial pose and shape estimation Θt\Theta_{t} and the visual feature ft\text{f}_{t} for a single frame to calculate its estimation Θt′\Theta_{t}^{\prime} for kk iterations. Notice that though all utilize the same regressor, MEVA uses it in a fundamentally different way–in , the regressor is initialized with mean SMPL pose Θmean\Theta_{mean}. At a sequence level, a regressor initialized uniformly with the mean pose Θmean\Theta_{mean} is trying to capture the overall motion in one pass, while in MEVA, the regressor is initialized with the computed poses from VME. Thus, the regressor is only tasked to do small cosmetic changes to the coarse estimation, adding back the fine details of the motion lost during our compression step. Similar to , the input visual features ft{f}_{t} to the regressor are encoded using a temporal visual encoder, so even though each frame’s estimation is calculated separately at this stage, the visual feature is already temporally correlated. The VME computes the overall coarse motion from videos by using a general model of motion, and the regressor jointly refines motion and human shape estimates, which amounts to adding back the person-specific motion details at a per frame level. We call this regressor the Motion Residual Regressor (MRR). MRR completes the overall framework of our proposed method, as shown in Fig. 2.

5 Training and Losses

Our framework is trained in two stages. At first, the motion VAE is pretrained. Then STE, VME, and MRR are trained jointly end-to-end. Using videos with various levels of annotation (2D joint positions, 3D joint positions, SMPL parameters), similar to in , the networks are trained with losses consisting of L2DL_{2D}, L3DL_{3D}, LSMPLL_{SMPL}, as long as respective data is available. Specifically:

For implementation details, please refer to the supplementary material.

Experiments

To demonstrate the effectiveness of our proposed method, we evaluate our method in terms of the overall accuracy and smoothness of the estimated motion on MPI-INF-3DPH , 3DPW , and human 3.6M . In the following sections, we will first describe the main datasets used to train and evaluate MEVA. Then in Section 4.2 we provide extensive evaluation results. Finally, in Section 4.3, we provide the ablation studies for our proposed method.

For the motion VAE, we use motion sequences from AMASS for training and sequences from 3DPW for evaluation. For training with videos, in addition to the train split of MPI-INF-3DPH , 3DPW , and human 3.6M , which have 3D joint annotation, we also use InstaVariety and PennAction which contain 2D joint annotation.

For training our motion VAE, we use AMASS . It is a recent dataset that contains a large sample of human motion sequences in SMPL parameters. These sequences are fitted from MoCap sequences using Mosh++ . There are in total 13k motion sequences with varying length. We use this dataset only for training our motion VAE.

For training with videos, 3DPW is the only dataset that contains paired SMPL parameters and video sequences (which provide direct supervision to MEVA). The videos from this dataset are mostly outdoors and in-the-wild. It uses paired IMU sensors and video input to compute the near ground truth SMPL pose and shape parameters. This is a relatively small dataset and we use the official split in for train, validation, and test. There are in total 60 videos with varying length (24 train, 24 test, 12 val). MPI-INF-3DHP is a dataset that contains 3D joint position annotation. It is captured using a multiview camera setup, and the 3D joint annotation is calculated through multiview methods . There are 8 subjects and 16 videos per subject, in total 128 videos with varying length. We use the official test and train split. H3.6M is a popular pose estimation dataset that contains 3D joint position annotations, captured indoors with MoCap markers. Notice that a number of previous works had access to a near-ground truth SMPL pose and shape parameters calculated by the Mosh method. However, this annotation has since been removed from public access due to legal issues. SMPL parameters provide the best supervision, as noted in , so for a fair comparison we retrain some of the state-of-the-art methods without such supervision. There are 840 videos in total across 7 subjects in H3.6M and we use the official train/test subject split ([S1, S5, S6, S7, S8] vs [S9, S11]). During preprocessing, we subsample every 5 frames from the dataset. The PennAction dataset contains human annotated ground truth 2D keypoints paired with video sequences. There are in total 2326 videos with varying length. InstaVariety dataset contains human annotated pseudo ground truth 2D keypoints paired with video sequences. The 2D keypionts are estimated using openpose. There are in total 28,272 videos with varying length.

2 Evaluation Results and Analysis

To best capture human motion, we utilize three popular metrics that measure the overall accuracy and smoothness of the motion. Mean per joint position error (MPJPE) and MPJPE after Procrustes Alignment (PA-MPJPE) measure the 3D joint positional discrepancy between the predicted and ground truth 3D joint positions in millimeters (mmmm), and are calculated after aligning the root position (human pelvis) of the 3D joint positions. Both MPJPE and PA-MPJPE serve as the accuracy indicator of the motion estimator. Acceleration error (ACC-ERR), proposed in , measures the difference between the predicted and ground truth 3D acceleration for each keypoint in mm/s2mm/s^{2}. ACC-ERR serves as the major smoothness indicator for the estimated motion sequences. Acceleration is calculated using the finite difference between individual frames. It is imperative to view these metrics jointly: a low position error indicates overall correctness in motion capture and a better acceleration error marks a smooth and natural estimation of human motion.

2.2 Generalization of Motion VAE

In this section, we study the genealizability of our learned motion VAE. We report the reconstruction error of the VAE on unseen motion sequences from the 3DPW dataset. Table 1 shows VAE motion reconstruction error over the different splits of 3DPW. The Motion VAE model is the best performing model that is trained with all three forms of data augmentation techniques. Detailed analysis about the effects of data augmentation can be found in 4.3.1. The result shows that our VAE generalizes well to unseen sequences and the learned subspace can represent the human motion space with reasonable quality.

2.3 Quantitative Results

The result in Table 2 shows that our method obtains state-of-the-art results on video motion estimation across all three test datasets. Overall, our method achieves comparable results in terms of position error (MPJPE and PA-MPJPE) while significantly improving the smoothness (acceleration error), signifying a smoother and more natural motion estimation without sacrificing accuracy. Notice that all methods in italics have access to SMPL parameter annotation to the H3.6M dataset, which has since been removed from the web due to legal reasons. The SMPL parameters provide the most direct supervision for the task, so the performance gain is significant especially on the H3.6M test set. For a more direct comparison, we retrain the previous state-of-the-art method, VIBE , using the official implementation with the exact same datasets as ours. On 3DPW, under the same training condition, MEVA outperforms VIBE on almost all three metrics while reducing the acceleration error by 54.3%, 59.3%, and 41.3 %, respectively. Even compared to VIBE trained with extra data, our model achieves comparable results in accuracy while sporting a great reduction in acceleration error, except for the H3.6M dataset. Note that the H3.6M dataset contains mainly indoor scenes with limited background variation and models trained with direct SMPL supervision tend to perform well on this dataset. Compared to HMMR , which is the state-of-the-art on smoothness, our model still achieves a smoother result (23.7% reduction in acceleration error) while improving greatly in MPJPE by 25.4%.

2.4 Qualitatively Results

As motions are best seen in videos, please refer to the supplementary video for qualitative results. Overall, our model achieves better acceleration error while preserving high joint position accuracy, resulting in an overall smooth and natural motion.

3 Ablation Experiments

As mentioned in Sec.3.3.2, data augmentation performed on the AMASS dataset significantly improves the generalizability of our motion VAE. Table 3 demonstrates the VAE reconstruction error on the unseen sequences from the 3DPW dataset (train/test/val), with various levels of data augmentation techniques. Overall, RR (random root) rotation is essential since the motion sequences in AMASS dataset, captured mostly in MoCap studio, have a single fixed initial root rotation. Model trained only on AMASS would suffer greatly when it encounters variation in root rotation. Changing the sampling frame rate (FR) and flip left and right (LR) also provide a significant boost to generalizability. Using all three augmentation techniques results in our best performing motion VAE.

3.2 Coarse motion vs fine motion retrieval

MEVA benefits from an explicit breakdown of coarse and fine motion retrieval, using a temporal compressive step that captures the overall motion in a given human motion sequence. Just how much coarse/smooth motion information is retrieved in the framework? Table 4 shows the result of MEVA if trained only using the VME (capturing only coarse motion). Notice that MEVA with only VME achieves result similar to HMMR , the previous state-of-the-art in producing low acceleration error motion estimation. An illustrative visualization of coarse and fine motion decomposition can be found in Fig.4.

3.3 Effects of the pretrained Motion VAE

MEVA benefits from using a pretrained motion VAE’s latent space. As argued in Sec.3.3, using a pretrained VAE provides a human motion subspace that assists in constraining the estimated motion sequences to be natural and plausible human motion. Table 4 shows the result of MEVA trained without using a pretrained VAE (not using the pretrained DvaeD_{vae}). In this case, the whole framework is trained end-to-end from scratch. Here we observe that the model performed relatively well in both accuracy and smoothness, demonstrating the power of our two-stage estimation framework. However, upon visual inspection, as shown in Fig.5, a few kinematically invalid human poses are estimated during the sequence, resulting in an overall accurate but flawed estimation.

Conclusion

We have shown that to achieve temporally smooth and accurate 3D human pose estimates, it is important to learn a compressive model that encodes the smoothness of general human motion while also learning an image-based regression model that can capture person-specific motion. We propose a two-stage model that first trains a Variational Autoencoder to model the general statistics of coarse/smooth human motion and then learns a person-specific motion refinement regression module to retain motions not captured by the general motion model. Through comprehensive experiments, we demonstrate that our method produces both smooth and accurate motion.

Acknowledgements: This project was sponsored in part by IARPA (D17PC00340), and JST AIP Acceleration Research Grant (JPMJCR20U1).

References

Qualitative Results

To best view our motion estimation and compare it with state-of-the-art, please refer to the supplementary video.

Specifically, in the supplementary video, we first show a visual demonstration of our two-stage decomposition of coarse and fine motion from a given video sequence. Next, we demonstrate the qualitative comparison between our algorithm and the prior state-of-the-art (VIBE) and show that our method achieves smoother, more natural, and accurate motion estimation. Finally, we will discuss the implementation details of our method.

Additional Ablation Studies

While our method has significantly reduced the acceleration error and achieves state-of-the-art accuracy, one can still apply postprocessing to existing sequences to further improve the prediction. To best study its effects, here we implement a simple average filter using spherical linear interpolation (slerp) in quaternion. Specifically, for each joint rotation in SMPL qiq^{i} at timestep tt, we apply slerp with a ratio of 0.5: qti=slerp(qti,qt+1i,0.5)q^{i}_{t}=slerp(q^{i}_{t},q^{i}_{t+1},0.5). Table 5 shows the result of applying averaging filtering as postprocessing on both VIBE and MEVA. From the results, it is clear that average filtering can help reduce the acceleration error of both VIBE and MEVA while slightly affecting accuracy. It is also conceivable that more sophisticated methods such as solving a constrained optimization problem can further improve results. Nonetheless, in the paper we only compare with feed-forward methods without any postprocessing, since postprocessing approaches are complementary to feed-forward methods and would be beneficial to all of them.

2 Effects of a long temporal window

MEVA uses a significantly longer temporal window (90 frames) than prior art (HMMR : 20 frames, VIBE: 16 frames). To show that our MEVA framework benefits more from this setting, we retrain VIBE with a 90 frames temporal window. As shown in Table 6, using the same size temporal window, MEVA produces better results on all three metrics and maintains a significant advantage in acceleration error. Notice that VIBE trained with a longer temporal window shows a slight improvement against the ones that use a shorter window, validating our intuition that a longer temporal window provides a more substantial context for motion estimation. Nonetheless, our two-stage decomposition method is more effective in utilizing a longer temporal window due to its separate motion compression and refinement stages.

3 Effects of STE and VME on MEVA

Here we take a further look into the effects of different components (STE, VME, and MRR) of our proposed method. Notice that without the Variational Motion Estimator (VME), our method will collapse into a single-stage estimator that only relies on the SMPL regressor, which has been studied extensively in prior art. Thus, here we only study the effects of Spatial Temporal Feature Extractor (STE) and Motion Refinement Regressor (MRR). Table 7 shows the results of our framework trained without STE or MRR. Without the STE, MEVA obtains high accuracy but suffers from high acceleration error. This indicates that STE produces correlated features that impart the necessary temporal consistency information to MRR. We reason that without STE, even although initialized with coarse estimation from VME, MRR will be biased by the input visual features and produce a temporally inconsistent refinement pose that negatively affects the overall estimation. On the other hand, without MRR, our method is reduced to one stage and only estimates the coarse motion. As shown in the result, using only VME will lead to an overly smoothed motion estimation and result in a higher acceleration error (underestimating movement also leads to high acceleration error).

Failure Modes

Although MEVA shows promising results in producing smooth and accurate human motion, there is still room for improvement.

MEVA processes videos using a sliding window: input video sequences are splited into chunks of 90 frames for processing. Due to the natural of this sliding window approach, inconsistency can sometimes be observed at a 3 second interval (videos are assumed to be at 30 fps). The explanation is as follows: the coarse motion estimated by VME can be quite different between each temporal window and MRR sometimes is unable to make enough adjustments to account for a smooth transition. Each temporal window also has their own STE, so the features from each window are longer correlated. Fig. 6 shows an instance of this behavior. For visual inspection, please refer to our supplement video.

2 Occluded body parts

Occluded body parts can still be challenging for MEVA. During occlusion, the lack of visual indicators will compel MEVA to rely on coarse motion estimation over the whole sequence and leads to a misscapture of detailed motion. Please refer to the supplementary video for an example.

3 Missing hands and face movement

Since the original SMPL model does not contain joints for the hand and face, all methods using SMPL do not capture hand movements and facial expressions. Moreover, there is not enough high quality 3D data that provides hand and face annotations. A recent work develops an enhanced SMPL model that jointly models body pose, hands, and face, but this model has not gained significant traction in the pose estimation community. We believe that capturing hands and face movement in motion estimation is an essential direction for future work.

Implementation Details

The motion VAE’s encoder, EvaeE_{vae}, is a bidirectional Gated Recurrent Unit (bi-GRU) with average pooling to obtain the temporal encoding hh of the overall input motion sequence MW∈RW×144M_{W}\in R^{W\times 144}. We pass the temporal encoding hh into a multilayer perceptron (MLP) with two hidden layers (1024, 512) and two heads to obtain the mean μ\mu and variance σ\sigma for the latent code zz. For the decoder DvaeD_{vae}, a forward GRU is used to decode the output motion sequence. At each time step, the GRU takes in the previous step estimation θt−1\theta_{t-1} and the current latent code z∈R1×Szz\in R^{1\times S_{z}} to output a 512 latent feature. The feature is then passed through an MLP with two hidden layers (1024, 512) to generate the reconstructed pose θ∈R1×144\theta\in R^{1\times 144}

2 Spatio-Temporal Feature Extractor

For a video input, we first preprocess the video frames using a pretrained ResNet-59 network . The feature extractor outputs fi∈R2048f_{i}\in R^{2048} for each frame. The extracted features within the same temporal window WW (we choose W=90W=90) are stacked together [ft]t=190∈R90×2048[f_{t}]^{90}_{t=1}\in R^{90\times 2048} and are encoded by STE into a sequence of temporally correlated features [ft′]t=190∈R90×2048[f^{\prime}_{t}]^{90}_{t=1}\in R^{90\times 2048}. STE is a 2 layer bi-GRU with hidden size 1024 that outputs a feature encoding at each timestep. EmotionE_{motion} that shares the same architecture as EfeatE_{feat} except for the final average pooling step to come up with a latent code z∈R1×512z\in R^{1\times 512} that represents the whole motion sequence.

3 Motion Residual Regressor

The MRR consists of 2 fully connected layers, each with 1024 neurons. It takes in per-frame features and a set of initializing parameters (pose, shape, and camera) and iterative refines its predictions (pose, shape, and camera) for kk iterations.