4D Human Body Capture from Egocentric Video via 3D Scene Grounding

Miao Liu, Dexin Yang, Yan Zhang, Zhaopeng Cui, James M. Rehg, Siyu Tang

Introduction

Continuous advancements in the capabilities of Augmented Reality (AR) headsets promise new trends of entertainment, communication, healthcare, and worker productivity, and point towards a revolution in how we interact with the world and communicate with each other. Egocentric vision is a key building block for these emerging capabilities, as AR experiences can benefit from an accurate understanding of the user’s perception, attention, and actions. Substantial progress has been made in understanding human-object interaction and social interaction from egocentric videos. However, future intelligent AR headsets should have the ability to capture the subtle nuances of second-person body pose and render an interactive 3D avatar that is grounded in the 3D scene as it is captured from an egocentric point of view. To this end, we introduce a novel task of 4D second-person full body capture from monocular egocentric videos. As shown in Fig. 4D Human Body Capture from Egocentric Video via 3D Scene Grounding, we seek to reconstruct a time series of 3D second-person body meshes that are temporally-consistent and grounded on the reconstructed 3D scene.

3D human body capture from videos is a key challenge in computer vision, which has received substantial attention over the years . However, none of the previous works considered the challenging setting of reconstructing 3D second-person human body from an egocentric videoWe note that another branch of prior work addresses the related but quite different task of predicting the 3D body pose of the camera-wearer from egocentric video .. The unique viewpoints and embodied camera motions of egocentric video create formidable technical obstacles to 3D body estimation, causing previous SOTA methods for video-based motion capture to fail. For example, the close interpersonal distances that characterize social interactions result in partial observation of the second-person as body parts move in and out of the frame. The egocentric camera motion creates an additional set of challenges, as the second-person motion is entangled with the embodied movement of the camera wearer.

To address these challenges, we propose a simple yet effective optimization-based method that jointly considers a time series of 2D observations and 3D scene information. Our key insight is that the 3D scene provides additional evidence for estimating partially observable human body models. Previous work used a 3D scanner to obtain high quality 3D scene reconstructions, however this approach is not scalable, and is infeasible for outdoor egocentric capture settings. In contrast, we are the first to show that Structure-from-Motion (SfM) can provide a valuable 3D scene context for partially observable body estimation. This is particularly challenging because SfM estimates are up to an unknown scale, and directly placing the 3D body meshes into the reconstructed 3D scene and enforcing human-scene contact will result in unrealistic human-scene interaction. To overcome this challenge, we carefully design the optimization method so that it not only encourages human-scene contact, but also estimates the relative scale between 3D human body and scene reconstruction.The prior work did not face the challenge of scale ambiguity because their 3D scene models came from a 3D scanner. We further unite the time series of body models with a temporal prior to recover more plausible global human motion even when the second-person body captured by the egocentric view is only partially observable.

Because existing egocentric datasets were not collected to address the problem of reconstructing the second-person body pose and shape in 4D, we have collected a new egocentric video dataset – EgoMoCap. Interactions between the first- and second-person in EgoMoCap were structured to yield a variety of interactions over a variety of interpersonal distances, which efficiently cover the variability in real-world social interactions in a compact number of clips. In contrast, previous social interaction datasets such as are naturalistic, but do not systematically cover the range of interpersonal distances needed for research in 4D capture. EgoMoCap is annotated with 2D human keypoints at the frame level. Using EgoMoCap, we compare our body capture approach with the previous state-of-the-art methods for human motion capture from monocular videos, and demonstrate that our method can address the challenging cases where the second-person human body is partially observable. Moreover, we demonstrate that our method can solve the relative scale between 3D scene reconstruction and 3D human body reconstruction from monocular videos, and thereby produce more realistic human-scene interactions. Detailed ablation studies highlight the benefits of our method. In summary, our work makes the following contributions:

• We introduce a novel problem of reconstructing time series of second-person poses and shapes from egocentric videos. To the best of our knowledge, we are the first to capture global human motion grounded in the 3D scene.

• We propose a simple yet effective optimization-based approach that jointly considers a time series of 2D observations and 3D scene context for accurate 4D human body capture. In addition, our approach addresses the scale ambiguity of 3D reconstruction from monocular videos.

• We conduct detailed experiments on our novel EgoMoCap dataset and show that our approach can more accurately reconstruct second-person human body, and encourage more realistic human-scene interaction.

Related Work

The most relevant works to ours are prior investigations of 4D human body reconstruction and human-scene interaction modeling. Our work is also related to recent efforts on reasoning about social interaction from egocentric videos. Specifically, we compare our EgoMoCap dataset with other egocentric human interaction datasets.

4D Human Body Reconstruction. A rich literature has addressed the topic of human body reconstruction. Previous approaches have demonstrated great success in inferring 3D human pose and shape from a single image. More closely-related to this work are prior efforts that leverage video of a moving peprson to infer a time series of 3D human body poses and shapes. Alldieck et al. used optical flow to estimate temporally-coherent human bodies from monocular videos. Tung et al. introduced a self-supervised learning method that uses optical flow, silhouettes, and keypoints to estimate SMPL human body parameters from two consecutive video frames. Dabral et al. presented a weakly supervised learning framework for learning 3D human body pose, and adopted a temporal network to harmonize sequences of 3D pose estimations. used a fully convolutional network to predict 3D human pose from 2D images sequences. Kocabas et al. proposed an adversarial learning framework to produce realistic and accurate human pose and motion from video sequences. Tripathi et al. also explored the knowledge distillation method for 3D human pose prediction. Shimada et al. used a physics engine to capture physically plausible and temporally stable 3D human motion. Those previous works assumed a fixed camera view and fully observable human body. Two works separately considered either the partial observation setting or moving camera setting for human body reconstruction. Rockwell et al. proposed a self-training pipeline for reconstructing 3D human body poses from truncated single image frames. Huang et al. enforced temporal coherence to reconstruct body pose from monocular videos with a moving camera. In contrast, we are the first to tackle the challenging task of reconstructing time-varying body models from egocentric video characterized by both partial observability and significant camera motion.

Other prior works addressed the question of body reconstruction from moving cameras by either combining video with IMU sensor data or leveraging multiple cameras . jointly optimize the camera pose and human body model by leveraging IMU sensory data. Wang et al. proposed to utilize multiple cameras for outdoor human motion capture. More recently, Guzov et al. proposed to estimate the full 3D human pose and location of the camera-wearer within large 3D scenes, by means of wearable sensors. In contrast to these works, we seek to estimate the second-person human motion grounded in the 3D scene using only monocular egocentric videos.

Human-Scene Interaction. Several prior works on human-scene interaction seek to reason about environment affordance . Our work is more relevant to previous efforts on using the environmental cues to estimate 3D human body models. Savva et al. proposed to learn a probabilistic model that captures how humans interact with the indoor scene from RGB-D sensors. Li et al. factorized estimating 3D person-object interactions into an optimal control problem, and used contact constraints to recover human motion and contact forces from monocular videos. Zhang et al. proposed an optimization-based framework that incorporates the scale loss to jointly reconstruct the 3D spatial arrangement and shape of humans and objects in the scene from a single image. Zhang et al. studied the problem of generating plausible human bodies grounded in 3D scenes. The most relevant work is PROX, from Hassan et al. , which introduced an optimization based method to use the 3D scene context for estimating more accurate human pose and shape from a single image. However, our work differs from PROX on several important aspects: First, PROX uses the Matterport 3D scanner to pre-scan the 3D scenes, whereas we use the Structure from Motion (SfM) to reconstruct the 3D scene geometry, resulting in a more general and scalable solution for 4D human pose and shape capture from a single egocentric camera. Second, we show that our contact term can address the challenging problem of the relative scale ambiguity between the estimated 3D human body and the reconstructed 3D scene from monocular videos, whereas the contact term used in PROX is only used to improve the physical plausibility. Third, the problem setting in PROX assumes a static camera with known 3D scene geometry, whereas we capture human motion in unconstrained environments with a moving egocentric camera, resulting in truncation and severe ego-motion.

Egocentric Social Interaction. Understanding human social interaction has been the subject of many recent efforts in egocentric vision . Several egocentric datasets have been proposed for the analysis of human social behavior. The NUS Dataset and JPL Dataset support more general human interaction classification tasks. Yonetani et al. collected a paired egocentric human interaction dataset to study human action and reaction. Park et al. introduced an RGB-D egocentric dataset – EgoMotion, for forecasting first-person walking trajectory. Ng et al. intorduced You2Me dataset to study the problem of egocentric body pose prediction. Fathi et al. presented an egocentric video dataset that captures conversational interactions within a social group, and therefore limited second-person motion is captured in this dataset. Recently, the Ego4D social dataset from was collected for understanding social communication behavior. In fact, none of those datasets were designed to study the second-person body pose and human-scene interaction. In prior datasets, the majority of captured second-person bodies are either largely occluded by objects or frequently truncated by the frustum, which makes their utilization for full body capture infeasible. In contrast, our EgoMoCap dataset focuses on outdoor social interaction scenarios where the second-person body has less occlusion and the Structure-from-Motion is more robust.

Method

We denote an input monocular egocentric video as x=(x1,...,xt)x=(x^{1},...,x^{t}) with its frame xtx^{t} indexed by time tt. We estimate the human body pose and shape at each time step from input xx. Due to the unique viewpoint of egocentric video, the captured second-person body is partially observable within a time window. In addition, the second-person body motion is entangled with the camera motion, creating additional barriers for the enforcement of temporal coherency. To address these challenges, we propose an optimization method that jointly considers the 2D observations of the entire video sequence and the 3D scene context in order to more accurately reconstruct the 4D human body in the presence of partial observability. We illustrate our method in Fig. 2. Specifically, we first recover the 3D human body at each time instant from the 2D observation of xtx^{t}. We then estimate Structure from Motion (SfM) to project a sequence of 3D body meshes into 3D world coordinates based on the recovered global camera motion, and further adopt a contact term to enforce appropriate human-scene interaction. In addition, we combine the 2D cues from entire video sequences for reconstructing temporally-coherent time series of body poses using a human dynamics prior. In the following sections, we introduce each component of our method.

We use the differentiable human body model SMPL-X to represent the body, hands, and facial expression. SMPL-X produces a body mesh of a fixed topology with Nb=10475N_{b}=10475 vertices, using a compact set of body configuration parameters. Specifically, the shape parameter β\beta encodes variations in height, volume, and body proportions; θ\theta encodes the 3D body pose, hand pose, and facial expression information; and γ\gamma denotes the body translation. Formally, the SMPL-X function is defined as Mb(β,θ,γ)M_{b}(\beta,\theta,\gamma). It outputs a 3D body mesh Mb=(Vb,Fb)M_{b}=(V_{b},F_{b}), where Vb∈RNb×3V_{b}\in R^{N_{b}\times 3} and FbF_{b} denote the body vertices and triangular faces. Similar to , we factorize fitting the SMPL-X model to each video frame as an optimization problem. Formally, we optimize (β,θ,γ)(\beta,\theta,\gamma) by minimizing:

where K is the intrinsic camera parameters; the shape prior term Eβ(β)E_{\beta}(\beta) is learned from SMPL-X model body shape training data and the pose prior term Eθ(θ)E_{\theta}(\theta) is learned from CMU MoCap dataset ; λβ\lambda_{\beta} and λθ\lambda_{\theta} denote the weights of Eβ(β)E_{\beta}(\beta) and Eθ(θ)E_{\theta}(\theta); EJE_{J} refers to the energy function that minimizes the weighted robust distance between the 2D projection of the body joints, hand joints and face landmarks, and the corresponding 2D joints estimation from OpenPose . EJE_{J} is given by:

where J(.)J(.) returns 3D joints location based on embedded shape parameters β\beta, and Rθγ(.)R_{\theta\gamma}(.) transforms the joints along the kinematic tree according to the pose θ\theta and body translation γ\gamma; ΠK\Pi_{K} is the 3D to 2D projection function based on intrinsic parameters KK; JestJ_{est} refers to the 2D joints estimated from OpenPose; wiw_{i} is the 2D joints detection confident score which accounts for the noise in 2D joint estimation; kik_{i} is the per-joint weights for annealed optimization as in ; ρJ\rho_{J} denotes a robust Geman-McClure error function that downweights outliers, which is given by: ρJ(e)=e2/(σJ2+e2)\rho_{J}(e)=e^{2}/(\sigma_{J}^{2}+e^{2}), where ee is the residual error, and σj\sigma_{j} is the robustness constant, which is chosen empirically.

2 Egocentric Camera Representation

To capture 4D second-person bodies that are grounded on the 3D scene from egocentric videos, we need to take the embodied camera motion into consideration. Here we elaborate the egocentric camera representation adopted in our method. Formally, we denote Tcb∈R4×4T_{cb}\in R^{4\times 4} as the transformation from the human body coordinate to the egocentric camera coordinate, and TwcT_{wc} as the transformation from the egocentric camera coordinate to the world coordinate. Note that Tcb∈R4×4T_{cb}\in R^{4\times 4} is derived from the translation parameter γ\gamma of SMPL-X model fitting introduced in Sec. 3.1, while TwcT_{wc} is returned from COLMAP Structure from Motion (SfM) . In order to utilize the 3D scene context and enforce the temporal coherency on reconstructed human body meshes, we project the 3D second-person body vertices VbV_{b} into world coordinate using human body to world transformation TwbT_{wb}, which is given by:

where V^bt\hat{V}_{b}^{t} refers to the body vertices at time step tt, represented in homogeneous coordinate.

3 Optimization with 3D Scene

3D Scene Representation. The structure of the 3D scene constrains and informs human behavior, and therefore 3D scene context can play an important role in 3D human body recovery. Human-scene interaction can be described in relation to 3D surfaces, and therefore we adopt a mesh representation for the 3D scene. Formally, we denote the 3D scene mesh as Ms=(Vs,Fs)M_{s}=(V_{s},F_{s}), where Vs∈RNs×3V_{s}\in R^{N_{s}\times 3} denotes the vertices of the scene representation, and FsF_{s} denotes the corresponding triangular faces. We use the dense environment reconstruction from COLMAP to obtain MsM_{s}. Specifically, COLMAP first reconstructs a sparse representation of the scene and the camera poses of the input images using SfM. It then calculates the depth and normal maps for all registered images, and fuses the depth and normal maps into a dense point cloud with normal information. Finally, Poisson Reconstruction is used to generate the 3D scene mesh representation MsM_{s}.

Human-Scene Contact. Note that the reconstructed 3D scene from the monocular video is up to a scale. To address this scale ambiguity, we design a novel energy function that not only encourages contact between the human body and 3D scene, but also estimates the relative scale between 3D scene mesh MsM_{s} and 3D body mesh MbM_{b}. Specifically, we make use of the annotation from , where a candidate set of SMPL-X mesh vertices Vc∈VbV_{c}\in V_{b} to contact with the world were provided. We then multiply an optimizable scale parameter S∈RS\in R to human body vertices VcV_{c} during optimization. Therefore, the energy function for enforcing human-scene contact is given by:

where ρc\rho_{c} is the robust Geman-McClure error function introduced before, and TwbT_{wb} is human body to world transformation introduced in Eq. 3. Note that the scale factor SS is shared across the video sequence. This is because we estimate a consistent 3D shape parameter θ\theta from the entire sequence by taking the median of all the shape parameters obtained from the per-frame SMPL-X model fitting.

4 Human Dynamics Prior

Fitting SMPL-X human body model to each video frame will incur notable temporal inconsistency. Due to drastic camera motion, this problem is further amplified under egocentric scenarios. Here, we propose to use the empirical human dynamics priors to enforce temporal coherency on human body models in the world coordinates. Formally, we have the following energy function:

where VwbiV_{wb}^{i} is the 3D human body vertices at time step ii, transformed in world coordinate as in Eq. 3; ρT\rho_{T} is another robust Geman-McClure error function that accounts for possible outliers; and wviw_{v}^{i} is the confidence score of 2D human keypoints estimation. As shown in Eq. 5, we design this energy function to focus on body parts that do not have reliable 2D observations (caused by the unique egocentric viewpoint). Notably, we assume a zero acceleration motion prior. This naive prior was proven to be effective in capturing human motion in the outdoor environment .

5 Optimization

Putting everything together, we have the following energy function for our optimization method:

where EMiE_{M}^{i} denotes the SMPL-X model fitting energy function for video frame xix^{i}; λC\lambda_{C} and λT\lambda_{T} represent the weights for human-scene contact term and human dynamic prior term, respectively. We optimize Eq. 6 using a gradient-based optimizer Adam w.r.t. SMPL-X body parameters β,θ,γ\beta,\theta,\gamma, scale parameter SS, and camera to world transformation TwcT_{wc}. Note that the SfM already provides an initialization of TwcT_{wc}, and incorporating TwcT_{wc} into the optimization can further smooth the global second-person human motion.

Note that EME_{M} performs model fitting at each time step, while ECE_{C} and ETE_{T} optimize a time series of body models. In addition, ECE_{C} and ETE_{T} seek to optimize human body parameters in world coordinates, and the scale ambiguity can cause the gradients of the contact term to shift the body global position in the wrong direction. Therefore, we propose a multi-stage optimization strategy. Specifically, we set λC\lambda_{C} and λT\lambda_{T} to be zero, so that the optimizer will only look at the 2D observations in the first stage. We then set λC\lambda_{C} to be 0.1, keeping λT\lambda_{T} as zero, and freezing TwcT_{wc}, so that the optimizer will focus on recovering the scale parameter SS. At the final stage, we set λT\lambda_{T} to 0.1 and enable the gradients of TwcT_{wc} to enforce temporal coherency.

Experiments

Datasets. To study the problem of second-person human body reconstruction, we present a new egocentric social interaction dataset – EgoMoCap. This dataset consists of 36 videos sequences from one-on-one interactions between 4 individuals. The camera wearer is equipped with a head-mounted camera, and the other participant is asked to interact with the camera wearer in a natural manner. The video sequences were recorded in 1920×1080 resolution at 60 fps using the GoPro Hero8 camera. This dataset captures 4 types of outdoor human social interactions: Greeting, Touring, Jogging Together, and Throw and Catch. We further annotate the captured second-person human bodies with 2D keypoints via Amazon Mechanical Turk (AMT).

Evaluation Metrics. For our experiments, we evaluate the human body reconstruction accuracy, motion smoothness, and the plausibility of the human-scene interaction.

• Human Body Reconstruction Accuracy: We acknowledge that the 3D ground truth of human bodies can be obtained from RGB-D data , or Motion Capture Systems . However, capturing the 3D human body ground truth in naturalistic outdoor social interaction setting remains a challenge. Therefore, we follow to evaluate the reconstruction quality using per-joint 2D projection error (PJE) on the image plane. Specifically, we report PJE-P on frames with partially observable second-person body, and PJE-U on frames with mostly untruncated second-person body (uniform sampled frames). Here, we evaluate human body poses, even though our method has the capacity of reconstructing 3D hands and faces, as human-scene contact used in our method has minor influence on facial expression and hand pose during social interaction.

• Motion Smoothness: We follow to adopt a physics-based metric that uses average magnitude of joint accelerations to measure the smoothness of the estimated pose sequence. Thus, a lower value indicates that the times series of body meshes have more consistent human motion. Note that the motion smoothness is evaluated on 3D human joints projected in world coordinate. For fair comparison, we normalize the scale factor when reporting the results.

• Plausibility of Human-Scene Interaction: To evaluate whether our method provides more realistic human-scene interaction, we transform the human body meshes into 3D world coordinates, render the results as video sequences, and further upload them to Amazon Mechanical Turk (AMT) for a user study. Specifically, we put the rendered results of all compared methods and our method side-by-side (sample videos can be found in supplementary materials), and ask the AMT worker to choose the instance that has the most realistic human-scene interaction.

2 Quantitative Results

In this section, we introduce our quantitative experimental results. We first present detailed ablation studies, and then compare our method with state-of-the-art for 3D human body reconstruction from monocular videos.

Ablation Study. We first analyze the functionality of the terms in Eq. 6. The results are summarized in Table 1. EME_{M} refers to the baseline method that performs per-frame fitting with 2D observations as in SMPLify-X . EME_{M} has undesirable performance on PJE-P, motion smoothness and human-scene interaction user study. In the second row EM+ECE_{M}+E_{C}, we report the method that makes use of both human scene contact term and 2D observations. Though adding the contact term alone leads to more realistic human-scene interaction, it compromises the performance on 2D projection error and motion smoothness by a notable margin. EM+ETE_{M}+E_{T} in the third row refers to the method that optimizes the 2D observations together with the human dynamic prior term ETE_{T}. Not surprisingly, ETE_{T} can significantly improve the motion smoothness. It is also worthy noting that using temporal prior alone can not improve the reconstruction quality. This suggests that simply enforcing temporal coherency without 3D scene grounding may cause the optimization method converging to sub-optimal equilibrium. In the last row, we present the results of our full optimization approach. Our method achieves the best performance on motion smoothness and plausibility of human-scene interaction. An interesting observation is that ours outperforms EM+ETE_{M}+E_{T} by a notable margin on motion smoothness. We speculate that this is because the physical human scene constraints narrow down the solution space of model fitting, and thereby leads to more optimal performance on temporal coherency. Note that, in the simple case when the figure is untruncated (PJE-U), EME_{M} gives good performance and slightly exceeds the accuracy of our full method. This is because PJE is a 2D metric, and therefore favors the method that adopts only 2D projection error as objective function during optimization. However, when the 2D observation can not be robustly estimated due to partial observation, our method outperforms other baselines by a significant margin (66.03 vs. 73.14 in PJE-P). Those results demonstrate that our method can address the challenge of partially observable human bodies, and estimate plausible global human motion grounded on the 3D scene.

Comparison to VIBE. Though many methods have addressed human body capture from monocular video. In Table 2, we compare our approach with a widely-used competitive method – VIBE . Since VIBE does not model the human-scene constraints, it provides unrealistic human-scene interaction. Moreover, the egocentric camera motion causes VIBE failing to capture coherent human motion. In contrast, our method outperforms VIBE on motion smoothness and human-scene interaction plausibility by a large margin. Though VIBE performs slightly better on PJE-U (22.45 vs. 24.03), it lags far behind of our method on PJE-P (75.91 vs. 66.03). We have to re-emphasize that the 2D projection error cannot reflect the true performance improvement of our method. This is because the 2D keypoints annotation is only available for visible human body parts, and therefore 2D PJE does not penalize the method that fits a wrong 3D body model to partially 2D observation. Take the VIBE result shown in the third row of Fig. 3 for an instance, the 2D projection error may have decent performance, yet the reconstructed 3D human body is completely wrong.

3 Discussion

Qualitative Results. As shown in Fig. 3, we first visualize our results by rendering the estimated body models from the egocentric viewpoint, so it can be directly imposed on the input video frames. Notably, both SMPLify-X and VIBE fail substantially for challenging cases where human body is partially-observable. Our method, on the other hand, makes uses of 3D scene context and harmonizes the 2D cues from the entire video sequence, and therefore successfully reconstructs the human body with only partial observation. We provide an additional visualization of the results of both EME_{M} baseline and our method in the world coordinate system in Fig. 4. By examining the SMPLify-X baseline results, we can observe an obvious mismatched scale between the 3D reconstruction of the human body and the 3D environment. In contrast, our method produces more plausible human body motion grounded in the 3D scene by resolving the relative scale between 3D scene reconstruction and 3D human body reconstruction from monocular videos. In the supplementary materials, we provide additional videos that demonstrate the benefits of our approach.

Limitations. A key issue of our method is the need to retrieve the camera trajectory and 3D scene only from monocular RGB videos via Structure from Motion (SfM). Therefore, our method has the same bottleneck as SfM: Challenging factors such as dynamic scenes, featureless surfaces, large illumination change, etc., may cause visual feature matching to fail. One promising direction for future work is to incorporate additional sensing modalities (IMU, depth estimation, and multiple cameras) to further stabilize the 3D scene reconstruction in challenging conditions. Another issue is that the naive human motion prior (zero acceleration) adopted in our method may result in unrealistic motions in some cases. More efforts in learning motion priors could potentially address this issue. In summary, we believe our efforts constitute an important step forward for a largely unexplored egocentric vision task, and we hope our work can motivate the community to make further investments.

Conclusion

We introduce a novel task of estimating a time series of 3D human body models for the second-person in an egocentric video, which are temporally-coherent and grounded on the 3D scene. To address the challenges of egocentric video, we propose an effective optimization-based method that exploits the 2D observations of the entire video sequence and human-scene contact for human motion capture. We conduct detailed experiments on our EgoMoCap dataset to demonstrate the benefits of our approach. We believe our work points to exciting research directions in egocentric social interaction analysis and 4D human body reconstruction.

Acknowledgments. Portions of this research were supported in part by National Science Foundation Award 2033413 and a gift from Facebook.

References