GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras

Ye Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani, Jan Kautz

Introduction

Recovering fine-grained 3D human meshes from monocular videos is essential for understanding human behaviors and interactions, which can be the cornerstone for numerous applications including virtual or augmented reality, assistive living, autonomous driving, etc. Many of these applications use dynamic cameras to capture human behaviors yet also require estimating human motions in global coordinates consistent with their surroundings. For instance, assistive robots and autonomous vehicles need a holistic understanding of human behaviors and interactions in the world to safely plan their actions even when they are moving. Therefore, our goal in this paper is to tackle the important task of recovering global human meshes from monocular videos captured by dynamic cameras.

However, this task is highly challenging for two main reasons. First, dynamic cameras make it difficult to estimate human motions in consistent global coordinates. Existing human mesh recovery methods estimate human meshes in the camera coordinates or even in the root-relative coordinates . Hence, they can only recover global human meshes from dynamic cameras by using SLAM to estimate camera poses . However, SLAM can often fail for in-the-wild videos due to moving and dynamic objects. It also has the problem of scale ambiguity, which often leads to camera poses that are inconsistent with the human motions. Second, videos captured by dynamic cameras often contain severe and long-term occlusions of humans, which can be caused by missed detection, complete obstruction by objects and other people, or the person going outside the camera’s field of view (FoV). These occlusions pose serious challenges to standard human mesh recovery methods, which rely on detections or visible parts to estimate human meshes. Only a few works have attempted to tackle the occlusion problem in human mesh recovery . However, these methods can only address partial occlusions of a person and fail to handle severe occlusions when the person is completely invisible for an extended period of time.

To tackle the above challenges, we propose Global Occlusion-Aware Human Mesh Recovery (GLAMR), which can handle severe occlusions and estimate human meshes in consistent global coordinates – even for videos recorded with dynamic cameras. We start by using off-the-shelf methods (e.g., KAMA or SPEC ) to estimate the shape and pose sequences (motions) of visible people in the camera coordinates. These methods also rely on multi-object tracking and re-identification, which provide occlusion information, and the motion of occluded frames is not estimated. To tackle potentially severe occlusions, we propose a deep generative motion infiller that autoregressively infills the local body motions of occluded people based on visible motions. The motion infiller leverages human dynamics learned from a large motion database, AMASS . Next, to obtain global motions, we propose a global trajectory predictor that can generate global human trajectories based on local body motions. It is motivated by the observation that the global root trajectory of a person is highly correlated with the local body movements. Finally, using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence such as 2D keypoints.

The contributions of this paper are as follows: (1) We propose the first approach to address long-term occlusions and estimate global 3D human pose and shape from videos captured by dynamic cameras; (2) We propose a novel generative Transformer-based motion infiller that autoregressively infills long-term missing motions, which considerably outperforms state-of-the-art motion infilling methods; (3) We propose a method to generate global human trajectories from local body motions and use the generated trajectories as anchors to constrain global motion and camera optimization; (4) Extensive experiments on challenging indoor and in-the-wild datasets demonstrate that our approach outperforms prior state-of-the-art methods significantly in tackling occlusions and estimating global human meshes.

Related Work

Camera-Relative Pose Estimation. 3D human mesh recovery from RGB images or videos is an ill-posed problem due to the depth ambiguity. Most existing methods simplify the problem by estimating human poses relative to the pelvis (root) of the human body . These methods assume an orthographic camera projection model and neglect the absolute 3D translation of the person w.r.t. the camera. To address the lack of translation, recent methods start to estimate human meshes in the camera coordinates . Several approaches recover the absolute translation of the person using an optimization framework . A few methods exploit various scene constraints during the optimization process to improve depth prediction . Alternatively, recent approaches use physics-based constraints to ensure the physical plausibility of the estimated poses . Iqbal et al. exploit a limb-length constraint to recover the absolute translation of the person using a 2.5D representation. Some approaches approximate the depth of the person using the bounding box size . HybrIK and KAMA employ inverse kinematics to estimate human meshes with absolute translations in the camera coordinates. Several methods directly predict the absolute depth of each person using a heatmap representation . Recently, SPEC learns to predict the camera parameters (pitch, yaw, FoV) from the image, which are used for absolute pose regression in the camera coordinates. THUNDR also adopts a similar strategy but uses known camera parameters. While these methods show impressive results, they cannot estimate global human motions from videos captured by dynamic cameras. In contrast, our approach can recover human meshes in consistent global coordinates for dynamic cameras and handle severe and long-term occlusions.

Global Pose Estimation. Most existing methods that estimate 3D poses in world coordinates rely on calibrated, synchronized, and static multi-view capture setups . Huang et al. use uncalibrated cameras but still assume time synchronization and static camera setups. Hasler et al. handle unsynchronized moving cameras but assume multi-view input and rely on audio stream for synchronization. More recently, Dong et al. propose to recover 3D poses from unaligned internet videos of different actors performing the same activity from unknown cameras. However, they assume that multiple viewpoints of the same pose are available in the videos. Different from these methods, our approach estimates human meshes in global coordinates from monocular videos recorded with dynamic cameras. Several methods rely on additional IMU sensors or pre-scanned environments to recover global human motions , which is unpractical for large-scale adoption. Recently, another line of work starts to focus on estimating accurate human-scene interaction . Liu et al. first obtain the camera poses and dense reconstruction of the scene from dynamic cameras using a SLAM algorithm, COLMAP . The camera poses are used for camera-to-world transformation, while the reconstructed scene is used to encourage human-scene contacts. However, SLAM can often fail for the in-the-wild videos and is prone to error propagation. In contrast, our approach does not require SLAM but instead uses global trajectory prediction to constrain the joint reconstruction of human motions and camera poses. Additionally, our approach can also handle severe and long-term occlusions common in dynamic camera setups.

Occlusion-Aware Pose Estimation. Most existing human pose estimation methods assume the person is fully visible in the images and are not robust to strong occlusions. Only a few methods address the occlusion problem in pose estimation . While these methods show impressive results under partial occlusions, they do not address severe and long-term occlusions when people are completely obstructed or outside the camera’s FoV for a long time. In contrast, our approach leverages deep generative human motion models to tackle severe and long-term occlusions.

Human Motion Modeling. Extensive research has studied 3D human dynamics for various tasks including motion prediction and synthesis . Recent human pose estimation methods start to leverage learned human dynamics models to improve the accuracy of estimated motions . Several motion infilling approaches are also proposed to generate complete motions from partially observed motions . Additionally, recent work on motion capture shows that global human translations can be predicted from 3D local joint positions . In contrast to prior work, our trajectory predictor does not require GT root orientations but can predict both global root translations and orientations. Furthermore, we also propose a novel generative autoregressive motion infiller that can use noisy poses as input instead of high-quality GT poses, and we demonstrate its effectiveness in tackling long-term occlusions in human pose estimation.

Method

As outlined in Fig. 1, our framework consists of four stages. In Stage I, we first use multi-object tracking (MOT) and re-identification algorithms to obtain the bounding box sequence of each person, which is input to a human mesh recovery method (e.g., KAMA or SPEC ) to extract the motion Q~i\boldsymbol{\widetilde{Q}}^{i} of each person (including translation) in the camera coordinates. The motion Q~i\boldsymbol{\widetilde{Q}}^{i} may be incomplete due to various occlusions (e.g., obstruction, missed detection, going outside FoV), where bounding boxes from MOT are missing for some frames. In Stage II (Sec. 3.1), we propose a generative motion infiller to tackle the occlusions in the estimated body motion Θ~i\boldsymbol{\widetilde{\Theta}}^{i} and produce occlusion-free body motion Θ^i\boldsymbol{\widehat{\Theta}}^{i}. In Stage III (Sec. 3.2), we propose a global trajectory predictor that uses the infilled body motion Θ^i\boldsymbol{\widehat{\Theta}}^{i} to generate the global trajectory (root translations and rotations) of each person and obtain their global motion Q^i\boldsymbol{\widehat{Q}}^{i}. In Stage IV (Sec. 3.3), we jointly optimize the global trajectories of all people and the camera parameters to produce global motions Qˇi\boldsymbol{\widecheck{Q}}^{i} consistent with the video evidence.

The task of the generative motion infiller M\mathcal{M} is to infill the occluded body motion Θ~i\boldsymbol{\widetilde{\Theta}}^{i} of each person to produce occlusion-free body motion Θ^i\boldsymbol{\widehat{\Theta}}^{i}. Here, we do not use the motion infiller M\mathcal{M} to infill other components in the estimated motion Q^i\boldsymbol{\widehat{Q}}^{i}, i.e., root trajectory (T~i,R~i\boldsymbol{\widetilde{T}}^{i},\boldsymbol{\widetilde{R}}^{i}) and shapes B~i\boldsymbol{\widetilde{B}}^{i}. This is because it is difficult to infill the root trajectory (T~i,R~i)(\boldsymbol{\widetilde{T}}^{i},\boldsymbol{\widetilde{R}}^{i}) using learned human dynamics, since it resides in the camera coordinates rather than a consistent coordinate system due to the dynamic camera. In Sec. 3.2, we will use the proposed global trajectory predictor to generate occlusion-free global trajectory (T^i,R^i)(\boldsymbol{\widehat{T}}^{i},\boldsymbol{\widehat{R}}^{i}) from the infilled body motion Θ^i\boldsymbol{\widehat{\Theta}}^{i}. The trajectory (T~i,R~i)(\boldsymbol{\widetilde{T}}^{i},\boldsymbol{\widetilde{R}}^{i}) from the pose estimator is not discarded and will be used in the global optimization (Sec. 3.3). We use linear interpolation to produce occlusion-free shapes B^i\boldsymbol{\widehat{B}}^{i}, which can be time-varying to be compatible with per-frame pose estimators such as KAMA.

Given a general occluded human body motion Θ~=(θ~1,…,θ~h)\boldsymbol{\widetilde{\Theta}}=(\boldsymbol{\widetilde{\theta}}_{1},\ldots,\boldsymbol{\widetilde{\theta}}_{h}) of hh frames and its visibility mask V=(V1,…,Vh)\boldsymbol{V}=(V_{1},\ldots,V_{h}) as input, the motion infiller M\mathcal{M} outputs a complete occlusion-free motion Θ^=(θ^1,…,θ^h)\boldsymbol{\widehat{\Theta}}=(\boldsymbol{\widehat{\theta}}_{1},\ldots,\boldsymbol{\widehat{\theta}}_{h}). The visibility mask V\boldsymbol{V} encodes the visibility of the occluded motion Θ~\boldsymbol{\widetilde{\Theta}}, where Vt=1V_{t}=1 if the body pose θ~t\boldsymbol{\widetilde{\theta}}_{t} is visible in frame tt and Vt=0V_{t}=0 otherwise. Since the human pose for occluded frames can be highly uncertain and stochastic, we formulate the motion infiller M\mathcal{M} using the conditional variational autoencoder (CVAE) :

where the motion infiller M\mathcal{M} corresponds to the CVAE decoder and z\boldsymbol{z} is a Gaussian latent code. We can obtain different occlusion-free motions Θ^\boldsymbol{\widehat{\Theta}} by varying z\boldsymbol{z}.

Autoregressive Motion Infilling. To ensure that the motion infiller M\mathcal{M} can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 2 (Left). The key idea is to use a sliding window of hh frames, where we assume the first hch_{\texttt{c}} frames of motion are already occlusion-free or infilled and serve as context, and we also use the last hlh_{\texttt{l}} frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions. Excluding the context and look-ahead frames, only the middle ho=h−hc−hlh_{\texttt{o}}=h-h_{\texttt{c}}-h_{\texttt{l}} frames of motion are infilled. We iteratively infill the motion using the sliding window and advance the window by hoh_{\texttt{o}} frames every step.

Motion Infiller Network. The overall network design of the CVAE-based motion infiller is outlined in Fig. 2 (Right). In particular, we employ a Transformer-based seq2seq architecture, which consists of three parts: (1) a context network that uses a Transformer encoder to encode the visible poses from the occluded motion Θ~\boldsymbol{\widetilde{\Theta}} into a context sequence, which serves as the condition for other networks; (2) a decoder network that uses the latent code z\boldsymbol{z} and context sequence to generate occlusion-free motion Θ^\boldsymbol{\widehat{\Theta}} via a Transformer decoder and a multilayer perceptron (MLP); (3) prior and posterior networks that generate the prior and posterior distributions for the latent code z\boldsymbol{z}. In the networks, we adopt a time-based encoding that replaces the position in the original positional encoding with the time index. Unlike prior CNN-based methods , our Transformer-based motion infiller does not require padding missing frames, but instead restricts its attention to visible frames to achieve effective temporal modeling.

Training. We train the motion infiller M\mathcal{M} using a large motion capture dataset, AMASS . To synthesize occluded motions Θ~\boldsymbol{\widetilde{\Theta}}, for any GT training motion Θ~′\boldsymbol{\widetilde{\Theta}}^{\prime} of hh frames, we randomly occlude HoccH_{\texttt{occ}} consecutive frames of motion where HoccH_{\texttt{occ}} is uniformly sampled from [Hlb,Hub][H_{\texttt{lb}},H_{\texttt{ub}}]. Note that we do not occlude the first hch_{\texttt{c}} frames which are reserved as context. We use the standard CVAE objective to train the motion infiller M\mathcal{M}:

where LKLzL_{\texttt{KL}}^{\boldsymbol{z}} is the KL divergence between the prior and posterior distributions of the CVAE latent code z\boldsymbol{z}.

2 Global Trajectory Predictor

After we obtain occlusion-free body motion Θ^i\boldsymbol{\widehat{\Theta}}^{i} for each person using the motion infiller, a key problem still remains: the estimated trajectory (T~i,R~i)(\boldsymbol{\widetilde{T}}^{i},\boldsymbol{\widetilde{R}}^{i}) of the person is still occluded and not in a consistent global coordinate system. To tackle this problem, we propose to learn a global trajectory predictor T\mathcal{T} that generates a person’s occlusion-free global trajectory (T^i,R^i)(\boldsymbol{\widehat{T}}^{i},\boldsymbol{\widehat{R}}^{i}) from the local body motion Θ^i\boldsymbol{\widehat{\Theta}}^{i}.

Given a general occlusion-free body motion Θ=(θ1,…,θm){\boldsymbol{\Theta}}=({\boldsymbol{\theta}}_{1},\ldots,{\boldsymbol{\theta}}_{m}) as input, the trajectory predictor T\mathcal{T} outputs its corresponding global trajectory (T,R)({\boldsymbol{T}},{\boldsymbol{R}}) including the root translations T=(τ1,…,τm){\boldsymbol{T}}=({\boldsymbol{\tau}}_{1},\ldots,{\boldsymbol{\tau}}_{m}) and rotations R=(γ1,…,γm){\boldsymbol{R}}=({\boldsymbol{\gamma}}_{1},\ldots,{\boldsymbol{\gamma}}_{m}). To address any potential ambiguity in the global trajectory, we also formulate the global trajectory predictor using the CVAE:

where the global trajectory predictor T\mathcal{T} corresponds to the CVAE decoder and v\boldsymbol{v} is the latent code for the CVAE. In Eq. (3), the immediate output of the global trajectory predictor T\mathcal{T} is an egocentric trajectory Ψ=(ψ1,…,ψm)\boldsymbol{{\Psi}}=(\boldsymbol{{\psi}}_{1},\ldots,\boldsymbol{{\psi}}_{m}), which by design can be converted to a global trajectory (T,R)(\boldsymbol{{T}},\boldsymbol{{R}}) using a conversion function EgoToGlobal.

Egocentric Trajectory Representation. The egocentric trajectory Ψ\boldsymbol{{\Psi}} is just an alternative representation of the global trajectory (T,R)(\boldsymbol{{T}},\boldsymbol{{R}}). It converts the global trajectory into relative local differences and represents rotations and translations in the heading coordinates (yy-axis aligned with the heading, i.e., the person’s facing direction). In this way, the egocentric trajectory representation is invariant of the absolute xyxy translation and heading. It is more suitable for the prediction of long trajectories, since the network only needs to output the local trajectory change of every frame instead of the potentially large global trajectory offset.

The conversion from the global trajectory to the egocentric trajectory is given by another function: Ψ=GlobalToEgo(T,R)\boldsymbol{{\Psi}}=\texttt{GlobalToEgo}(\boldsymbol{{T}},\boldsymbol{{R}}), which is the inverse of the function EgoToGlobal. In particular, the egocentric trajectory ψt=(δxt,δyt,zt,δϕt,ηt)\boldsymbol{{\psi}}_{t}=(\delta x_{t},\delta y_{t},z_{t},\delta\phi_{t},\boldsymbol{\eta}_{t}) at time tt is computed as:

where τtxy\boldsymbol{\tau}_{t}^{xy} is the xyxy component of the translation τt\boldsymbol{\tau}_{t}, τtz\boldsymbol{\tau}_{t}^{z} is the zz component (height) of τt\boldsymbol{\tau}_{t}, γtϕ\boldsymbol{\gamma}_{t}^{\phi} is the heading angle of the rotation γt\boldsymbol{\gamma}_{t}, ToHeading is a function that converts translations or rotations to the heading coordinates defined by the heading γtϕ\boldsymbol{\gamma}_{t}^{\phi}, and ηt\boldsymbol{\eta}_{t} is the local rotation. As an exception, (δx0,δy0)(\delta x_{0},\delta y_{0}) and δϕ0\delta\phi_{0} are used to store the initial xyxy translation τ0xy\boldsymbol{\tau}_{0}^{xy} and heading τ0ϕ\boldsymbol{\tau}_{0}^{\phi}. These initial values are set to the GT during training and arbitrary values during inference (as the trajectory can start from any position and heading). The inverse process of Eq. (5)-(7) defines the inverse conversion EgoToGlobal used in Eq. (4), which accumulates the egocentric trajectory to obtain the global trajectory. To correct potential drifts in the trajectory, in Sec. 3.3, we will optimize the global trajectory of each person to match the video evidence, which also solves the trajectory’s starting point (δx0,δy0,δϕ0)(\delta x_{0},\delta y_{0},\delta\phi_{0}). More details about the egocentric trajectory are given in Appendix D.

Network and Training. The trajectory predictor adopts a similar network design as the motion infiller with one main difference: we use LSTMs for temporal modeling instead of Transformers since the output of each frame is the local trajectory change in our egocentric trajectory representation, which mainly depends on the body motion of nearby frames and does not require long-range temporal modeling. We will show in Sec. 4.2 that the egocentric trajectory and use of LSTMs instead of Transformers are crucial for accurate trajectory prediction. Please refer to Appendix D for the detailed network architectures. We use the standard CVAE objective to train the trajectory predictor T\mathcal{T}:

where τt′\boldsymbol{{\tau}}^{\prime}_{t} and γt′\boldsymbol{\gamma}^{\prime}_{t} denote the GT translation and rotation, ⊖\ominus computes the relative rotation, ∥⋅∥a\|\cdot\|_{a} computes the rotation angle, and LKLvL_{\texttt{KL}}^{\boldsymbol{v}} is the KL divergence between the prior and posterior distributions of the CVAE latent code v\boldsymbol{v}. We again use AMASS to train the trajectory predictor T\mathcal{T}.

3 Global Optimization

After using the generative motion infiller and global trajectory predictor, we have obtained an occlusion-free global motion Q^i=(T^i,R^i,Θ^i,B^i)\widehat{\boldsymbol{Q}}^{i}=(\widehat{\boldsymbol{T}}^{i},\widehat{\boldsymbol{R}}^{i},\widehat{\boldsymbol{\Theta}}^{i},\widehat{\boldsymbol{B}}^{i}) for each person in the video. However, the global trajectory predictor generates trajectories for each person independently, which may not be consistent with the video evidence. To tackle this problem, we propose a global optimization process that jointly optimizes the global trajectories of all people and the extrinsic camera parameters to match the video evidence such as 2D keypoints. The final output of the global optimization and our framework is Qˇi=(Tˇi,Rˇi,Θˇi,Bˇi)\widecheck{\boldsymbol{Q}}^{i}=(\widecheck{\boldsymbol{T}}^{i},\widecheck{\boldsymbol{R}}^{i},\widecheck{\boldsymbol{\Theta}}^{i},\widecheck{\boldsymbol{B}}^{i}) where (Θˇi,Bˇi)=(Θ^i,B^i)(\widecheck{\boldsymbol{\Theta}}^{i},\widecheck{\boldsymbol{B}}^{i})=(\widehat{\boldsymbol{\Theta}}^{i},\widehat{\boldsymbol{B}}^{i}), i.e., we directly use the occlusion-free body motion and shapes from the previous stages.

Optimization Variables. The first set of variables we optimize is the egocentric representation {Ψˇi}i=1N\{\widecheck{\boldsymbol{\Psi}}^{i}\}_{i=1}^{N} of the global trajectories {(Tˇi,Rˇi)}i=1N\{(\widecheck{\boldsymbol{T}}^{i},\widecheck{\boldsymbol{R}}^{i})\}_{i=1}^{N}. We adopt the egocentric representation since it allows corrections of the translation and heading at one frame to propagate to all future frames. Therefore, it enables optimizing the trajectories of occluded frames since they will impact future visible frames under the egocentric trajectory representation. We will empirically demonstrate its effectiveness in Sec. 4.2.

Energy Function. The energy function we aim to minimize is defined as

where we use five energy terms with their corresponding coefficients λ2D,λtraj,λreg,λcam,λpen\lambda_{\texttt{2D}},\lambda_{\texttt{traj}},\lambda_{\texttt{reg}},\lambda_{\texttt{cam}},\lambda_{\texttt{pen}}.

where VtiV_{t}^{i} is person ii’s visibility at frame tt, Π\Pi is the camera projection with extrinsics Ct\boldsymbol{C}_{t} and approximated intrinsics K\boldsymbol{K}, and Xˇti\widecheck{\boldsymbol{X}}_{t}^{i} is computed using the SMPL joint function J\mathcal{J} from the optimized global pose qˇti=(τˇti,γˇti,θˇti,βˇti)∈Qˇi\widecheck{\boldsymbol{q}}_{t}^{i}=(\widecheck{\boldsymbol{\tau}}_{t}^{i},\widecheck{\boldsymbol{\gamma}}_{t}^{i},\widecheck{\boldsymbol{\theta}}_{t}^{i},\widecheck{\boldsymbol{\beta}}_{t}^{i})\in\widecheck{\boldsymbol{Q}}^{i}.

The second term EtrajE_{\texttt{traj}} measures the difference between the optimized global trajectory (Tˇi,Rˇi)(\widecheck{\boldsymbol{T}}^{i},\widecheck{\boldsymbol{R}}^{i}) viewed in the camera coordinates and the trajectory (T~i,R~i)(\widetilde{\boldsymbol{T}}^{i},\widetilde{\boldsymbol{R}}^{i}) output by the pose estimator (e.g., KAMA ) in Stage I:

where the function Γ(⋅,Ct)\Gamma(\cdot,\boldsymbol{C}_{t}) transforms the global rotation γˇti\widecheck{\boldsymbol{\gamma}}_{t}^{i} or translation τˇti\widecheck{\boldsymbol{\tau}}_{t}^{i} to the camera coordinates defined by Ct\boldsymbol{C}_{t}, and wtw_{t} is a weighting factor for the translation term.

The third term EregE_{\texttt{reg}} regularizes the egocentric trajectory Ψˇi\widecheck{\boldsymbol{\Psi}}^{i} to stay close to the output Ψ^i\widehat{\boldsymbol{\Psi}}^{i} of the trajectory predictor:

where ∘\circ denotes the element-wise product and wψ\boldsymbol{w}_{\psi} is a weighting vector for each element inside the egocentric trajectory. As an exception, we do not regularize each person’s initial xyxy position and heading (δxˇ0i,δyˇ0i,δϕˇ0i)⊂ψˇ0i(\delta\widecheck{x}^{i}_{0},\delta\widecheck{y}^{i}_{0},\delta\widecheck{\phi}^{i}_{0})\subset\widecheck{\boldsymbol{\psi}}_{0}^{i} as they need to be inferred from the video.

The fourth term EcamE_{\texttt{cam}} measures the smoothness of the camera parameters C\boldsymbol{C} and the uprightness of the camera:

where ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle denotes the inner product, Cty\boldsymbol{C}_{t}^{y} is the +y+y vector of the camera Ct\boldsymbol{C}_{t}, and Y\boldsymbol{Y} is the global up direction. Ctγ\boldsymbol{C}_{t}^{\gamma} and Ctτ\boldsymbol{C}_{t}^{\tau} denote the rotation and translation of the camera Ct\boldsymbol{C}_{t}.

The final term EpenE_{\texttt{pen}} is an signed distance field (SDF)-based inter-person penetration loss adopted from .

Experiments

Datasets. We employ the following datasets in our experiments: (1) AMASS , which is a large human motion database with 11000+ human motions. We use AMASS to train and evaluate the motion infiller and trajectory predictor. (2) 3DPW , which is an in-the-wild human motion dataset that uses videos and wearable IMU sensors to obtain GT poses, even when the person is occluded. We evaluate our approach using the test split of 3DPW. (3) Dynamic Human3.6M is a new benchmark for human pose estimation with dynamic cameras that we create from the Human3.6M dataset . We simulate dynamic cameras and occlusions by cropping each frame with a small view window that oscillates around the person (see Fig. 4). More details are provided in Appendix A.

Evaluation Metrics. We use the following metrics for evaluation: (1) G-MPJPE and G-PVE, which extend the mean per joint position error (MPJPE) and per-vertex error (PVE) by computing the errors in the global coordinates. As errors in estimated global trajectories accumulate over time in our dynamic camera setting, we follow standard evaluations for open-loop reconstruction (e.g., SLAM and inertial odometry ) to compute errors using a sliding window (10 seconds) and align the root translation and rotation with the GT at the start of the window. (2) PA-MPJPE, which is the Procrustes-aligned MPJPE for evaluating estimated body poses. For invisible poses, since there can be many plausible poses beside the GT, we follow prior work to compute the best PA-MPJPE out of multiple samples for our probabilistic approach. (3) Accel, which computes the mean acceleration error of each joint and is commonly used to measure the jitter in estimated motions . (4) FID, which is an extension of the original Frechet Inception Distance that calculates the distribution distance between estimated motions and the GT. FID is a standard metric in motion generation literature to evaluate the quality of generated motions . Following prior work , we compute FID using the well-designed kinetic motion feature extractor in the fairmotion library .

Implementation Details. Thorough details about the entire framework are provided in Appendix A to E.

Baselines. Since no prior methods can estimate global motions from dynamic cameras and address long-term occlusions, we design various baselines by combining state-of-the-art human mesh recovery methods (KAMA or SPEC ), motion infilling methods, and SLAM-based camera estimation (OpenSfM ). In particular, we use the estimated camera parameters to convert estimated motions from the camera coordinates to the global coordinates. For motion infilling, we use (1) linear interpolation, (2) last pose, i.e., replicating the last visible pose, and (3) a state-of-the-art CNN-based motion infilling method, ConvAE .

The results on Dynamic Human3.6M and 3DPW are summarized in Table 1 and 2 respectively. We only report G-MPJPE and G-PVE on Dynamic Human3.6M since they require accurate GT trajectories, which 3DPW does not provide. It is evident that our approach, GLAMR, outperforms the baselines in almost all metrics. In particular, GLAMR achieves significantly lower G-MPJPE and G-PVE, which demonstrates its strong ability to reconstruct global human motions. Furthermore, GLAMR attains considerably lower FID and PA-MPJPE (with ten samples) for occluded (invisible) poses. The lower FID means GLAMR can infill more humanlike motions, and the lower PA-MPJPE also shows GLAMR’s probabilistic motion samples can cover the GT better. Finally, while GLAMR achieves almost the same PA-MPJPE for visible poses as the best method, it yields much smoother motions (smaller acceleration error). This is because our motion infiller leverages human dynamics learned from a large motion dataset to produce motions.

Qualitative Results. Fig. 3 and 4 show qualitative comparisons of GLAMR against the strong baseline, KAMA + Linear Interpolation. Additionally, we provide abundant qualitative results on the project page.

2 Evaluation of Key Components

Benchmarking Motion Infiller. We evaluate the proposed generative motion infiller on the test split of the AMASS dataset . We compare against three motion infilling baselines: linear interpolation, replicating the last pose, and ConvAE . As shown in Table 3, our generative motion infiller achieves significantly better PA-MPJPE for both the sampled motions (with five samples) and reconstructed motion for the infilled frames. Our approach also achieves considerably better FID, reducing the FID of ConvAE by half, which indicates that the infilled motions by our approach are much closer to real human motions.

Benchmarking Trajectory Predictor. We also evaluate our global trajectory predictor against two variants on the AMASS test set: (1) “Transformer”, which replaces the LSTMs in the trajectory predictor with Transformers; (2) “Ours w/o Ego Trajectory”, which does not use the egocentric trajectory but instead directly outputs the 6-DoF global trajectory. As shown in Table 4, both variants lead to worse global trajectory prediction (higher best-of-five G-MPJPE and G-PVE). We believe the reasons are: (1) the positional encoding in Transformers may not generalize well to longer motions compared to the LSTMs in our approach; (2) directly predicting the 6-DoF global trajectory offsets instead of egocentric trajectories from local body motions is also hard to generalize since the global offsets can be large.

Ablations for Global Optimization. We further perform ablation studies on the effect of key components in our global optimization. Specifically, we design two variants: (1) “Ours w/o Trajectory Predictor”, which does not use our trajectory predictor to generate the global human trajectories and uses camera parameters from OpenSfM to obtain global trajectories instead; (2) “Ours w/o Opt Ego Trajectory”, which does not employ the egocentric trajectory representation and directly optimizes the 6-DoF root trajectory instead. As shown in Table 5, both variants lead to significantly worse global trajectory reconstruction with large increases in G-MPJPE, G-PVE, and Accel. This demonstrates that both the global trajectory predictor and egocentric trajectory representation are vital in our approach.

Discussion and Limitations

In this paper, we proposed an approach for 3D human mesh recovery in consistent global coordinates from videos captured by dynamic cameras. We first proposed a novel Transformer-based generative motion infiller to address severe occlusions that often come with dynamic cameras. To resolve ambiguity in the joint reconstruction of global human motions and camera poses, we proposed a new solution by predicting global human trajectories from local body motions. Finally, we proposed a global optimization framework to refine the predicted trajectories, which serve as anchors for camera optimization. Our method achieves SOTA results on challenging datasets and marks a significant step towards global human mesh recovery in the wild.

As the first paper on this new problem, our method has a few limitations: propagation of errors in multiple stages, limited body shape estimation, not being real-time, not including scene information, etc. A detailed discussion is provided in Appendix H. We believe these limitations are exciting avenues for future work to explore.

References

Appendix A Details for the Datasets

AMASS is a large human motion database with 11000+ human motions. We use AMASS to train and evaluate the motion infiller and trajectory predictor. Specifically, we use the Transitions, SSM, and HumanEva subsets for testing and all other subsets for training.

3DPW is an in-the-wild human motion dataset that consists of 60 videos recorded with dynamic cameras in diverse environments. The GT 3D poses are obtained using wearable IMU sensors. Since non-optical sensors are used to obtain GT data, the dataset also provides body pose information when the persons go outside the FoV of the camera. 3DPW also provides the global trajectories of people in the dataset. However, the global trajectories are quite inaccurate since they are estimated from IMU data. Therefore, we do not use 3DPW to evaluate global trajectory reconstruction in the paper. Since we do not use 3DPW for training, we use sequences from the entire 3DPW dataset for visualization. We use the official 3DPW test split to report quantitative results in the paper.

Dynamic Human3.6M is a new benchmark for global human pose estimation with dynamic cameras that we create from the Human3.6M dataset . We simulate dynamic cameras and occlusions by cropping each frame with a view window of 300×600300\times 600 that horizontally oscillates around the person’s bounding box center with a period of 4.8 seconds and a magnitude of 200 pixels. In this way, we synthesize large camera motions and severe occlusions where the person is occluded for almost half of the time, which makes it very challenging for existing 3D human pose and shape estimation methods. Additionally, since Human3.6M provides accurate global human trajectories and human poses, we use Dynamic Human3.6M to evaluate global trajectory reconstruction and pose estimation for occluded frames. We follow the standard protocol and use the official test split (subjects 9 and 11) for evaluation. Please refer to the [supplementary video](https://youtu.be/wpObDXcYueo) for an example sequence of the Dynamic Human3.6M dataset. Code for generating Dynamic Human3.6M are available here for users who have downloaded the original Human3.6M dataset .

Appendix B Implementation Details for Preprocessing

3D Multi-Object Tracking and Re-identification. We use DeepSORT with ResNet-50 in the MMTracking package for 3D multi-object tracking (MOT) and re-identification. We use the GT tracks to evaluate our approach and the baselines, following the standard protocol for human pose estimation.

Initial Human Pose and Shape Estimation. As mentioned in the main paper, we use KAMA or SPEC to provide the initial human pose and shape estimation from the bounding boxes extracted by 3D MOT. We choose these two methods since both KAMA and SPEC estimate 3D human poses in the camera coordinates with absolute root translations, while many state-of-the-art human pose estimation methods do not provide the root translations. We also use HRNet to extract 2D human keypoints from the video, which are used in the proposed global optimization framework.

Appendix C Implementation Details for Generative Motion Infiller

Hyperparameters and Training. The dimension of the latent code z\boldsymbol{z} is 128. The sliding window size hh of the autoregressive motion infilling is 50. Both the number of context frames hch_{\texttt{c}} and the number of look-ahead hlh_{\texttt{l}} frames are 10. When synthesizing occluded motions, for any GT training motion of h=50h=50 frames, we randomly occlude HoccH_{\texttt{occ}} consecutive frames of motion where HoccH_{\texttt{occ}} is uniformly sampled from $.Notethatwedonotoccludethefirst. Note that we do not occlude the firsth_{\texttt{c}}=10$ frames which are reserved as context. The KL divergence term in Eq. (2) uses a weighting factor of 0.001. We train the networks for 2000 epochs with a batch size of 1024 where each epoch uses a total of 10 million frames of motion. For optimization, we use the Adam optimizer with a learning rate of 0.001 and clip the gradient if its norm is larger than 5. We use PyTorch to implement and train the networks.

Appendix D Implementation Details for Global Trajectory Predictor

Heading Coordinate and Egocentric Trajectory Representation. The heading vector of a person points towards where the person is facing and is parallel to the ground. We obtain the heading vector by aligning the zz-axis of the person’s root coordinate with the world zz-axis and use the resulting yy-axis of the aligned root coordinate as the heading vector. This way of obtaining the heading is more stable than using the yaw of the Euler angle representation, which suffers from singularities and can be quite unstable. The heading coordinate is defined by first placing the world coordinate at the root position of the person and then rotating the world coordinate around the zz-axis (vertical) to align the yy-axis with the heading vector. By definition, representing and predicting human trajectories in the heading coordinate allows the predicted trajectory to be invariant of the person’s absolute xyxy translation and heading. In the egocentric trajectory representation ψt=(δxt,δyt,zt,δϕt,ηt)\boldsymbol{{\psi}}_{t}=(\delta x_{t},\delta y_{t},z_{t},\delta\phi_{t},\boldsymbol{\eta}_{t}), we use absolute height ztz_{t} since the height of a person relative to the ground does not vary a lot and is highly correlated with the body motion of the person. For the local rotation ηt\boldsymbol{\eta}_{t}, we adopt the 6D rotation representation to avoid discontinuity.

Network Architecture. The detailed network architecture of the CVAE-based global trajectory predictor is illustrated in Fig. 6. We use two bidirectional LSTM layers with hidden dimension 256 for all the LSTM blocks in the networks. We use two hidden layers (512, 256) with ReLU activations for all the token-wise MLPs. For the input poses, we first convert them to 3D joint positions using the SMPL joint function without global rotations and translations. This is because we find that using 3D joint positions leads to better performance than using joint rotations directly. In both the prior and posterior networks, token-wise mean pooling is used to produce a single feature from a sequence of tokens, which is then used to produce the parameters of the prior or posterior distribution of the latent code v\boldsymbol{v}.

Hyperparameters and Training. The dimension of the latent code v\boldsymbol{v} is 128. The KL divergence term in Eq. (8) uses a weighting factor of 0.001. We train the networks for 2000 epochs with a batch size of 256 where each epoch uses a total of 2 million frames of motion. The training sequence length is 100 frames For optimization, we use the Adam optimizer with a learning rate of 0.0001 and clip the gradient if its norm is larger than 5. We use PyTorch to implement and train the networks.

Appendix E Implementation Details for Global Optimization

Initialization. We initialize the egocentric trajectories using the output from the global trajectory predictor. For the camera, we approximate the camera intrinsic parameters K\boldsymbol{K} using the dimensions of the image [w,h][\texttt{w},\texttt{h}] where we assume the principal point is at the image center [w/2,h/2][\texttt{w}/2,\texttt{h}/2]. Note that the camera intrinsics are kept fixed during the optimization process. For the camera extrinsic parameters C\boldsymbol{C}, we initialize them from the persons’ global trajectories using the following equations:

Hyperparameters and Optimization. The optimization loss coefficients (λ2D,λtraj,λreg,λcam,λpen)(\lambda_{\texttt{2D}},\lambda_{\texttt{traj}},\lambda_{\texttt{reg}},\lambda_{\texttt{cam}},\lambda_{\texttt{pen}}) in Eq. (9) are set to (1, 100000, 100, 10000, 100000) for 3DPW and (1, 100000, 100, 10000, 0) for Human3.6M. We do not use the inter-person penetration loss for Human3.6M since it only has one person in each video. The weighting factor wtw_{t} for the translation term in Eq. (12) is set to 0 since the translation estimated by the pose estimator can be quite noisy. The trajectory regularization weighting factor wψ\boldsymbol{w}_{\psi} in Eq. (13) is set to (3,10,10000,5,10000) for each element in the egocentric trajectory ψt=(δxt,δyt,zt,δϕt,ηt)\boldsymbol{{\psi}}_{t}=(\delta x_{t},\delta y_{t},z_{t},\delta\phi_{t},\boldsymbol{\eta}_{t}), where we use large weights to penalize changes in height ztz_{t} and local rotation ηt\boldsymbol{\eta}_{t}. The global optimization is also implemented in PyTorch , where we use the Adam optimizer with a learning rate of 0.001 to optimize the global trajectories and camera extrinsics.

Computation Time. The overall processing time for a 1-min scene is around 5 mins with 500 optimization iterations, which is much faster than using OpenSfM (>30>30 mins).

Appendix F Evaluation of Global Optimization on 3DPW

We also perform experiments on 3DPW with and without our global optimization framework to study the importance of global optimization when there are multiple people in the video. Although 3DPW does not provide accurate GT human trajectories in the global coordinates, the relative translations and rotations between people in 3DPW are quite accurate. Therefore, we compute the relative translations and rotations between pairs of humans and calculate their errors w.r.t. the ground truth. These metrics, i.e., relative translation and rotation errors, serve as an alternative way to evaluate global reconstruction quality. As shown in Table 6, using global optimization can greatly reduce the relative translation and rotation errors between humans, which means our global optimization framework can greatly help to reconstruct the spatial relationships of humans in the video.

Appendix G Effect of Sliding Window Length.

As shown in Fig. 7, when increasing the window length hh (with context hch_{\texttt{c}} and look-ahead hlh_{\texttt{l}} being 0.2h0.2h), the reconstruction error increases because it is harder for the latent code z\boldsymbol{z} to encode a longer window which contains more motion variations than a shorter window. In the meantime, the sample error first drops and then increases since there is a trade-off: a longer window provides more context for better inference, but it also puts more burden on the latent code as indicated by the increasing reconstruction error.

Motion Infilling without Visible Pose. In the extreme case, when there is no visible pose (hc=hl=0h_{\texttt{c}}=h_{\texttt{l}}=0), our motion infiller can still produce plausible motions sampled from the prior learned from the training motion datasets. In this case, the motion infiller essentially becomes an unconditional VAE model.

Appendix H Discussion of Limitations

As the first paper on this new problem, our method has a few limitations that are important for future research to address. First, our approach has five stages that are sequentially dependent. Therefore, errors in early stages can propagate to late stages, which may lead to inaccurate global pose estimation. Future work could integrate these stages together to form an end-to-end learnable framework. Second, like many works in human mesh recovery, our approach can only recover the SMPL parameters which omit the fine details of human meshes such as clothing. Integrating neural articulated shapes such as into our approach could potentially address this problem. Third, our approach is not real-time due to the batch processing and global optimization. Future work could explore a causal version of our approach where only a small window around the incoming frame is optimized, which could substantially improve computational efficiency. Finally, the generative motion infiller and global trajectory predictor in our approach operate for each person independently. Therefore, the generated motions and trajectories may not capture potentially complex and nuanced interactions between occluded people such as hugging or dancing. Future work could address this limitation by employing new generative models that produce interaction-aware motions of multiple people.

Appendix I Discussion of Potential Negative Impact

With its strong ability to reconstruct global human motions and tackle severe occlusions, our method marks a significant step towards global human mesh recovery in the wild. However, misuse of this technology could lead to potential privacy concerns and the propagation of misinformation. For instance, combined with advanced neural rendering approaches , the reconstructed global human motion of our approach could be used to fabricate videos of human actions that are indistinguishable from real ones. To address this issue, future research should continue to study the detection of synthesized videos with realistic human motion.