RoHM: Robust Human Motion Reconstruction via Diffusion
Siwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler, Petr Kadlecek, Siyu Tang, Federica Bogo
Introduction
In this paper, we tackle the problem of 3D human motion reconstruction from monocular RGB(-D) videos in real scenarios – i.e., in the presence of noise and occlusions. Reconstructing 3D human motion is crucial for many applications, ranging from augmented and virtual reality to robotics. Many methods in the literature tackle the problem by training deep neural networks to directly regress 3D body pose and shape from monocular input . However, these approaches commonly suffer from two major shortcomings: 1) they estimate only local motion, representing pose in body-root relative coordinates without plausible global motion, in world coordinates consistent over time; 2) they lack robustness when the body undergoes occlusions, in either the spatial or temporal dimension.
In such scenarios, optimization-based methods have shown better performance. For instance, HuMoR and PhaseMP explicitly address scenarios with noisy and incomplete input by combining data-driven motion priors with application-agnostic optimizers such as L-BFGS . Still, these methods may fail under heavy occlusions (Fig. 1). Furthermore, test-time optimization is time-consuming, prone to local minima, and requires significant manual tuning .
To overcome these limitations, we propose to leverage the iterative, generative nature of diffusion models. Diffusion models were initially proposed for 2D generation tasks ; recently, they achieved compelling results in 3D human motion generation from input such as text and action labels , music , and sparse (noise-free) keypoints . In particular, these models proved effective in learning and modelling long-term motion dependencies over time , which go beyond the single-frame conditioning considered in . Furthermore, their iterative denoising process poses them as a data-driven alternative to the iterative minimization of optimization-based techniques. However, so far diffusion-based approaches have mostly focused on synthesizing motion from user input, rather than reconstructing motion from monocular videos exhibiting noise and occlusions – where reconstructions need to match image evidence whenever available. Here, we explore how to leverage diffusion models in such reconstruction scenarios.
Given a monocular video, we obtain initial, per-frame 3D pose estimates using off-the-shelf regressors and/or per-frame optimization (akin to ). These estimates are inaccurate and incomplete, with implausible motions. From them, our goal is to reconstruct a smooth and complete 3D motion. This is a complex task, requiring us to address different problems (motion denoising and infilling) in different solution spaces (local and global motion). We observe that previously proposed diffusion-based motion models do not work well in this scenario: they expect noise-free, user-provided keypoints as input, model global and local motion together, and cannot ensure alignment between reconstruction and image evidence. Inspired by , we therefore propose to decompose the problem. We leverage motions from the AMASS dataset to learn two diffusion models conditioned on noisy input, one for global trajectory reconstruction and one for local motion prediction– showing the benefits of decoupling the two spaces. Still, this formulation ignores the correlations between global and local motion space. While addresses this by predicting trajectory based only on infilled local motion, in our scenario we require the estimated trajectory to remain faithful to the input. Drawing inspiration from , we therefore propose a flexible conditioning strategy for trajectory reconstruction, exploiting both input data and denoised local motion. We show that this module, combined with an iterative scheme at inference time, improves both local and global motion quality. Finally, to further encourage physically plausible motions that match image evidence, we guide sampling at test time with physics-based and image-based scores.
In summary, our contributions are: 1) RoHM, a novel diffusion-based approach for Robust Human Motion reconstruction in consistent global coordinates from monocular sequences with noise and occlusions; 2) a flexible conditioning strategy to capture inter-dependencies between root trajectory and local pose; 3) various applications enabled by RoHM, including motion reconstruction, denoising, spatial and temporal infilling. Extensive experiments on three widely used datasets show that RoHM achieves superior accuracy and realism compared to state-of-the-art optimization-based methods, while being 30 times faster at inference time. Code and models will be released for research purposes.
Related Work
Regression-based methods. Many approaches in the literature focus on 3D human shape and pose reconstruction from a single image , recently also considering robustness to occlusions . Dealing with occlusions is even more challenging for monocular human motion reconstruction, requiring one to model plausible dynamics over time for non-visible body parts. Many approaches train neural networks to predict 3D motion from RGB videos and are not easily adaptable to different input modalities . Some methods introduce adversarial priors , leverage multi-view cues or learn denoising models to achieve robustness. Still, most of them estimate only local motion in the camera frame without recovering global trajectory, thus suffering from jitter and motion artifacts. estimate plausible global trajectories from per-frame local features, but are not robust to occlusions . In contrast, our method robustly recovers both global and local motion and can be applied to different inputs.
Optimization-based methods. These methods typically fit a parametric body model to observations (such as body keypoints, depth, silhouettes etc.) by iteratively minimizing an objective function . To regularize motion, such function contains one or more temporal priors. Some approaches define hand-crafted priors encouraging motion smoothness (e.g., on body joint velocity or acceleration) , or constrain motion in a low-dimensional space ; however, they may produce over-smooth motions and foot skating . Lately, the availability of large-scale motion capture datasets such as AMASS made it possible to learn powerful motion models from data . LEMO learns two separate, fully deterministic priors for motion smoothness and infilling. HuMor and PhaseMP propose generative priors, modelling the distribution of state transitions between frames via Conditional Variational Autoencoders (CVAEs) . Learning transitions only between pairs of frames, these methods struggle when occlusions span long time intervals (see Sec. 5). Optimization-based methods tend to match input data more closely than regression methods, thanks to their iterative minimization; but they are in general slower, prone to local minima, and require significant manual tuning . In contrast, we propose to leverage the iterative nature of diffusion models.
Human motion models. A variety of approaches has been proposed for motion tracking and synthesis, including mixtures-of-Gaussians , Gaussian processes , pose embeddings , VAEs , and normalizing flows . These methods may not generalize well to unseen motions and body-scene interactions. Physics-based methods address these challenges by enforcing physics laws via physics simulation. However, simulators are computationally expensive, non-differentiable, and may introduce discrepancies between predicted motions and observed input.
Diffusion models for human motion. Diffusion models have demonstrated compelling results for human motion synthesis conditioned on input such as text and action labels , music , environment , and (noise-free) 3D joints . Given their ability to model long-range spatio-temporal relationships, they have been applied to motion forecasting and infilling . use diffusion to synthesize lower body motion given head and wrist positions and rotations. These methods do not focus on scenarios with noisy input or body occlusion varying over time. Recently, diffusion has been leveraged for 3D body pose estimation , even under severe body truncations , but without considering the temporal dimension. In contrast, we leverage diffusion models to achieve robustness against varying ranges of occlusion and noise over long temporal sequences.
Preliminaries
SMPL-X is a parametric model that represents the human body as a function , parameterized by global translation , global orientation , body pose , body shape , hand pose and facial expression . The function returns a triangle mesh with 10,475 vertices. In the following we do not use and and consider only the main body parameters, with a total of body joints. We use SMPL-X instead of SMPL to leverage the gender-neutral annotations from AMASS .
Conditional Diffusion Models. We adopt the Denoising Diffusion Probabilistic Models (DDPMs) formulation . At the core of it are a forward diffusion process and an inverse, iterative denoising process. Let be the real motion distribution. The forward diffusion process is a Markov chain adding i.i.d. Gaussian noise at each step :
where defines the variance at each step according to a pre-defined schedule and is the identity matrix. shows that can be directly sampled from :
Starting from Gaussian noise , the inverse diffusion process reconstructs by iteratively denoising over steps. In practice, we train a denoiser neural network to remove the added Gaussian noise based on condition signal and step : (cf. ). Leveraging Eq. 2, the denoiser can be trained by sampling noise step and optimizing the simple objective :
Note that conditioning is crucial in our setup, where we want to exploit input observations whenever available.
Method
where the dot indicates the first derivative (velocity). For rotations, we adopt the 6D representation from . For , we select two joints per foot and set the corresponding contact label as 1 if the joint is in contact with the ground, else 0 . For each frame , we define a local coordinate system such that local joint positions are relative to the current frame pelvis joint, projected on the ground . This over-parameterized representation allows explicit modelling of both 3D skeleton joints and SMPL-X meshes, such that they can benefit the one from the other. Together with and , we define root joint and local joint visibility masks and , respectively (1 denotes visible, 0 otherwise). They are obtained by randomly masking joints at training time and computing joint visibility at test time (see Sec. 4.4 and Supp. Mat.).
We tackle the problem of denoising and infilling global and local motion by using two networks, TrajNet and PoseNet.
denotes root trajectory at each diffusion denoising step and denotes element-wise matrix multiplication. The trajectory representation for TrajNet is parameterized as , excluding first derivatives to avoid global drifting caused by inaccurate velocities.
Architectures. TrajNet adopts a U-Net encoder-decoder structure built upon . An extra conditioning encoder maps the conditional trajectory signal into multi-layer features, which are concatenated with U-Net encoder features at each intermediate layer. PoseNet employs a transformer encoder structure akin to . The conditioning signal is encoded via an MLP (Multilayer Perceptron) and fed into the transformer. See Supp. Mat. for more details.
2 Controlling Global Motion Reconstruction
Inspired by , we introduce TrajControl, an auxiliary module for fine-tuning TrajNet with additional control signal from local body pose (Fig. 2 right). Specifically, we freeze the parameters of our pre-trained TrajNet and clone the encoder layers to a trainable duplicate, , which serves as the TrajControl module. Iteratively, we feed the output of PoseNet into , thus improving TrajNet output based on denoised, complete local motion. In turn, the improved TrajNet output can be leveraged to further refine local body motion. We detail this iterative scheme in Sec. 4.3.
Architecture. is connected to the frozen pre-trained TrajNet with zero convolution layers (1D convolution layer with kernel size 1, initialized from zero bias and zero weight), to ensure a smooth start for the fine-tuning. During fine-tuning, we update only the weights of . In this way, the fine-tuned TrajNet can still be conditioned on noisy trajectory only, when clean local motion is not available.
3 Inference
For any subsequent iteration , TrajNet is conditioned on estimated trajectory and local pose , output of PoseNet at iteration , as the additional control signal; PoseNet is conditioned on current TrajNet output and local pose predicted at iteration :
Algorithm 1 summarizes the approach. Note that in each ‘iteration’ (), we sample from TrajNet and PoseNet running all diffusion denoising steps. is set to 100 for TrajNet, and 1000 for PoseNet. Empirically, we find that iterations are sufficient to obtain accurate results. While one could still run this iterative inference between TrajNet and PoseNet without using TrajControl, we show in Sec. 5.5 that this leads to degraded results.
Score-guided sampling. Besides embedding condition signals in the decoder architecture, diffusion models enable also test-time conditioning via classifier-based guidance, which has been already leveraged for image generation , trajectory prediction , motion generation , and human mesh recovery . In a similar spirit, we enhance physical plausibility and accuracy of reconstructed motions by guiding PoseNet sampling with two scores, penalizing foot skating and enforcing 2D joint reprojection consistency:
with as a linear combination of and . The guidance is modulated by , the variance of a pre-scheduled Gaussian distribution , and by the scaling weights and .
4 Training
We train our diffusion denoisers and using in Eq. 3, plus losses enforcing consistency with the ground truth in terms of 3D joint position () and 3D joint velocity (), and penalizing foot skating ():
where is the ground-truth motion; \color[rgb]{0.43,0.71,0.88}\boldsymbol{R} refers to ground-truth root trajectory for PoseNet, and predicted root trajectory for TrajNet. denotes ground-truth foot contact labels, and denotes the predicted foot joint velocities. The overall loss is defined as:
PoseNet and TrajNet are trained separately on the AMASS dataset , with each sequence trimmed into short clips of frames. For both training the “vanilla” TrajNet and fine-tuning TrajNet with TrajControl, we exclude , and only compute and for the root joint. We utilize ground-truth local pose as the control input of TrajControl to fine-tune TrajNet.
Experiments
AMASS is a large-scale motion capture dataset collecting high-quality 3D human pose and shape annotations. We use the official SMPL-X neutral body annotations for training and evaluation. We downsample each sequence to 30fps. As in , we use TCD_handMocap, TotalCapture, and SFU for testing and the remaining subsets for training.
PROX collects monocular RGB-D videos of people interacting with various 3D indoor scenes. Since the dataset does not provide ground-truth annotations, we use a subset of sequences to evaluate physical plausibility as in .
EgoBody collects sequences of people interacting with each other in various 3D indoor environments, capturing multi-modal input streams with both head-mounted (first-person) and external (third-person) cameras. The dataset provides ground-truth SMPL/SMPL-X annotations. We manually select a set of third-person RGB sequences (around 24k frames) exhibiting severe human-scene occlusions, and use them for evaluation.
2 Evaluation Metrics
Accuracy. We adopt the Mean Per-Joint Position Error in to evaluate predicted body joint accuracy in the pelvis-aligned coordinate system (MPJPE) and in the global coordinate system (GMPJPE). We report numbers for full-body (all), visible (vis) and occluded (occ) body joints separately, considering the 22 SMPL-X main body joints. Furthermore, we measure foot-ground contact binary classification accuracy (Contact) for the 4 foot joints as in .
Physical Plausibility. We report additional metrics to assess motion and scene-interaction plausibility. When ground-truth motion is available (AMASS and EgoBody), we report the acceleration error (Accel) as the difference in acceleration between predicted and ground-truth 3D joints ; for PROX, we report the norm of mean per-joint acceleration (Accel). Both metrics are in . Foot skating ratio (Skating) measures how often feet slide on the floor. We define sliding as happening when the velocity of all foot joints exceeds 10cm/s, toe joints’ height above the ground is lower than 10cm, and ankle joints’ height is lower than 15cm. The mean ground penetration distance (Dist) measures to what extent left and right toe joints penetrate into the ground, measured in .
3 Motion from 3D Observations
Experimental setup. To evaluate RoHM’s robustness to noisy and occluded data, we run two sets of experiments on AMASS: (1) motion denoising + infilling (Occ-L.), and (2) motion denoising + in-betweening (Occ-10%). Given a SMPL-X motion sequence from our AMASS test set, in (1) we mask out all lower body joint parameters (both positions and rotations), simulating scenarios commonly occurring when people move in a 3D scene; in (2), we mask out an entire subsequence of frames ( of the original sequence), thus requiring the model to generate the in-between motion. In both setups, we add Gaussian noise to SMPL-X pose and translation parameters and use the resulting noisy and occluded 3D motion data as input for our model. We consider different, increasing noise levels: noise level corresponds to Gaussian noise with standard deviation of for . Note that the noise, defined on SMPL-X parameters, will accumulate along the kinematic tree.
Baselines. We compare RoHM with VPoser-t, HuMoR , and an adapted version of MDM (‘MDM++’). As in , VPoser-t is an optimization-based method using VPoser and 3D joint smoothing . It serves as the initialization stage for HuMoR. We cannot directly apply MDM/PriorMDM to our setup, since they require noise-free visible joints for infilling and do not support explicit conditioning on noisy observations – resulting in both pose and trajectory drifting. We therefore train an adapted and improved version, which allows conditioning on the initial corrupted motion, using the same data augmentation as RoHM (see Supp. Mat. for details).
Results. Tab. 1 reports results on the AMASS test set. Our approach demonstrates significantly superior performance, in both accuracy and physical plausibility. Specifically, when compared to HuMoR, we achieve 48% reduction in GMPJPE for occluded body parts in both Occ-L. and Occ-10% setups. The reduced acceleration errors suggest RoHM can recover more realistic motion dynamics. This facilitates also accurate foot contact label predictions, leading to 44% improvement in foot skating over HuMoR and fewer foot-ground inter-penetrations (Fig. 4, row 3). MDM++ performs similarly to us in foot-ground collisions, but reconstruction accuracy and other plausibility metrics are noticeably inferior in comparison. In scenarios with larger noise (e.g., level 7), baselines struggle to recover plausible lower body motions, often leading to legs floating in the air (see Fig. 4, row 1-2). This is particularly evident for VPoser-t, which therefore exhibits a low skating ratio, as the skating score only considers frames with foot-ground contact. Fig. 3 (left, middle) compares robustness to noise of our method and HuMoR. Increasing input noise levels correspond to a substantial decline in performance for HuMoR, while RoHM shows more robustness. Note that our method is only trained with a noise level of 2.
4 Motion from RGB(-D) Videos
We compare the performance of RoHM against baselines on motion reconstruction from RGB/RGB-D videos on PROX, and from RGB videos on EgoBody.
Baselines. We compare our method against (1) a per-frame human mesh regressor, CLIFF (RGB only) and (2) four optimization-based methods leveraging motion priors: VPoser-t , HuMoR , LEMO (RGB-D only) and PhaseMP (RGB only)Since PhaseMP code is unavailable, we only compare with it in the PROX-RGB setup using the results kindly provided by PhaseMP authors.. Note that methods of type (2) are currently the ones reporting the best results for robust monocular motion reconstruction. For reference, we also include as a baseline our initialization stage (‘Ours-init’).
Results. Tab. 2 reports physical plausibility results obtained on PROX. CLIFF and Ours-init (RGB/RGB-D) are per-frame methods, producing noticeable motion jitter and foot skating. VPoser-t simply enforces 3D joint smoothness and struggles to recover realistic motion dynamics. LEMO tackles noise and occlusions separately, generalizing less well to such complex scenarios. HuMoR and PhaseMP model motion transitions between frames but are less effective in capturing longer-range temporal correlations – producing implausible results under heavy occlusions. In contrast, RoHM reconstructs smooth motions with improved foot-ground interactions (Skating and Dist), and realistic motions for occluded body parts, as shown in Fig. 5. Notably, our method starts from a much more challenging initialization (Ours-init) compared to HuMoR and PhaseMP (VPoser-t), with more severe jitterings and foot skatings, as can be observed by comparing rows 2 and 6 in Tab. 2.
Factoring out the impact of the initialization stage, Tab. 3 presents quantitative results on EgoBody. We start from the same initialization as HuMoR (VPoser-t) and consistently outperform the baselines across all metrics. Qualitative results are shown in Fig. 5 (right). This indicates that RoHM can generate more plausible motions, even in the highly occluded, challenging scenarios of this dataset.
Efficiency. Our approach exhibits significantly reduced runtime compared to HuMoR, being 30 times faster factoring out the initialization stage (tested with the same settings, see details in Supp. Mat.).
5 Ablation Study
We perform ablation studies on AMASS (Fig. 3 right, with respect to different noise levels) and PROX (Tab. 4). Our iterative inference scheme leveraging TrajControl effectively alleviates foot skating by closing the gap between PoseNet and TrajNet, particularly in the presence of large noise (see Fig. 3 right, and ‘w/o TC’ versus ‘Ours’ in Tab. 4). Iterating between TrajNet and PoseNet twice as described in Sec. 4.3 without TrajControl (‘w/o TC itr=2’ in Tab. 4) improves motion plausibility to some extent but is still sub-optimal. Test-time score guidance further improves result-observation alignment (see Fig. 6) and alleviates foot skating for all setups. As expected, test-time guidance slightly impacts motion smoothness – an aspect compensated by the iterative inference scheme.
Conclusion
We proposed RoHM, an approach for robust human motion reconstruction. Differently from previous work relying on test-time optimization , the approach learns how to reconstruct motion from data using diffusion models. The approach decouples the problem of recovering global and local motion by learning two models and conditioning them on available image evidence; a flexible control module captures correlations between global and local dynamics, leveraged by an iterative inference scheme to refine motion plausibility. Experiments on three publicly available datasets show that the approach can reconstruct more realistic and accurate motions than state-of-the-art baselines, especially in challenging scenarios exhibiting noise and occlusions.
Limitations and Future Work. In its current formulation, RoHM does not work online at real-time framerates. In the future, we plan to evaluate accuracy-efficiency tradeoffs using different architectures (e.g., ). Moreover, RoHM does not consider 3D environment constraints to model interactions between body and 3D scene geometry; adapting RoHM to further incorporate scene conditioning is an exciting avenue for future research. Finally, while here we focused on full-body reconstruction, future work should extend RoHM to model also facial expressions and articulated hand poses over time.
Acknowledgements. Siyu Tang is partially funded by the SNSF project grant 200021 204840. We sincerely thank Korrawe Karunratanakul, Marko Mihajlovic, Tony Tung, Yuting Ye, Artsiom Sanakoyeu, and Yuhua Chen for insightful discussions.
References
Appendix A Architecture Details
The detailed architecture of our model is illustrated in Fig. S1.
TrajControl models pose-trajectory correlations and further refines root trajectory (Sec. 4.2 in the main paper), based on the denoised and infilled local body pose from the previous inference iteration . Namely, upon completing the training of the vanilla TrajNet, the U-Net encoder, along with its weights, is duplicated to serve as the TrajControl encoder, to encode pose information. The intermediate pose features are added to the U-Net decoder via zero convolution layers (1x1 convolution with weights and bias initialized from zero). The TrajControl module is fine-tuned while keeping other TrajNet components frozen. This ensures that the vanilla TrajNet can process input even when only a corrupted trajectory is provided.
Appendix B Implementation Details
TrajNet undergoes training in four stages. In the initial two stages, the model is trained with noise level and , respectively, without occlusions. In the third stage, the noise level is raised to , with and 10% of the frames entirely masked out. Upon completion of these stages, the training for the vanilla TrajNet is concluded. In the last stage, TrajControl is fine-tuned to incorporate additional control from body pose, with a noise level of and no occlusion masks.
PoseNet follows a two-stage training process. In the first stage, the model is trained with a noise level of . To synthesize occlusion masks, we randomly mask out 1-6 joints in the initial 500 epochs. Afterwards, a mixed occlusion scheme is applied: with 0.5 probability, occlusion masks from PROX pseudo ground truth are used; with 0.3 probability, all lower body parts are masked out; with 0.2 probability, all upper body parts are masked out; with 0.1 probability, the full body is masked out in 30% of the frames. In the second stage, we continue the mixed occlusion scheme and increase the noise level to .
Training weights. For both PoseNet and TrajNet, weights and are set to 100 and 1000, respectively. is set to 0 during the first training stage, and 0.1 during the second training stage in PoseNet.
B.2 Motion Initialization
For RGB-D sequences on PROX, we additionally perform a per-frame optimization step to incorporate depth observations. More precisely, for each frame, we optimize the SMPL-X body parameters by minimizing the following objective function:
penalizes the 2D joint distances between the optimized 2D SMPL-X body joints projected onto the RGB image, and detections from OpenPose . penalizes the 3D Chamfer distance between the human point cloud obtained from the depth frame and SMPL-X surface points visible from the camera as in . and denote priors that regularize SMPL-X body pose and shape. s denote the corresponding weights. This approach is akin to VPoser-t but excludes the 3D joint smoothness term, working per-frame.
On EgoBody , to conduct a quantitative comparison with the baselines while factoring out the influence of various initialization strategies, we employ VPoser-t for initialization as in HuMoR . Regarding the input OpenPose 2D detections for our method and baseline methods, instead of raw detections, we use a manually post-processed version provided by EgoBody, where the detections for most occluded joints are masked out.
It is worth highlighting that our approach can be combined with various initialization strategies (both optimization- and regression-based), ensuring flexibility for different applications and inputs.
B.3 Inference Details
Occlusion masks for reconstruction from RGB(-D) videos. To obtain joint occlusion masks for inference on PROX and EgoBody, given the initialized 3D body, we identify a body joint as occluded if it fulfills two conditions: (1) the confidence score of the corresponding 2D joint detection is below 0.2; and (2) the depth of the joint is greater than the depth of the scene vertex which is projected on the same 2D pixel in the image plane as the body joint, from the camera view. The depth of the joint is determined by rendering the 3D body mesh obtained from initialization from the camera view.
Score-guided sampling. In Eq. (14) in the main paper, we set to 3e5. is set to 1e5 for experiments on PROX and EgoBody, and to 3e6 for experiments on AMASS. The score-guided sampling is enabled for the last 100 denoising steps for PoseNet. Furthermore, as the modulation variance diminishes towards the end of the diffusion denoising steps, we skip the last 20 denoising steps for PoseNet for experiments on PROX and EgoBody; this ensures stronger gradient guidance for 2D alignment with image observations.
Runtime. To assess the runtime difference between our method and HuMoR , we omit the initialization stage and focus solely on the inference/test-optimization stage for both methods. For RGB-D input, employing an NVIDIA A100 GPU with a batch size of 10, and with a sequence length of 144 frames, our method completes the inference in 59 seconds, while HuMoR requires 30 minutes for the entire test-time optimization. We use the default configurations of the official HuMoR code.
Appendix C Baseline MDM++ Details
For motion infilling and in-between tasks, at each denoising step, MDM and PriorMDM replace denoised joints with visible input joints, when they are available. This assumes clean motion for visible body parts as input, and therefore cannot handle noisy scenarios like the ones we consider. Moreover, we observe that the relative trajectory representation in , which only considers trajectory velocities, results in severe global trajectory drifting and deviation from the input, due to accumulated errors in the estimated trajectory velocities. To address these limitations and enable denoising together infilling and in-between tasks, we adapt the original MDM formulation to obtain MDM++, as explained below.
However, addressing both denoising and infilling tasks in two different spaces (root trajectory and local pose) within one single model remains very challenging. Indeed, MDM++ still exhibits degraded reconstruction accuracy and motion plausibility compared to RoHM, as shown in Tab. 1 in the main paper.
Appendix D Limitations and Failure Cases
We show example failure cases in Fig. S2. As it is common for learning-based approaches, our method can struggle to generalize to out-of-distribution test cases – such as shapes and poses that are rarely seen in the training data. For instance, the first two columns of Fig. S2 show subjects that are relatively tall, and the last two columns show the rare poses.
Another limitation lies in the model’s dependence on both the 3D scene mesh and 2D joint detections to determine if a joint is occluded. This reliance becomes problematic when the 3D scene mesh is unavailable or when 2D joint detections are unreliable. A potential solution could involve learning an occlusion classifier based on the initial 3D body pose and image inputs to identify joint occlusions. We consider this avenue a promising direction for future exploration.