Pose-NDF: Modeling Human Pose Manifolds with Neural Distance Fields

Garvita Tiwari, Dimitrije Antic, Jan Eric Lenssen, Nikolaos Sarafianos, Tony Tung, Gerard Pons-Moll

Introduction

Realistic and accurate human motion capture and generation is essential for understanding human behavior and human interaction in the scene . Human motion capturing systems, like marker-based systems , IMU-based methods , or reconstruction from RGB/RGB-D data , often suffer from artifacts like skating, self-intersections and jitters and produce non-realistic human poses, especially in the presence of noisy data and occlusion. To make the results applicable in fields like 3D scene understanding, human motion generation, or AR/VR applications, it is often required to apply exhaustive manual or automatic cleaning procedures.

In recent years, learned data priors to post-process such non-realistic human poses has become increasingly popular. Prior human pose models mainly focus on learning a joint distribution of individual joints in pose space or recently in a latent space, using VAEs . They have demonstrated to greatly improve the plausibility of poses after model fitting. However, VAE-based methods, such as VPoser or HuMoR make a Gaussian assumption on the space of possible poses, which leads to several limitations: 1) They have the tendency of producing more likely poses that lie near the mean of the computed Gaussian. Those poses however, might not be the correct ones. 2) Distances between individual human poses are not preserved in the VAE latent space. Hence, taking small steps towards the Gaussian mean might result in large steps in pose space. 3) VAEs have been shown to fold a manifold into a Gaussian distribution , exposing dead regions without any data points in the outer parts of the distribution. Thus, they produce non-plausible samples that are far from the input when traversed in outer regions, as we demonstrate in our experiments.

To alleviate these issues, we present \blah, a human pose prior that models the full manifold of plausible poses in high-dimensional pose space. We represent the manifold as a surface, where plausible poses lie on the manifold, hence having a zero distance, and non-plausible poses lie outside of it, having a non-zero distance from the surface. We propose to learn this manifold using a high-dimensional neural field, analogously to representing 3D shapes using neural distance fields . This formulation preserves distances between poses and allows to traverse the pose space along the negative gradient of the distance function, which points to the direction of maximum distance decrease. Using gradient descent in pose space from an initial potentially non-plausible pose, we always find the closest point on the manifold of plausible poses.

An overview of our method is given in Fig. 1. We formulate the problem of learning the pose manifold as a surface completion task in n-dimensional space. In order to learn a pose manifold, there are two key challenges: a) the input space is high-dimensional, and b) the input space is not Euclidean, as it is for 3-dimensional implicit surfaces . Instead, the pose space is given as SO(3)KSO(3)^{K}, in which a single pose can be represented by KK elements of the rotation group SO(3)SO(3), describing the orientations of joints in a human body model. To represent group elements, we opted for a quaternion representation, as they are continuous, have an easy-to-compute distance, and are subject to an efficient gradient descent algorithm. We map a given pose to a distance by applying a hierarchical implicit neural function, which encodes the pose based on the kinematic structure of the human body. We train our model using the AMASS dataset , where each sample from the dataset is treated as a point on the manifold. The learned neural field representation can be used to project any pose onto the manifold, similar to . We leverage this property and use \blah for diverse pose generation, pose interpolation, as a pose prior for 3D pose estimation from images , and motion denoising , improving on state-of-the-art methods in all areas. In summary our contributions are:

A novel high-dimensional neural field representation in SO(3)KSO(3)^{K}, \blah, which represents the manifold of plausible human poses.

improves the state of the art in human body fitting from images by acting as a pose prior. It outperforms other human pose priors, such as VPoser and the human motion prior HuMoR on motion denoising.

Our method is as fast or faster than current state-of-the-art methods, is fully differentiable and the distance from the manifold can be leveraged for finding the optimal step size during optimization.

generates more diverse samples than previous methods with Gaussian assumptions, which are biased towards generating more likely poses.

Related Work

Our method is a human pose prior build as neural field in high-dimensional space. Thus, we review related work in both of these areas.

Pose and Motion Priors. Human pose and motion priors are crucial for preserving the realism of models estimated from captured data and to estimate human pose from images and videos . Further, they can be powerful tools for data generation. Initial work along this direction mainly focused on learning constraints for joint limits in Euler angles or swing and twist representations , to avoid twists and bends beyond certain limits. A next iteration of methods fits a Gaussian Mixture Model (GMM) to a pose dataset and uses the GMM-based prior for downstream tasks like image-based 3D pose estimation or registration of 3D scans . Additionally, simple statistical models, such as PCA, have been proposed . With the rise of deep learning and GANs , adversarial training has been used to bring predicted poses close to real poses and for motion prediction . However these are task specific models, HMR models p(θ∣I)p(\boldsymbol{\theta}|I) and requires an image II. HP-GAN models p(θt∣θt−1)p(\boldsymbol{\theta}_{t}|\boldsymbol{\theta}_{t-1}) and requires pose parameters θt−1\boldsymbol{\theta}_{t-1} for previous frame/time. Therefore they cannot be used as a prior for other tasks.

More recent work uses VAEs to learn pose priors , which can be used for generating pose samples, as prior in pose estimation, or 3d human reconstruction from images or sparse/occluded data. Some works propose VAE-based human motion models. HuMoR proposes to learn a distribution of possible pose transitions in motion sequences using a conditional VAE. ACTOR learns an action conditioned VAE-Transformer prior. Further work designs pose representations along the hierarchy of human skeletons and uses it for character animation . Concurrent work learns a human pose prior using GANs and highlights the shortcomings of Gaussian assumption based models like VPoser . A VPoser decoder is used as generator (mapping z→θz\rightarrow\theta) and an HMR -like discriminator is used to train the model. As described in Sec. 1, our approach follows a different paradigm than the VAE and GAN-based methods, as we directly model the manifold of plausible poses in high-dimensional space, which leads to a distance-preserving representation.

Before the rise of deep learning, modeling partial pose spaces as implicit functions was common, e.g. as fields on a single shoulder joint quaternion or an elbow joint quaternion, conditioned on the shoulder joint . However, those ignore the real part of the quaternion, leading to ambiguities in representation, are not differentiable, and are limited to 2 joints in the human body model. In contrast, our method uses a fully differentiable neural network, which learns an implicit surface in higher dimension, taking all human joints and all four components of each quaternion into account.

Neural Fields. Neural fields for surface modeling have received increasing interest over the recent years. They have been used to model fields in 2D or 3D, representing images or partial differentiable equations , signed or unsigned distances from static 3D shapes , pose-conditioned distance field , radiance fields and more recently for human-object and hand-object interactions. For a more detailed overview of neural fields please refer to . Neural fields have recently been brought to higher dimensions to model surfaces in Euclidean spaces . In this work, we apply the concept to the high-dimensional, non-Euclidean space of SO(3)KSO(3)^{K}, modeling the unsigned distance to manifolds of plausible human body poses in pose space.

Method

such that the value of ff represents the unsigned distance to the manifold, similar to neural fields-based 3D shape learning . Without loss of generality, we use the SMPL body model , resulting in poses θ\boldsymbol{\theta} with K=21K=21 joints.

where the individual elements of summation are a metric on SO(3)SO(3) and wiw_{i} is the weight associated with each joint based on their position in the kinematic structure of the SMPL body model (i.e. early joints in the chain have higher weights). It should be noted that the double cover property of unit quaternions, that is, the quaternions q\mathbf{q} and −q-\mathbf{q} represent the same SO(3)SO(3) element, does not lead to additional challenges. We simply train the network to be point symmetric by applying sign flip augmentation on input quaternions.

2 Hierarchical Implicit Neural Function

Formally, for a given pose θ={θ1,...,θK}\boldsymbol{\theta}=\{\boldsymbol{\theta}_{1},...,\boldsymbol{\theta}_{K}\}, where θk\boldsymbol{\theta}_{k} is the pose for joint kk, and a function τ(k)\tau(k), mapping the index of each joint to its parent joints index, we encode each pose using an MLP as:

3 Loss functions

More details about training data, network architecture is provided in the supplementary material.

4 Projection Algorithm

where d(θ,S)d(\boldsymbol{\theta},\mathcal{S}) is the distance (Eq. 2) of θ\boldsymbol{\theta} to the closest point in S\mathcal{S}. We find θ^\hat{\boldsymbol{\theta}} by applying gradient descent on the 33-sphere, using gradient information ∇θf(θ)\nabla_{\boldsymbol{\theta}}f(\boldsymbol{\theta}) and distances f(θ)f(\boldsymbol{\theta}), obtained from the implicit neural function ff. One step is given as:

followed by a re-projection to the sphere (i.e. vector normalization) after several iterations. This algorithm is guaranteed to converge to local minima on the sphere, which in our case, assuming a correctly learned distance function, is the nearest point on the pose manifold.

Experiments and Results

In this section we evaluate \blah and show the different use cases of our pose model, which include the ability to serve as a prior in denoising motion sequences or recovery from partial observations (Sec. 4.2), prior for recovering plausible poses from images (Sec. 4.3) using an optimization-based method, pose generation (Sec. 4.4) and pose interpolation (Sec. 4.5). We demonstrate that the \blah method outperforms the state-of-the-art VAE-based human pose prior methods. We also show the advantages of our distance field formulation over VAEs or Gaussian assumption models (Sec. 4.6). Before turning to the results, we explain training and implementation details of \blah in Sec. 4.1.

2 Denoising Mocap Data

Human motion capture has been done using diverse setups ranging from RGB, RGB-D to IMU based capture systems. The data captured from these sources often produce artifacts like jitters, unnaturally rigid joints or weird bends at some joints, or positions with only partial observations. Prior work improves the quality of captured motion sequences by using an optimization-based method, with the goal of recovering the captured data and preserving the realism of human poses. A robust and expressive human pose prior is key to preserve the realism of optimized poses, along with preserving the original data. Following HuMoR , we demonstrate the effectiveness of our pose manifold for: 1) motion denoising and 2) fitting to partial data.

We follow the same experimental setup as , but only deal with human poses and thus, remove the terms corresponding to human-scene contact and translation of root joint. In total, we find the pose parameters θ^t\hat{\theta}^{t} at frame t as:

Results. We compare motion denoising between HuMoR (TestOpt), Eq. (7) with VPoser prior , and Eq. (7) with \blah prior in Table 1. \blah achieves the lowest error in all settings. For mocap datasets like AMASS and HPS the motion is realistic, but can have small artifacts and jitter. Thus, an ideal motion/pose prior should not change the overall pose of these examples, but only fix these local artifacts. We observe that, numerically, VPoser and \blah-based optimization do not change the input pose significantly. However HuMoR changes the pose and this change increases with an increasing number of frames. This is because HuMoR is a motion-based prior (conditioned on the previous pose) and, hence, over time the correction in pose accumulates and makes the output pose significantly different from the input.

For the “Noisy AMASS” data, \blah-based optimization outperforms prior work. We visualise the denoising results in Fig. 2, and observe that the \blah-based method produces realistic and close to GT results. We further compare results of a sequence with HuMoR in Fig. 3. HuMoR results in large deviations from the input/GT, due to accumulation of correction over time.

Fitting to partial data. We use the test set of AMASS and randomly create occluded poses (e.g. missing arm or legs or shoulder joint) and quantitatively compare with HuMoR and VPoser in Table 2. We use Eq. (7) for VPoser and \blah-based optimization. We only optimize for the occluded joints and for our model, we initialize the occluded joint pose randomly (close to 0). For HuMoR, we use the TestOpt provided in their paper. We evaluate on three different type of occlusions: 1) occluded left leg, 2) occluded left arm and 3) occluded right shoulder and upper arm. For the occluded leg case, VPoser and our prior-based method perform better. We believe this is because the majority of the poses in both AMASS training and test are upright with nearly straight legs and hence VPoser is biased towards these poses. For our method, it highly depends on initialization. Since we have used an initialization close to rest position, our optimization method generates smaller error for occluded legs but higher errors for occluded arms and shoulders, as they usually are more far away from the rest pose. For HuMoR, the motion generated is realistic and plausible, but in some cases results in large deviation from ground truth, because the correction in input pose accumulates over the time.

3 3D pose Estimation from Images

We now show that \blah can also be used as a prior in optimization-based 3D pose estimation from images . We use the objective function proposed in SMPLify-X , see Eq. (10). Since we are working with a SMPL body only (without hands or faces), we remove the respective loss and prior terms. Thus, we find the desired pose θ^\hat{\boldsymbol{\theta}} and shape β^\hat{\boldsymbol{\beta}} as:

with data term LJ\mathcal{L}_{J}, bending term Lα\mathcal{L}_{\alpha}, shape regularizer Lβ\mathcal{L}_{\boldsymbol{\beta}}, and prior term Lθ\mathcal{L}_{\boldsymbol{\theta}}. The data term and the bending term are given as:

Results. We use the EHF dataset for quantitative evaluation and compare our work with the state-of-the-art priors VPoser and GAN-S . A \blah prior term slightly improves on the VPoser and GAN-S based optimization (Tab. 3). We observe that the neural network based model ExPose outperforms all optimization-based results. However, we show that such methods can benefit from an optimization-based refinement step. We refine the ExPose output using Eq (10) with \blah as prior and compare this refinement with no-prior and other priors (Tab. 3). With no prior, the optimization objective only minimizes the joint projection loss, resulting in unrealistic poses. In contrast, GAN and \blah improve the result (qualitatively and quantitatively), generating realistic poses, while \blah outperforms the GAN prior. Finally, in Fig. 4 we show qualitative results of optimization-based 3D pose estimation on in-the-wild images from 3DPW , LSP and MS-COCO datasets.

4 Pose Generation

We evaluate our model on the task of pose generation. Due to our distance field formulation, we can generate diverse poses by sampling a random point from SO(3)KSO(3)^{K} and projecting it onto the manifold (Sec. 3.4). We compare the results of our model with sampling from the state-of-the-art pose prior VPoser , GMM and GAN-S in Fig. 5. We use Average Pairwise Distance (APD) , to quantify the diversity of generated poses. APD is defined as mean joint distance between all pairs of samples. We randomly sample 500 poses for each GMM, VPoser, GAN-S and \blah, which results in APD values of 48.24, 23.13, 27.52, 32.31 (in cm), respectively. We see that numerically, the GMM produces very large variance, but also results in unrealistic poses, as seen in Fig 5 (top-left). \blah generates more diverse poses than VPoser while producing only plausible poses. We also calculate the percentage of self-intersecting faces in generated poses, to evaluate one aspect of realism in poses. \blah generates poses with less self-intersecting faces (0.89%\textbf{0.89}\%), as compared to the GAN-S (1.43%\textbf{1.43}\%) and VPoser (2.10%\textbf{2.10}\%).

5 Pose Interpolation

Results: We compare the results of \blah with those from VPoser and GAN-S interpolation. For VPoser , we project the start and end pose into the latent space and perform linear interpolation using the latent vectors. For GAN-S , we use the spherical interpolation in latent space, as suggested in the work. We qualitatively evaluate the interpolation quality by calculating mean per-vertex distance between consecutive frames. Smaller value means smooth interpolation. We observe that \blah-based interpolation has a mean per-vertex distance of 2.72 ± 2.16, GAN-S has 2.71 ± 2.45 and VPoser has 2.53 ± 4.62, which shows that Pose-NDF and GAN-S based interpolation is smooth and the distance in input space is not entirely preserved in case of VAEs. We compare VPoser based interpolation with \blah in Fig. 6) and observe large jumps in VPoser interpolation. This behaviour is not observed in GAN-S and Pose-NDF based interpolation. Since the VAE learns a compact latent representation of poses, the distance between two input poses is not preserved in the latent space.

6 Pose-NDF vs. Gaussian Assumption models

Prior work uses VAE-based models as pose/motion prior, which follow a Gaussian assumption in the latent space. This has three major limitations, as mentioned in Sec. 1. Conversely, \blah learns the manifold directly in the pose-space without such assumptions and, hence, overcomes these limitations.

We report the cumulative error based on deviation from the mean pose. We evaluate on AMASS Noisy (60 and 120 frames) and report cumulative error for samples with σ,2σ,3σ\sigma,2\sigma,3\sigma for both \blah and VPoser motion denoising. We obtain per-vertex error of 8.18,8.20,8.21\textbf{8.18},\textbf{8.20},\textbf{8.21} cmcm for \blah and 8.35,9.11,9.13\textbf{8.35},\textbf{9.11},\textbf{9.13} cmcm for VPoser, and 10.08,11.38,16.86\textbf{10.08},\textbf{11.38},\textbf{16.86} cmcm for HuMoR which reflects that VPoser and HuMoR perform well for poses close to the mean but the error increases for samples deviating from mean pose. Since the Gaussian distribution is unbounded, it produces dead regions, without any data points in these parts of distribution. Hence sampling in these regions might result in completely unrealistic poses for GMM and VPoser (Fig. 5). Lastly, since we learn the manifold in pose space, the distance between individual poses is preserved and leads to smoother interpolation compared to VPoser (see Sec. 4.5).

Conclusion

We introduced a novel human pose prior model represented by a scalar neural distance field that describes a manifold of plausible poses as zero level set in SO(3)KSO(3)^{K}. The method extends the idea of classic 3D shape representation using neural fields to higher the dimensions of human poses and maps quaternion-based poses to an unsigned distance value, representing the distance to the pose manifold. The resulting network can be used to project arbitrary poses to the pose manifold, opening applications in several areas. We comprehensively evaluate the performance of our model in diverse pose sampling, pose estimation from images, and motion denoising. We show that our model is able to generate poses with much more diversity than prior VAE-based works and improves state-of-the-art results in reconstruction from images and motion estimation.

Special thanks to the RVH team and reviewers, their feedback helped improve the manuscript and Andrey Davydov, for providing the code for GAN-based pose prior. This work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans), German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039A and a Facebook research award. Gerard Pons-Moll is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645.

References