PINA: Learning a Personalized Implicit Neural Avatar from a Single RGB-D Video Sequence

Zijian Dong, Chen Guo, Jie Song, Xu Chen, Andreas Geiger, Otmar Hilliges

Introduction

Making immersive AR/VR a reality requires methods to effortlessly create personalized avatars. Consider telepresence as an example: the remote participant requires means to simply create a detailed scan of themselves and the system must then be able to re-target the avatar in a realistic fashion to a new environment and to new poses. Such applications impose several challenging constraints:

To address these requirements, we introduce a novel method for learning Personalized Implicit Neural Avatars (PINA) from only a sequence of monocular RGB-D video.

Existing methods do not fully meet these criteria. Most state-of-the-art dynamic human models burov2021dsfn; loper2015smpl; CAPE:CVPR:20; bhatnagar2020ipnet represent humans as a parametric mesh and deform it via linear blend skinning (LBS) and pose correctives. Sometimes learned displacement maps to capture details of tight-fitting clothing are used burov2021dsfn. However, the fixed topology and resolution of meshes limit the type of clothing and dynamics that can be captured. To address this, several methods saito2019pifu; chibane20ifnet propose to learn neural implicit functions to model static clothed humans. Furthermore, several methods that learn a neural avatar for a specific outfit from watertight meshes Saito:CVPR:2021; chen2021snarf; tiwari2021neural; deng2020nasa; Peng_2021_ICCV; liu2021neural; habermann2021 have been proposed. These methods either require complete full-body scans with accurate surface normals and registered poses Saito:CVPR:2021; chen2021snarf; tiwari2021neural; deng2020nasa or rely on complex and intrusive multi-view setups Peng_2021_ICCV; liu2021neural; habermann2021.

Learning an animatable avatar from a monocular RGB-D sequence is challenging since raw depth images are noisy and only contain partial views of the body (Fig. 1, left). At the core of our method lies the idea to fuse partial depth maps into a single, consistent representation and to learn the articulation-driven deformations at the same time. To do so, we formulate an implicit signed distance field (SDF) in canonical space. To learn from posed observations, the inverse mapping from deformed to canonical space is required. We follow SNARF chen2021snarf and locate the canonical correspondences via optimization. A key challenge brought on by the monocular RGB-D setting is to learn from incomplete point clouds. Inspired by rigid learned SDFs for objects gropp2020implicit, we propose a point-based supervision scheme that enables learning of articulated non-rigid shapes (i.e. clothed humans). Transforming the spatial gradient of the SDF into posed space and comparing it to surface normals from depth images leads to the learning of geometric details. Training is formulated as a global optimization that jointly optimizes the canonical SDF, the skinning fields and the per-frame pose. PINA learns animatable avatars without requiring any additional supervision or priors extracted from large datasets of clothed humans.

In detailed ablations, we shed light on the key components of our method. We compare to existing methods in the reconstruction and animation tasks, showing that our method performs best across several datasets and settings. Finally, we demonstrate the ability to capture and animate different humans in a variety of clothing styles qualitatively.

a method to fuse partial RGB-D observations into a canonical, implicit representation of 3D humans; and

to learn an animatable SDF representation directly from partial point clouds and normals; and

a formulation which jointly optimizes shape, per-frame pose and skinning weights.

The code and video will be available on the project page: https://zj-dong.github.io/pina/.

Related Work

A large body of literature utilizes explicit surface representations (particularly polygonal meshes) for human body modeling anguelov2005scape; hasler2009statistical; joo2018total; pavlakos2019expressive; xu2020ghum; osman2020star; bhatnagar2019mgn. These works typically leverage parametric models for minimally clothed human bodies pavlakos2019expressive; dong2021shape; song2020lgd; kocabas2020vibe; bogo2016keep; fang2021reconstructing (e.g. SMPL loper2015smpl) and use a displacement layer on top of the minimally clothed body to model clothing guo2021human; alldieck2019learning; alldieck2018video; ma2020learning; neophytou2014layered; yang2018analyzing. Recently, DSFN burov2021dsfn proposes to embed MLPs into the canonical space of SMPL to model pose-dependent deformations. However, such methods depend upon SMPL learned skinning for deformation and are upper-bounded by the expressiveness of the template mesh. During animation or reposing, the surface deformations of parametric human models rely on the skinning weights trained from minimally clothed body scans loper2015smpl; ma2020learning; Zhang_2017_CVPR. These methods suffer from artifacts during reposing as they rely on skinning weights of a naked body for animation which may be incorrect for points on the surface of the garment. In contrast, our method represents clothed humans as a flexible implicit neural surface and jointly learns shape and a neural skinning field from depth observations. Other methods gundogdu2019garnet; guan2012drape; patel20tailornet; santesteban2021self; Bertiche_2021_ICCV leverage physical simulation to drape garments onto the SMPL model. These approaches are a promising direction towards more realistic cloths deformation compared to previous template-based methods with fixed skinning weights. These methods are orthogonal to ours: we focus on acquiring a surface representation of body and clothing from the raw inputs, without assuming prior knowledge about the subject.

Implicit Human Models from 3D Scans

Implicit neural representations park2019deepsdf; mescheder2019occupancy; chen2019learning can handle topological changes better bozic2021neural; park2021hypernerf and have been used to reconstruct clothed human shapes huang2020arch; li2020monocular; saito2019pifu; saito2020pifuhd; he2020geo; peng2021neural; raj2021anr; chen2022gdna; xiu2022icon. Typically, based on a learned prior from large-scale datasets, they recover the geometry of clothed humans from images saito2019pifu; saito2020pifuhd; zheng2021pamir; xiu2022icon or point clouds chibane20ifnet. However, these reconstructions are static and cannot be reposed. Follow-up work huang2020arch; bhatnagar2020ipnet attempts to endow static reconstructions with human motion based on a generic deformation field which tends to output unrealistic animated results. To model pose-dependent clothing deformations, SCANimate Saito:CVPR:2021 proposes to transform scans to canonical space in a weakly supervised manner and to learn the implicit shape model conditioned on joint-angle rotations. Follow-up works further improve the generalization ability to unseen poses and accelerate the training process via a displacement network tiwari2021neural, deform the shape via a forward warping field chen2021snarf; zheng2021avatar or leverage prior information from large-scale human datasets wang2021metaavatar. However, all of these methods require complete and registered 3D human scans for training, even if they sometimes can be fine-tuned on RGB-D data. In contrast, PINA is able to learn a personalized implicit neural avatar directly from a short monocular RGB-D sequence without requiring large-scale datasets of clothed human 3D scans or other priors.

Reconstructing Clothed Humans from RGB-D Data

One straightforward approach to acquiring a 3D human model from RGB-D data is via per-frame reconstruction chibane20ifnet; bhatnagar2020ipnet. To achieve this, IF-Net chibane20ifnet learns a prior to reconstruct an implicit function of a human and IP-Net bhatnagar2020ipnet extends this idea to fit SMPL to this implicit surface for articulation. However, since input depth observations are partial and noisy, artifacts appear in unseen regions. Real-time performance capture methods incrementally fuse observations into a volumetric SDF grid. DynamicFusion newcombe2015dynamicfusion extends earlier approaches for static scene reconstruction newcombe2011kinectfusion to non-rigid objects. BodyFusion BodyFusion and DoubleFusion DoubleFusion build upon this concept by incorporating an articulated motion prior and a parametric body shape prior. Follow-up work bozic2021neural; li2021posefusion; burov2021dsfn leverages a neural network to model the deformation or to refine the shape reconstruction. However, it is important to note that such methods only reconstruct the surface, and sometimes the pose, for tracking purposes but typically do not allow for the acquisition of skinning information which is crucial for animation. In contrast, our focus differs in that we aim to acquire a detailed avatar including its surface and skinning field for reposing and animation.

Method

We introduce PINA, a method for learning personalized neural avatars from a single RGB-D video, illustrated in Fig. 2. At the core of our method lies the idea to fuse partial depth maps into a single, consistent representation of the 3D human shape and to learn the articulation-driven deformations at the same time via global optimization.

We parametrize the 3D surface of clothed humans as a pose-conditioned implicit signed-distance field (SDF) and a learned deformation field in canonical space (Sec. 3.1). This parametrization enables the fusion of partial and noisy depth observations. This is achieved by transforming the canonical surface points and the spatial gradient into posed space, enabling supervision via the input point cloud and its normals. Training is formulated as global optimization (Sec. 3.2) to jointly optimize the per-frame pose, shape and skinning fields without requiring prior knowledge extracted from large datasets. Finally, the learned skinning field can be used to articulate the avatar (Sec. 3.3).

We model the human avatar in canonical space and use a neural network fsdff_{\text{sdf}} to predict the signed distance value for any 3D point xc\mathbf{x}_{c} in this space. To model pose-dependent local non-rigid deformations such as wrinkles on clothes, we concatenate the human pose p\mathbf{p} as additional input and model fsdff_{\text{sdf}} as:

The pose parameters (p\mathbf{p}) are defined consistently to the SMPL skeleton loper2015smpl and npn_{p} is their dimensionality. The canonical shape S\mathcal{S} is then given by the zero-level set of fsdff_{\text{sdf}}:

In addition to signed distances, we also compute normals in canonical space. We empirically find that this resolves high-frequency details better than calculating the normals in the posed space. The normal for a point xc\mathbf{x}_{c} in canonical space is computed as the spatial gradient of the signed distance function at that point (attained via backpropagation):

To deform the canonical shape into novel body poses we additionally model deformation fields. To animate implicit human shapes in the desired body pose p\mathbf{p}, we leverage linear blend skinning (LBS). The skeletal deformation of each point in canonical space is modeled as the weighted average of a set of bone transformations B\mathbf{B}, which are derived from the body pose p\mathbf{p}. We follow chen2021snarf and define the skinning field in canonical space using a neural network fwf_{w} to model the continuous LBS weight field:

Here, nbn_{b} denotes the number of joints in the transformation and wc={wc1,...,wcnb}=fw(xc)\mathbf{w}_{c}=\{w_{c}^{1},...,w_{c}^{n_{b}}\}=f_{w}(\mathbf{x}_{c}) represents the learned skinning weights for xc\mathbf{x}_{c}.

Skeletal Deformation

Given the bone transformation matrix Bi\mathbf{B}_{i} for joint i∈{1,...,nb}i\in\{1,...,n_{b}\}, a canonical point xc\mathbf{x_{c}} is mapped to the deformed point xd\mathbf{x}_{d} as follows:

The normal of the deformed point xd\mathbf{x_{d}} in posed space is calculated analogously:

where Ri\mathbf{R}_{i} is the rotation component of Bi\mathbf{B}_{i}.

To compute the signed distance field SDF(xd)SDF(\mathbf{x}_{d}) in deformed space, we need the canonical correspondences xc∗\mathbf{x}_{c}^{*}.

Correspondence Search

For a deformed point xd\mathbf{x}_{d}, we follow chen2021snarf and compute its canonical correspondence set Xc={xc1,...,xck}\mathbf{\mathcal{X}}_{c}=\{\mathbf{x}_{c}^{1},...,\mathbf{x}_{c}^{k}\}, which contains kk canonical candidates satisfying Eq. 5, via an iterative root finding algorithm. Here, kk is an empirically defined hyper-parameter of the root finding algorithm (see Supp. Mat for more details).

Note that due to topological changes, there exist one-to-many mappings when retrieving canonical points from a deformed point, i.e., the same point xd\mathbf{x}_{d} may correspond to multiple different valid xc\mathbf{x}_{c}. Following Ricci et al. ricci1973constructive , we composite these proposals of implicitly defined surfaces into a single SDF via the union (minimum) operation:

The canonical correspondence xc∗\mathbf{x}_{c}^{*} is then given by:

2 Training Process

Defining our personalized implicit model in canonical space is crucial to integrating partial observations across all depth frames because it provides a common reference frame. Here, we formally describe this fusion process. We train our model jointly w.r.t. body poses and the weights of the 3D shape and skinning networks.

Given an RGB-D sequence with N input frames, we minimize the following objective:

Loni\mathcal{L}_{\text{on}}^{i} represents an on-surface loss defined on the human surfaces for frame ii. Loffi\mathcal{L}_{\text{off}}^{i} represents the off-surface loss which helps to carve free-space and Leiki\mathcal{L}_{\text{eik}}^{i} is the Eikonal regularizer which ensures a valid signed distance field. Θ\mathbf{\Theta} is the set of optimized parameters which includes the shape network weights Θsdf\mathbf{\Theta_{\text{sdf}}}, the skinning network weights Θw\mathbf{\Theta}_{w} and the pose parameters pi\mathbf{p}_{i} for each frame.

To calculate Loni\mathcal{L}_{\text{on}}^{i}, we first back-project the depth image into 3D space to obtain partial point clouds Poni\mathbf{\mathcal{P}}_{\text{on}}^{i} of human surfaces for each frame ii. For each point xd\mathbf{x}_{d} in Poni\mathbf{\mathcal{P}}_{\text{on}}^{i}, we additionally calculate its corresponding normal ndobs\mathbf{n}_{d}^{\text{obs}} from the raw point cloud using principal component analysis of points in a local neighborhood. Loni\mathcal{L}_{\text{on}}^{i} is then defined as

Here, NC(xd)=ndobs(xd)−nd(xd)NC(\mathbf{x}_{d})=\mathbf{n}_{d}^{\text{obs}}(\mathbf{x}_{d})-\mathbf{n}_{d}(\mathbf{x}_{d}).

We add two additional terms to regularize the optimization process. Loffi\mathcal{L}_{\text{off}}^{i} complements Loni\mathcal{L}_{\text{on}}^{i} by randomly sampling points Poffi\mathbf{\mathcal{P}}_{\text{off}}^{i} that are far away from the body surface. For any point xd\mathbf{x}_{d} in Poffi\mathbf{\mathcal{P}}_{\text{off}}^{i}, we calculate the signed distance between this point and an estimated body mesh (see initialization section below). This signed distance SDFbody(xd)SDF_{body}(\mathbf{x}_{d}) serves as pseudo ground truth to force plausible off-surface SDF values. Loffi\mathcal{L}_{\text{off}}^{i} is then defined as:

Following IGR gropp2020implicit, we leverage Leiki\mathcal{L}_{\text{eik}}^{i} to force the shape network fsdff_{\text{sdf}} to satisfy the Eikonal equation in canonical space:

Implementation

The implicit shape network and blend skinning network are implemented as MLPs. We use positional encoding mildenhall2020nerf for the query point xc\mathbf{x}_{c} to increase the expressive power of the network. We leverage the implicit differentiation derived in chen2021snarf to compute gradients during iterative root finding.

Initialization

We initialize body poses by fitting SMPL model loper2015smpl to RGB-D observations. This is achieved by minimizing the distances from point clouds to the SMPL mesh and jointly minimizing distances between the SMPL mesh and the corresponding surface points obtained from a DensePose guler2018densepose model. Please see Supp. Mat. for details.

Optimization

Given a sequence of RGB-D video, we deform our neural implicit human model for each frame based on the 3D pose estimate and compare it with its corresponding RGB-D observation. This allows us to jointly optimize both shape parameters Θsdf\mathbf{\Theta}_{\text{sdf}}, Θw\mathbf{\Theta}_{w} and pose parameters pi\mathbf{p}_{i} of each frame and makes our model robust to noisy initial pose estimates. We follow a two-stage optimization protocol for faster convergence and more stable training: First, we pretrain the shape and skinning networks in canonical space based on the SMPL meshes obtained from the initialization process. Then we optimize the shape network, skinning network and poses jointly to match the RGB-D observations.

3 Animation

To generate animations, we discretize the deformed space at a pre-defined resolution and estimate SDF(xd)SDF(\mathbf{x}_{d}) for every point xd\mathbf{x}_{d} in this grid via correspondence search (Sec 3.1). We then extract meshes via MISE mescheder2019occupancy.

Experiments

We first conduct ablations on our design choices. Next, we compare our method with state-of-the-art approaches on the reconstruction and animation tasks. Finally, we demonstrate personalized avatars learned from only a single monocular RGB-D video sequence qualitatively.

We first conduct experiments on two standard datasets with clean scans projected to RGB-D images to evaluate our performance on both reconstruction and animation. To further demonstrate the robustness of our method to real-world sensor noise, we collect a dataset with a single Kinect including various challenging garment styles.

This dataset contains textured 3D scan sequences. Following burov2021dsfn, we obtain monocular RGB-D data by rendering the scans and use them for our reconstruction task, comparing to the ground truth scans.

CAPE Dataset CAPE:CVPR:20:

This dataset contains registered 3D meshes of people wearing different clothes while performing various actions. It also provides corresponding ground-truth SMPL parameters. Following Saito:CVPR:2021, we conduct animation experiments on CAPE. To adapt CAPE to our monocular depth setting, we acquire single-view depth inputs by rendering the meshes. The most challenging subject (blazer) is used for evaluation where 10 sequences are used for training and 3 unseen sequences are used for evaluating the animation performance. Note that our method requires RGB-D for initial pose estimation since CAPE does not provide texture, we take the ground-truth poses for training (same for the baselines).

Real Data:

To show robustness and generalization of our method to noisy real-world data, we collect RGB-D sequences with an Azure Kinect at 3030 fps (each sequence is approximately 2-3 minutes long). We use the RGB images for pose initialization. We learn avatars from this data and animate avatars with unseen poses aist-dance-db; AMASS:ICCV:2019; CAPE:CVPR:20.

Metrics:

We consider volumetric IoU, Chamfer distance (cm) and normal consistency for evaluation.

2 Ablation Study

The initial pose estimate from a monocular RGB-D video is usually noisy and can be inaccurate. To evaluate the importance of jointly optimizing pose and shape, we compare our full model to a version without pose optimization. Results: Tab. 1 shows that joint optimization of pose and shape is crucial to achieve high reconstruction quality and globally accurate alignment (Chamfer distance and IoU). It is also important to recover details (normal consistency). As shown in Fig. 3, unnatural reconstructions such as the artifacts on the head and trouser leg can be corrected by pose optimization.

Deformation Model:

The deformation of the avatar can be split into pose-dependent deformation and skeletal deformation via LBS with the learned skinning field. To model pose-dependent deformations such as cloth wrinkles, we leverage a pose-conditioned shape network to represent the SDF in canonical space. Results: Fig. 4 shows that without pose features, the network cannot represent dynamically changing surface details of the blazer, and defaults to a smooth average. This is further substantiated by a 70%\% increase in Chamfer distance, compared to our full method.

To show the importance of learned skinning weights, we compare our full model to a variant with a fixed shape network (w/ SMPL weights). Points are deformed using SMPL blend weights at the nearest SMPL vertices. Results: Tab. 2 indicates that our method outperforms the baseline in all metrics. In particular, the normal consistency improved significantly. This is also illustrated in Fig. 4, where the baseline (w/ SMPL weights) is noisy and yields artifacts. This can be explained by the fact that the skinning weights of SMPL are defined only at the mesh vertices of the naked body, thus they can’t model complex deformations.

3 Reconstruction Comparisons

Although not our primary goal, we also compare to several reconstruction methods, including IP-Net bhatnagar2020ipnet, CAPE CAPE:CVPR:20 and DSFN burov2021dsfn. The experiments are conducted on the BUFF Zhang_2017_CVPR dataset. The RGB-D inputs are rendered from a sequence of registered 3D meshes. IP-Net relies on a learned prior from twindom; treedys. It takes the partial 3D point cloud from each depth frame as input and predicts the implicit surface of the human body. The SMPL+D model is then registered to the reconstructed surface. DSFN models per-vertex offsets to the minimally-clothed SMPL body via pose-dependent MLPs. For CAPE, we follow the protocol in DSFN burov2021dsfn and optimize the latent codes based on the RGB-D observations.

Results:

Tab. 3 summarizes the reconstruction comparison on BUFF. We observe that our method leads to better reconstructions in all three metrics compared to current SOTA methods. A qualitative comparison is shown in Fig. 5. Compared to methods based on implicit reconstruction, i.e., IP-Net, our method reconstructs person-specific details better and generates complete human bodies. This is due to the fact that IP-Net reconstructs the human body frame-by-frame and can’t leverage information across the sequence. In contrast, our method solves the problem via global optimization. Compared to methods with explicit representations, i.e., CAPE and DSFN, our method reconstructs details (hair, trouser leg) that geometrically differ from the minimally clothed human body better. We attribute this to the flexibility of implicit shape representations.

4 Animation Comparisons

We compare animation quality on CAPE CAPE:CVPR:20 with IP-Net bhatnagar2020ipnet and SCANimate Saito:CVPR:2021 as baselines. IP-Net does not natively fuse information across the entire depth sequence (discussed in Sec 4.3). For a fair comparison, we feed one complete T-pose scan of the subject as input to IP-Net and predict implicit geometry and leverage the registered SMPL+D model to deform it to unseen poses. For SCANimate, we create two baselines. The first baseline (SCANimate 3D) is learned from complete meshes and follows the original setting of SCANimate. Note that in this comparison ours is at a disadvantage since we only assume monocular depth input without accurate surface normal information. Therefore, we also compare to a variant (SCANimate 2.5D) which operates on equivalent 2.5D inputs.

Results:

Tab. 4 shows the quantitative results. Our method outperforms IP-Net and SCANimate (2.5D), and achieves comparable results to SCANimate (3D) which is trained on complete and noise-free 3D meshes. Fig. 6 shows that the clothing deformation of the blazer is unrealistic when animating IP-Net. This may be due to overfitting to the training data. Moreover, the animation is driven by skinning weights that are learned from minimally-clothed human bodies. As seen in Fig. 6, SCANimate also leads to unrealistic animation results for unseen poses. This is because the deformation field in SCANimate depends on the pose of the deformed object, which limits generalization to unseen poses chen2021snarf. Furthermore, we find that this issue is amplified in SCANimate (2.5D) with partial point clouds. In contrast, our method solves this problem well via jointly learned skinning field and shape in canonical space.

5 Real-world Performance

To demonstrate the performance of our method on noisy real-world data, we show results on additional RGB-D sequences from an Azure Kinect in Fig. 7. More specifically, we learn a neural avatar from an RGB-D video and drive the animation using unseen motion sequences from aist-dance-db; AMASS:ICCV:2019; CAPE:CVPR:20. Our method is able to reconstruct complex cloth geometries like hoodies, high collar and puffer jackets. Moreover, we demonstrate reposing to novel out-of-distribution motion sequences including dancing and exercising.

Conclusion

In this paper, we presented PINA to learn personalized implicit avatars for reconstruction and animation from noisy and partial depth maps. The key idea is to represent the implicit shape and the pose-dependent deformations in canonical space which allows for fusion over all frames of the input sequence. We propose a global optimization that enables joint learning of the skinning field and surface normals in canonical representation. Our method learns to recover surface details and is able to animate the human avatar in novel unseen poses. We compared the method to explicit and neural implicit state-of-the-art baselines and show that we outperform all baselines in all metrics. Currently, our method does not model the appearance of the avatar. This is an exciting direction for future work. We discuss potential negative societal impact and limitations in the Supp. Mat.

Acknowledgements: Zijian Dong was supported by ELLIS. Xu Chen was supported by the Max Planck ETH Center for Learning Systems.

References