gDNA: Towards Generative Detailed Neural Avatars

Xu Chen, Tianjian Jiang, Jie Song, Jinlong Yang, Michael J. Black, Andreas Geiger, Otmar Hilliges

Introduction

The ability to easily create diverse high-quality virtual humans with full control over their pose has many applications in movie production, games, VR/AR, architecture, and computer vision. While modern computer graphics techniques achieve photorealism, they typically require a lot of expertise and extensive manual effort. Our goal is to make 3D human avatars widely accessible by learning a generative model of people. Towards this goal, we propose the first method that can generate 1) diverse 3D virtual humans with 2) various identities and shapes, appearing in 3) different clothing styles and poses, with 4) realistic and stochastic high-frequency details such as wrinkles in garments.

Generative modeling of 3D rigid objects has recently seen rapid progress, fueled by continuous and resolution-independent neural 3D representations . However, modeling clothed humans and their articulation is more difficult due to the complex interaction of garments, their topology, and pose-driven deformations. Recent work leverages neural implicit surfaces to learn high-quality articulated avatars for a single subject but these methods are not generative, i.e. they cannot synthesize novel human identities and shapes. Generative models of clothing exist that augment SMPL by predicting displacements from the body mesh (CAPE ), or by draping an implicit garment representation on a T-posed body (SMPLicit ), and relying on SMPL’s learned skinning for reposing. We show empirically that holistic modeling of identity, shape, articulation and clothing leads to higher fidelity generation and animation of virtual humans and to higher accuracy in fitting to 3D scans.

Taking a step towards fully generative modeling of detailed neural avatars, we propose gDNA, a method that synthesizes 3D surfaces of novel human shapes, with control over the clothing style and pose, and that produces realistic high-frequency details of the garments (Fig. 1). To leverage raw (posed) 3D scans, we build a multi-subject implicit generative representation. We build upon SNARF , a recent method for learning single-subject articulation-dependent effects that has been shown to generalize well to unseen poses. SNARF requires many poses of a single subject for training. In contrast, our multi-subject method can be learned from very few posed scans (1-3) of many different subjects. This is achieved via the addition of a latent space for the conditional generation of shape and skinning weights for clothed humans. Furthermore, a learned warping field yields accurate deformations, using the same skinning field, independent of body size.

Clothing wrinkles are produced by an underlying stochastic process. To capture these effects, we propose a method that learns the underlying statistics of 3D clothing details via an adversarial loss. Previous mesh-based approaches formulate this in UV-space , which is not directly applicable to implicit surfaces due to the lack of mesh connectivity. To learn high-frequency details, we first predict a 3D normal field, conditioned on the coarse shape features. To backpropagate the adversarial loss to the 3D normal field we establish 3D-2D correspondences by augmenting forward skinning with an implicit surface renderer. We show that adversarial training leads to significantly improved fidelity of 3D geometric details, see Fig. 9.

Trained from posed scans only, we demonstrate the first method that can generate a large variety of 3D clothed human shapes with detailed wrinkles under pose control. The generated samples can be reposed via the learned skinning weights. We evaluate gDNA quantitatively, qualitatively, and through a perceptual study; gDNA strongly outperforms baselines. Furthermore, we show that gDNA can be used for fitting and re-animation of 3D scans, outperforming the state of the art (SOTA). In summary, we contribute:

The first method to generate a large variety of animatable 3D human shapes in detailed garments; that

learns from raw posed 3D scans without requiring canonical shapes, detailed surface registration, or manually defined skinning weights, and

a technique to significantly improve the geometric detail in clothing deformation, based on recovering the underlying statistics of cloth deformation.

Related Work

2D and 3D Generative Models: Most modern methods for synthesizing natural images leverage generative adversarial networks (GANs) or variational auto-encoders (VAEs) . These methods have achieved a high level of photorealism and can yield impressive results on the task of synthesizing 2D images of humans . However, such methods reason in 2D and hence 3D consistency cannot be guaranteed nor is extracting 3D geometry from such approaches straightforward.

Several methods for the task of learning rigid 3D shapes exist. Early methods rely on voxel or point cloud representations. More recently, several methods represent object shapes by learning an implicit function using neural networks . Such representations have also been proposed for the task of generative modeling of 3D shapes . However, these methods are typically not easily extended to non-rigid clothed humans. In this paper, we study the problem of 3D implicit generative modeling of non-rigid human shape.

3D Human Models: Parametric 3D human body models can synthesize 3D human shapes from a set of low-dimensional control parameters by deforming a template mesh. This idea has also been extended to model clothed humans . However, geometric expressivity is limited due to the fixed mesh topology and the bounded resolution of the template mesh.

To overcome the topology and resolution limitations of meshes, other representations, including point clouds , implicit surfaces , and radiance fields , have been explored. In particular, neural implicit surface representations have emerged as a powerful tool to model 3D (clothed) human shapes due to their topological flexibility and resolution independence. Recent work uses implicit surfaces to learn human avatars for a single subject, wearing a specific garment. These methods model clothing details such as wrinkles as a deterministic function of the body poses. However, due to hysteresis and complex material properties, garment folds and wrinkles are stochastic and existing methods struggle to capture these effects. In contrast, we propose a multi-subject generative model of 3D humans that provides separate control over poses, garments and can synthesize realistic geometric details.

CAPE and SMPLicit are generative models of clothing only, based on meshes and implicit surfaces respectively. Both methods are purely additive, that is they drape an implicit garment over the SMPL body or predict the displacement parameters of a SMPL+D template mesh . We experimentally show that this leads to lower fidelity in generated samples and higher error when fitting to 3D scans. NPMs provide a latent space of multiple subjects for fitting to RGB-D depth maps or 3D scans.

A common problem of all aforementioned approaches is the specific training data requirements these models impose: They either require synthetic data in canonical space , or precise registration of a template mesh to posed scans . The former are rare and suffer from a domain gap, while the latter is challenging to attain. Our method overcomes this issue by requiring only a few training samples of each subject in posed space. We show that our method learns complex shape and clothing details and models realistic deformation even from such limited data.

Adversarial Training of Clothing Details: Adversarial loss formulations have been used to learn detailed cloth wrinkles by optimizing 2D representations such as UV normal maps or depth images . It is noteworthy that implicit surfaces lack a notion of connectivity and therefore, incorporating 2D representations that have been designed to augment explicitly parametrized meshes is not straightforward. In contrast, we propose a formulation that leverages a 2D adversarial loss computed with posed images to optimize a 3D implicit representation in canonical space. Finally, our focus is the generation of human shapes appearing in varied clothing styles and diverse identities while previous methods focus on reconstruction or single garment pose-dependent wrinkle enhancement .

Method

Our goal is to build a model that generates diverse 3D clothed humans with varying identities and fine-grained geometric details in arbitrary poses. Our model is learned from a sparse set of static scans without assuming surface correspondences. Our method is summarized Fig. 2.

First, we formulate a pose- and body-size-independent canonical representation of clothed human shapes (Section 3.1). Second, to learn the canonical shape and deformation properties from very few posed scans of each of the subjects, we extend a single-subject differentiable forward skinning method to multiple subjects via a latent space of shape, articulation and garment (Section 3.2). Finally, to learn rich yet stochastic geometric details, we learn a detailed 3D normal field via a 2D adversarial loss formulation. To achieve this, we augment the forward skinning module with an implicit surface renderer (Section 3.3). Training details are discussed in Section 3.4.

Our method is based on neural implicit representations, leveraging their topological flexibility and resolution independence. We model the clothed human shape and geometric clothing details jointly.

Coarse Shape: We model the shape in canonical space as the τ=0.5\tau=0.5 level set of a neural occupancy function:

This occupancy network also outputs a feature vector f\mathbf{f} of dimension LfL_{\mathbf{f}} for each surface point. This feature carries coarse shape information and is used to predict fine details.

We combine a 3D CNN-based feature generator and a locally conditioned MLP to model O\mathcal{O}. A 3D style-based generator, illustrated in Fig. 3 first produces a 3D feature volume conditioned on zshape\mathbf{z}_{\text{shape}} via adaptive instance normalization . The final occupancy is obtained via trilinear sampling of the feature volume and by feeding the feature and the 3D coordinate into an MLP.

2 Multi-Subject Forward Skinning

We additionally model the deformation properties and define the body size (β\boldsymbol{\beta}) and pose (θ\boldsymbol{\theta}) parameters to be consistent with SMPL, enabling use of existing datasets (e.g. AMASS ) for animation. The body size parameter β\boldsymbol{\beta} is a 10-dimensional vector, and the body pose parameter θ\boldsymbol{\theta} represents the joint angles of SMPL’s skeleton.

Single-Subject Skinned Representation: To animate implicit human shapes in controllable body poses θ\boldsymbol{\theta}, recent work generalizes mesh-based linear blend skinning algorithms to neural implicit surfaces. The skeletal deformation of each 3D point is modeled as the weighted average of a set of bone transformations, with weights at each point predicted by an MLP. A key difference is whether this skinning weight field is defined in canonical space or in posed space. We follow Chen et al. who define the skinning field in canonical space:

where nbn_{b} denotes the number of bones and the weights w={w1,…,wnb}\mathbf{w}=\{w_{1},\dots,w_{n_{b}}\} of each point x\mathbf{x} are enforced to satisfy wi≥0w_{i}\geq 0 and ∑iwi=1\sum_{i}w_{i}=1 by a softmax activation function. As shown in , defining the skinning weights field in canonical space is desirable because the skinning weights are then pose-independent, thus easier to learn and enabling generalization to out-of-distribution poses.

Multi-Subject Skinned Representation: We extend this forward skinning idea to multiple subjects. Since the skinning weight field is defined in canonical space, the model can aggregate information over multiple training instances. Importantly, this enables us to learn skinning from one or a few poses of multiple subjects, instead of requiring many poses of the same subject.

To achieve this, we decouple the effects originating from the body size variation β\boldsymbol{\beta} and the clothed human shape zshape\mathbf{z}_{\text{shape}}. We model the skinning field in a body-size-neutral space, analogously to the canonical surface representations. To capture diverse clothed human shapes, we condition the field on the latent shape code zshape\mathbf{z}_{\text{shape}}:

We then model body size change with an additional warping field. Given a point x^\hat{\mathbf{x}} in β\boldsymbol{\beta}-size space, the warping field maps it back to the mean size by predicting its canonical correspondence x\mathbf{x} (see Fig. 4):

In this formulation, β\boldsymbol{\beta} captures body shape variations analogously to SMPL, e.g. body height. Therefore, the canonical shape network only needs to model the remaining shape variations beyond SMPL, e.g. clothing and hair, controlled by zshape\mathbf{z}_{\text{shape}}. The final resized canonical surface is defined by:

Given the target body pose θ\boldsymbol{\theta}, a point x^\hat{\mathbf{x}} in β\boldsymbol{\beta}-size space is transformed to posed space x′\mathbf{x}^{\prime} via

where Bi(β,θ)\boldsymbol{B}_{i}(\boldsymbol{\beta},\boldsymbol{\theta}) are the bone transformation matrices obtained from the parametric skeleton of SMPL.

Implicit Differentiable Forward Skinning: While our model learns a canonical representation, its supervision is provided in posed space. Given a point x′\mathbf{x}^{\prime} in posed space we need to determine its correspondence in canonical space x\mathbf{x} to compare the predicted occupancy and normals to ground-truth. We first find the correspondence x^∗\hat{\mathbf{x}}^{*} of x′\mathbf{x}^{\prime} in resized canonical space and then map x^∗\hat{\mathbf{x}}^{*} to canonical space x∗\mathbf{x}^{*}. An overview is provided in Fig. 4. While the goal is to determine x′↦x^\mathbf{x}^{\prime}\mapsto\hat{\mathbf{x}}, we only have direct access to the inverse mapping defined by forward skinning Eq. (8), which is not invertible. Following , we determine the correspondence numerically by finding the root of the equation:

using Broyden’s method . Subsequently, the canonical correspondence x∗\mathbf{x}^{*} is given by:

We can now determine the occupancy at x′\mathbf{x}^{\prime} as o′=O(x∗,zshape)o^{\prime}=\mathcal{O}(\mathbf{x}^{*},\mathbf{z}_{\text{shape}}) and the normal n′\mathbf{n}^{\prime} as

where Ri\mathbf{R}_{i} denotes the rotational component of Bi\boldsymbol{B}_{i}.

For convenient future reference, we define the occupancy field O′\mathcal{O}^{\prime} and normal function N′\mathcal{N}^{\prime} in posed space as:

3 Implicit Surface Rendering

Geometric clothing details are challenging to learn due to their stochastic nature. In 2D image generation tasks, GANs have achieved impressive results on learning high fidelity local textures. We propose to learn better geometric details N\mathcal{N} using an adversarial loss. Towards this goal, we augment the forward skinning module with an implicit renderer to establish direct correspondences between 2D projections of 3D points in posed space and corresponding 3D points in canonical space, enabling end-to-end training.

Implicit Rendering with Skinning: Given a pixel p\mathbf{p} in the 2D posed normal map, its correspondence in deformed 3D space x′\mathbf{x}^{\prime} can be determined by the intersection between the ray through p\mathbf{p} and the forward skinned surface:

where rd\mathbf{r}_{d} and rc\mathbf{r}_{c} denote the ray direction and origin, and tt is the scalar distance along the ray. Following , we determine the intersection point x′\mathbf{x}^{\prime} by finding the first change of occupancy O′\mathcal{O}^{\prime} along the ray using the Secant method. We also obtain the canonical correspondence point x\mathbf{x} of p\mathbf{p} via forward skinning. Solving the 3D canonical correspondence for each pixel, yields the 2D normal map II:

4 Training

We train our method via a set of posed scans and their corresponding SMPL parameters θ,β\boldsymbol{\theta},\boldsymbol{\beta}. We follow the auto-decoding framework of , and assign one shape code zshape\mathbf{z}_{\text{shape}} and one detail code zdetail\mathbf{z}_{\text{detail}} to each training sample. These are initialized to be zero and optimized jointly with the network weights. To enable sampling, we fit a Gaussian distribution to the latent codes after training.

We split training into two stages: We first train the coarse shape, skinning, and warping networks and then train the normal network. This two-stage training is essential. Otherwise, the normal supervision will be back-propagated to wrong locations in canonical space due to wrong correspondences before training of shape and skinning converges.

For the first stage, we use the binary cross entropy loss LBCE\mathcal{L}_{\text{BCE}} between predicted occupancy O′(x′,zshape,β,θ)\mathcal{O}^{\prime}(\mathbf{x}^{\prime},\mathbf{z}_{\text{shape}},\boldsymbol{\beta},\boldsymbol{\theta}) and ground-truth ogto_{\text{gt}}. Following , we add auxiliary losses Lbone\mathcal{L}_{\text{bone}} and Ljoint\mathcal{L}_{\text{joint}} to guide learning during early iterations:

where xbone\mathbf{x}_{\text{bone}} are randomly sampled points on canonical bones, xjoint\mathbf{x}_{\text{joint}} are randomly sampled canonical joints, and wjoint, target\mathbf{w}_{\text{joint, target}} is a vector that is 0.50.5 for the neighboring bones and elsewhere (for details see Sup. Mat.). To ensure that the warping field changes body size consistently, we enforce the warping field to warp SMPL vertices v(β)\mathbf{v}(\boldsymbol{\beta}) to the corresponding location in the neutral shape v(β0)\mathbf{v}(\boldsymbol{\beta}_{0}):

Finally, we regularize the latent code to be close to the origin of the latent space via Lreg,shape=∥zshape∥22\mathcal{L}_{\text{reg,shape}}=\|\mathbf{z}_{\text{shape}}\|_{2}^{2}.

The normal prediction network is trained subsequently. Here we penalize differences between the predicted and GT normal ngt′\mathbf{n}^{\prime}_{\text{gt}} for randomly sampled surface points:

In addition, we apply non-saturating adversarial losses Ladv=−log⁡(1+exp⁡(D(I)))\mathcal{L}_{\text{adv}}=-\log(1+\exp(D(I))) with R1R_{1} gradient penalty on the predicted 2D normal maps II and the real normal maps rendered from the posed scans IrealI_{\text{real}}. DD is a jointly trained discriminator (see Sup. Mat. for details). We further regularize zdetail\mathbf{z}_{\text{detail}} with Lreg,detail=∥zshape∥22\mathcal{L}_{\text{reg,detail}}=\|\mathbf{z}_{\text{shape}}\|_{2}^{2}.

5 Inference

We generate human avatars by randomly sampling zshape\mathbf{z}_{\text{shape}} and zdetail\mathbf{z}_{\text{detail}} from the estimated Gaussian distribution. We then extract meshes in resized canonical space using MISE from the implicit representation S^(zshape,β)\hat{\mathcal{S}}(\mathbf{z}_{\text{shape}},\boldsymbol{\beta}) and predict the vertex normal with our normal field. Finally, we pose the meshes to desired poses θ\boldsymbol{\theta} following Eq. (8).

Experiments

Our main goal is to generate 3D human avatars. Since we are the first to tackle this problem setting, we compare our method to carefully designed ablative baselines, enabling analysis of each component of our method. We also evaluate the expressiveness of our model by fitting it to unseen scans and compare the accuracy to SOTA 3D human shape modeling methods. We outline the evaluation protocols in the following and refer the readers to Sup. Mat. for details.

3D Scans: We train our model on commercial scans .

SIZER: Following , we use the SIZER dataset to evaluate fitting. This dataset contains 3D scans of humans in 21 garments, including shirts, T-shirts, coats and pants.

Fréchet Inception Distance (FID): To evaluate generation quality, we compute FID between 2D normal maps of training scans and those of randomly generated 3D shapes.

User Preference: We conduct a perceptual study among 44 subjects and report how often participants preferred a particular method over ours.

Surface Distance: To evaluate fitting accuracy, we measure the one-directional Chamfer distance between predicted surfaces and the target scans, following SMPLicit .

NPMs : NPMs learn the latent space of human shapes and deformation from ground-truth canonical shapes and vertex displacements, obtained from synthetic 3D animations and real scans with registered surfaces .

SMPLicit : SMPLicit learns a generative model of 3D garments, and drapes these over the SMPL T-pose. This model is trained with a collection of 3D synthetic garments.

Random Generation of Canonical Shapes: We show random samples generated by our method in Fig. 5 (top). While trained with posed scans only, our method learns plausible canonical shapes with surface details.

Disentangled Pose and Shape: The generated shapes can be reposed as desired, even to poses far beyond the training pose distribution (cf. Fig. 5 bottom and Fig. 1).

Interpolation: Interpolating the shape and details codes, yields smooth transitions of shapes and details between two very different samples, as shown in Fig. 6.

Disentangled Shape and Details: Our disentangled formulation allows us to generate diverse clothing details for the same coarse shape. Fig. 7 shows results with the same coarse shape zshape\mathbf{z}_{\text{shape}} but different details codes zdetail\mathbf{z}_{\text{detail}}. While the coarse shape remains the same, gDNA generates varied plausible wrinkles that match the underlying coarse shape.

Extrapolation Beyond Training Distribution: To further illustrate generalization, we show the training samples with the most similar pose and latent code to the generated sample in Fig. 8. The nearest neighbors are noticeably different from our generation, demonstrating that our method generalizes and is able to generate novel shapes in novel poses.

2 Ablation Study

We now ablate our design choices. The results are summarized in Tab. 1 and Fig. 9.

Canonical Space Modeling: We verify the necessity to model shapes in canonical space and joint learning of skinning weights. Towards this goal, we implement a baseline that generates posed shapes directly given the latent code and the body pose as input. As shown in Fig. 9 (first row), the individual samples lack details as the baseline must capture a large shape space caused by the pose change. Since the method does not reason about articulation, the sampled shapes suffer from invalid pose configurations, leading to high FID values as shown in Tab. 1 (Pose ONet).

Adversarial Learning: The adversarial loss plays an important role in improving the perceptual realism of the generated samples, as evidenced by the FID improvement from Detailed Normal (w/o Adversarial) to Ours in Tab. 1. The normals estimated directly from the occupancy field suffer from artifacts on the surface (Fig. 9 (second row)). Training without adversary leads to overly smooth geometry (Fig. 9 (third row)), as the reconstruction loss induces a bias that averages out details. In contrast, our method produces realistic high-frequency details (Fig. 9 (bottom)). Notably, in 21.3% of the cases, users consider our generated shapes to be even more realistic than real scans.

3 Comparison with SOTA on Model Fitting

While our main goal is to generate clothed human shapes, our model can be fit to raw observations, just like existing 3D parametric human or clothing models. We consider two recent SOTA methods, i.e. NPMs and SMPLicit . We follow SMPLicit and fit ours and the baselines to scans from the SIZER dataset.

Accuracy: While not designed for fitting, our method achieves better accuracy than previous special purpose methods, as demonstrated in Tab. 2. Our method captures the person identity and clothing shapes more faithfully than NPMs and SMPLicit, and our results exhibit more details such as wrinkles (Fig. 10 top). Since the model is trained directly from posed scans, disentangling pose and shape, it learns about real clothing details and can reproduce them.

Reposing Scans: During fitting we also recover skinning weights. This enables reposing of the shape as demonstrated in Fig. 10 bottom.

Conclusion

We propose gDNA, a generative model of 3D clothed humans that can produce a large variety of clothed people with detailed wrinkles and explicit pose control. Using implicit multi-subject forward skinning enables learning from only a few posed scans per subject. To model the stochastic details of garments, we exploit a 2D adversarial loss to update a 3D normal field. We demonstrate that gDNA can be used in various applications such as animation and 3D fitting, outperforming state-of-the-art methods.

Limitations: Learning loose clothing (e.g. skirts) from deformed observations remains challenging due to the topology ambiguity and the large pose-dependent non-linear cloth deformation. Please refer to Sup. Mat. for more discussions about limitations and societal impact.

Acknowledgements: Xu Chen was supported by the Max Planck ETH Center for Learning Systems. Andreas Geiger was supported by the DFG EXC number 2064/1 - project number 390727645. We thank Alex Zicong Fan, Marcel C. Bühler, Priyanka Patel, Qianli Ma, Sai Kumar Dwivedi, Thomas Langerak and Yuliang Xiu for their feedback, Garvita Tiwari for her suggestions about the SIZER dataset, and Tsvetelina Alexiadis for her help with the user study.

Disclosure: MJB has received research gift funds from Adobe, Intel, Nvidia, Meta/Facebook, and Amazon. MJB has financial interests in Amazon, Datagen Technologies, and Meshcapade GmbH. MJB’s research was performed solely at, and funded solely by, the Max Planck.

References