SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar Networks

Shunsuke Saito, Jinlong Yang, Qianli Ma, Michael J. Black

Introduction

Parametric models of 3D human bodies are widely used for the analysis and synthesis of human shape, pose, and motion. While existing models typically represent “minimally clothed” bodies , many applications require realistically clothed bodies. Our goal is to make it easy to produce a realistic 3D avatar of a clothed person that can be reposed and animated as easily as existing models like SMPL . In particular, the model must support clothing that moves and deforms naturally, with detailed 3D wrinkles, and the rendering of realistically textured images.

To that end, we introduce SCANimate (Skinned Clothed Avatar Networks for animation), which creates high-quality animatable clothed humans, called Scanimats, from raw 3D scans. SCANimate has the following properties: (1) we learn an articulated clothed human model directly from raw scans, completely eliminating the need for surface registration of a custom template or synthetic clothing simulation data, (2) our parametric model retains the complex and detailed deformations of clothing present in the original scans such as wrinkles and sliding effects of garments with arbitrary topology, (3) a Scanimat can be animated directly using SMPL pose parameters, and (4) our approach predicts pose-dependent clothing deformations based on local pose parameters, providing generalization to unseen poses.

Recent data-driven approaches have shown promise for learning parametric models of clothed humans from real-world observations . However, these approaches typically limit the supported clothing types and topology because they require accurate surface registration of a common template mesh to 3D training scans . Concurrent work by Ma et al. learns clothing deformation without surface registration, yet it is unclear if the method works on raw scans with noise and holes. Learning from real-world observations is essentially challenging because raw 3D scans are un-ordered point clouds with missing data, changing topology, multiple clothing layers, and sliding motions between the body and garments. Although one can learn from synthetic data generated by physics-based clothing simulation , the results are less realistic, the data preparation is time consuming and non-trivial to scale to the real-world clothing.

To address these issues, SCANimate learns directly from raw scans of people in clothing. Body scanning is becoming common, and scans can be obtained from a variety of devices. Scans contain high-frequency details, capture varied clothing topology, and are inherently realistic. To make learning from scans possible, we make several contributions: canonicalization, implicit skinning fields, cycle consistency, and implicit shape learning.

Canonicalization and Implicit Skinning Fields. The first step involves transforming the raw scans to a common pose so we can learn to model pose-dependent surface deformations (e.g. bulging, stretching, wrinkling, and sliding), i.e. pose “correctives”. But we are not seeking a traditional “registration” of the scans to a common mesh topology, since this is, in general, not feasible with clothed bodies. Instead, we learn continuous functions of 3D space that allow us to transform posed scans to a canonical pose and back again.

Cycle Consistency. Despite the desirable properties of canonicalization, learning the skinning function is ill-posed since we do not have ground truth training data that specifies the weights. To address this, we exploit two key observations. First, as demonstrated in previous work , fitting a parametric human body model such as SMPL to 3D scans is more tractable than surface registration. We leverage SMPL’s skinning weights, which are defined only on the body surface, to regularize our more general skinning function. Second, the transformations between the posed space and the canonical space should be cycle-consistent. Namely, inverse LBS and forward LBS together should form an identity mapping as illustrated in Fig. 3, which provides a self-supervision signal for training the skinning function. After training the skinning function, we obtain the canonicalized scans (all in the same pose).

Learning Implicit Pose Correctives. Given the canonicalized scans, we learn a model that captures the pose-dependent deformations. However a problem remains: the original raw scans often contain holes, and so do the canonicalized scans. To deal with this and with the arbitrary topology of clothing, we use an implicit surface representation . As multiple canonicalized scans will miss different regions, with this approach, they complement each other, while retaining details present in the original inputs. Furthermore, unlike traditional approaches , where pose-dependent deformations are conditioned on entire pose parameters, we spatially filter out irrelevant pose features from the input conditions by leveraging the learned skinning weights. In this way, we effectively prune long-range spurious correlations between garment deformations and body joints, achieving plausible pose correctives for unseen poses even from a small number of training scans. The resulting learned Scanimat can be easily reposed and animated with SMPL pose parameters.

In summary, our main contributions are (1) the first end-to-end trainable framework to build a high-quality parametric clothed human model from raw scans, (2) a novel weakly-supervised formulation with geometric cycle-consistency that disentangles articulated deformations from the local pose correctives without requiring ground-truth training data, and (3) a locally pose-aware implicit surface representation that models pose-dependent clothing deformation and generalizes to unseen poses. Our results show that SCANimate is superior to existing solutions in terms of generality and accuracy. Furthermore, we perform an extensive study to evaluate the technical contributions that are critical for success. The code and example Scanimats can be found at https://scanimate.is.tue.mpg.de.

Related Work

Parametric Models for Human Bodies and Clothing. Parametric body models learn statistical body shape variations and pose-dependent shape correctives that capture non-linear body deformation and compensate for linear blend skinning artifacts . While these approaches achieve high-fidelity and intuitive control of human body shape and pose, they only focus on bodies without clothing. Similar ideas have been extended to model clothed bodies by introducing additional garment layers or adding displacements or transformations to the base human body mesh . These parametric clothed human models decompose garment deformations into articulated deformations and local deformations such that pose correctives only focus on non-rigid local deformations. Thus, it is essential to obtain the inverse skinning transformation by using the surface registration of a well-defined template or using synthetic simulation data . However, these requirements limit the applicability of the approaches to fairly simple clothing, with a fixed topology, and without complex interactions between garments and the body.

In contrast, our work uses a weakly supervised approach to build a parametric clothed human model from raw scans without the requirement of a template and surface registration. We canonicalize posed scans and learn an implicit surface with arbitrary topology conditioned on pose parameters by leveraging a fitted human body model to the scan data . Moon et al. similarly propose a weakly supervised method for learning a fine-grained hand model from scan data by deforming a fitted base hand model ; the approach is non-trivial to extend to human clothing with varying topology.

The most related work to ours is Neural Articulated Shape Approximation (NASA) , where the composition of occupancy networks articulated by the fitted SMPL model are directly learned from posed scans in the same spirit as structured implicit functions . Concurrent work, LEAP , extends a similar framework to a multi-subject setting. Through an extensive study in Sec. 4.1, we find that the compositional implicit functions proposed in are more prone to artifacts and less generalizable to unseen poses than our LBS-based formulation.

Pose Canonicalization via Inverse LBS. The key to successful canonicalization is learning transformations in the form of skinning weights in a continuous space. Learning skinning weights for varied topologies has become possible using neural networks with graph convolutions . Given a neutral-posed template, these networks predict skinning weights together with a skeleton or pose-dependent deformations . While they predict skinning weights on a neutral-posed template in a fully supervised manner, our problem requires learning skinning weights, not only on the surface mesh, but in both the canonical and posed space without ground-truth skinning weights.

Extending LBS skinning weights from an underlying body model to the continuous space is used in the data preparation step of ARCH and LoopReg . However, in these approaches, the skinning weights are uniquely determined by the underlining body and not learnable. We argue, and experimentally demonstrate, in Sec. 4.1 that jointly learning skinning weights leads to visually pleasing canonicalization while maximizing the reproducibility of input scans by the reconstructed parametric model. Inspired by recent unsupervised methods using cycle consistency , we leverage geometric cycle consistency between the canonical space and posed space to learn skinning weights in a weakly supervised manner without requiring any ground-truth training data. Concurrent work, FTP , proposes a similar idea but is limited to body modeling; instead, we extend the traditional LBS to the entire 3D space and enable clothing surface modeling.

Reconstructing Clothed Humans. Reconstructing humans from depth maps , images , or video is also extensively studied. While many works focus on the minimally clothed human body , recent approaches show promise in reconstructing clothed human models from RGB inputs using the SMPL mesh with displacements , external garment layers , depth maps , voxels , or implicit functions . However, these approaches do not learn, or infer, pose-dependent deformation of garments, and simply apply articulated deformations to the reconstructed shapes. This results in unrealistic pose-dependent deformations that lack garment specific wrinkles. Our work differs by focusing on learning pose-dependent clothing deformation from scans.

Method

Figure 2 shows an overview of our pipeline. The input is a set of raw 3D scans of a person in clothing, together with fitted minimally clothed body models. Here we use the SMPL model fit to the scans to obtain body joints and blend skinning weights, which we exploit in learning. Given the input, we first learn bidirectional transformations between the posed space and canonical space by predicting skinning weights as a function of space coordinates (Sec. 3.1). To address the lack of ground truth correspondence of the scan data, we leverage geometric cycle consistency to learn continuous skinning functions. The raw scans are canonicalized with the learned bidirectional transformations. We further learn a locally pose-aware signed distance function, parameterized by a neural network, from canonicalized scans using implicit geometric regularization (Sec. 3.2). For implementation details, including hyper parameters and network architectures, see Appendix A.

Implicit Skinning Fields. In contrast to traditional applications, where the skinning weights for each point are predefined, either by artists or by automatic methods , skinning weights on the raw scan data are not known a priori. Fortunately, we can learn them in a weakly supervised manner, such that all the scans can be decomposed into articulated deformations and non-rigid deformations.

To this end, we introduce two neural networks called the forward skinning net and the inverse skinning net:

where zis\boldsymbol{z}^{s}_{i} represents a latent embedding, and Θ1\Theta_{1} and Θ2\Theta_{2} are the learnable parameters of the multilayer perceptrons (MLP), which we omit below for notational brevity. The forward skinning net predicts LBS skinning weights of queried 3D locations in the canonical space. Similarly, the inverse skinning net predicts skinning weights in the posed space of each training scan. Notably, this continuous representation is advantageous over other alternatives including fully connected networks and graph convolutional networks as it does not depend on a fixed number of vertices or predefined topology. Empirically we observe that jointly learning zis\boldsymbol{z}^{s}_{i} in an auto-decoding fashion leads to superior performance compared to taking pose parameters as input; see Appendix B for discussion.

By combining Eq. 1 and 2, we can compute the mappings between the canonical and posed spaces via:

Note that these functions are differentiable.

Learning Skinning. To successfully train gc(⋅)g^{c}(\cdot) and gs(⋅)g^{s}(\cdot) without ground truth weights on the scans, we leverage two key observations: (1) the regions close to the human body model are highly correlated with the nearest body parts where ground-truth skinning weights are available; (2) any points in the posed space should be mapped back to the same points after reapplying LBS to the canonicalized points. To utilize (1), we use the underlying SMPL body model’s LBS skinning weights as guidance for the canonical and posed space. More specifically, gs(⋅)g^{s}(\cdot) and gc(⋅)g^{c}(\cdot) at points on the scans are loosely guided by the nearest neighbor point on the body model and its SMPL skinning weights, propagating skinning weights from body models to the input scans.

Most importantly, observation (2) plays a central role in the success of the weakly supervised learning. It allows us to formulate cycle consistency constraints, updating both gc(⋅)g^{c}(\cdot) and gs(⋅)g^{s}(\cdot) such that wrongly associated skinning weights that break the cycle consistency are highly penalized. Our evaluation in Sec. 4.1 shows that the cycle consistency constraints are critical to decompose articulated deformations. Note that the jointly learned gc(⋅)g^{c}(\cdot) is used to learn and animate the pose-aware clothed human model (see Sec. 3.2).

Our final objective function is defined as:

where EBE_{B} and ESE_{S} are body-guided loss functions, ECE_{C} is based on cycle consistency, and ERE_{R} is a regularization term. EBE_{B} ensures gc(⋅)g^{c}(\cdot) and gp(⋅)g^{p}(\cdot) predict SMPL skinning weights on the body surface by

Note that this nearest neighbor assignment is also used in for training data preparation. However, in Sec. 4.1, we show that this alone is prone to inaccurate assignments, causing severe artifacts.

We facilitate cycle consistency with two terms. EC′E_{C^{\prime}} directly constrains the consistency of skinning weights between the canonical space and the posed space, and EC′′E_{C^{\prime\prime}} facilitates cycle consistency on the vertices of the posed meshes as follows:

Notice that cycle consistency can hold only if we start from the posed space since points in the canonical space can be mapped to the same location in case of self-intersection.

Lastly, our regularization term consists of a sparsity constraint ESpE_{Sp}, a smoothness term ESmE_{Sm}, and a statistical regularization on the latent code EZE_{Z} as follows:

where e=(e1,e2)\boldsymbol{e}=(\boldsymbol{e}_{1},\boldsymbol{e}_{2}), E\mathbf{E} is the set of edges on the triangulated scans and we mask out concave regions C\mathbf{C} so that skinning weights are not propagated across merged body parts due to self-intersection ( See Appendix A for details.).

After training, we canonicalize all the scans by applying the inverse LBS transform (Eq. 3) to all vertices on the scans. By eliminating triangles with large distortion (see Appendix A for details), we obtain the canonical scans used to learn a pose-aware parametric clothed human model.

2 Locally Pose-aware Implicit Shape Learning

Given the canonicalized partial scans together with the learned skinning weights, we learn a parametric clothed human model with pose-aware deformations. To this end, we base our shape representation on an implicit surface representation as it supports arbitrary topology with fine details. However, real scans have holes and such partial scans cause difficulty obtaining ground truth occupancy labels since the meshes are not water-tight. To handle partial scans as input, we learn a signed distance function fΦ(x)f_{\Phi}(\boldsymbol{x}) based on a multilayer perceptron (for brevity, we omit the network parameters Φ\Phi), using implicit geometric regularization (IGR) by minimizing the following objective function:

where ELSE_{LS} ensures the zero level-set of the predicted SDF lies on the given points with its surface normal aligned with that of the input scans, n(x)\boldsymbol{n}(\boldsymbol{x}). EIGRE_{IGR} is the Eikonal regularization term that regularizes the function ff to satisfy the Eikonal equation ∥∇xf(⋅)∥=1\left\|\nabla_{\boldsymbol{x}}f(\cdot)\right\|=1. EOE_{O} regularizes off-surface SDF values from being close to the level-set surface as in . Remarkably, this formulation does not require ground truth signed distance for non-surface points and naturally fills in the missing regions by leveraging the inductive bias of multilayer perceptrons as shown in .

where gc(⋅)g^{c}(\cdot) is the skinning network learned in Sec. 3.1, WW is the weight map that converts skinning weights into pose attention weights , and ∘\circ denotes element-wise product. Specifically, if we want a 3D point that is skinned to the nthn^{th} joint with non-zero skinning weights to pay attention to the mthm^{th} joint, Wm,nW_{m,n} and Wn,mW_{n,m} are set to 1, otherwise, they are set to 0. The weight map is essential because the movement of one joint will be propagated to regions associated with neighboring body joints (e.g. raising the shoulders lifts up an entire T-shirt). In this paper, we set Wn,m=1W_{n,m}=1 when nthn^{th} joint is within 44-ring neighbors of mthm^{th} joint in the kinematic tree. By reducing spurious correlations, our formulation significantly reduces over-fitting artifacts given a set of unseen poses, demonstrating better generalization ability even with a small number of input scans (see Sec. 4.1).

Experimental Results

For evaluation and comparison with baseline methods, we use the CAPE dataset , which includes raw 3D scan sequences and SMPL model fits.

We evaluate generalization to unseen poses with both pose interpolation (denoted as Int. in tables) and extrapolation tasks (denoted as Ex. in tables). The motion sequences are randomly split into training (80%) and test (20%) sets, where the test sequences are used to evaluate extrapolation. For the training sequences, we choose every 10th frame starting from the first frame as training scans and every 10th frame with 5 frame strides from the training sequences for the interpolation evaluation. We perform Marching Cubes to the predicted implicit surface in canonical space as in Eq 18 and then pose it by forward LBS in Eq. 1 to get the resulting meshes. For quantitative evaluation, we use scan-to-mesh distance Ds2mD_{s2m} (cm) and surface normal consistency DnD_{n}, where a nearest neighbor vertex on the resulting meshes is used to compute the average L2 norm.

In addition, we conduct a perceptual study to assess the plausibility score, PP, of generated garment shapes and deformations. Workers on Amazon Mechanical Turk (AMT) are given a pair of side-by-side images or videos showing a rendered result from our approach and another approach; the left-right order of the results is randomized. The task is to choose the result with the most realistic clothing. We continue this NN times and compute the probability of the other approach being favored P=M/NP=M/N, where MM is how many times the users chose the other method over ours. In other words, we set our approach as baseline with a constant score P=0.5P=0.5; for other approaches, if P<0.5P<0.5, ours achieves higher fidelity. The perceptual score for image and video pairs is denoted as PiP_{i} and PvP_{v}, respectively. While PiP_{i} focuses on the plausibility of static clothing, PvP_{v} reveals the temporal consistency and realism of pose-dependent clothing deformations. Note that we provide only the perceptual scores for the extrapolation task as numerical evaluation is difficult due to the stochasticity of clothing deformations.

1 Evaluation

Canonicalization. The goal of canonicalization is to disentangle articulated deformations from other non-rigid deformations for effective shape learning. We choose two baseline approaches to replace our canonicalization module. The first, as used by , copies skinning weights on the clothed scans from the nearest neighbor body vertex. The other approach is based on weighted correspondences by interpolating skinning weights from the k-nearest neighbors (we use k=6k=6) in the spirit of . This reduces the impact of a wrong clothing-body association that limits the performance of single nearest neighbor assignment.

Figure 4 shows that the two baseline methods break the cycle consistency with wrong associations of the skinning weights, resulting in noticeable artifacts. The inaccurate canonicalization results are propagated to the parametric model learning, substantially degrading the quality of reconstructed avatars as shown in Tab. 1. Our approach with cycle consistency successfully normalizes the input scans into a canonical pose while retaining coherent geometric details, enabling the parametric modeling of clothed avatars.

Locally Pose-aware Shape Learning. We evaluate our local pose representation using the learned skinning weights for pose-dependent shape learning and compare against commonly used global pose conditioning . To this end, we replace the second input of Eq. 18 with the global pose parameter, θ\theta, as a baseline. To assess the generalization ability, both models are trained on 100%, 50%, 10% and 5% of the original training set.

Table 2 shows that our local pose conditioning achieves better reconstruction accuracy and fidelity for both interpolation and extrapolation. Note that the performance of global pose conditioning drastically degrades when the training data is reduced to less than 10%, suffering from severe overfitting. In contrast, our approach keeps roughly equivalent reconstruction accuracy even when only 5% of the original training data is used, exhibiting few noticeable artifacts (see Fig. 5).

Comparison with SoTA. We compare the proposed method with two state-of-the-art methods that also learn an articulated parametric human model with pose correctives from real world scans . CAPE learns pose-dependent deformations on a fixed mesh topology using graph convolutions , but requires surface registration for training. NASA , on the other hand, can be learned without registration but needs to determine occupancy values. We train both methods using registered CAPE data. Table 3 shows that our approach achieves superior reconstruction accuracy and perceptual realism, while Fig. 6 illustrates limitations of the prior methods. As CAPE relies on a template mesh with a fixed topology, the reconstructions are not only less detailed but also fail to capture topological changes such as the lifting up of the jacket. While NASA can model pose-dependent shapes using articulated implicit functions, discontinuites and ghosting artifacts are visible, as the implicit functions of each body part are learned independently, which limits generalization to unseen poses. In contrast, our approach can produce highly detailed and globally coherent pose-dependent deformations without template-registration.

Learning a Fully Textured Avatar. We extend our pose-aware shape modeling to appearance modeling by predicting texture fields ; see Appendix A for details. Figure 7 shows that high-resolution texture can be modeled without 2D texture mapping, which illustrates another advantage of eliminating the template-mesh requirement.

Discussion and Future Work

We introduced SCANimate, a fully automatic framework to create high-quality avatars (Scanimats), with realistic clothing deformations, driven by pose parameters, that are directly learned from raw 3D scans. Our experiments show that decomposing articulated deformations from scanned data is now possible in a weakly supervised manner by combining body-guided supervision with cycle-consistency regularization. Previously, the difficulty of accurate and coherent surface registration limited the field from analysing and modeling complex clothing deformations involving multiple garments from real-world observations. Our approach enables, for the first time, learning of physically plausible clothing deformations from raw scans, unlocking the possibility of realistic avatar learning from data.

Limitations and Future Work. The current representation works well for clothing that is topologically similar to the body. The method may fail for clothing, like skirts, that deviates significantly from the body; see Appendix B for an example. Clothing wrinkles tend to be stochastic; that is, for a specific pose, they may differ depending on the preceding sequence of poses. The current model, however, is deterministic. Future work should factor the surface texture into albedo, shape, and lighting enabling more realistic relighting of Scanimats. Additionally, an adversarial texture loss could improve visual quality. Here we model a person in a single garment. Learning a generative model with clothing variety should be possible but will require training data of varied clothing in varied poses. Most exciting is the idea of fitting Scanimats to, or even learning them from, images or videos. Finally, extending this approach to model hand articulation and facial expressions should be possible using expressive body models like SMPL-X .

Acknowledgements: Q. Ma was partially funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 276693517 SFB 1233. Disclosure: MJB has received research gift funds from Adobe, Intel, Nvidia, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, Max Planck. MJB has financial interests in Amazon, Datagen Technologies, and Meshcapade GmbH.

References

Appendix A Implementation Details

A.2 Training Procedure

Our training consists of three stages. First, we pretrain gΘ1c(⋅)g^{c}_{\Theta_{1}}(\cdot) and gΘ2s(⋅)g^{s}_{\Theta_{2}}(\cdot) with the following relative weights: λB=10.0\lambda_{B}=10.0, λS=1.0\lambda_{S}=1.0, λC′=0.0\lambda_{C^{\prime}}=0.0, λC′′=0.0\lambda_{C^{\prime\prime}}=0.0, λSp=0.001\lambda_{Sp}=0.001, λSm=0.0\lambda_{Sm}=0.0, and λZ=0.01\lambda_{Z}=0.01. After pretraining, we jointly train gΘ1c(⋅)g^{c}_{\Theta_{1}}(\cdot) and gΘ2s(⋅)g^{s}_{\Theta_{2}}(\cdot) using the proposed cycle consistency constraint with the following weights: λB=10.0\lambda_{B}=10.0, λS=1.0\lambda_{S}=1.0, λC′=1.0\lambda_{C^{\prime}}=1.0, λC′′=1.0\lambda_{C^{\prime\prime}}=1.0, λSp=0.001\lambda_{Sp}=0.001, λSm=0.1\lambda_{Sm}=0.1, and λZ=0.01\lambda_{Z}=0.01. We multiply λC′′\lambda_{C^{\prime\prime}} by 1010 for the second half of the training iterations. For the two stages above, we use 68906890 points of the SMPL vertices and 80008000 points uniformly sampled on the scan data, which is dynamically updated at every iteration.

Once the training of the skinning networks is complete, we fix the weights of gΘ1c(⋅)g^{c}_{\Theta_{1}}(\cdot), gΘ2s(⋅)g^{s}_{\Theta_{2}}(\cdot), and {zis}\left\{\boldsymbol{z}^{s}_{i}\right\}, and train the geometry module fΦ(⋅)f_{\Phi}(\cdot) with the following hyper parameters: λigr=1.0\lambda_{igr}=1.0, λo=0.1\lambda_{o}=0.1, and α=100\alpha=100. To compute ELSE_{LS}, we uniformly sample 50005000 points on the scan surface at each iteration. We compute EIGRE_{IGR} by combining 20002000 points within a bounding box and 1000010000 points perturbed with the standard deviation of 1010cm from the surface geometry, half of which is sampled from the scans and the remaining from the SMPL body vertices. Note that EOE_{O} uses only 20002000 points sampled from the bounding box to avoid overly penalizing zero crossing near the surface.

We train each stage with the Adam optimizer with learning rates of 0.0040.004, 0.0010.001, and 0.0010.001, respectively. They are decayed by the factor of 0.10.1 at 1/21/2 and 3/43/4 of the training iterations. The first stage runs for 80 epochs and the second for 200 epochs.

A.3 Texture Inference

A.4 Other Details

We exclude concave regions from the smoothness constraint to avoid propagating incorrect skinning weights at the self-intersection regions. We detect them by computing the mean curvature on the surface of scans with the threshold of 0.20.2. Note that while we empirically find our detection algorithm is sufficient for our training data, utilizing external information such as body part labels is possible when available to improve robustness.

Obtaining Canonical Body.

The canonicalized body Bic\mathbf{B}^{c}_{i} in Eq. 5 is a body model of the subject in a canonical pose with pose dependent deformations. We obtain the pose correctives by activating pose-aware blend shapes in the SMPL model given the body pose θ\theta at frame ii.

Removing Distorted Triangles.

When the input scans are canonicalized, triangle edges that belong to self-intersection regions are highly distorted. As these regions must be separated in the canonical pose, we remove all triangles for which any edge length is larger than its initial length multiplied by 44.

Appendix B Discussions

We provide additional discussion to clarify technical details in the main text, including choice of latent autodecoding and the similarity of forward and inverse skinning networks, and discussion over failure cases.

The purpose of learning gs(⋅,z)g^{s}(\cdot,\boldsymbol{z}) is to stably canonicalize raw scans. To this end, we use auto-decoding z\boldsymbol{z} as in for the following advantages. Auto-decoding self-discovers the latent embedding z\boldsymbol{z} such that the loss function is minimized, allowing the network to better distinguish each scan regardless of the similarity in the pose parameters. Thus, z\boldsymbol{z} can implicitly encode not only pose information but also anything necessary to distinguish each frame. Furthermore, due to no dependency on pose parameters, auto-decoding is more robust to the fitting error of the underlying body model. As a baseline we replace autodecoding by regressing skinning weights on pose parameters of a fitted SMPL body. We use the energy function EcanoE_{cano} in Eq. 4 without the term of EZE_{Z} to evaluate the performance of the two. While pose regression results in 0.043, autodecoding achieves a much lower local minimum at 0.025, showing superior performance against the baseline.

B.2 CAPE Dataset Limitation

Some frames of the CAPE dataset contain erroneous body fitting around the wrists and ankles, as shown in the right inset figure, resulting in unnecessary distortions around the regions. Due to the smoothness regularization in our method, such a distortion can be propagated to the nearby regions, and hence a larger region may be discarded. However, the proposed shape learning method complements such a missing region from other canonicalized scans, and our reconstructed Scanimats do not suffer from the small errors in pose fitting.

B.3 Combining Skinning Networks

As in Eq. 2, gcg^{c} and gsg^{s} are formulated separately. This is in accordance with the idea of predicting skinning weights for both forward and backward transformations. However, if one considers the skinning networks in another point of view, particularly when regarding them as mappings from 3D space coordinates conditioned on different frames to skinning weights, it is clear that gcg^{c} is a special case of gsg^{s}. Thus in practical implementation, one can either set up two networks corresponding to gcg^{c} and gsg^{s}, respectively, or set up a single networks in an autodecoder manner with a single common latent vector zc\boldsymbol{z}^{c} for all the forward skinning weights prediction and per-frame latent vectors zis\boldsymbol{z}^{s}_{i} for inverse skinning weights prediction in each posed frame.

B.4 Failure Cases

As mentioned in the main paper, while the current pipeline performs well for clothing that is topologically similar to the body, the method may fail for clothing, like skirts, whose topology may deviate significantly. Fig. B.1 shows a failure case of canonicalizing a person with a skirt synthetically generated using a physics-based simulation. The SMPL-guided initialization of skinning weights fails recovering from poor local minima. We leave for future work a garment-specific tuning of hyperparameters and more robust training schemes for various clothing types.

Appendix C Additional Qualitative Results

Fig. C.1, an extended figure of Fig. 5 in the main paper, shows more qualitative comparison on pose encoding with different sizes of training data.

Comparison with the SoTA methods.

Fig. C.2, an extended figure of the main paper Fig. 6, shows more qualitative comparison with the SoTA methods.

Textured Scanimats.

Fig. C.3, an extended figure of Fig. 1, shows more examples of textured Scanimats.