ICON: Implicit Clothed humans Obtained from Normals

Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. Black

Introduction

Realistic virtual humans will play a central role in mixed and augmented reality, forming a key foundation for the “metaverse” and supporting remote presence, collaboration, education, and entertainment. To enable this, new tools are needed to easily create 3D virtual humans that can be readily animated. Traditionally, this requires significant artist effort and expensive scanning equipment. Therefore, such approaches do not scale easily. A more practical approach would enable individuals to create an avatar from one or more images. There are now several methods that take a single image and regress a minimally clothed 3D human model . Existing parametric body models, however, lack important details like clothing and hair . In contrast, we present a method that robustly extracts 3D scan-like data from images of people in arbitrary poses and uses this to construct an animatable avatar.

We base our approach on implicit functions (IFs), which go beyond parametric body models to represent fine shape details and varied topology. IFs allow recent methods to infer detailed shape from an image . Despite promising results, state-of-the-art (SOTA) methods struggle with in-the-wild data and often produce humans with broken or disembodied limbs, missing details, high-frequency noise, or non-human shape; see Fig. 2 for examples.

The issues with previous methods are twofold: (1) Such methods are typically trained on small, hand-curated, 3D human datasets (e.g. Renderpeople ) with very limited pose, shape and clothing variation. (2) They typically feed their implicit-function module with features of a global 2D image or 3D voxel encoder, but these are sensitive to global pose. While more, and more varied, 3D training data would help, such data remains limited. Hence, we take a different approach and improve the model.

Specifically, our goal is to reconstruct a detailed clothed 3D human from a single RGB image with a method that is training-data efficient and robust to in-the-wild images and out-of-distribution poses. Our method, called ICON, stands for Implicit Clothed humans Obtained from Normals. ICON replaces the global encoder of existing methods with a more data-efficient local scheme; Fig. 3 shows a model overview. ICON takes as input an RGB image of a segmented clothed human and a SMPL body estimated from the image . The SMPL body is used to guide two of ICON’s modules: one infers detailed clothed-human surface normals (front and back views), and the other infers a visibility-aware implicit surface (iso-surface of an occupancy field). Errors in the initial SMPL estimate, however, might misguide inference. Thus, at inference time, an iterative feedback loop refines SMPL (i.e., its 3D shape, pose, and translation) using the inferred detailed normals, and vice versa, leading to a refined implicit shape with better 3D details.

We evaluate ICON quantitatively and qualitatively on challenging datasets, namely AGORA and CAPE , as well as on in-the-wild images. Results show that ICON has two advantages w.r.t. the state of the art: (1) Generalization. ICON’s locality helps it generalize to in-the-wild images and out-of-distribution poses and clothes better than previous methods. Representative cases are shown in Fig. 2; notice that, although ICON is trained on full-body images only, it can handle images with out-of-frame cropping, with no fine tuning or post processing. (2) Data efficacy. ICON’s locality helps it avoid spurious correlations between pose and surface shape. Thus, it needs less data for training. ICON significantly outperforms baselines in low-data regimes, as it reaches SOTA performance when trained with as little as 12%12\% of the data.

We provide an example application of ICON for creating an animatable avatar; see ICON: Implicit Clothed humans Obtained from Normals for an overview. We first apply ICON on the individual frames of a video sequence, to obtain 3D meshes of a clothed person in various poses. We then use these to train a poseable avatar using a modified version of SCANimate . Unlike 3D scans, which SCANimate takes as input, our estimated shapes are not equally detailed and reliable from all views. Consequently, we modify SCANimate to exploit visibility information in learning the avatar. The output is a 3D clothed avatar that moves and deforms naturally; see ICON: Implicit Clothed humans Obtained from Normals-right and Fig. 7(b).

ICON takes a step towards robust reconstruction of 3D clothed humans from in-the-wild photos. Based on this, fully textured and animatable avatars with personalized pose-aware clothing deformation can be created directly from video frames. Models and code are available at https://icon.is.tue.mpg.de.

Related work

Mesh-based statistical models. Mesh-based statistical body models are a popular explicit representation for 3D human reconstruction. This is not only because such models capture the statistics across a human population, but also because meshes are compatible with standard graphics pipelines. A lot of work estimates 3D body meshes from an RGB image, but these have no clothing. Other work estimates clothed humans, instead, by modeling clothing geometry as 3D offsets on top of body geometry . The resulting clothed 3D humans can be easily animated, as they naturally inherit the skeleton and surface skinning weights from the underlying body model. An important limitation, though, is modeling clothing such as skirts and dresses; since these differ a lot from the body surface, simple body-to-cloth offsets are insufficient. To address this, some methods use a classifier to identify cloth types in the input image, and then perform cloth-aware inference for 3D reconstruction. However, such a remedy does not scale up to a large variety of clothing types. Another advantage of mesh-based statistical models, is that texture information can be easily accumulated through multi-view images or image sequences , due to their consistent mesh topology. The biggest limitation, though, is that the state of the art does not generalize well w.r.t. clothing-type variation, and it estimates meshes that do not align well to input-image pixels.

Deep implicit functions. Unlike meshes, deep implicit functions can represent detailed 3D shapes with arbitrary topology, and have no resolution limitations. Saito et al. introduce deep implicit functions for clothed 3D human reconstruction from RGB images and, later , they significantly improve 3D geometric details. The estimated shapes align well to image pixels. However, their shape reconstruction lacks regularization, and often produces artifacts like broken or disembodied limbs, missing details, or geometric noise. He et al. add a coarse-occupancy prediction branch, and Li et al. and Dong et al. use depth information captured by an RGB-D camera to further regularize shape estimation and provide robustness to pose variation. Li et al. speed up inference through an efficient volumetric sampling scheme. A limitation of all above methods is that the estimated 3D humans cannot be reposed, because implicit shapes (unlike statistical models) lack a consistent mesh topology, a skeleton, and skinning weights. To address this, Bozic et al. infer an embedded deformation graph to manipulate implicit functions, while Yang et al. also infer a skeleton and skinning fields.

Statistical models & implicit functions. Mesh-based statistical models are well regularized, while deep implicit functions are much more expressive. To get the best of both worlds, recent methods combine the two representations. Given a sparse point cloud of a clothed person, IPNet infers an occupancy field with body/clothing layers, registers SMPL to the body layer with inferred body-part segmentation, and captures clothing as offsets from SMPL to the point cloud. Given an RGB image of a clothed person, ARCH and ARCH++ reconstruct 3D human shape in a canonical space by warping query points from the canonical to the posed space, and projecting them onto the 2D image space. However, to train these models, one needs to unpose scans into the canonical pose with an accurately fitted body model; inaccurate poses cause artifacts. Moreover, unposing clothed scans using the “undressed” model’s skinning weights alters shape details. For the same RGB input, Zheng et al. condition the implicit function on a posed and voxelized SMPL mesh for robustness to pose variation and reconstruct local details from the image pixels, similar to PIFu . However, these methods are sensitive to global pose, due to their 3D convolutional encoder. Thus, for training data with limited pose variation, they struggle with out-of-distribution poses and in-the-wild images.

Positioning ICON w.r.t. related work. ICON combines the statistical body model SMPL with an implicit function, to reconstruct clothed 3D human shape from a single RGB image. SMPL not only guides ICON’s estimation, but is also optimized “in the loop” during inference to enhance its pose accuracy. Instead of relying on the global body features, ICON exploits local body features that are agnostic to global pose variations. As a result, even when trained on heavily limited data, ICON achieves state-of-the-art performance and is robust to out-of-distribution poses. This work links monocular 3D clothed human reconstruction to scan/depth based avatar modeling algorithms .

Method

ICON is a deep-learning model that infers a 3D clothed human from a color image. Specifically, ICON takes as input an RGB image with a segmented clothed human (following the suggestion of PIFuHD’s repository ), along with an estimated human body shape “under clothing” (SMPL), and outputs a pixel-aligned 3D shape reconstruction of the clothed human. ICON has two main modules (see Fig. 3) for: (1) SMPL-guided clothed-body normal prediction and (2) local-feature based implicit surface reconstruction.

We train the normal networks, GN\mathcal{G}^{\text{N}}, with the following loss:

where Lpixel=∣Nvc−N^vc∣\mathcal{L}_{\text{pixel}}=|\mathcal{N}^{\text{c}}_{\text{v}}-\widehat{\mathcal{N}}^{\text{c}}_{\text{v}}|, v={front,back}\text{v}=\{\text{front},\text{back}\}, is a loss (L1) between ground-truth and predicted normals (the two GN\mathcal{G}^{\text{N}} in Fig. 3 have different parameters), and LVGG\mathcal{L}_{\text{VGG}} is a perceptual loss weighted by λVGG\lambda_{\text{VGG}}. With only Lpixel\mathcal{L}_{\text{pixel}}, the inferred normals are blurry, but adding LVGG\mathcal{L}_{\text{VGG}} helps recover details.

Refining SMPL. Intuitively, a more accurate SMPL body fit provides a better prior that helps infer better clothed-body normals. However, in practice, human pose and shape (HPS) regressors do not give pixel-aligned SMPL fits. To account for this, during inference, the SMPL fits are optimized based on the difference between the rendered SMPL-body normal maps, Nb\mathcal{N}^{\text{b}}, and the predicted clothed-body normal maps, N^c\widehat{\mathcal{N}}^{\text{c}}, as shown in Fig. 4. Specifically we optimize over SMPL’s shape, β\beta, pose, θ\theta, and translation, tt, parameters to minimize:

where LN_diff\mathcal{L}_{\text{N}\text{\_diff}} is a normal-map loss (L1), weighted by λN_diff\lambda_{\text{N\_diff}}; LS_diff\mathcal{L}_{\text{S}\text{\_diff}} is a loss (L1) between the silhouettes of the SMPL body normal-map Sb\mathcal{S}^{\text{b}} and the human mask S^c\widehat{\mathcal{S}}^{\text{c}} segmented from I\mathcal{I}. We ablate LN_diff\mathcal{L}_{\text{N}\text{\_diff}}, LS_diff\mathcal{L}_{\text{S}\text{\_diff}} in Appx

Refining normals. The normal maps rendered from the refined SMPL mesh, Nb\mathcal{N}^{\text{b}}, are fed to the GN\mathcal{G}^{\text{N}} networks. The improved SMPL-mesh-to-image alignment guides GN\mathcal{G}^{\text{N}} to infer more reliable and detailed normals N^c\widehat{\mathcal{N}}^{\text{c}}.

Refinement loop. During inference, ICON alternates between: (1) refining the SMPL mesh using the inferred N^c\widehat{\mathcal{N}}^{\text{c}} normals and (2) re-inferring N^c\widehat{\mathcal{N}}^{\text{c}} using the refined SMPL. Experiments show that this feedback loop leads to more reliable clothed-body normal maps for both (front/back) sides.

2 Local-feature based implicit 3D reconstruction

Given the predicted clothed-body normal maps, N^c\widehat{\mathcal{N}}^{\text{c}}, and the SMPL-body mesh, M\mathcal{M}, we regress the implicit 3D surface of a clothed human based on local features FP\mathcal{F}_{\text{P}}:

where Fs\mathcal{F}_{\text{s}} is the signed distance from a query point P to the closest body point Pb∈M\text{P}^{\text{b}}\in\mathcal{M}, and Fnb\mathcal{F}_{\text{n}}^{\text{b}} is the barycentric surface normal of Pb\text{P}^{\text{b}}; both provide strong regularization against self occlusions. Finally, Fnc\mathcal{F}_{\text{n}}^{\text{c}} is a normal vector extracted from N^frontc\widehat{\mathcal{N}}^{\text{c}}_{\text{front}} or N^backc\widehat{\mathcal{N}}^{\text{c}}_{\text{back}} depending on the visibility of Pb\text{P}^{\text{b}}:

where π(P)\pi(\text{P}) denotes the 2D projection of the 3D point P.

Please note that FP\mathcal{F}_{\text{P}} is independent of global body pose. Experiments show that this is key for robustness to out-of-distribution poses and efficacy w.r.t. training data.

We feed FP\mathcal{F}_{\text{P}} into an implicit function, IF\mathcal{IF}, parameterized by a Multi-Layer Perceptron (MLP) to estimate the occupancy at point P, denoted as o^(P)\widehat{o}(\text{P}). A mean squared error loss is used to train IF\mathcal{IF} with ground-truth occupancy, o(P)o(\text{P}). Then the fast surface localization algorithm is used to extract meshes from the 3D occupancy inferred by IF\mathcal{IF}.

Experiments

We compare ICON primarily with PIFu and PaMIR . These methods differ from ICON and from each other w.r.t. the training data, the loss functions, the network structure, the use of the SMPL body prior, etc. To isolate and evaluate each factor, we re-implement PIFu and PaMIR by “simulating” them based on ICON’s architecture. This provides a unified benchmarking framework, and enables us to easily train each baseline with the exact same data and training hyper-parameters for a fair comparison. Since there might be small differences w.r.t. the original models, we denote the “simulated” models with a “star” as:

PIFu∗\text{PIFu}^{*} : {f2D(I,N)}\{f_{\text{2D}}(\mathcal{I},\mathcal{N})\} →O\rightarrow\mathcal{O},

PaMIR∗\text{PaMIR}^{*} : {f2D(I,N),f3D(V)}\{f_{\text{2D}}(\mathcal{I},\mathcal{N}),f_{\text{3D}}(\mathcal{V})\} →O\rightarrow\mathcal{O},

ICON : {N,γ(M)}\{\mathcal{N},\gamma(\mathcal{M})\} →O\rightarrow\mathcal{O},

where f2Df_{\text{2D}} denotes the 2D image encoder, f3Df_{\text{3D}} denotes the 3D voxel encoder, V\mathcal{V} denotes the voxelized SMPL, O\mathcal{O} denotes the entire predicted occupancy field, and γ\gamma is the mesh-based local feature extractor described in Sec. 3.2. The results are summarized in Tab. 2-A, and discussed in Sec. 4.3-A. For reference, we also report the performance of the original PIFu , PIFuHD , and PaMIR ; our “simulated” models perform well, and even outperform the original ones.

2 Datasets

Several public or commercial 3D clothed-human datasets are used in the literature, but each method uses different subsets and combinations of these, as shown in Tab. 1.

Training data. To compare models fairly, we factor out differences in training data as explained in Sec. 4.1. Following previous work , we retrain all baselines on the same 450450 Renderpeople scans (subset of AGORA). Methods that require the 3D body prior (i.e., PaMIR, ICON) use the SMPL-X meshes provided by AGORA. ICON’s GN\mathcal{G}^{\text{N}} and IF\mathcal{IF} modules are trained on the same data.

Testing data. We evaluate primarily on CAPE , which no method uses for training, to test their generelizability. Specifically, we divide the CAPE dataset into the “CAPE-FP” and “CAPE-NFP” sets that have “fashion” and “non-fashion” poses, respectively, to better analyze the generalization to complex body poses; for details on data splitting please see Appx To evaluate performance without a domain gap between train/test data, we also test all models on “AGORA-50” , which contains 5050 samples from AGORA that are different from the 450450 used for training.

Generating synthetic data. We use the OpenGL scripts of MonoPort to render photo-realistic images with dynamic lighting. We render each clothed-human 3D scan (I\mathcal{I} and Nc\mathcal{N}^{\text{c}}) and their SMPL-X fits (Nb\mathcal{N}^{\text{b}}) from multiple views by using a weak perspective camera and rotating the scan in front of it. In this way we generate 138,924138,924 samples, each containing a 3D clothed-human scan, its SMPL-X fit, an RGB image, camera parameters, 2D normal maps for the scan and the SMPL-X mesh (from two opposite views) and SMPL-X triangle visibility information w.r.t. the camera.

3 Evaluation

We use 3 evaluation metrics, described in the following: “Chamfer” distance. We report the Chamfer distance between ground-truth scans and reconstructed meshes. For this, we sample points uniformly on scans/meshes, to factor out resolution differences, and compute average bi-directional point-to-surface distances. This metric captures large geometric differences, but misses smaller geometric details.

“P2S” distance. CAPE has raw scans as ground truth, which can contain large holes. To factor holes out, we additionally report the average point-to-surface (P2S) distance from scan points to the closest reconstructed surface points. This metric can be viewed as a 1-directional version of the above metric.

“Normals” difference. We render normal images for reconstructed and ground-truth surfaces from fixed viewpoints (Sec. 4.2, “generating synthetic data”), and calculate the L2 error between them. This captures errors for high-frequency geometric details, when Chamfer and P2S errors are small.

A. ICON -vs- SOTA. ICON outperforms all original state-of-the-art (SOTA) methods, and is competitive to our “simulated” versions of them, as shown in Tab. 2-A. We use AGORA’s SMPL-X ground truth (GT) as a reference. We notice that our re-implemented PaMIR∗\text{PaMIR}^{*} outperform the SMPL-X GT for images with in-distribution body poses (“AGORA-50” and “CAPE-FP”), However, this is not the case for images with out-of-distribution poses (“CAPE-NFP”). This shows that, although conditioned on GT SMPL-X fits, PaMIR∗\text{PaMIR}^{*} is still sensitive to global body pose due to its global feature encoder, and fails to generalize to out-of-distribution poses. On the contrary, ICON generalizes well to out-of-distribution poses, because its local features are independent from global pose (see Sec. 3.2).

B. Body-guided normal prediction. We evaluate the conditioning on SMPL-X-body normal maps, Nb\mathcal{N}^{\text{b}}, for guiding inference of clothed-body normal maps, N^c\widehat{\mathcal{N}}^{\text{c}} (Sec. 3.1). Table 2-B shows performance with (“ICON”) and without (“ICONN†\text{ICON}_{\text{N}^{\dagger}}”) conditioning. With no conditioning, errors on “CAPE” increase slightly. Qualitatively, guidance by body normals heavily improves the inferred normals, especially for occluded body regions; see Fig. 5. We also ablate the effect of the body-normal feature (Sec. 3.2), Fnb\mathcal{F}_{\text{n}}^{\text{b}}, by removing it; this worsens results, see “ICON w/o Fnb\mathcal{F}_{\text{n}}^{\text{b}}” in Tab. 2-B.

C. Local-feature based implicit reconstruction. To evaluate the importance of our “local” features (Sec. 3.2), FP\mathcal{F}_{\text{P}}, we replace them with “global” features produced by 2D convolutional filters. These are applied on the image and the clothed-body normal maps (“ICONenc(I,N^c)\text{ICON}_{\text{enc}(\mathcal{I},\widehat{\mathcal{N}}^{\text{c}})}” in Tab. 2-C), or only on the normal maps (“ICONenc(N^c)\text{ICON}_{\text{enc}(\widehat{\mathcal{N}}^{\text{c}})}” in Tab. 2-C). We use a 2-stack hourglass model , whose receptive field expands to 46%46\% of the image size. This takes a large image area into account and produces features sensitive to global body pose. This worsens reconstruction performance for out-of-distribution poses, such as in “CAPE-NFP”. For an evaluation of PaMIR’s receptive field size, see Appx

We compare ICON to state-of-the-art (SOTA) models for a varying amount of training data in Fig. 6. The “Dataset scale” axis reports the data size as the ratio w.r.t. the 450450 scans of the original PIFu methods ; the left-most side corresponds to 5656 scans and the right-most side corresponds to 3,7093,709 scans, i.e., all the scans of AGORA and THuman . ICON consistently outperforms all methods. Importantly, ICON achieves SOTA performance even when trained on just a fraction of the data. We attribute this to the local nature of ICON’s point features; this helps ICON generalize well in the pose space and be data efficient.

D. Robustness to SMPL-X noise. SMPL-X estimated from an image might not be perfectly aligned with body pixels in the image. However, PaMIR and ICON are conditioned on this estimation. Thus, they need to be robust against various noise levels in SMPL-X shape and pose. To evaluate this, we feed PaMIR∗\text{PaMIR}^{*} and ICON with ground-truth and perturbed SMPL-X, denoted with (✓) and (✓) in Tab. 2-A,D. ICON conditioned on perturbed (✓) SMPL-X produces larger errors w.r.t. conditioning on ground truth (✓). However, adding the body refinement module (“ICON +BR”) of Sec. 3.1, refines SMPL-X and improves performance. As a result, “ICON +BR” conditioned on noisy SMPL-X (✓) performs comparably to PaMIR∗\text{PaMIR}^{*} conditioned on ground-truth SMPL-X (✓); it is slightly worse/better for in-/out-of-distribution poses.

Applications

We collect 200200 in-the-wild images from Pinterest that show people performing parkour, sports, street dance, and kung fu. These images are unseen during training. We show qualitative results for ICON in Fig. 7(a) and comparisons to SOTA in Fig. 2; for more results see our video and Appx

To evaluate the perceived realism of our results, we compare ICON to PIFu∗\text{PIFu}^{*}, PaMIR∗\text{PaMIR}^{*}, and the original PIFuHD in a perceptual study. ICON, PIFu∗\text{PIFu}^{*} and PaMIR∗\text{PaMIR}^{*} are trained on all 3,7093,709 scans of AGORA and THuman (“8x” setting in Fig. 6). For PIFuHD we use its pre-trained model. In the study, participants were shown an image and either a rendered result of ICON or of another method. Participants were asked to choose the result that best represents the shape of the human in the image. We report the percentage of trails in which participants preferred the baseline methods over ICON in Tab. 3; p-values correspond to the null-hypothesis that two methods perform equally well. For details on the study, example stimuli, catch trials, etc. see Appx

2 Animatable avatar creation from video

Given a sequence of images with the same subject in various poses, we create an animatable avatar with the help of SCANimate . First, we use ICON to reconstruct a 3D clothed-human mesh per frame. Then, we feed these meshes to SCANimate. ICON’s robustness to diverse poses enables us to learn a clothed avatar with pose-dependent clothing deformation. Unlike raw 3D scans, which are taken with multi-view systems, ICON operates on a single image and its reconstructions are more reliable for observed body regions than for occluded ones. Thus, we reformulate the loss of SCANimate to downweight occluded regions depending on camera viewpoint. Results are shown in ICON: Implicit Clothed humans Obtained from Normals and Fig. 7(b); for animations see the video on our webpage.

Conclusion

We have presented ICON, which robustly recovers a 3D clothed person from a single image with accuracy and realism that exceeds prior art. There are two keys: (1) Regularizing the solution with a 3D body model while optimizing that body model iteratively. (2) Using local features to eliminate spurious correlations with global pose. Thorough ablation studies validate these choices. The quality of results is sufficient to build a 3D avatar from monocular image sequences.

Limitations and future work. Due to the strong body prior exploited by ICON, loose clothing that is far from the body may be difficult to reconstruct; see Fig. 7. Although ICON is robust to small errors of body fits, significant failure of body fits leads to reconstruction failure. Because it is trained on orthographic views, ICON has trouble with strong perspective effects, producing asymmetric limbs or anatomically improbable shapes. A key future application is to use images alone to create a dataset of clothed avatars. Such a dataset could advance research in human shape generation , be valuable to fashion industry, and facilitate graphics applications.

Possible negative impact. While the quality of virtual humans created from images is not at the level of facial “deep fakes”, as this technology matures, it will open up the possibility for full-body deep fakes, with all the attendant risks. These risks must also be balanced by the positive use cases in entertainment, tele-presence, and future metaverse applications. Clearly regulation will be needed to establish legal boundaries for its use. In lieu of societal guidelines today, we have made our code available with an appropriate license.

Disclosure. https://files.is.tue.mpg.de/black/CoI_CVPR_2022.txt

Acknowledgments. We thank Yao Feng, Soubhik Sanyal, Hongwei Yi, Qianli Ma, Chun-Hao Paul Huang, Weiyang Liu, and Xu Chen for their feedback and discussions, Tsvetelina Alexiadis for her help with perceptual study, Taylor McConnell assistance, Benjamin Pellkofer for webpage, and Yuanlu Xu’s help in comparing with ARCH and ARCH++. This project has received funding from the European Union’s Horizon 20202020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No.860768860768 (CLIPE project).

We provide more details for the method and experiments, as well as more quantitative and qualitative results, as an extension of Sec. 3, Sec. 4 and Sec. 5 of the main paper.

Appendix A Method & Experiment Details

Dataset size. We evaluate the performance of ICON and SOTA methods for a varying training-dataset size (Figs. 6 and 9). For this, we first combine AGORA (3,1093,109 scans) and THuman (600600 scans) to get 3,7093,709 scans in total. This new dataset is 88x times larger than the 450450 Renderpeople (“450450-Rp”) scans used in . Then, we sample this “88x dataset” to create smaller variations, for 1/81/8x, 1/41/4x, 1/21/2x, 11x, and 88x the size of “450450-Rp”.

Dataset splits. For the “88x dataset”, we split the 3,1093,109 AGORA scans into a new training set (3,0343,034 scans), validation set (2525 scans) and test set (5050 scans). Among these, 1,8471,847 come from Renderpeople (see Fig. 9(a)), 622622 from AXYZ , 242242 from Humanalloy , 398398 from 3DPeople , and we sample only 600600 scans from THuman (see Fig. 9(b)), due to its high pose repeatability and limited identity variants (see Tab. 1), with the “select-cluster” scheme described below. These scans, as well as their SMPL-X fits, are rendered after every 1010 degrees rotation around the yaw axis, to totally generate (3109 \mboxAGORA+600 \mboxTHuman+150 \mboxCAPE)×36=138,924(3109\text{~{}\mbox{AGORA}}+600\text{~{}\mbox{THuman}}+150\text{~{}\mbox{CAPE}})\times 36=138,924 samples.

Dataset distribution via “select-cluster” scheme. To create a training set with a rich pose distribution, we need to select scans from various datasets with poses different from AGORA. Following SMPLify , we first fit a Gaussian Mixture Model (GMM) with 88 components to all AGORA poses, and select 2K THuman scans with low likelihood. Then, we apply M-Medoids (n_cluster=50\text{n\_cluster}=50) on these selections for clustering, and randomly pick 1212 scans per cluster, collecting 50×12=60050\times 12=600 THuman scans in total; see Fig. 9(b). This is also used to split CAPE into “CAPE-FP” (Fig. 9(c)) and “CAPE-NFP” (Fig. 9(d)), corresponding to scans with poses similar (in-distribution poses) and dissimilar (out-of-distribution poses) to AGORA ones, respectively.

Perturbed SMPL. To perturb SMPL’s pose and shape parameters, random noise is added to θ and β\theta\text{ and }\beta by:

where μ∈\mu\in, sθ=0.15s_{\theta}=0.15 and sβ=0.5s_{\beta}=0.5. These are set empirically to mimic the misalignment error typically caused by off-the-shell HPS during testing.

Discussion on simulated data. The wide and loose clothing in CLOTH3D++ demonstrates strong dynamics, which would complement commonly used datasets of commercial scans. Yet, the domain gap between CLOTH3D++ and real images is still large. Moreover, it is unclear how to train an implicit function from multi-layer non-watertight meshes. Consequently, we leave it for future research.

A.2 Refining SMPL (Sec. 3.1)

To statistically analyze the necessity of LN_diff\mathcal{L}_{\text{N\_diff}} and LS_diff\mathcal{L}_{\text{S\_diff}} in Eq. 4, we do a sanity check on AGORA’s validation set. Initialized with different pose noise, sθs_{\theta} (Eq. 8), we optimize the {θ,β,t}\{\theta,\beta,t\} parameters of the perturbed SMPL by minimizing the difference between rendered SMPL-body normal maps and ground-truth clothed-body normal maps for 2K iterations. As Fig. 9 shows, LN_diff+LS_diff\mathcal{L}_{\text{N\_diff}}+\mathcal{L}_{\text{S\_diff}} always leads to the smallest error under any noise level, measured by the Chamfer distance between the optimized perturbed SMPL mesh and the ground-truth SMPL mesh.

A.3 Perceptual study (Tab. 3)

Reconstruction on in-the-wild images. We perform a perceptual study to evaluate the perceived realism of the reconstructed clothed 3D humans from in-the-wild images. ICON is compared against 33 methods, PIFu , PIFuHD , and PaMIR . We create a benchmark of 200200 unseen images downloaded from the internet, and apply all the methods on this test set. All the reconstruction results are evaluated on Amazon Mechanical Turk (AMT), where each participant is shown pairs of reconstructions from ICON and one of the baselines, see Fig. 11. Each reconstruction result is rendered in four views: front, right, back and left. Participants are asked to choose the reconstructed 3D shape that better represents the human in the given color image. Each participant is given 100100 samples to evaluate. To teach participants, and to filter out the ones that do not understand the task, we set up 11 tutorial sample, followed by 1010 warm-up samples, and then the evaluation samples along with catch trial samples inserted every 1010 evaluation samples. Each catch trial sample shows a color image along with either (1) the reconstruction of a baseline method for this image and the ground-truth scan that was rendered to create this image, or (2) the reconstruction of a baseline method for this image and the reconstruction for a different image (false positive), see Fig. 10(c). Only participants that pass 70%70\% out of 1010 catch trials are considered. This leads to 2828 valid participants out of 3636 ones. Results are reported in Tab. 3.

Normal map prediction. To evaluate the effect of the body prior for normal map prediction on in-the-wild images, we conduct a perceptual study against prediction without the body prior. We use AMT, and show participants a color image along with a pair of predicted normal maps from two methods. Participants are asked to pick the normal map that better represents the human in the image. Front- and back-side normal maps are evaluated separately. See Fig. 12 for some samples. We set up 22 tutorial samples, 10 warm-up samples, 100100 evaluation samples and 1010 catch trials for each subject. The catch trials lead to 2020 valid subjects out of 2424 participants. We report the statistical results in Tab. 5. A chi-squared test is performed with a null hypothesis that the body prior does not have any influence. We show some results in Fig. 13, where all participants unanimously prefer one method over the other. While results of both methods look generally similar on front-side normal maps, using the body prior usually leads to better back-side normal maps.

A.4 Implementation details (Sec. 4.1)

Network architecture. Our body-guided normal prediction network uses the same architecture as PIFuHD , originally proposed in , and consisting of residual blocks with 44 down-sampling layers. The image encoder for PIFu∗\text{PIFu}^{*}, PaMIR∗\text{PaMIR}^{*}, and ICONenc\text{ICON}_{\text{enc}} is a stacked hourglass with 22 stacks, modified according to . Tab. 6 lists feature dimensions for various methods; “total dims” is the neuron number for the first MLP layer (input). The number of neurons in each MLP layer is: 1313 (77 for ICON), 512512, 256256, 128128, and 11, with skip connections at the 33rd, 44th, and 55th layers.

Training details. For training GN\mathcal{G}^{\text{N}} we do not use THuman due to its low-quality texture (see Tab. 1). On the contrary, IF\mathcal{IF} is trained on both AGORA and THuman. The front-side and back-side normal prediction networks are trained individually with batch size of 1212 under the objective function defined in Eq. 3, where we set λVGG=5.0\lambda_{\text{VGG}}=5.0. We use the ADAM optimizer with a learning rate of 1.0×10−41.0\times 10^{-4} until convergence at 8080 epochs.

Test-time details. During inference, to iteratively refine SMPL and the predicted clothed-body normal maps, we perform 5050 iterations (each iteration takes ∼460\sim 460 ms on a Quadro RTX 50005000 GPU) and set λN=2.0\lambda_{\text{N}}=2.0 in Eq. 4. We conduct an experiment to show the influence of the number of iterations (#iterations) on accuracy, see Tab. 8.

The resolution of the queried occupancy space is 2563256^{3}. We use rembg https://github.com/danielgatis/rembg to segment the humans in in-the-wild images, and use Kaolin https://github.com/NVIDIAGameWorks/kaolin to compute per-point the signed distance, Fs\mathcal{F}_{\text{s}}, and barycentric surface normal, Fnb\mathcal{F}_{\text{n}}^{\text{b}}.

Discussion on receptive field size. As Tab. 8 shows, simply reducing the size of receptive field of PaMIR does not lead to better performance. This shows that our informative 3D features as in Eq. 6 and normal maps N^c\widehat{\mathcal{N}}^{\text{c}} also play important roles for robust reconstruction. A more sophisticated design of smaller receptive field may lead to better performance and we would leave it for future research.

Appendix B More Quantitative Results (Sec. 4.3)

Table 4 compares several ICON variants conditioned on perturbed SMPL-X meshes. For the plot of Fig. 6 of the main paper (reconstruction error w.r.t. training-data size), extended quantitative results are shown in Tab. 9.

Appendix C More Qualitative Results (Sec. 5)

Figures 14, 15 and 16 show reconstructions for in-the-wild images, rendered from four different view points; normals are color coded. Figure 17 shows reconstructions for images with out-of-frame cropping. Figure 18 shows additional representative failures. The video on our website shows animation examples created with ICON and SCANimate.

References