ARCH++: Animation-Ready Clothed Human Reconstruction Revisited

Tong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, Tony Tung

Introduction

Digital humans have become an increasingly important building block for numerous AR/VR applications, such as video games, social telepresence and virtual try-on. Towards truly immersive experiences, it is crucial for these avatars to obtain higher level of realism that goes beyond the uncanny valley . Building a photorealistic avatar involves many manual works by artists or expensive capture systems under controlled environments , limiting access and increasing cost. Therefore, it is vital to revolutionize reconstruction techniques with minimal prerequisite (e.g., a selfie) for future digital human applications.

Recent human models reconstructed from a single image combine category-specific data prior with image observations . Among which, template-based approaches nevertheless suffer from lack of fidelity and difficulty supporting clothing variations; while non-parametric reconstruction methods , e.g., using implicit surface functions, do not provide intuitive ways to animate the reconstructed avatar despite impressive fidelity. In the recent work ARCH , the authors propose reconstructing non-parametric human model using pixel-aligned implicit functions in a canonical space, where all reconstructed avatars are transformed to a common pose. To do so, a parametric human body model is exploited to determine the transformations. By transferring skinning weights, which encode how much each vertex is influenced by the transformation of each body joint, from the underling body model, the reconstruction results are ready to animate. However, we observe that the advantages of a parametric body model and pixel-aligned implicit functions are not fully exploited.

In this paper we introduce ARCH++, which revisits the major steps of animatable avatar reconstruction from images and addresses the limitations in the formulation and representation of the prior work. First, current implicit function based methods mainly use hand-crafted features as the 3D space representation, which suffers from depth ambiguity and lacks human body semantic information. To address this, we propose an end-to-end geometry encoder based on PointNet++ , which expressively describes the underlying 3D human body. Second, we find the unposing process to obtain the canonical space supervision causes topology change (e.g., removing self-intersecting regions) and consequently the articulated reconstruction fails to obtain the same level of accuracy in the original posed space. Therefore, we present a co-supervising framework where occupancy is jointly predicted in both the posed and canonical spaces, with additional constraints on the cross-space consistency. This way, we benefit from both: supervision in the posed space allows the prediction to retain all the details of the original scans; while canonical space reconstruction can ensure the completeness of a reconstructed avatar. Last, image-based avatar reconstruction often suffers from degraded geometry and texture in the occluded regions. To make the problem more tractable, we first infer surface normals and texture of the occluded regions in the image domain using image translation networks, and then refine the reconstructed surface with a moulding-inpainting scheme.

In the experiments, we evaluate ARCH++ on photorealistically rendered synthetic images as well as in-the-wild images, outperforming prior works based on implicit functions and other design choices on public benchmarks.

The contributions of ARCH++ include: 1) a point-based geometry encoder for implicit functions to directly extract human shape and pose priors, which is efficient and free from quantization errors; 2) we are the first to point out and study the fundamental issue of determining target occupancy space: posed-space fidelity vs. canonical-space completeness. Albeit ignored before, we outline the pros and cons of different spaces, and propose a co-supervising framework of occupancy fields in joint spaces; 3) we discover image-based surface attribute estimation could address the open problem of view-inconsistent reconstruction quality. Our moulding-inpainting surface refinement strategy generates 360∘ photorealistic 3D avatars. 4) our method demonstrates enhanced performance on the brand new task of image-based animatable avatar reconstruction.

Related Work

Template-based reconstruction utilizes parametric human body models, e.g., SCAPE and SMPL to provide strong prior on body shape and pose to address ill-posed problems including body estimation under clothing and image-based human shape reconstruction . While these works primarily focus on underling body shapes without clothing, the template-based representations are later extended to modeling clothed humans with displacements from the minimal body , or external clothing templates , from 3D scans , videos , and a single image . As these approaches build clothing shapes on a body template mesh, the reconstructed models can be easily driven by pose parameters of the parametric body model. To address the lack of details with limited mesh resolutions, recent works propose to utilize 2D UV maps . However, as a clothing topology can significantly deviate from the underling body mesh and its variation is immense, these template-based solutions fail to capture clothing variations in the real world.

Non-parametric capture is widely used to capture highly detailed 3D shapes with an arbitrary topology from multi-view systems under controlled environments . Recent advances of deep learning further push the envelope by supporting sparse view inputs , and even monocular input . For single-view clothed human reconstruction, direct regression methods demonstrate promising results, supporting various clothing types with a wide range of shape representations including voxels , two-way depth maps , visual hull , and implicit functions . In particular, pixel-aligned implicit functions (PIFu) and its follow-up works demonstrate impressive reconstruction results by leveraging neural implicit functions and fully convolutional image features. Unfortunately, despite its high-fidelity results, non-parametric reconstructions are not animation-ready due to missing body part separation and articulation. Recently, IF-Net exploits partial point cloud inputs and learns implicit functions using latent voxel features. Compared with image-based avatar reconstruction, completion from points can leverage directly provided strong shape and pose cues, and thus skip learning them from complex images.

Hybrid approaches combine template-based and non-parametric methods and allow us to leverage the best of both worlds, namely structural prior and support of arbitrary topology. Recent work shows that using SMPL model as guidance significantly improves robustness of non-rigid fusion from RGB-D inputs. For single-view human reconstruction, Zheng et al. first introduce a hybrid approach of a template-model (SMPL) and a non-parametric shape representation (voxel and implicit surface ). These approaches, however, choose an input view space for shape modeling with reconstructed body parts potentially glued together, making the reconstruction difficult to animate as in the aforementioned non-parametric methods. The most relevant work to ours is ARCH , where the reconstructed clothed humans are ready for animation as pixel-aligned implicit functions are modeled in an unposed canonical space. However, such framework fundamentally leads to sub-optimal reconstruction quality. We achieve significant improvement on accuracy and photorealism by addressing the hand-crafted spatial encoding for implicit functions, the lack of supervision in the original posed space, and the limited fidelity of occluded regions.

Proposed Methods

Our proposed framework, ARCH++, uses a coarse-to-fine scheme, i.e., initial reconstruction by learning joint-space implicit surface functions (see Fig. 2), and then mesh refinement in both spaces (see Fig. 3).

Semantic-Aware Geometry Encoder. The spatial feature representation of a query point is critical for deep implicit function. While the pixel-aligned appearance feature via Stack Hourglass Network has already demonstrated its effectiveness in detailed clothed human reconstruction by prior works , an effective design of point-wise spatial encoding has not yet been well studied. The extracted geometry features should be informed of the semantics of the underlying 3D human body, which provide strong priors to regularize the overall dressed people shape.

The spatial encoding methods used previously include hand-crafted features (e.g., RBF ) and latent voxel features . The former is constructed based on Euclidean distances between a query point and the body joints, ignoring the shapes. The voxel-based features capture both shape and pose priors of a parametric body mesh. Compared with the hand-crafted features, the end-to-end learned voxel features are better informed of the underlying body structures but often constrained by GPU memory sizes and suffer from quantization errors due to low spatial resolution. To effectively encode the shape and pose priors without losing any precision, we propose a novel semantic-aware geometry encoder that extracts point-wise spatial encodings. Essentially a parametric body mesh can be sampled into a point cloud and fed into PointNet++ to learn point-based spatial features, which have several advantages over both hand-crafted RBF features and voxel-based ones. Our method encodes both shape and pose priors from parametric shapes without computation overhead and quantization errors caused by the mesh voxelization process. Additional detailed statistical comparisons on points v.s. voxels in representing 3D shapes are reported in .

Given a parametric body mesh estimated and deformed by , we use a PointNet++ based semantic-aware geometry encoder to learn the underlying 3D human body prior. We sample N0N_{0} (e.g., 7324) points from the body mesh surfaces and feed them into the geometry encoder for spatial feature learning, that is,

where B(⋅)\mathcal{B}(\cdot) indicates the differentiable bilinear sampling operation, and π(⋅)\pi(\cdot) means weak perspective camera projection from the query point pbp_{b} to the image plane of II.

Joint-Space Occupancy Estimator. While most non-parametric and hybrid methods use the posed space as the learning and inference target space, ARCH instead reconstructs the clothed human mesh directly in a canonical space where humans are in a normalized A-shape pose. Different choices of the target space have pros and cons. The posed space is naturally aligned with the input pixel evidence and therefore the reconstructions have high data fidelity leveraging the direct image feature correspondences. Thus, many works choose to reconstruct a clothed human mesh in its original posed space (e.g., PIFu(HD) , Geo-PIFu , PaMIR ). However, in many situations the human can demonstrate complex poses with self-intersection (e.g., hands in the pocket, crossed arms) and cause a ”glued” mesh that is difficult to articulate. Meanwhile, canonical pose reconstruction offers us a rigged mesh that is animation ready (via its registered A-shape parametric mesh ). The problem of using the canonical space as the target space is that when we warp the mesh into its posed space there could be artifacts like intersecting surfaces and distorted body parts (see Fig. 6). Thus, the reconstruction fidelity of the warping obtained canonical-to-posed space mesh will degenerate. To maintain both input image fidelity and reconstruction surface completeness, we propose to learn the joint-space occupancy distributions.

We use a joint-space defined occupancy map OO to implicitly represent the 3D clothed human under both its original posed space and a rigged canonical space:

where oa,obo_{a},o_{b} denote the occupancy for points pap_{a} and pbp_{b}. A point in the posed space is pbp_{b} and its mapped counterpart in the canonical space is pa=SemDF(pb)p_{a}=\text{SemDF}(p_{b}). The semantic deformation mapping (SemDF) between the original posed and the rigged canonical spaces is enabled by nearest neighbor-based skinning weights matching between pbp_{b} and the estimated underlying parametric body mesh .

where θ,β\theta,\beta are network weights of the MLP-based deep implicit surface functions. To reconstruct avatars from the dense occupancy estimations in two spaces, we use Marching Cube to extract the isosurface at oa=τo_{a}=\tau and ob=τo_{b}=\tau (i.e., τ=0\tau=0), respectively.

The network outputs oa,obo_{a},o_{b} are supervised by the ground truth joint-space occupancy o^a,o^b\hat{o}_{a},\hat{o}_{b}, depending on whether a posed space query point pbp_{b} and its corresponding canonical space point pap_{a} are inside the clothed human meshes or not. Though pa,pbp_{a},p_{b} are a pair of mapped points their ground-truth occupancy values are not the same in all cases. For example, a point outside and close to the hand of a parametric body could has o^b>0\hat{o}_{b}>0 and o^a<0\hat{o}_{a}<0 if the original mesh in posed space has self-contact (e.g., hands in the pocket). Namely, the SemDF defines a dense correspondence mapping between the two spaces but their occupancy values are not necessarily the same. Therefore, naively learning the distribution in one space and then warping the reconstruction into another pose can cause mesh artifacts (see Fig. 6). This motivates us to model two space occupancy distributions jointly in order to maintain both canonical space mesh completeness and posed space reconstruction fidelity.

2 Mesh Refinement

We further refine the reconstructed meshes in joint spaces by adding geometric surface details and photorealistic textures. As illustrated in Fig. 3, we propose a moulding-inpainting scheme to utilize the front and back side normals and textures estimated in the image space. This is based on the observation that direct learning and inference of dense normal/color fields using deep implicit functions as usually leads to over-smooth blur patterns and block artifacts (see Fig. 5). In contrast, image space estimation of normal and texture maps produces sharp results with fine-scale details, and is robust to human pose and shape changes. These benefits are from well-designed 2D convolutional deep networks (e.g., Pix2Pix ) and advanced (adversarial) image generation training schemes like GAN, with perceptual losses. The image-space estimated normal (and texture) maps could be used in two different ways. They can be used as either direct inputs into the Stack Hourglass as additional channels of the single-view image, or moulding-based front and back side mesh refinement sampling sources. In the experiments, we conduct ablation studies on these two schemes (i.e., early direct input, late surface refinement) and demonstrate that our moulding-based refinement is better at maintaining fine-scale surface details across different views (see Fig. 8).

where α\alpha is the angle between the unrefined normal and the forward camera raycast, and α′{\alpha}^{\prime} is the normalized value of α\alpha. Again, B(⋅)\mathcal{B}(\cdot) indicates the bilinear sampling operation. The indicator function χ(⋅)\chi(\cdot) determines the blending weights of sampled normals from the front and the back sides:

This simple yet effective fusion scheme creates a normal-refined mesh with negligible blending boundary artifacts. With the refined surface normals we can further apply Poisson Surface Reconstruction to update the mesh topology but in practice we find this unnecessary since the moulding-refined avatar can already satisfy various AR/VR and novel-view rendering applications. This bump rendering idea is also used in DeepHuman but they only refine meshes using the front views. We further conduct the texture refinement in a similar manner but use the refined normals to help determine the linear blending weights of boundary vertices. Our moulding-based front/back normal and texture refinement method yields clothed human meshes that look photorealistic at different viewpoints with full-body surface details (e.g., clothes wrinkles, hairs).

Canonical Space. The reconstructed canonical space avatar is rigged and thus can be warped back to its posed space and then refined via the same pipeline described above. However, a unique challenge for canonical avatar refinement is that mesh reconstructions in this space might contain unseen surfaces under the posed space. For example, in the third row of Fig. 5, the folded arm is in contact with the chest in the posed space but unfolded in the canonical space. Therefore, we do not have direct normal/texture correspondences of the chest regions of the canonical mesh. To address this problem, we render the front and the back side images of the canonical mesh with incomplete normal and texture, and treat it as an inpainting task. This problem has been well studied using deep neural networks and patch matching based methods . We use PatchMatch for its robustness. As demonstrated in the last two columns of Fig. 5, compared to directly regressing point-wise normal and texture, our inpainting-based results obtain sharper details and fewer artifacts.

Training Losses

The training process involves learning deep networks for two goals: joint-space occupancy estimation with Lo\mathcal{L}_{o}, and normal/texture estimation with Ln\mathcal{L}_{n} and Lt\mathcal{L}_{t}. Specifically, Lo\mathcal{L}_{o} is the occupancy regression loss of our joint-space deep implicit functions, and Ln,Lt\mathcal{L}_{n},\mathcal{L}_{t} are image translation losses of the normal, texture estimation networks.

where Loocc(oa),Loocc(ob)\mathcal{L}_{o}^{occ}(o_{a}),\mathcal{L}_{o}^{occ}(o_{b}) denote the Smooth L1L1-Loss between the estimated occupancy values and their ground truth in the canonical and the posed spaces, respectively. Locon(oa,ob)\mathcal{L}_{o}^{con}(o_{a},o_{b}) is a contrastive loss regularizing the occupancy consistency between the two spaces, that is,

where λ1\lambda_{1} and λ2\lambda_{2} are two parameters to adjust the penalty of inconsistent joint-space groundtruth pairs. Those pairs usually exist around the self-intersecting regions and need to be down-weighted due to the errors in canonical space supervision. Empirically, we set λ1=0.1\lambda_{1}=0.1 and λ2=0.3\lambda_{2}=0.3.

where Lrec(⋅)\mathcal{L}^{rec}(\cdot) denotes the L1L1 distance reconstruction loss, Ladv(⋅)\mathcal{L}^{adv}(\cdot) means the generative adversarial loss and Lvgg(⋅)\mathcal{L}^{vgg}(\cdot) is the VGG-perceptual loss proposed by . In the experiments, we found that the generative adversarial loss Ladv(⋅)\mathcal{L}^{adv}(\cdot) counteracts to performance in the normal map estimation task and thus we only enforce this loss term upon the back side texture map. One explanation is that the normal map space is more constrained and has fewer variations than the texture map, and therefore adversarial training does not fully show its effectiveness in this case.

Experiments

In this section, we present the experimental settings, result comparisons and ablation studies of ARCH++.

We implement our framework using PyTorch and conduct the training with one NVIDIA Tesla V100 GPU. The proposed deep neural networks are trained with RMSprop optimizer with a learning rate starting from 1e-4. We use an exponential learning rate scheduler to update it every 3 epochs by multiplying with the factor 0.10.1 and terminate the training after 12 epochs.

2 Datasets

We adopt the dataset setting from . Our training dataset consists of 375 3D scans from RenderPeople dataset and 205 3D scans from AXYZ dataset . These watertight human meshes have various clothes styles as well as body shapes and poses. Our testing set includes 37 scans from RenderPeople dataset , 192 scans from AXYZ dataset, 26 scans from BUFF dataset , and 2D images from Internet public domains, representing clothed people with a large variety of complex clothes. The subjects in the training dataset are mostly in standing pose, while the subjects in the test dataset contain various poses including sitting, twisted and standing, as well as self-glued and separated limbs. We use Blender and 38 environment maps to render each scan under different natural lighting conditions. For each 3D scan, we generate 360 images by rotating a camera around the mesh with a step size of 1 degree. These RenderPeople images are used to train both the occupancy estimation and the image translation networks.

We generate ground truth clothed human meshes in the canonical pose using the method introduced in . Note that the warping process between the posed and the canonical spaces inevitably contain model noises (e.g., self-contact region artifacts, skinning weights nearest neighbor discontinuities), which motivates our joint-space co-supervision and reconstruction scheme.

3 Results and Comparisons

We use the same metrics as for quantitative evaluation of the reconstructed meshes. We report the average point-to-surface Euclidean distance (P2S) and the Chamfer distance in centimeters, as well as the L2L2 normal re-projection errors. The two state of the art methods for our main comparisons are PIFuHD and ARCH , both are built upon PIFu with improvements in different aspects. PIFuHD ingests high-resolution images in a sliding window manner to achieve rich surface reconstruction details. ARCH leverages nearest neighbor-based linear blend skinning weights and hand-crafted RBF features to reconstruct animatable avatars in a canonical space. In addition to these two most related methods, we also include multiple prior methods and report the benchmark results on the RenderPeople and the BUFF datasets in Tab. 2. ARCH++ [Ours] results outperform the second best method ARCH by large gaps.

The visual comparisons in Fig. 5 and Fig. 2 further explain the advantages of our improvement. PIFuHD suffers from shape distortions due to lacking shape and pose priors provided by the end-to-end geometry encoder. Note that PIFuHD is incapable of reconstruct canonical space avatars and lacks texture estimation. ARCH reconstructions tend to be over smooth and blurry. Its recovered mesh normal and texture also have several block artifacts. Additionally, both methods fail to hallucinate plausible back-side surface details like clothes wrinkles, hairs, etc. In comparison, our approach achieves photorealistic and animatable reconstructions in joint spaces and across different viewpoints. We further show our results on Internet images in Fig. 9.

4 Ablation Studies

Joint Space Reconstruction. To further understand the impact of the proposed methods, we present ablation studies in Tab. 1. The first three rows demonstrate the effectiveness of joint-space co-supervision, achieving balanced performances on both the posed and the canonical space mesh reconstructions. Choosing the posed space as the reconstruction target space (e.g., PIFu, PIFuHD, Geo-PIFu, PaMIR) can cause missing surfaces and topology distortions in the posed-to-canonical space warped meshes (see Fig. 6). Meanwhile, choosing the canonical space as the target space (e.g., ARCH) can cause self-intersecting meshes with broken manifold as well as body part un-natural deformations in the canonical-to-posed space warped meshes. In contrast, our co-supervision and joint-space inference methods achieve both reconstruction fidelity in the posed space and body mesh completeness in the canonical space.

Geometry Encoding. As shown in Tab. 1, we observe further error reduction leveraging the end-to-end learned point-wise spatial encodings. The prior method ARCH uses handcrafted RBF features that only model the pose prior of parametric body mesh skeletons, ignoring the mesh shape. In comparison, our point-based features are informed of both pose and shape priors of the underlying parametric body model w.r.t. a clothed human mesh, and thus improve the surface reconstruction quality. We further implement the learned volumetric spatial feature encodings used in Geo-PIFu and PaMIR as an alternative encoder and inject into our framework for direct comparisons. The results are shown in Tab. 3 and Fig. 7. While both types of end-to-end spatial features outperform the hand crafted RBF features, our point-based feature extraction method does not suffer from computation overhead and mesh quantization errors of the voxel-based approach.

Normal Refinement. While single-image based direct inference of human meshes with rich surface details at both the front and the back side remains an open question, some empirical observations and prior works indicate that normal estimation is a relatively easier task and can help refine the reconstructions. In Tab. 4 and Fig. 8 we experiment on three principle ways of leveraging the estimated normals for mesh reconstructions with refined surface details. Among these normal refinement methods, our front/back-side image space normal regression and moulding-based surface refinement approach outperforms other variants. Object-space normal regression is adopted in ARCH and is based on learning deep implicit functions of spatial normal fields. It fails to generate rich back side details and sometimes causes block artifacts as shown in the fourth row of Fig. 5. Image-space input is used in PIFuHD. It concatenates the color image input with estimated image-space normal maps and feeds them into Stack Hourglass for feature extraction. While this method achieves the same level of quantitative performance as our mesh refinement approach, its visual results are not as sharp as ours at both the front and the back sides. A degenerated case of our mesh refinement method is studied before in DeepHuman where they only estimate front-view normal maps and therefore lack reconstruction details at the back side.

Conclusion

In this paper, we revisit the major components in existing deep implicit function based 3D avatar reconstruction. Our method ARCH++ produces results which have high-level fidelity and are animation-ready for many AR/VR applications. We conduct a series of comparisons with and analysis on the state of the art to validate our findings. For future works, we plan to incorporate environment information (e.g., lighting, affordance) to further understand the body pose and appearance, and address current limitations.

Acknowledgements. We would like to thank Minh Vo and Nikolaos Sarafianos for the discussions and synthetic data creation.

References

Supplementary

Appendix A Implementation Details

In this section, we provide the implementation details of our proposed method.

During both training and test time, the input images to the network are normalized with regard to the human body scale. In particular, we re-scale the image based on the 3D skeleton estimation of the subject. The image is resized then centered, such that the pelvis of the person is aligned with the center of the image. Each pixel represents 11cm length using an orthographic scene projection. In this way we ensure proper scaling of the body parts, which allows us to capture the variations of different heights of people.

A.2 Network Architectures

Semantic-Aware Geometry Encoder is based on PointNet++ , which consists of 3 Set Abstraction (SA) layers. The configurations of each layer are SA(2048, 0.1, 16, 3, ), SA(512, 0.2, 32, 32, ), SA(128, 0.4, 64, 64, ). The explanation of each argument is (furthest point sampling size, point neighborhood radius, point neighborhood size limit, input feature channel, MLP output channels list). Namely, the multi-scale point set sizes of Eq. (1) in the main paper are: N1=2048,N2=512,N3=128N_{1}=2048,N_{2}=512,N_{3}=128. When extracting spatially-aligned features for any given query point, we leverage the point Feature Propagation (FP) layers defined at the aforementioned 3 different point set scales: FP(32, ), FP(64, ), FP(128, ). The explanation of each argument is (input feature channel, MLP output channels list). Therefore, the dimensions of our spatially-aligned geometry features fgf_{g} in Eq. (3) of the main paper are 96=32∗396=32*3. Please refer to for further details.

Pixel-Aligned Appearance Encoder adopts the architecture from Stack Hourglass Network . The layer configuration is the same as PIFu(HD) and ARCH , which is composed of a 4-stack model and each stack uses 2 residual blocks. The output latent image feature length is 256. Therefore, the dimensions of our pixel-aligned appearance features faf_{a} in Eq. (3) of the main paper are 256256. Please refer to for further details.

Joint-Space Occupancy Estimator is a two-branch multilayer perceptron (MLP). Each branch takes the spatially-aligned geometry features fgf_{g} and pixel-aligned appearance features faf_{a} described above, and estimates one-dimension occupancy oao_{a} or obo_{b} using Tanh activation. oao_{a} is the occupancy probability in the canonical space and obo_{b} is the one in the posed space. Similar to , we design the MLP with four fully-connected layers and the numbers of hidden neuron sizes are (1024,512,256,128)(1024,512,256,128). Each layer of MLP has skip connections from the input features.

Normal and Texture Image Translation both use a network architecture designed by using 9 residual blocks with 4 downsampling layers. The same network is also used in PIFuHD for normal maps estimation. In our work we extend this architecture to back side texture inference by adding GAN losses.

A.3 Hyper-Parameters

When training the joint-space occupancy losses Lo\mathcal{L}_{o}, we use 0.5, 0.5, 0.05 to weight the canonical/posed space occupancy estimation losses Loocc\mathcal{L}_{o}^{occ} as well as the contrastive regularizer Locon\mathcal{L}_{o}^{con}. When supervising the normal and texture image translation losses Ln,Lt\mathcal{L}_{n},\mathcal{L}_{t}, we set the weights of the L1L1 reconstruction losses Lrec\mathcal{L}^{rec} and the perceptual losses Lvgg\mathcal{L}^{vgg} to 5.0 and 1.0, respectively. Particularly, the GAN losses (i.e. generator, discriminator) used in back-side texture hallucination is weighted by 0.1.

A.4 Computation Cost

Empirically, when using a single Tesla V100 GPU for training and one batch of 44 images (each image with 20480 pairs of query points), the forward pass takes around 1.4s and the backward propagation takes around 0.6s.

Appendix B Inference on Images in the Wild

To perform inference on in-the-wild images, our staged pipeline involves person instance segmentation, parametric 3D human body estimation and the proposed 3D avatar reconstruction. The total pipeline takes around 55 seconds to reconstruct a fully colored animation-ready avatar from an unconstrained photo (RGB image) using one Tesla V100 GPU. In comparison, PaMIR takes over 4040 seconds. Note our current implementation lets all modules run sequentially and the intermediate results (e.g., masks, 3D human body parameters) are mostly exchanged through CPU memory and file IO, which leaves room for optimization. For our proposed avatar reconstruction framework alone, we could run 1) Semantic-Aware Geometry Encoder, 2) Pixel-Aligned Appearance Encoder and 3) Normal and Texture Refinement Networks in parallel to greatly boost the efficiency.

Similar to most existing implicit surface function based methods, our occupancy estimation module also requires the person segmentation mask. Such mask is used to remove redundant and erroneous estimated occupancy in the background regions and serves as a visual hull prior similar to multiple view stereo. In this paper, we utilize one state of the art semantic instance segmentation , which is able to generate per-person segmentation mask. Note we are able to handle multiple people in the same image with such a detection-and-segmentation method. We set a minimum detection score 0.50.5 and minimum bounding box size 100×100100\times 100 to filter out people instances with too small resolution and guarantees the proper scale used for our proposed avatar reconstruction approach.

B.2 Parametric 3D Human Body Estimation

Underlying 3D human body serves as an important semantic cue for our approach. In this paper, we adopt a similar way as ARCH to estimate the parametric 3D human body from the input image. We first run DenseRaC to obtain the initial estimation of 3D human pose and shape parameters. Furthermore, we observe that when re-projecting such estimated 3D human body back to the input image, the body landmarks (e.g., joints, face, hands, feet) do not align with the input image well. We thus implement an additional optimization script using pytorch to compute the offsets between the re-projected body landmarks and detected body landmarks from OpenPose and back-propagate to the estimated 3D human pose and shape parameters (see Fig.10). The optimization is run over 200 iterations and we obtain better-aligned 3D human body in this way.

B.3 3D Avatar Reconstruction

Given the intermediate results obtained from the modules above, we are able to run our proposed approach and obtain the jointly reconstructed avatars in both the original posed space and the canonical space.

Appendix C Extended Experiments

In this section, we show some interesting conclusions we obtained along the way and extended experiments as well as comparisons (e.g., user studies, applications of avatar animations and video-based fusion, failure cases). Note we remove the background for all images for better visualization of the inputs.

In Fig. 11 we show that reconstruction results of PIFuHD fail to capture the underlying correct body shapes and poses. PIFuHD and our method are trained using the same set of RenderPeople clothed human meshes, which consist of mostly upstanding poses. While obtaining large-scale ground truth clothed avatars with various poses and shapes is still an open problem, we can leverage parametric body shape estimation networks (e.g. DenseRaC and HMR ) whose training data is easier to obtain. This motivates our design of learning both semantic-aware geometry features and pixel-aligned appearance features. The geometry features encode shape and pose priors of the underlying parametric body mesh, while the appearance features provide image evidence for fine-scale clothing wrinkles and surface details reconstruction.

C.2 Moulding-based Surface Refinement Obtains Better Consistency across Views

For previous methods like PIFu , ARCH and PIFuHD , we often observe the reconstructed surface details only look plausible from the input camera view. Once we change the camera view to preview the reconstructed avatar from other view points, the rendered results contain fewer surface details and are less realistic. Such quality inconsistency limits the applicability of the prior works to AR/VR applications that require free viewpoint rendering.

Based on the aforementioned observations, we conclude that such phenomenon is caused by the suppression of occluded region hallucination. Although PIFuHD generates detailed back-side surface compared to PIFu and ARCH by leveraging inferred normal maps, its final rendering quality remains less sharp than our moulding-based refinement approach. Besides the results shown in the main paper, we provide more results with zoom-in in Fig. 12 to further demonstrate the improvement on reconstruction details. Moreover, the normal refinement step enhances photorealistic rendering results by enabling fine-grained shading effects. In Fig. 13, we show the rendered images with and without the refined normals using the same mesh textures. With the normal refinement, the rendered images show more plausible clothing wrinkles at different views than the ones without normal refinement.

C.3 GAN Improves Back-Side Texture Estimation

In our experiments, we found that GAN losses help to enhance the realism of back-side texture maps. In Fig. 14, we demonstrate that the estimated textures with GAN training contain more plausible texture details and better lighting effect than those without GAN.

C.4 Estimating Shaded Textures Preserves More Details

Notably PIFu and ARCH all choose the albedo color space to predict. This requires the neural network to implicitly learn to compensate light/shading from the given shaded input images. However our current data scale seems insufficient to capture such a complicated space and might produce erroneous reconstructed textures. We believe our best strategy is to “reconstruct the texture as similar/compatible as the input image“. We conduct ablation studies on learning different color spaces and the results are shown in Fig. 15. It can be observed that estimating back-side shaded textures, in comparison with albedos, preserves more details and overall generates similar and consistent color space to the input images.

C.5 Inpainting Is a Needed Step for Avatar Reconstruction

Compared with PIFu and ARCH, we propose to utilize the image space characteristics to better refine the reconstructed geometry. Our insight is that such features are naturally encoded in the image and there are already lots of powerful generative models in the literature which can solve similar tasks. However, one issue when trying to estimate the surface normals and textures lies in the missing surfaces cannot be ray traced from either the front or the back side, e.g., the occluded clothes by the arms in Fig. 16. As a result, we believe such missing surfaces could be formulated and solved as an inpainting task. As shown in the figure, we are able to inpaint those missing surfaces (marked as gray regions) using the context. We could also observe that the final reconstructed avatar in the canonical space looks more complete and realistic than the one obtained by ARCH via implicit color field interpolation.

C.6 More Normal Refinement Results

We show more qualitative results on the testing sets (e.g., DeepHuman, AXYZ, RenderPeople and Unsplash) for people in different camera views, poses and clothes in Fig. 18. It can be observe that our normal refinement results achieve high fidelity.

C.7 User Study on Reconstruction Photorealism Shows Superior Quality of Our Method

We set up a user study to further evaluate the photorealism of our method against other state of the art. In our study, we randomly pick 3030 examples from our testing set (i.e., RenderPeople, AXYZ, Unsplash) and run the comparison methods to obtain the results. For each example, we showed the input image and side-by-side rendered reconstruction results from two approaches in both front and back views to the participants. We numerate all pairs of approaches following a “similarity judgment” design . For each pair, the participants were asked to choose the more realistic result (“Which reconstructed avatar looks more real given the input image?”). They could choose either results or indicate that they were equally real. To avoid biases and learning effects, we randomized the order of pairs as well as the position of the results while ensuring that for each participant the same number of techniques were shown on the left and the right sides. To evaluate the results, we attribute the choice of one technique with +1+1 and the other with −1-1. Averaging over all results in a quality score for each pair of techniques. Using a t-test, we determine the probability of the drawn sample to come from a zero-mean distribution—zero-mean would indicate both techniques to be of equal quality.

We recruited 22 participants from universities and research institutes. Most participants are with medium to high experiences for computer vision and graphics. The results of the user study are summarized in Fig. 17. We conduct two sessions with one studying the normal reconstruction quality and the other one studying the texture reconstruction quality. We compute the p-value to indicates statistical significant difference between the methods according to a t-test (all significant results achieved at least p value <.001<.001). All three pair-wise comparisons showed significant results:

For normal reconstruction, Ours vs. PIFuHD (mean =0.39697=0.39697, std =0.34625=0.34625, t(21)=4.02t(21)=4.02, p<.001p<.001), Ours vs. ARCH (mean =0.83636=0.83636, std =0.15535=0.15535, t(21)=25.25t(21)=25.25, p<.001p<.001), and PIFuHD vs. ARCH (mean =0.32121=0.32121, std =0.27826=0.27826, t(21)=5.41t(21)=5.41, p<.001p<.001), with ours being significantly better than other state of the art and PIFuHD being significantly better than ARCH.

For texture reconstruction, Ours vs. PIFu (mean =0.49091=0.49091, std =0.21513=0.21513, t(21)=10.70t(21)=10.70, p<.001p<.001), Ours vs. ARCH (mean =0.74545=0.74545, std =0.21615=0.21615, t(21)=16.18t(21)=16.18, p<.001p<.001), and ARCH vs. PIFu (mean =0.22424=0.22424, std =0.21134=0.21134, t(21)=4.98t(21)=4.98, p<.001p<.001), with ours being significantly better than other state of the art and ARCH being significantly better than PIFu.

C.8 Applications

When multi-view inputs are available (e.g., monocular videos), our method naturally supports canonical space normal and texture fusion to recover a photorealistic and animatable avatar thanks to the shared canonical space among different poses. For mesh vertices that are co-visible under multiple viewpoints we apply a simple yet effective normal-based linear blending scheme. For each vertex, the normal/texture fusion weights w.r.t. one visible view is determined by the angle between the (unrefined)normal of that vertex and the camera direction. This is similar to Eq. (6) and (7) of the main paper. Here we further extend them to multiple views. Namely, one image that is facing towards the surface is weighted higher than another image of a large viewing angle. As show in Fig. 19, we fuse the normal and texture from multiple frames, and further generate an animation sequence.

C.9 Failure cases

As shown in Fig. 20, there are some typical failure cases due to strong directional lighting and challenging poses. To tackle these issues we plan to add more lighting augmentation for the back-side texture estimation Pix2Pix module and also increase the size of our training scan set. For example, compared with the widely used image classification and detection datasets like ImageNet and COCO , our training dataset is relatively small consisting of only hundreds of 3D scans. Building a large-scale and high-quality clothed human mesh dataset with sufficient clothes types and human pose/shape variations is critical for pushing research works in image-based photorealistic avatar reconstruction.