Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

Thiemo Alldieck, Mihai Zanfir, Cristian Sminchisescu

3.25ex plus1ex minus.2ex3pt plus 1pt minus 1pt

Introduction

We present PHORHUM, a method to photorealistically reconstruct the 3D geometry and appearance of a dressed person as photographed in a single RGB image. The produced 3D scan of the subject not only accurately resembles the visible body parts but also includes plausible geometry and appearance of the non-visible parts, see fig. 1. 3D scans of people wearing clothing have many use cases and demand is currently rising. Applications like immersive AR and VR, games, telepresence, virtual try-on, free-viewpoint photo-realistic visualization, or creative image editing would all benefit from accurate 3D people models. The classical way to obtain models of people is to automatically scan using multi-camera set-ups, manual creation by an artist, or a combination of both as often artists are employed to ‘clean up’ scanning artifacts. Such approaches are difficult to scale, hence we aim for alternative, automatic solutions that would be cheaper and easier to deploy.

Prior to us, many researchers have focused on the problem of human digitization from a single image . While these methods sometimes produce astonishingly good results, they have several shortcomings. First, the techniques often produce appearance estimates where shading effects are baked-in, and some methods do not produce color information at all. This limits the usefulness of the resulting scans as they cannot be realistically placed into a virtual scene. Moreover, many methods rely on multi-step pipelines that first compute some intermediate representation, or perceptually refine the geometry using estimated normal maps. While the former is at the same time impractical (since compute and memory requirements grow), and potentially sub-optimal (as often the entire system cannot be trained end-to-end to remove bias), the latter may not be useful for certain applications where the true geometry is needed, as in the case of body measurements for virtual try-on or fitness assessment, among others. In most existing methods color is exclusively estimated as a secondary step. However, from a methodological point of view, we argue that geometry and surface color should be computed simultaneously, since shading is a strong cue for surface geometry and cannot be disentangled.

Our PHORHUM model specifically aims to address the above-mentioned state of the art shortcomings, as summarised in table 1. In contrast to prior work, we present an end-to-end solution that predicts geometry and appearance as a result of processing in a single composite network, with inter-dependent parameters, which are jointly estimated during a deep learning process. The appearance is modeled as albedo surface color without scene specific illumination effects. Furthermore, our system also estimates the scene illumination which makes it possible, in principle, to disentangle shading and surface color. The predicted scene illumination can be used to re-shade the estimated scans, to realistically place another person in an existing scene, or to realistically composite them into a photograph. Finally, we found that supervising the reconstruction using only sparse 3D information leads to perceptually unsatisfactory results. To this end, we introduce rendering losses that increase the perceptual quality of the predicted appearance. Our contributions can be summarised as follows:

We present an end-to-end trainable system for high quality human digitization

Our method computes, for the first time, albedo and shading information

Our rendering losses significantly improve the visual fidelity of the results

Our results are more accurate and feature more detail than current state-of-the-art

Related Work

Reconstructing the 3D shape of a human from a single image or a monocular video is a wide field of research. Often 3D shape is a byproduct of 3D human pose reconstruction and is represented trough parameters of a statistical human body model . In this review, we focus on methods that go beyond and reconstruct the 3D human shape as well as garments or hairstyle. Early pioneering work is optimization-based. Those methods use videos of moving subjects and integrate information over time in order to reconstruct the complete 3D shape . The advent of deep learning questioned the need for video. First, hybrid reconstruction methods based on a small number of images have been presented . Shortly after, approaches emerged to predict 3D human geometry from a single image. Those methods can be categorized by the used shape representation: voxel-based techniques predict whether a given segment in space is occupied by the 3D shape. A common limitation is the high memory requirement resulting in shape estimates of limited spatial resolution. To this end, researchers quickly adopted alternative representations including visual hulls , moulded front and back depth maps , or augmented template meshes . Another class of popular representations consists of implicit function networks (IFNs). IFNs are functions over points in space and return either whether a point is inside or outside the predicted shape or return its distance to the closest surface . Recently IFNs have been used for various 3D human reconstruction tasks and to build implicit statistical human body models . Neural radiance fields are a related class of representations specialized for image synthesis that have also been used to model humans . Saito et al. were the first to use IFNs for monocular 3D human reconstruction. They proposed an implicit function conditioned on pixel-aligned features . Other researchers quickly adopted this methodology for various use-cases . ARCH and ARCH++ also use pixel-aligned features but transform information into a canonical space of a statistical body model. This process results in animatable reconstructions, which comes, however, at the cost of artifacts that we will show. In this work, we also employ pixel-aligned features but go beyond the mentioned methods in terms of reconstructed surface properties (albedo and shading) and in terms of the quality of the 3D geometry. Also related is H3D-Net , a method for 3D head reconstruction, which uses similar rendering losses as we do, but requires three images and test-time optimization. In contrast, we work with a monocular image, purely feed-forward.

Method

Our goal is to estimate the 3D geometry S\mathcal{S} of a subject as observed in a single image I\mathbf{I}. Further, we estimate the unshaded albedo surface color and a per-image lighting model. S\mathcal{S} is defined as the zero-level-set of the signed distance function (SDF) ff represented using a neural network,

where θ\boldsymbol{\theta} is the superset of all learnable parameters. The surface S\mathcal{S} is parameterized by pixel aligned features z\boldsymbol{z} (cf. ) computed from the input image I\mathbf{I} using the feature extractor network GG

where bb defines pixel access with bilinear interpolation and π(x)\pi(\boldsymbol{x}) defines the pixel location of the point x\boldsymbol{x} projected using camera π\pi. ff returns the signed distance dd of the point x\boldsymbol{x} w.r.t. S\mathcal{S} and additionally its albedo color a\boldsymbol{a}

where γ\gamma denotes basic positional encoding as defined in . In the sequel, we will use dxd_{\boldsymbol{x}} for the estimated distance at x\boldsymbol{x} and ax\boldsymbol{a}_{\boldsymbol{x}} for the color component, respectively.

To teach the model to decouple shading and surface color, we additionally estimate the surface shading using a per-point surface shading network

where nx=∇xdx\boldsymbol{n}_{\boldsymbol{x}}=\nabla_{\boldsymbol{x}}d_{\boldsymbol{x}} is the estimated surface normal defined by the gradient of the estimated distance w.r.t. x\boldsymbol{x}. l(I;θ)=ll(\mathbf{I};\boldsymbol{\theta})=\boldsymbol{l} is the illumination model estimated from the image. In practice, we use the bottleneck of GG for l\boldsymbol{l} and further reduce its dimensionality. The final shaded color is then c=s∘a\boldsymbol{c}=\boldsymbol{s}\circ\boldsymbol{a} with ∘\circ denoting element-wise multiplication. We now define the losses we use to train ff, GG, and ss.

We create training examples by rendering scans of humans and drawing samples from the raw meshes – please see §3.2 for details. We define losses based on sparse 3D supervision and losses informed by ray-traced image patches.

Geometry and Color Losses. Given a ground truth mesh M\mathcal{M} describing the surface S\mathcal{S} as observed in an image I\mathbf{I} and weights λ∗\lambda_{*} we define losses as follows. The surface is supervised via samples O\mathcal{O} taken from the mesh surface M\mathcal{M} and enforcing their distance to return zero and the distance gradient to follow their corresponding ground truth surface normal nˉ\bar{\boldsymbol{n}}

Moreover, we supervise the sign of additional samples F\mathcal{F} taken around the surface

where ll are inside/outside labels, ϕ\phi is the sigmoid function, and BCE is the binary cross-entropy. kk determines the sharpness of the decision boundary and is learnable. Following , we apply geometric regularization such that ff approximates a SDF with gradient norm 11 everywhere

Finally, we supervise the albedo color with the ‘ground truth’ albedo aˉ\bar{\boldsymbol{a}} calculated from the mesh texture

Following , we apply La\mathcal{L}_{a} not only on but also near the surface. Since albedo is only defined on the surface, we approximate the albedo for points near the surface with the albedo of their nearest neighbor on the surface.

Rendering losses. The defined losses are sufficient to train our networks. However, as we show in the sequel, 2D rendering losses help further constrain the problem and increase the visual fidelity of the results. To this end, during training, we render random image patches of the surface S\mathcal{S} with random strides and fixed size using ray-tracing. First, we compute the rays R\mathcal{R} corresponding to a patch as defined by π\pi. We then trace the surface using two strategies. First, to determine if we can locate a surface along a ray, we query ff in equal distances along every ray r\boldsymbol{r} and compute the sign of the minimum distance value

where o\boldsymbol{o} is the camera location. We then take the subset RS⊂R\mathcal{R}_{\mathcal{S}}\subset\mathcal{R} of the rays containing rays where σ≤0.5\sigma\leq 0.5 and l=0l=0, i.e. we select the rays which located a surface where a surface is expected. Hereby, the inside/outside labels ll are computed from pixel values of the image segmentation mask M\mathbf{M} corresponding to the rays. For the subset RS\mathcal{R}_{\mathcal{S}}, we exactly locate the surface using sphere tracing. Following , we make the intersection point x^\hat{\boldsymbol{x}} at iteration tt differentiable w.r.t. to the network parameters without having to store the gradients of sphere tracing

In practice, we trace the surface both from the camera into the scene and from infinity back to the camera. This means, we locate both the front surface and the back surface. We denote the intersection points x^f\hat{\boldsymbol{x}}^{f} for the front side and x^b\hat{\boldsymbol{x}}^{b} for the back side, respectively. Using the above defined ray set RS\mathcal{R}_{\mathcal{S}} and intersection points x^\hat{\boldsymbol{x}}, we enforce correct surface colors through

where ground truth albedo colors aˉ\bar{\boldsymbol{a}} are taken from synthesized unshaded images Af\mathbf{A}^{f} and Ab\mathbf{A}^{b}. The back image Ab\mathbf{A}^{b} depicts the backside of the subject and is created by inverting the Z-buffer during rendering. We explain this process in more detail in §3.2. Additionally, we also define a VGG-loss LVGG\mathcal{L}_{\text{VGG}} over the rendered front and back surface patches, enforcing that structure is similar to the unshaded ground-truth images. Finally, we supervise the shading using

with p\boldsymbol{p} being the pixel color in the image I\mathbf{I} corresponding to the ray r\boldsymbol{r}. We found it also useful to supervise the shading on all pixels of the image I={p0,…,pN}\mathcal{I}=\{\boldsymbol{p}_{0},\dots,\boldsymbol{p}_{N}\} using ground truth normals nˉ\bar{\boldsymbol{n}} and albedo aˉ\bar{\boldsymbol{a}}

The final loss is a weighted combination of all previously defined losses L∗\mathcal{L}_{*}. In §4.3, we ablate the usage of the rendering losses and the shading estimation network.

2 Dataset

We train our networks using pairs of meshes and rendered images. The meshes are scans of real people from commercial websites and our own captured data. We employ high dynamic range images (HDRI) for realistic image-based lighting and as backgrounds. Additionally to the shaded images, we also produce an alpha mask and unshaded albedo images. In the absence of the true surface albedo, we use the textures from the scans. Those are uniformly lit but may contain small and local shading effects, e.g. from small wrinkles. As mentioned earlier, we produce not only a front side albedo image, but also one showing the back side. We obtain this image by inverting the Z-buffer during rendering. This means, not that the first visible point along each camera ray is visible, but the last passed surface point. See fig. 3 for an example of our training images. Furthermore, we produce normal maps used for evaluation and to supervise shading. Finally, we take samples by computing 3D points on and near the mesh surface and additionally sample uniformly in the bounding box of the whole dataset. For on-surface samples, we compute their corresponding albedo colors and surface normals, and for near and uniform samples we compute inside/outside labels by casting randomized rays and checking for parity.

We use 217217 scans of people in different standing poses, wearing various outfits, and sometimes carrying bags or holding small objects. The scans sources allow for different augmentations: we augment the outfit colors for 100 scans and repose 38 scans. In total we produce a dataset containing ≈190\approx 190K images, where each image depicts a scan rendered with a randomly selected HDRI backdrop and with randomized scan placement. Across the 217217 scans some share the same identity. We strictly split test and train identities and create a test-set containing 20 subjects, each rendered under 5 different light conditions.

3 Implementation Details

We now present our implementation and training procedure. Our networks are trained with images of 512×512512\times 512px resolution. During training we render 32×3232\times 32px patches with stride ranging from zero to three. We discard patches that only include background. Per training example we draw random samples for supervision from the surface and the space region around it. Concretely, we draw each 512 samples from the surface, near the surface and uniformly distributed over the surrounding space. The samples are projected onto the feature map using a projective camera with fixed focal length.

Experiments

We present quantitative evaluation results and ablation studies for geometric and color reconstruction on our own dataset. We also show qualitative results for real images.

At inference time, we take as input an RGB image of a person in a scene. Note that we do not require the foreground-background mask of the person. However, in practice we use a bounding box person detector to center the person and crop the image – a step that can also be performed manually. We use Marching Cubes to generate our reconstructions by querying points in a 3D bounding box at a maximum resolution of 5123512^{3}. We first approximate the bounding box of the surface by probing at coarse resolution and use Octree sampling to progressively increase the resolution as we get closer to the surface. This allows for very detailed reconstructions of the surface geometry with a small computational overhead, being made possible by the use of signed distance functions in our formulation.

Camera Model.

Different from other methods in the literature, we deviate from the standard orthographic camera model and instead use perspective projection, due to its general validity. A model assuming an orthographic camera would in practice produce incorrect 3D geometry. In fig. 5 one can see the common types of errors for such models. The reconstructed heads are unnaturally large, as they extend in depth away from the camera. In contrast, our reconstructions are more natural, with correct proportions between the head and the rest of the body.

Competing Methods.

We compare against other single-view 3D reconstructions methods that leverage pixel-aligned image features. PIFu is the pioneering work and learns an occupancy field. PIFuHD , a very parameter-heavy model, builds upon PIFu with higher resolution inputs and leverages a multi-level architecture for coarse and fine grained reconstruction. It also uses offline estimated front and back normal maps as additional input. GeoPIFu is also a multi-level architecture, but utilizes latent voxel features as a coarse human shape proxy. ARCH and ARCH++ transform information into the canonical space of a statistical body model. This sacrifices some of the reconstruction quality for the ability to produce animation-ready avatars. For PIFu, ARCH, ARCH++, an off-the-shelf detector is used to segment the person in the image, whereas PHORHUM (us) and PIFuHD use the raw image. The results of ARCH and ARCH++ have been kindly provided by the authors.

Due to the lack of a standard dataset and the non-availability of training scripts of most methods, all methods have been trained with similar but different datasets. All datasets are sufficiently large to enable generalization across various outfits, body shapes, and poses. Please note that our dataset is by far the smallest with only 217 scans. All other methods use >400>400 scans.

3 Ablations

We now ablate two main design choices of our method: first, the rendering losses, and second, shading estimation. In tab. 4.1, we report metrics for our method trained without rendering losses (w/o rendering) and without shading estimation (w/o shading). Furthermore, in fig. 6 we show visual examples of results produced by our model variant trained without rendering losses.

While only using 3D sparse supervision produces accurate geometry, the albedo estimation quality is, however, significantly decreased. As evident in fig. 6 and also numerically in tab. 4.1, the estimated albedo contains unnatural color gradient effects. We hypothesize that due to the sparse supervision, where individual points are projected into the feature map, the feature extractor network does not learn to understand structural scene semantics. Here our patch-based rendering losses help, as they provide gradients for neighboring pixels. Moreover, our rendering losses could better connect the zero-level-set of the signed distance function with the color field, as they supervise the color at the current zero-level-set and not at the expected surface location. We plan to structurally investigate these observations, and leave these for future work.

Estimating the shading jointly with the 3D surface and albedo does not impair the reconstruction accuracy. On the contrary, as evident in tab. 4.1, this helps improve albedo reconstruction. This is in line with our hypothesis that shading estimation helps the networks to better decouple shading effects from albedo. Finally, shading estimating makes our method a holistic reconstruction pipeline.

Discussion and Conclusions

The limitations of our method are sometimes apparent when the clothing or pose of the person in the input image deviates too much from our dataset distribution, see fig. 8. Loose, oversized, and non-Western clothing items are not well covered by our training set. The backside of the person sometimes does not semantically match the front side. A larger, more geographic and culturally diverse dataset would alleviate these problems, as our method does not make any assumptions about clothing style or pose.

Application Use Cases and Model Diversity.

The construction of our model is motivated by the breadth of transformative, immersive 3D applications, that would become possible, including clothing virtual apparel try-on, immersive visualisation of photographs, personal AR and VR for improved communication, special effects, human-computer interaction or gaming, among others. Our models are trained with a diverse and fair distribution, and as the size of this set increases, we expect good practical performance.

Conclusions. We have presented a method to reconstruct the three-dimensional (3D) geometry of a human wearing clothing given a single photograph of that person. Our method is the first one to compute the 3D geometry, surface albedo, and shading, from a single image, jointly, as prediction of a model trained end-to-end. Our method works well for a wide variation of outfits and for diverse body shapes and skin tones, and reconstructions capture most of the detail present in the input image. We have shown that while sparse 3D supervision works well for constraining the geometry, rendering losses are essential in order to reconstruct perceptually accurate surface color. In the future, we would like to further explore weakly supervised differentiable rendering techniques, as they would support, long-term, the construction of larger and more inclusive models, based on diverse image datasets of people, where accurate 3D surface ground truth is unlikely to be available.

Supplementary Material

In this supplementary material, we detail our implementation by listing the values of all hyper-parameters. Further, we report inference times, demonstrate how we can repose our reconstructions, conduct further comparisons, and show additional results.

In this section, we detail our used hyper-parameters and provide timings for mesh reconstruction via Marching Cubes .

When training the network, we minimize a weighted combination of all defined losses:

Further, we have defined the weights λg1\lambda_{g_{1}}, λg2\lambda_{g_{2}}, λa1\lambda_{a_{1}}, and λa2\lambda_{a_{2}} inside the definitions of Lg\mathcal{L}_{g} and La\mathcal{L}_{a}. During all experiments, we have used the following empirically determined configuration: λe=0.1\lambda_{e}=0.1, λl=0.2\lambda_{l}=0.2, λr=1.0\lambda_{r}=1.0, λc=1.0\lambda_{c}=1.0, λs=50.0\lambda_{s}=50.0, λVGG=1.0\lambda_{\text{VGG}}=1.0, λg2=1.0\lambda_{g_{2}}=1.0, λa1=0.5\lambda_{a_{1}}=0.5, λa2=0.3\lambda_{a_{2}}=0.3 Additionally we found it beneficial to linearly increase the surface loss weight λg1\lambda_{g_{1}} from 1.01.0 to 15.015.0 over the duration of 100k interactions.

A.2 Inference timings

To create a mesh we run Marching Cubes over the distance field defined by ff. We first approximate the bounding box of the surface by probing at coarse resolution and use Octree sampling to progressively increase the resolution as we get closer to the surface. This allows us to extract meshes with high resolution without large computational overhead. We query ff in batches of 64364^{3} samples up to the desired resolution. The reconstruction of a mesh in a 2563256^{3} grid takes on average 1.211.21s using a single NVIDIA Tesla V100. Reconstructing a very dense mesh in a 5123512^{3} grid takes on average 5.725.72s. Hereby, a single batch of 64364^{3} samples takes 142.1142.1ms. In both cases, we query the features once which takes 243243ms. In practise, we also query ff a second time for color at the computed vertex positions which takes 56.556.5ms for meshes in 2563256^{3} and 223.3223.3ms for 5123512^{3}, respectively. Meshes computed in 2563256^{3} and 5123512^{3} grids contain about 100k and 400k vertices, respectively. Note that we can create meshes in arbitrary resolutions and our reconstructions can be rendered through sphere tracing without the need to generate an explicit mesh.

Appendix B Additional Results

In the sequel, we show additional results and comparisons. First, we demonstrate how we can automatically rig our reconstructions using a statistical body model. Then we conduct further comparisons on the PeopleSnapshot Dataset . Finally, we show additional qualitative results.

In fig. 10, we show examples of rigged and animated meshes created using our method. For rigging, we fit the statistical body model GHUM to the meshes. To this end, we first triangulate joint detections produced by an off-the-shelf 2D human keypoint detector on renderings of the meshes. We then fit GHUM to the triangulated joints and the mesh surface using ICP. Finally, we transfer the joints and blend weights from GHUM to our meshes. We can now animate our reconstructions using Mocap data or by sampling GHUM’s latent pose space. By fist reconstructing a static shape that we then rig in a secondary step, we avoid reconstruction errors of methods aiming for animation ready reconstruction in a single step .

B.2 Comparisons on the PeopleSnapshot Dataset

We use the public PeopleSnapshot dataset for further comparisons. The PeopleSnapshot dataset contains of people rotating in front of the camera while holding an A-pose. The dataset is openly available for research purposes. For this comparison we use only the first frame of each video. We compare once more with PIFuHD and additionally compare with the model-based approach Tex2Shape . Tex2Shape does not estimate the pose of the observed subject but only its shape. The shape is represented as displacements to the surface of the SMPL body model . In fig. 10 we show the results of both methods side-by-side with our method. Also in this comparison our method produces the most realistic results and additionally also reconstructs the surface color.

B.3 Qualitative Results

We show further qualitative results in fig. 11. Our methods performs well on a wide range of subjects, outfits, backgrounds, and illumination conditions. Further, despite never being trained on this type of data, our method performs extremely well on image of people with solid white background. In fig. 12 we show a number of examples. This essentially means, matting the image can be performed as a pre-processing step to boost the performance of our method in cases where the model has problems identifying foreground regions.

References