H3D-Net: Few-Shot High-Fidelity 3D Head Reconstruction

Eduard Ramon, Gil Triginer, Janna Escur, Albert Pumarola, Jaime Garcia, Xavier Giro-i-Nieto, Francesc Moreno-Noguer

Introduction

Recent learning based methods have shown impressive results in reconstructing 3D shapes from 2D images. These approaches can be roughly split into two main categories: model-based and model-free . The former incorporate prior knowledge obtained from training data to limit the space of feasible solutions, making these approaches well suited for few-shot and one-shot shape estimation. However, most model-based methods produce shapes that usually lack geometric detail and cannot handle arbitrary topology changes.

On the other hand, model-free approaches based on discrete representations like e.g. voxels, meshes or point-clouds, have the flexibility to represent a wider spectrum of shapes, although at the cost of being computationally tractable only for small resolutions or being restricted to fixed topologies. These limitations have been overcome by neural implicit representations , which can represent both geometry and appearance as a continuum, encoded in the weights of a neural network. have shown the success of such representations in learning detail-rich 3D geometry directly from images, with no 3D ground truth supervision. Unfortunately, the performance of these methods is currently conditioned to the availability of a large number of input views, which leads to a time consuming inference.

In this work we introduce H3D-Net, a hybrid scheme that combines the strengths of model-based and model-free representations by incorporating prior knowledge into neural implicit models for category-specific multi-view reconstruction. We apply this approach to the problem of few-shot full head reconstruction. In order to build the prior, we first use several thousands of raw incomplete scans to learn a space of Signed Distance Functions (SDF) representing 3D head shapes . At inference, this learnt shape prior is used to initalize and guide the optimization of an Implicit Differentiable Renderer (IDR) that, given a potentially reduced number of input images, estimates the full head geometry. The use of the learned prior enables faster convergence during optimization and prevents it from being trapped into local minima, yielding 3D shape estimates that capture fine details of the face, head and hair from just three input images (see. Fig. H3D-Net: Few-Shot High-Fidelity 3D Head Reconstruction).

We exhaustively evaluate our approach on a mid-resolution Multiview-Stereo (MVS) public dataset and on a high-resolution dataset we collected with a structured-light scanner, consisting of 10 3D full-head scans. The results show that we consistently outperform current state-of-the-art, both in a few-shot setting and when many input views are available. Importantly, the use of the prior also makes our approach very efficient, achieving competitive results in terms of accuracy about 20×\times faster than IDR . Our key contributions can be summarized as follows:

We introduce a method for reconstructing high quality full heads in 3D from small sets of in-the-wild images.

Our method is the first to use implicit functions for reconstructing 3D humans heads from multiple images and also to rival parametric and non-parametric models in 3D accuracy at the same time.

We devise a guided optimization approach to introduce a probabilistic shape prior into neural implicit models.

We collect and will release a new dataset containing high-resolution 3D full head scans, images, masks and camera poses for evaluation purposes, which we dub H3DS.

Related work

Model-based. 3D Morphable Models (3DMMs) have become the de facto representation used for few-shot 3D face reconstruction in-the-wild given that they lead to light-weight, fast and robust systems. Adopting 3DMMs as a representation, the 3D reconstruction problem boils down to estimating the small set of parameters that best represent a target 3D shape. This makes it possible to obtain 3D reconstructions from very few images and even a single input . Nevertheless, one of the main limitations of morphable models is their lack of expressiveness, specially for high frequencies. This issue has been addressed by learning a post processing that transfers the fine details from the image domain to the 3D geometry . Another limitation of 3DMMs is their inability to represent arbitrary shapes and topologies. Thus, they are not suitable for reconstructing full heads with hair, beard, facial accessories and upper body clothing.

Model-free. Model-free approaches build upon more generic representations, such as voxel-grids or meshes, in order to gain expressiveness and flexibility. Voxel-grids have been extensively used for 3D reconstruction and concretely for 3D face reconstruction . Their main limitation is that memory requirements grow cubically with resolution, and octrees have been proposed to address this issue. On the other hand, meshes are a more efficient representation for surfaces than voxel-grids, and are suitable for graphics applications. Meshes have been proposed for 3D face reconstruction in combination with graph neural networks . However, similarly to 3DMMs, meshes are also usually restricted to fixed topologies and are not suitable for reconstructing other elements beyond the face itself.

Implicit representations. Recently, implicit representations have been proposed to jointly address the memory limitations of voxel grids and the topological rigidity of meshes. Implicit representations model surfaces as a level-set of a coordinate-based continuous function, e.g. a signed distance function or an occupancy function. These functions, usually implemented as multi-layer perceptrons (MLPs), can theoretically express any shape with infinite resolution and a fixed memory footprint. Implicit methods can be divided in those that, at inference time, perform a single forward pass of a previously trained model , and those that overfit a model to a set of input images through an optimization process using implicit differentiable rendering . In the later, given that the inference is an optimization process, the obtained 3D reconstructions are more accurate. However, they are slow and require an important number of multi-view images, failing in few-shot setups as those we consider in this work.

Priors for implicit representations. Building priors for implicit representations has been addressed with two main purposes. The first consists in speeding up convergence of methods that perform an optimization at inference time using meta-learning techniques . The second is to find a space of implicit functions that represent a certain category using auto-decoders . However, have been used to solve tasks using 3D supervision, and it is still an open problem how to use these priors when the supervision signal is generated from 2D images.

As done with morphable models, implicit shape models can be used to constrain image-based 3D reconstruction systems to make them more reliable. Drawing inspiration from this idea, in this work we leverage implicit shape models to guide the optimization-based implicit 3D reconstruction method towards more accurate and robust solutions, even under few-shot in-the-wild scenarios.

Method

In order to approximate F\mathcal{F}, we propose to optimize a previously learnt probabilistic model Fz,θ0\mathcal{F}_{\mathbf{z},\theta_{0}}, that represents a prior distribution over 3D head SDFs. z\mathbf{z} and θ0\theta_{0} are a latent vector encoding specific shapes and the learnt parameters of an auto-decoder , respectively. Building on DeepSDF , we learn these parameters from thousands of incomplete scans. We describe this process in section 3.1.

At test time, the reconstruction process is reduced to finding the optimal parameters {z∗,θ∗}\{\mathbf{z}^{*},\theta^{*}\} such that Fz∗,θ∗∼F\mathcal{F}_{\mathbf{z}^{*},\theta^{*}}\sim{\mathcal{F}}. To that end, we compose the prior model Fz,θ0\mathcal{F}_{\mathbf{z},\theta_{0}}, which we also refer to as geometry network, with a rendering network Gϕ:(x,n,v)→c\mathcal{G}_{\phi}:(\mathbf{x},\mathbf{n},\mathbf{v})\rightarrow\mathbf{c} that models the RGB radiance emitted from a surface point x\mathbf{x} with normal n\mathbf{n} in a viewing direction v\mathbf{v}, and minimize a photometric error w.r.t. the input images Iv\mathbf{I}_{v}, as in . Moreover, we propose a two-step optimization schedule that prevents the reconstruction process from getting trapped into local minima and, as we shall see in the results section, leads to much more accurate, robust and realistic reconstructions. We describe the reconstruction step in section 3.2.

Given a set of MM scenes with associated raw 3D point clouds, we use the DeepSDF framework to learn a prior distribution of signed distance functions representing 3D heads, Fz,θ0\mathcal{F}_{\mathbf{z},\theta_{0}}. While the original DeepSDF formulation requires watertight meshes as training data to use signed distances as supervision, we use the Eikonal loss to learn directly from raw, and potentially incomplete, surface point clouds. In addition, Fourier features are used to overcome the spectral bias of MLPs towards low frequencies in low dimensional tasks . We illustrate the training and inference process of the prior model in figure 2-left.

For each scene, indexed by i=1,…,Mi=1,\ldots,M, we sample a subset of points Ps(i)\mathcal{P}_{\rm s}^{(i)} on the surface, and another set Pv(i)\mathcal{P}_{\rm v}^{(i)} uniformly taken from a volume containing the scene, and minimize the following objective:

where λ0\lambda_{0} and λ1\lambda_{1} are hyperparameters and LSurf(i)\mathcal{L}_{\rm Surf}^{(i)} accounts for the SDF error at surface points:

LEmb(i)\mathcal{L}_{\rm Emb}^{(i)} enforces a zero-mean multivariate-Gaussian distribution with spherical covariance σ2\sigma^{2} over the space of latent vectors:

Finally, LEik(i)\mathcal{L}_{\rm Eik}^{(i)} regularizes Fzi,θF_{\mathbf{z}_{i},\theta} with the Eikonal loss to ensure that it approximates a signed distance function by keeping its gradients close to unit norm:

This regularization across the whole volume is necessary given that our meshes are not watertight and only a subset of surface points is available as ground truth. .

After training, we have obtained the parameters θ0\theta_{0} that represent a space of human head SDFs. We can now draw signed distance functions of heads from Fz,θ0F_{\mathbf{z},\theta_{0}} by sampling the latent space z\mathbf{z}. We use this pre-trained model as the prior for the 3D reconstruction schedule described in the following section.

2 Prior-aided 3D Reconstruction

Given a new scene, for which no 3D information is provided at this point, we aim to approximate the SDF that implicitly encodes the surface of the head by only supervising in the image domain. To that end, we compose the previously learnt geometry probabilistic model Fz,θ0\mathcal{F}_{\mathbf{z},\theta_{0}} with the rendering network Gϕ\mathcal{G}_{\phi}, and supervise on the photometric error to find the optimal parameters z∗\mathbf{z}^{*}, θ∗\theta^{*} and ϕ∗\phi^{*}. The reconstruction process is illustrated in figure 2-right.

For every pixel coordinate pp of each input image Iv\mathbf{I}_{v}, we march a ray r={c0+tv∣t≥0}\mathbf{r}=\{\mathbf{c}_{0}+t\mathbf{v}|t\geq 0\}, where c0\mathbf{c}_{0} is the position of the associated camera Cv\mathbf{C}_{v}, and v\mathbf{v} the viewing direction. The intersection point xi\mathbf{x}_{i} between the ray r\mathbf{r} and the surface Sz,θ={x∣Fz,θ(x)=0}\mathcal{S}_{\mathbf{z},\theta}=\{\mathbf{x}|\mathcal{F}_{\mathbf{z},\theta}(\mathbf{x})=0\} can be efficiently found using sphere tracing . This intersection point can be made differentiable w.r.t z\mathbf{z} and θ\theta without having to store the gradients corresponding to all the forward passes of the geometry network, as shown in and generalized by . The following expression is exact in value and first derivatives:

Here zk\mathbf{z}_{k} and θk\theta_{k} denote the parameters of Fz,θ\mathcal{F}_{\mathbf{z},\theta} at iteration kk, and xs\mathbf{x}_{s} represents the intersection point made differentiable w.r.t. the geometry network parameters.

Next, we evaluate the mapping Gϕ\mathcal{G}_{\phi} at xs\mathbf{x}_{s}, n=∇xFz,θ(xs)\mathbf{n}=\nabla_{\mathbf{x}}\mathcal{F}_{\mathbf{z},\theta}(\mathbf{x}_{s}) and v\mathbf{v} to estimate the color c\mathbf{c} for the pixel pp in the image Iv\mathbf{I}_{v}:

Finally, in order to optimize the surface parameters z\mathbf{z} and θ\theta, and the rendering parameters ϕ\phi, we minimize the following loss :

where β0\beta_{0} and β1\beta_{1} are hyperparameters. We next describe each component of this loss. Let P\mathcal{P} be a mini-batch of pixels from view vv, PRGB\mathcal{P}_{\rm RGB} the subset of pixels whose associated ray intersects Sz,θ\mathcal{S}_{\mathbf{z},\theta} and which have a nonzero mask value, and PMask=P∖PRGB\mathcal{P}_{\rm Mask}=\mathcal{P}\setminus\mathcal{P}_{\rm RGB}. The LRGB(v)\mathcal{L}_{\rm RGB}^{(v)} is the photometric error, computed as:

LMask(v)\mathcal{L}_{\rm Mask}^{(v)} accounts for silhouette errors:

where sv,α=sigmoid(−αmin⁡t≥0Fz,θ(rt))s_{v,\alpha}={\rm sigmoid}(-\alpha\min_{t\geq 0}\mathcal{F}_{\mathbf{z},\theta}(\mathbf{r}_{t})) is the estimated silhouette, CE\rm CE is the binary cross-entropy and α\alpha is a hyperparameter. Lastly, LEik\mathcal{L}_{\rm Eik} encourages Fz,θ\mathcal{F}_{\mathbf{z},\theta} to approximate a signed distance function as in equation 4.

Instead of jointly optimizing all the parameters {z,θ,ϕ}\{\mathbf{z},\theta,\phi\} to minimize L\mathcal{L} we introduce a two-step optimization schedule which is more appropriate for auto-decoders like DeepSDF. We begin by initializing the geometry network Fz,θ\mathcal{F}_{\mathbf{z},\theta} with the previously learnt prior for human head SDFs, Fz,θ0\mathcal{F}_{\mathbf{z},\theta_{0}}, and a randomly sampled z0\mathbf{z}_{0} such that ∥z0∥<ϵ\left\lVert\mathbf{z}_{0}\right\rVert<\epsilon to stay near the mean of the latent space. In a first phase, we only optimize z\mathbf{z} and ϕ\phi as arg min⁡z,ϕL\operatornamewithlimits{arg\,min}_{\mathbf{z},\phi}\mathcal{L}, which is equivalent to the standard auto-decoder inference. By doing so, the resulting surface Sz∗,θ\mathcal{S}_{z^{*},\theta} is forced to stay within the learnt distribution of 3D heads. Once the geometry and the radiance mappings have reached an equilibrium, i.e. the optimization has converged, we unfreeze the decoder parameters θ\theta to fine-tune the whole model as arg min⁡z,θ,ϕL\operatornamewithlimits{arg\,min}_{\mathbf{z},\theta,\phi}\mathcal{L}.

In section 5, we empirically prove that by using this optimization schedule instead of optimizing all the parameters at once, the obtained 3D reconstructions are more accurate and less prone to artifacts, specially in few-shot setups.

Implementation details

Our implementation of the prior model closely follows the one proposed in , with the addition that we apply a positional encoding γF\gamma_{\mathcal{F}} to the input coordinates x\mathbf{x} with 6 log-linear spaced frequencies. The encoded 3D coordinates are concatenated with the z\mathbf{z} latent vector of size 256 and set as the input to the decoder. The decoder is a MLP of 8 layers with 512 neurons in each layer and single skip connection from the input of the decoder to the output of the 4th layer. We use Softplus as activation function in every layer except the last, where no activation is used. The prior model is trained for 100 epochs using Adam with standard parameters, learning rate of 10−410^{-4} and learning rate step decay of 0.5 every 15 epochs. The training takes approximately 50 minutes for a small dataset (500 scenes) and 10 hours for a large one (10,000 scenes).

The 3D reconstruction network is composed by the prior model described above and a mapping Gϕ\mathcal{G}_{\phi} that is split into two sub-networks Qρ\mathcal{Q}_{\rho} and Rη\mathcal{R}_{\eta} as shown in figure 2. Qρ\mathcal{Q}_{\rho} is a MLP implemented exactly as the decoder of the prior model, except for the input layer, which takes in a 3 dimensional vector, and the output layer, which outputs a dd-dimensional vector. As in , Rη\mathcal{R}_{\eta} is a smaller MLP composed by 4 layers, each 512 neurons wide, no skip connections, and ReLU activations except in the output layer which is tanh. We also apply the positional encodings γQ\gamma_{\mathcal{Q}} and γR\gamma_{\mathcal{R}} to xs\mathbf{x}_{s} with 6 and 4 log-linear spaced frequencies respectively. Each scene is trained for 2000 epochs using Adam with fixed learning rate of 10−410^{-4} and learning rate step decay of 0.5 at epochs 1000 and 1500. The scene reconstruction process takes approximately 25 minutes for scenes of 3 views and 4 hours and 15 minutes for scenes of 32.

All the experiments for both prior and reconstruction models have been performed using a single Nvidia RTX 2080Ti.

Experiments

In this section, we evaluate quantitatively and qualitatively our multi-view 3D reconstruction method. We empirically demonstrate that the proposed solution surpasses the state of the art in the few-shot and many-shot scenarios for 3D face and head reconstruction in-the-wild.

Prior training. In order to train the geometry prior, we use an internal dataset made of 3D head scans from 10,000 individuals. The dataset is perfectly balanced in gender and diverse in age and ethnicity. The raw data is automatically processed to remove internal mesh faces and non-human parts such as background walls. Finally, all the scenes are aligned by registering a template 3D model with non-rigid Iterative Closest Point (ICP) .

3DFAW . We evaluate our method in the 3DFAW dataset. This dataset provides videos recorded in front, and around, the head of a person in static position as well as mid-resolution 3D ground truth of the facial region. We select 5 male and 5 female scenes and use them to evaluate only the facial region.

H3DS. We introduce a new dataset called H3DS, the first dataset containing high resolution full head 3D textured scans and 360º images with associated ground truth camera poses and ground truth masks. The 3D geometry has been captured using a structured light scanner, which leads to more precise ground truth geometries than the ones from 3DFAW , which were generated using Multi-View Stereo (MVS). The dataset consists of 10 individuals, 50% man and 50% woman. We use this dataset to evaluate the accuracy of the different methods in both the full head and the facial regions. We plan to release, maintain and progressively grow the H3DS dataset.

2 Experiments setup

We use the 3DMM-based methods MVFNet and DFNRMVS , and the model-free method IDR as baselines to compare against H3D-Net.

In the few-shot scenario (3 views), all the methods are evaluated on the 3DFAW and H3DS datasets. To benchmark our method when more than 3 views are available, we compare it against IDR on the H3DS dataset.

The evaluation criteria have been the same for all methods and in all the experiments. The predicted 3D reconstruction is roughly aligned with the ground truth mesh using manually annotated landmarks, and then refined with rigid ICP . Then, we compute the unidirectional Chamfer distance from the predicted reconstruction to the ground truth. All the distances are computed in millimeters.

We report metrics in two different regions, the face and the full head. For the finer evaluation in the face region, we cut both the reconstructions and the ground truth using a sphere of 95 mm radius and with center at the tip of the nose of the ground truth mesh, and refine the alignment with ICP as in . Then, we compute the Chamfer distance in this sub-region. For the full head evaluation, the ICP alignment is performed using an annotated region that includes the face, the ears, and the neck, since it is a region visible in all view configurations (3, 4, 8, 16 and 32). These configurations are defined by their yaw angles as follow: V3={0,±45}\mathcal{V}_{3}=\{0,\pm 45\}, V4={±45,±90}\mathcal{V}_{4}=\{\pm 45,\pm 90\} and VN={360Ni}i=1N\mathcal{V}_{N}=\{\frac{360}{N}i\}_{i=1}^{N} for N=8,16,32N={8,16,32}. In this case, the Chamfer distance is computed for all the vertices of the reconstruction.

3 Ablation study

We conduct an ablation study on the H3DS dataset in the few-shot scenario (3 views) and show the numerical results in table 1, and the qualitative results in figure 3. First, we reconstruct the scenes without prior and without schedule (a). In this case, the geometry network is initialized using geometric initialization , representing a sphere of radius one at the beginning of the optimization. Then, we initialize the geometry network with two different priors, a small one trained on 500 subjects (b), and a large one trained on 10,000 subjects (c), and perform the reconstructions without schedule. As it can be observed, initializing the geometry network with a previously learnt prior leads to smoother and more plausible surfaces, specially when more subjects have been used to train it. It is important to note that the benefits of the initialization are not only due to a better initial shape, but also to the ability of the initial weights to generalize to unseen shapes, which is greater in the large prior model. Finally, we initialize the geometry network with the large prior and use the proposed optimization schedule during the reconstruction process. It can be observed how the resulting 3D heads resemble much more to the ground truth in terms of finer facial details.

Given the notable effect that the number of samples has in the learnt prior representations and in the resulting 3D reconstructions as well, we visualize latent space interpolations in figure 4. To that end, we optimize the latent vector for two ground truth 3D scans as shown in figure 2-left-bottom in order minimize the loss 1. Then, we interpolate between the two optimal latent vectors. As it can be observed, the 3D reconstructions resulting from the interpolation in z\mathbf{z} space of the large prior model are more detailed and plausible than the ones from the small prior model, suggesting that the later achieves poorer generalization.

4 Quantitative results

Quantitative results in terms of surface error are reported in table 2. Remarkably, H3D-Net outperforms both 3DMM-based methods in the few-shot regime, and the model-free method IDR when the largest number of views (32) are available. It is worth noting how the enhancement due to the prior is more significant as the number of views decreases. Nevertheless, the prior does not prevent the model from becoming more accurate when more views are available, which is a current limitation of model-based approaches.

We also analyze the trade-off between the optimization time and the accuracy in IDR and H3D-Net for the case of 32 views, which we illustrate in figure 5. It can be observed that, despite reaching similar errors asymptotically, in average our method achieves the best performance attained by IDR much faster. In particular, we report convergence gains of 20×\times for the facial region error and 4×\times for the full head. Moreover, the smaller variance observed in H3D-Net (blue) indicates that it is a more stable method.

5 Qualitative results

Quantitative results show improvements over the baselines in both few-shot and many-shot setups. Here, we study how this is translated into the reconstructed 3D shape.

In figure 6, we qualitatively evaluate the three baselines and H3D-Net for the case of 3 input views. As expected, IDR is the worst performing model in this scenario, generating reconstructions with artifacts and with no resemblance to human faces. On the other hand, 3DMM-based models achieve more plausible shapes, but they struggle to capture fine details. H3D-Net, in contrast, is able to capture much more detail and reduce significantly the errors over the whole face and, concretely, in difficult areas such as the nose, the eyebrows, the cheeks, and the chin.

We also evaluate the impact that varying the number of available views has on the reconstructed surface, and compare our method to IDR . As shown in figure 7, H3D-Net is able to obtain surfaces with less error (greener) with far fewer views, which is consistent with the quantitative results reported in table 2. Notably, it can also be observed that, even when errors are numerically similar (first and third columns), the reconstructions from H3D-Net are much more realistic. In addition, H3D-Net improvements are especially notable within the face region. We attribute this to the fact that training data used to build the prior model is more rich in this area, whereas training examples frequently present holes in other parts of the head.

Conclusions

In this work we have presented H3D-Net, a method for high-fidelity 3D head reconstruction from small sets of in-the-wild images with associated head masks and camera poses. Our method combines a pre-trained probabilistic model, which represents a distribution of head SDFs, with an implicit differentiable renderer that allows direct supervision in the image domain. By constraining the reconstruction process with the prior model, we are able to robustly recover detailed 3D human heads, including hair and shoulders, from only three input images. After a thorough quantitative and qualitative evaluation, our experiments show that our method outperforms both model-based methods in the few-shot setup and model-free methods when a large number of views are available. One limitation of our method is that it still requires several minutes to generate 3D reconstructions. An interesting direction for future work is to use more efficient representations for SDFs in order to speed up the optimization process. We also find it promising to introduce texture priors, which could reduce the convergence time and reduce the final overall error.

Acknowledgments

This work has been partially funded by the Spanish government with the projects MoHuCo PID2020-120049RB-I00, DeeLight PID2020-117142GB-I00 and Maria de Maeztu Seal of Excellence MDM-2016-0656, and by the Government of Catalonia under the industrial doctorate 2017 DI 028.

References