GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction

Baris Gecer, Stylianos Ploumpis, Irene Kotsia, Stefanos Zafeiriou

Introduction

Estimation of the 3D facial surface and other intrinsic components of the face from single images (e.g., albedo, etc.) is a very important problem at the intersection of computer vision and machine learning with countless applications (e.g., face recognition, face editing, virtual reality). It is now twenty years from the seminal work of Blanz and Vetter which showed that it is possible to reconstruct shape and albedo by solving a non-linear optimization problem that is constrained by linear statistical models of facial texture and shape. This statistical model of texture and shape is called a 3D Morphable Model (3DMM). Arguably the most popular publicly available 3DMM is the Basel model built from 200 people . Recently, large scale statistical models of face and head shape have been made publicly available .

For many years 3DMMs and its variants were the methods of choice for 3D face reconstruction . Furthermore, with appropriate statistical texture models on image features such as Scale Invariant Feature Transform (SIFT) and Histogram Of Gradients (HOG), 3DMM-based methodologies can still achieve state-of-the-art performance in 3D shape estimation on images captured under unconstrained conditions . Nevertheless, those methods can reconstruct only the shape and not the facial texture. Another line of research in decouples texture and shape reconstruction. A standard linear 3DMM fitting strategy is used for face reconstruction followed by a number of steps for texture completion and refinement. In these papers , the texture looks excellent when rendered under professional renderers (e.g., Arnold), nevertheless when the texture is overlaid on the images the quality significantly drops Please see the supplementary materials for a comparison with ..

In the past two years, a lot of work has been conducted on how to harness Deep Convolutional Neural Networks (DCNNs) for 3D shape and texture reconstruction. The first such methods either trained regression DCNNs from image to the parameters of a 3DMM or used a 3DMM to synthesize images and formulate an image-to-image translation problem using DCNNs to estimate the depthThe depth was afterwards refined by fitting a 3DMM and then changing the normals by using image features. . The more recent unsupervised DCNN-based methods are trained to regress 3DMM parameters from identity features by making use of differentiable image formation architectures and differentiable renderers .

The most recent methods such as use both the 3DMM model, as well as additional network structures (called correctives) in order to extend the shape and texture representation. Even though the paper shows that the reconstructed facial texture has indeed more details than a texture estimated from a 3DMM , it is still unable to capture high-frequency details in texture and subsequently many identity characteristics (please see the Fig. 4). Furthermore, because the method permits the reconstructions to be outside the 3DMM space, it is susceptible to outliers (e.g., glasses etc.) which are baked in shape and texture. Although rendering networks (i.e. trained by VAE ) generates outstanding quality textures, each network is capable of storing up to few individuals whom should be placed in a controlled environment to collect ∼20{\sim}20 millions of images.

In this paper, we still propose to build upon the success of DCNNs but take a radically different approach for 3D shape and texture reconstruction from a single in-the-wild image. That is, instead of formulating regression methodologies or auto-encoder structures that make use of self-supervision , we revisit the optimization-based 3DMM fitting approach by the supervision of deep identity features and by using Generative Adversarial Networks (GANs) as our statistical parametric representation of the facial texture.

In particular, the novelties that this paper brings are:

We show for the first time, to the best of our knowledge, that a large-scale high-resolution statistical reconstruction of the complete facial surface on an unwrapped UV space can be successfully used for reconstruction of arbitrary facial textures even captured in unconstrained recording conditionsIn the very recent works, it was shown that it is feasible to reconstruct the non-visible parts a UV space for facial texture completion and that GANs can be used to generate novel high-resolution faces. Nevertheless, our work is the first one that demonstrates that a GAN can be used as powerful statistical texture prior and reconstruct the complete texture of arbitrary facial images..

We formulate a novel 3DMM fitting strategy which is based on GANs and a differentiable renderer.

We devise a novel cost function which combines various content losses on deep identity features from a face recognition network.

We demonstrate excellent facial shape and texture reconstructions in arbitrary recording conditions that are shown to be both photorealistic and identity preserving in qualitative and quantitative experiments.

History of 3DMM Fitting

Our methodology naturally extends and generalizes the ideas of texture and shape 3DMM using modern methods for representing texture using GANs, as well as defines loss functions using differentiable renderers and very powerful publicly available face recognition networks . Before we define our cost function, we will briefly outline the history of 3DMM representation and fitting.

The first step is to establish dense correspondences between the training 3D facial meshes and a chosen template with fixed topology in terms of vertices and triangulation.

Traditionally 3DMMs use a UV map for representing texture. UV maps help us to assign 3D texture data into 2D planes with universal per-pixel alignment for all textures. A commonly used UV map is built by cylindrical unwrapping the mean shape into a 2D flat space formulation, which we use to create an RGB image IUV\mathbf{I}_{UV}. Each vertex in the 3D space has a texture coordinate tcoordt_{coord} in the UV image plane in which the texture information is stored. A universal function exists, where for each vertex we can sample the texture information from the UV space as T=P(IUV,tcoord)\mathbf{T}=\mathcal{P}(\mathbf{I}_{UV},t_{coord}).

In order to define a statistical texture representation, all the training texture UV maps are vectorized and Principal Component Analysis (PCA) is applied. Under this model any test texture T0\mathbf{T}^{0} is approximated as a linear combination of the mean texture mt\mathbf{m}_{t} and a set of bases Ut\mathbf{U}_{t} as follows:

where pt\mathbf{p}_{t} is the texture parameters for the text sample T0\mathbf{T}^{0}. In the early 3DMM studies, the statistical model of the texture was built with few faces captured in strictly controlled conditions and was used to reconstruct the test albedo of the face. Since, such texture models can hardly represent faces captured in uncontrolled recording conditions (in-the-wild). Recently it was proposed to use statistical models of hand-crafted features such as SIFT or HoG directly from in-the-wild faces. The interested reader is referred to for more details on texture models used in 3DMM fitting algorithms.

The recent 3D face fitting methods still make use of similar statistical models for the texture. Hence, they can naturally represent only the low-frequency components of the facial texture (please see Fig. 4).

1.2 Shape

The method of choice for building statistical models of facial or head 3D shapes is still PCA . Assuming that the 3D shapes in correspondence comprise of NN vertexes, i.e. s=[x1T,…,xNT]T=[x1,y1,z1,…,xN,yN,zN]T\mathbf{s}={\left[\mathbf{x}_{1}^{\mathsf{T}},\ldots,\mathbf{x}_{N}^{\mathsf{T}}\right]}^{\mathsf{T}}={\left[x_{1},y_{1},z_{1},\ldots,x_{N},y_{N},z_{N}\right]}^{\mathsf{T}}. In order to represent both variations in terms of identity and expression, generally two linear models are used. The first is learned from facial scans displaying the neutral expression (i.e., representing identity variations) and the second is learned from displacement vectors (i.e., representing expression variations). Then a test facial shape S(ps,e)\mathbf{S}(\mathbf{p}_{s,e}) can be written as

2 Fitting

3D face and texture reconstruction by fitting a 3DMM is performed by solving a non-linear energy based cost optimization problem that recovers a set of parameters p=[ps,e,pt,pc,pl]\mathbf{p}=[\mathbf{p}_{s,e},\mathbf{p}_{t},\mathbf{p}_{c},\mathbf{p}_{l}] where pc\mathbf{p}_{c} are the parameters related to a camera model and pl\mathbf{p}_{l} are the parameters related to an illumination model. The optimization can be formulated as:

where I0\mathbf{I}^{0} is the test image to be fitted and W\mathbf{W} is a vector produced by a physical image formation process (i.e., rendering) controlled by p\mathbf{p}. Finally, RegReg is the regularization term that is mainly related to texture and shape parameters.

Various methods have been proposed for numerical optimization of the above cost functions . A notable recent approach is which uses handcrafted features (i.e., H\mathbf{H}) for texture representation simplified the cost function as:

where ∣∣a∣∣A2=aTAa||\mathbf{a}||_{\mathbf{A}}^{2}=\mathbf{a}^{T}\mathbf{A}\mathbf{a}, A\mathbf{A} is the orthogonal space to the statistical model of the texture and pr\mathbf{p}^{r} is the set of reduced parameters pr={ps,e,pc}\mathbf{p}^{r}=\{\mathbf{p}_{s,e},\mathbf{p}_{c}\}. The optimization problem in Eq. 4 is solved by Gauss-Newton method. The main drawback of this method is that the facial texture in not reconstructed.

In this paper, we generalize the 3DMM fittings and introduce the following novelties:

We use a GAN on high-resolution UV maps as our statistical representation of the facial texture. That way we can reconstruct textures with high-frequency details.

We replace physical image formation stage with a differentiable renderer to make use of first order derivatives (i.e., gradient descent). Unlike its alternatives, gradient descent provides computationally cheaper and more reliable derivatives through such deep architectures (i.e., above-mentioned texture GAN and identity DCNN).

Approach

We propose an optimization-based 3D face reconstruction approach from a single image that employs a high fidelity texture generation network as statistical prior as illustrated in Fig. 2. To this end, the reconstruction mesh is formed by 3D morphable shape model; textured by the generator network’s output UV map; and projected into 2D image by a differentiable renderer. The distance between the rendered image and the input image is minimized in terms of a number of cost functions by updating the latent parameters of 3DMM and the texture network with gradient descent. We mainly formulate these functions based on rich features of face recognition network for smoother convergence and landmark detection network for alignment and rough shape estimation.

The following sections introduce firstly our novel texture model that employs a generator network trained by progressive growing GAN framework. After describing the procedure for image formation with differentiable renderer, we formulate our cost functions and the procedure for fitting our shape and texture models onto a test image.

Although conventional PCA is powerful enough to build a decent shape and texture model, it is often unable to capture high frequency details and ends up having blurry textures due to its Gaussian nature. This becomes more apparent in texture modelling which is a key component in 3D reconstruction to preserve identity as well as photo-realism.

GANs are shown to be very effective at capturing such details. However, they suffer from preserving 3D coherency of the target distribution when the training images are semi-aligned. We found that a GAN trained with UV representation of real textures with per pixel alignment avoids this problem and is able to generate realistic and coherent UVs from 99.9%99.9\% of its latent space while at the same time generalizing well to unseen data.

In order to take advantage of this perfect harmony, we train a progressive growing GAN to model distribution of UV representations of 10,000 high resolution textures and use the trained generator network

as texture model that replaces 3DMM texture model in Eq. 1.

While fitting with linear models, i.e. 3DMM, is as simple as linear transformation, fitting with a generator network can be formulated as an optimization that minimizes per-pixel Manhattan distance between target texture in UV space Iuv\mathbf{I}_{uv} and the network output G(pt)\mathcal{G}(\mathbf{p}_{t}) with respect to the latent parameter pt\mathbf{p}_{t}, i.e. min⁡pt∣G(pt)−Iuv∣\min_{\mathbf{p}_{t}}|\mathcal{G}(\mathbf{p}_{t})-\mathbf{I}_{uv}|.

2 Differentiable Renderer

Following , we employ a differentiable renderer to project 3D reconstruction into a 2D image plane based on deferred shading model with given camera and illumination parameters. Since color and normal attributes at each vertex are interpolated at the corresponding pixels with barycentric coordinates, gradients can be easily backpropagated through the renderer to the latent parameters.

A 3D textured mesh at the center of Cartesian origin $isprojectedonto2Dimageplanebyapinholecameramodelwiththecamerastandingatis projected onto 2D image plane by a pinhole camera model with the camera standing at[x_{c},y_{c},z_{c}],directedtowards, directed towards[x_{c}^{\prime},y_{c}^{\prime},z_{c}^{\prime}]andwiththefocallengthand with the focal lengthf_{c}.Theilluminationismodelledbyphongshadinggiven1)directlightsourceat3Dcoordinates. The illumination is modelled by phong shading given 1) direct light source at 3D coordinates[x_{l},y_{l},z_{l}]withcolorvalueswith color values[r_{l},g_{l},b_{l}],and2)colorofambientlighting, and 2) color of ambient lighting[r_{a},g_{a},b_{a}]$.

Finally, we denote the rendered image given geometry (ps,e\mathbf{p}_{s,e}), texture (pt\mathbf{p}_{t}), camera (pc=[xc,yc,zc,xc′,yc′,zc′,fc]\mathbf{p}_{c}=[x_{c},y_{c},z_{c},x_{c}^{\prime},y_{c}^{\prime},z_{c}^{\prime},f_{c}]) and lighting parameters (pl=[xl,yl,zl,rl,gl,bl,ra,ga,ba]\mathbf{p}_{l}=[x_{l},y_{l},z_{l},r_{l},g_{l},b_{l},r_{a},g_{a},b_{a}] by the following:

where we construct shape mesh by 3DMM as given in Eq. 2 and texture by GAN generator network as in Eq. 5. Since our differentiable renderer supports only color vectors, we sample from our generated UV map to get vectorized color representation as explained in Sec. 2.1.1.

Additionally, we render a secondary image with random expression, pose and illumination in order to generalize identity related parameters well with those variations. We sample expression parameters from a normal distribution as pe^∼N(μ=0,σ=0.5)\hat{\mathbf{p}_{e}}\sim\mathcal{N}(\mu=0,\sigma=0.5) and sample camera and illumination parameters from the Gaussian distribution of 300W-3D dataset as p^c∼N(μc^,σc^)\hat{\mathbf{p}}_{c}\sim\mathcal{N}(\hat{\mu_{c}},\hat{\sigma_{c}}) and pl^∼N(μl^,σl^)\hat{\mathbf{p}_{l}}\sim\mathcal{N}(\hat{\mu_{l}},\hat{\sigma_{l}}). This rendered image of the same identity as IR\mathbf{I}^{\mathcal{R}} (i.e., with same ps\mathbf{p}_{s} and pt\mathbf{p}_{t} parameters) is expressed by the following:

3 Cost Functions

Given an input image I0\mathbf{I}^{0}, we optimize all of the aforementioned parameters simultaneously with gradient descent updates. In each iteration, we simply calculate the forthcoming cost terms for the current state of the 3D reconstruction, and take the derivative of the weighted error with respect to the parameters using backpropagation.

We formulate an additional identity loss on the rendered image I^R\hat{\mathbf{I}}^{\mathcal{R}} that is rendered with random pose, expression and lighting. This loss ensures that our reconstruction resembles the target identity under different conditions. We formulate it by replacing IR\mathbf{I}^{\mathcal{R}} by I^R\hat{\mathbf{I}}^{\mathcal{R}} in Eq. 8 and it is denoted as L^id\hat{\mathcal{L}}_{id}.

3.2 Content Loss

Face recognition networks are trained to remove all kinds of attributes (e.g. expression, illumination, age, pose) other than abstract identity information throughout the convolutional layers. Despite their strength, the activations in the very last layer discard some of the mid-level features that are useful for 3D reconstruction, e.g. variations that depend on age. Therefore we found it effective to accompany identity loss by leveraging intermediate representations in the face recognition network that are still robust to pixel-level deformations and not too abstract to miss some details. To this end, normalized euclidean distance of intermediate activations, namely content loss, is minimized between input and rendered image with the following loss term:

3.3 Pixel Loss

3.4 Landmark Loss

The alignment error is achieved by point-to-point euclidean distances between detected landmark locations of the input image and 2D projection of the 3D reconstruction landmark locations that is available as meta-data of the shape model. Since landmark locations of the reconstruction heavily depend on camera parameters, this loss is great a source of information the alignment of the reconstruction onto input image and is formulated as following:

4 Model Fitting

We first roughly align our reconstruction to the input image by optimizing shape, expression and camera parameters by: min⁡prE(pr)=λlanLlan\min_{\mathbf{p}^{r}}\mathcal{E}(\mathbf{p}^{r})=\lambda_{lan}\mathcal{L}_{lan}. We then simultaneously optimize all of our parameters with gradient descent and backpropagation so as to minimize weighted combination of above loss terms in the following:

where we weight each of our loss terms with λ\lambda parameters. In order to prevent our shape and expression models and lighting parameters from exaggeration to arbitrarily bias our loss terms, we regularize those parameters by Reg({ps,e,pl})Reg(\{\mathbf{p}_{s,e},\mathbf{p}_{l}\}).

While the proposed approach can fit a 3D reconstruction from a single image, one can take advantage of more images effectively when available, e.g. from a video recording. This often helps to improve reconstruction quality under challenging conditions, e.g. outdoor, low resolution. While state-of-the-art methods follow naive approaches by averaging either the reconstruction or features-to-be-regressed before making a reconstruction, we utilize the power of iterative optimization by averaging identity reconstruction parameters (ps,pt\mathbf{p}_{s},\mathbf{p}_{t}) after every iteration. For an image set I={I0,I1,…,Ii,…,Ini}\mathbf{I}=\{\mathbf{I}^{0},\mathbf{I}^{1},\dots,\mathbf{I}^{i},\dots,\mathbf{I}^{n_{i}}\}, we reformulate our parameters as p=[ps,pei,pt,pci,pli]\mathbf{p}=[\mathbf{p}_{s},\mathbf{p}_{e}^{i},\mathbf{p}_{t},\mathbf{p}_{c}^{i},\mathbf{p}_{l}^{i}] in which we average shape and texture parameters by the following:

Experiments

This section demonstrates the excellent performance of the proposed approach for 3D face reconstruction and shape recovery. We verify this by qualitative results in Figures GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction, 3, qualitative comparisons with the state-of-the-art in Sec. 4.2 and quantitative shape reconstruction experiment on a database with ground truth in Sec. 4.3.

For all of our experiments, a given face image is aligned to our fixed template using 68 landmark locations detected by an hourglass 2D landmark detection . For the identity features, we employ ArcFace network’s pretrained models. For the generator network G\mathcal{G}, we train a progressive growing GAN with around 10,000 UV maps from at the resolution of 512×512512\times 512. We use the Large Scale Face Model for 3DMM shape model with ns=158n_{s}=158 and the expression model learned from 4DFAB database with ne=29n_{e}=29. During fitting process, we optimize parameters using Adam Solver with 0.01 learning rate. And we set our balancing factors as the following: λid:2.0,λ^id:2.0,λcon:50.0,λpix:1.0,λlan:0.001,λreg:{0.05,0.01}\lambda_{id}:2.0,\hat{\lambda}_{id}:2.0,\lambda_{con}:50.0,\lambda_{pix}:1.0,\lambda_{lan}:0.001,\lambda_{reg}:\{0.05,0.01\}. The Fitting converges in around 30 seconds on an Nvidia GTX 1080 TI GPU for a single image.

2 Qualitative Comparison to the State-of-the-art

Fig. 4 compares our results with the most recent face reconstruction studies on a subset of MoFA test-set. The first four rows after input images show a comparison of our shape and texture reconstructions to and the last three rows show our reconstructed geometries without texture compared to . All in all, our method outshines all others with its high fidelity photorealistic texture reconstructions. Both of our texture and shape reconstructions manifest strong identity characteristics of the corresponding input images from the thickness and shape of the eyebrows to wrinkles around the mouth and forehead.

3 3D shape recovery on MICC dataset

We evaluate the shape reconstruction performance of our method on MICC Florence 3D Faces dataset (MICC) in Table 1. The dataset provides 3D scans of 53 subjects as well as their short video footages under three difficulty settings: ’cooperative’, ’indoor’ and ’outdoor’. Unlike which processes all the frames in a video, we uniformly sample only 5 frames from each video regardless of their zoom level. And, we run our method with multi-image support for these 5 frames for each video separately as shown in Eq. 13. Each test mesh is cropped at a radius of 9595mm around the tip of the nose according to in order to evaluate the shape recovery of the inner facial mesh. We perform dense alignment between each predicted mesh and its corresponding ground truth mesh, by implementing an iterative closest point (ICP) method . As evaluation metric, we follow to measure the error by average symetric point-to-plane distance.

Table 1 reports the normalized point-to-plain errors in millimeters. It is evident that we have improved the absolute error compared to the other two state-of-the-art methods by 36%36\%. Our results are shown to be consistent across all different settings with minimal standard deviation from the mean error.

4 Ablation Study

Fig. 5 shows an ablation study on our method where the full model reconstructs the input face better than its variants, something that suggests that each of our components significantly contributes towards a good reconstruction. Fig. 5(c) indicates albedo is well disentangled from illumination and our model capture the light direction accurately.

While Fig. 5(d-f) shows each of the identity terms contributes to preserve identity, Fig. 5(h) demonstrates the significance identity features altogether. Still, overall reconstruction utilizes pixel intensities to capture better albedo and illumination as shown in Fig. 5(g). Finally, Fig. 5(i) shows the superiority of our textures over PCA-based ones.

Conclusion

In this paper, we revisit optimization-based 3D face reconstruction under a new perspective, that is, we utilize the power of recent machine learning techniques such as GANs and face recognition network as statistical texture model and as energy function respectively.

To the best of our knowledge, this is the first time that GANs are used for model fitting and they have shown excellent results for high quality texture reconstruction. The proposed approach shows identity preserving high fidelity 3D reconstructions in qualitative and quantitative experiments.

Baris Gecer is funded by the Turkish Ministry of National Education. Stefanos Zafeiriou acknowledges support by EPSRC Fellowship DEFORM (EP/S010203/1) and a Google Faculty Award.

References

Appendix A Experiments on LFW

In order to evaluate identity preservation capacity of the proposed method, we run two face recognition experiments on Labelled Faces in the Wild (LFW) dataset . Following , we feed real LFW images and rendered images of their 3D reconstruction by our method to a pretrained face recognition network, namely VGG-Face. We then compute the activations at the embedding layer and measure cosine similarity between 1) real and rendered images and 2) renderings of same/different pairs.

In Fig. 6 and 7, we have quantitatively showed that our method is better at identity preservation and photorealism (i.e., as the pretrained network is trained by real images) than other state-of-the-art deep 3D face reconstruction approaches .

Appendix B More Qualitative Results

Figures 8, 9, 10, and 11 illustrate the reconstructions of our method under different settings in comparison to the other state-of-the-art methods. Please see figure captions for detailed explanation.