Neural Face Editing with Intrinsic Image Disentangling

Zhixin Shu, Ersin Yumer, Sunil Hadap, Kalyan Sunkavalli, Eli Shechtman, Dimitris Samaras

Introduction

Understanding and manipulating face images in-the-wild is of great interest to the vision and graphics community, and as a result, has been extensively studied in previous work. This ranges from techniques to relight portraits , to edit or exaggerate expressions , and even drive facial performance . Many of these methods start by explicitly reconstructing face attributes like geometry, texture, and illumination, and then edit these attributes to edit the image. However, reconstructing these attributes is a challenging and often ill-posed task; previous techniques deal with this by either assuming richer data (e.g., RGBD video streams) or a strong prior on the reconstruction that is adapted to the particular editing task that they seek to solve (e.g., low-dimensional geometry ). As a result, these techniques tend to be both costly and not generalize well to the large variations in facial identity and appearance that exist in images-in-the-wild.

In this work, our goal is to learn a compact, meaningful manifold of facial appearance, and enable face edits by walking along paths on this manifold. The remarkable success of morphable face models – where face geometry and texture are represented using low-dimensional linear manifolds – indicates that this is possible for facial appearance. However, we would like to handle a much wider range of manipulations including changes in viewpoint, lighting, expression, and even higher-level attributes like facial hair and age – aspects that cannot be represented using previous models. In addition, we would like to learn this model without the need for expensive data capture .

To this end, we build on the success of deep learning – especially unsupervised auto-encoder networks – to learn “good” representations from large amounts of data . Trivially applying such approaches to our problem leads to representations that are not meaningful, making the subsequent editing challenging. However, we have (approximate) models for facial appearance in terms of intrinsic face properties like geometry (surface normals), material properties (diffuse albedo), and illumination. We leverage this by designing the network to explicitly infer these properties and introducing an in-network forward rendering model that reconstructs the image from them. Merely introducing these factors into the network is not sufficient; because of the ill-posed nature of the inverse rendering problem, the learnt intrinsic properties can be arbitrary. We guide the network by imposing priors on each of these intrinsic properties; these include a morphable model-driven prior on the geometry, a Retinex-based prior on the albedo, and an assumption of low-frequency spherical harmonics-based lighting model . By combining these constraints with adversarial supervision on image reconstruction, and weak supervision on the inferred face intrinsic properties, our network is able to learn disentangled representations of facial appearance.

Since we work with natural images, faces appear in front of arbitrary backgrounds, where the physical constraints of the face do not apply. Therefore, we also introduce a matte layer to separate the foreground (i.e., the face) from the image background. This enables us to provide optimal reconstruction pathways in the network specifically designed for faces, without distorting the background reconstruction.

Our network naturally exposes low-dimensional manifold embeddings for each of the intrinsic properties, which in turn enables direct and data-driven semantic editing from a single input image. Specifically, we demonstrate direct illumination editing with explicit spherical harmonics lighting built into the network, as well as latent space manifold traversal for semantically meaningful expression edits such as smiling, and more structurally global edits such as aging. We show that by constraining physical properties that do not affect the target edits, we can achieve significantly more realistic results compared to other learning-based face editing approaches.

Our main contributions are: (1) We introduce an end-to-end generative network specifically designed for the understanding and editing of face images in the wild; (2) We encode the image formation and shading processes as in-network layers enabling the disentangling in the the latent space, of physically based rendering elements such as shape, illumination, and albedo; (3) We introduce statistical loss functions (such as batchwise white shading (BWS) corresponding to color consistency theory ) to improve disentangling latent representations.

Related Work

Face Image Manipulation. Face modeling and editing is an extensively studied topic in vision and graphics. Blanz and Vetter showed that facial geometry and texture can be approximated by a low-dimensional morphable face model. This model and its variants have been used for a variety of tasks including relighting , face attribute editing , expression editing , authoring facial performances , and aging . Another class of techniques uses coarse geometry estimates to drive image-based editing tasks . Each of these works develops techniques that are specifically designed for their application and often can not be generalized to other tasks. In contrast, our work aims to learn a general manifold for facial appearance that can support all these tasks.

Intrinsic decompositions. Barrow and Tanenbaum proposed the concept of decomposing images into their physical intrinsic components such as surface normals, surface shading, etc. Barron and Mallik extended this decomposition assuming a Lambertian rendering model with low-frequency illumination and made use of extensive priors on geometry, albedo, and illumination. This rendering model has also been used in face relighting and shape-from-shading-based face reconstruction . We use a similar rendering model in our work, but learn a face-specific appearance model by training a deep network with weak supervision.

Neural Inverse Rendering. Generative network architectures have shown to be effective for image manipulation. Kulkarni et al. utilized a variational auto-encoder (VAE) for synthesizing novel variations of the input image where the objects pose and lighting conditions are altered. Yang et al. demonstrated novel view synthesis for the object in a given image, where view specific properties were disentangled in latent space utilizing a recurrent network. In contrast, Tatarchenko et al. used an autoencoder style network for the same task, where transformations were encoded through a secondary input stream and mixed with the input image in the latent space. Recently, Yan et al. used a VAE variant and layered representations to generate images with specific semantic attributes. We adopt their background-foreground disentangling scheme through an in-network matte layer.

Face Representation Learning. Face representation learning is generally performed with a standard convolutional neural network trained for a recognition or labeling task . Such approaches often need a significantly large amount of data since the network is treated as a black box. Synthetically boosting the dataset using normalizations and augmentations has proven useful. Most recently, Masi et al. used face fitting using morphable models similar to our approach, but used the resulting 3D faces to generate more data for traditional recognition network training. Even though such learned representations are powerful, especially in recognition, they are not straight forward to utilize for face editing.

Recently, Gardner et al. demonstrated face editing through a standard recognition network. Since the network does not have a natural generation pathway, they use a two step optimization procedure (one in latent space, and one at low level feature space) to reconstruct the edited image. This, combined with the fact that they use a global latent space, leads to unintended changes and artifacts. On the other hand, our generative auto-encoder style network allows for a physically meaningful latent space disentangling, thereby solving both problems: we constrain semantic edits to their corresponding latent representation, and our decoder generates the editing result in a single forward pass.

We formulate the face generation process as an end-to-end network where the face is physically grounded by explicit in-network representations of its shape, albedo, and lighting. Fig. 2 shows the overall network structure. We first introduce the foreground Shading Layer and the Image Formation Layer (Sec. 2.1), followed by two alternative in-network face representations (Fig. 2(a)-(b) and Sec. 2.2) that are compatible with in-network image formation. Finally, we introduce in-network matting (Sec. 2.3) which further disentangles the learning process of the foreground and background for face images in the wild.

From a graphics point of view, we assume a given face image IfgI_{fg} is the result of a rendering process, frenderingf_{\text{rendering}} where the inputs are an albedo map AeA_{e}, a normal map NeN_{e}, and illumination/lighting LL:

We assume Lambertian reflectance and adopt Retinex Theory to separate the albedo (i.e. reflectance) from the geometry and illumination:

in which ⊙\odot denotes the per-element product operation in the image space, and SeS_{e} represents a shading map rendered by NeN_{e} and LL:

If Eqs. 2 and 3 are differentiable, they can be realized as in-network layers in an auto-encoder network (Fig. 2(a)). This allows us to represent the image with disentangled latent variables for physically meaningful factors in the image formation process: the albedo latent variable ZAeZ_{A_{e}} , the normals variable ZNeZ_{N_{e}} and the lighting variable ZLZ_{L}. We show that this is advantageous over the traditional approach of a single latent variable that encodes the combined effect of all image formation factors. Each of the latent variables allows us access to a specific manifold, where semantically relevant edits can be performed while keeping irrelevant latent variables fixed. For instance, one can trivially perform image relighting by only traversing the lighting manifold given by ZLZ_{L} or changing only the albedo (e.g., to grow a beard) by traversing ZAeZ_{A_{e}}.

Computing shading from geometry (NeN_{e}) and illumination (LL) is nontrivial under unconstrained conditions, and might result in fshading(⋅,⋅)f_{\text{shading}}(\cdot,\cdot) being a discontinuous function in a significantly large region of the space it represents. Therefore, we further assume distant illumination, LL, that is represented by spherical harmonics s.t. the Lambertian shading function, fshading(⋅,⋅)f_{\text{shading}}(\cdot,\cdot) has an analytical form and is differentiable.

Following previous work , lighting LL is represented by a 9-dimensional spherical harmonics coefficient vector. For a given pixel, ii, with normal ni=[nx,ny,nz]⊤\mathbf{n_{i}}=[n_{x},n_{y},n_{z}]^{\top}, the shading is rendered as:

We provide the formulas for the partial derivatives ∂Sei∂nx\frac{\partial S_{e}^{i}}{\partial n_{x}}, ∂Sei∂ny\frac{\partial S_{e}^{i}}{\partial n_{y}},∂Sei∂nz\frac{\partial S_{e}^{i}}{\partial n_{z}} and ∂Sei∂Lj\frac{\partial S_{e}^{i}}{\partial L_{j}} in the supplementary material. Using these two differential rendering modules fshadingf_{\text{shading}} and fimage-formationf_{\text{image-formation}}, we can now implement the rendering modules within the network as shown in Figure 2.

2 In-Network Face Representation

The formulation introduced in the previous section requires the image formation and shading variables to be defined in the image coordinate system. This can be achieved with an explicit per-pixel representation of the face properties: NeN_{e}, AeA_{e}. Figure 2(a) depicts the module where the explicit normals and albedo are represented by their latent variables ZNeZ_{N_{e}}, ZAeZ_{A_{e}}. Note that the lighting, LL, is independent of the face representation; we represent it using spherical harmonics coefficients, i.e., ZL=LZ_{L}=L is directly used by the shading layer whose forward process is given by Eqn. 4.

Implicit Representation.

Even though the explicit representation helps disentangle certain properties and relates edits more intuitively to the latent variable manifolds (i.e. relighting), it might not be satisfactory in some cases. For instance, pose and expression edits might change both the explicit per-pixel normals, as well as the per-pixel albedo in the image space. We therefore introduce an implicit representation, where the parametrization is over the face coordinate system rather than the image coordinate system. This will allow us to further constrain pose and expression changes to the shape (i.e. normal) space only.

To address this, we introduce an alternative network architecture where the explicit representation depicted in the module in Fig. 2(a) is replaced with Fig. 2(b). Here, UVUV represents the per-pixel face space uv-coordinates, NiN_{i} and AiA_{i} represent the normal and albedo maps in the face uv-coordinate system, and ZUVZ_{UV}, ZNiZ_{N_{i}}, and ZAiZ_{A_{i}} represent the corresponding latent variables respectively. This is akin to the standard UV-mapping process in computer graphics. Facial features are aligned in this space (eyes correspond to eyes, mouths to mouths, etc.), and as a result the network has to learn a smaller space of variation, leading to sharper, more accurate reconstructions. Note that even though the network only uses the explicit latent variables at test time, we have auxiliary decoder stacks for all implicit variables to encourage disentangling of these variables during training. The implementation and training details will be explained in Sec. 3.2.

3 In-Network Background Matting

To further encourage the physically based representations of albedo, normals and lighting to concentrate on the face region, we disentangle the background from the foreground with a matte layer similar to the work by Yan et al. . The matte layer computes the compositing of the foreground face onto the background:

The matting layer also enables us to utilize efficient skip layers where unpooling layers in the decoder stack can use pooling switches from the corresponding encoder stack of the input image (grey links from the input encoder to background and mask decoders in Figure 2). The skip connection between the encoder and the decoder, allow for the details of the background to be preserved to a greater extent. Such skip connections bypass the bottleneck ZZ and therefore allow only partial information flow through ZZ during training.

For the foreground face region we chose to “filter” all the information through the bottleneck ZZ without any skip connections in order to gain full control over the latent manifolds for editing, at the expense of some detail loss.

Implementation

The convolutional encoder stack (Fig. 2) is composed of three convolutions with 32∗3×332*3\times 3, 64∗3×364*3\times 3 and 64∗3×364*3\times 3 filter sets. Each convolution is followed by max-pooling and a ReLU nonlinearity. We pad the filter responses after each pooling layer so that the final output of the convolutional stack is a set of filter responses with size 64∗8×864*8\times 8 for an input image 3∗64×643*64\times 64.

ZIiZ_{I_{i}} is a latent variable vector of 128×1128\times 1 which is fully connected to the last encoder stack downstream as well as the individual latent variables for background ZbgZ_{bg}, mask ZmZ_{m}, light ZLZ_{L}, and the foreground representations. For the explicit foreground representation, it is directly connected to ZNeZ_{N_{e}} and ZAeZ_{A_{e}} (Fig. 2(a)), whereas for the implicit representation it is connected to ZUVZ_{UV}, ZNiZ_{N_{i}}, and ZAiZ_{A_{i}} (Fig. 2(b)). All individual latent representations are 128×1128\times 1 vectors except for ZLZ_{L} which represents the light LL directly and is thus a 27×127\times 1 vector (three 9×19\times 1 concatenated vectors representing the spherical harmonics of the RGB components).

All decoder stacks for upsampling per-pixel (explicit or implicit) values are strictly symmetric to the encoder stack. As described in Sec. 2.3, the decoder stacks for the mask and background have skip connections to the input encoder stack at corresponding layers. The implicit normals NiN_{i} and implicit albedo AiA_{i} share weights in the decoder, since we have supervision of the implicit normals only.

2 Training

We use “in-the-wild” face images for training. Hence, we only have access to the image itself (denoted by I∗I^{*}), and do not have ground-truth data for either illumination, normal map, or the albedo. The main loss function is therefore on the reconstruction of the image IiI_{i} at the output IoI_{o}:

where Erecon=∣∣Ii−Io∣∣2E_{\text{recon}}=||I_{i}-I_{o}||^{2}. EadvE_{\text{adv}} is given by the adversarial loss, where a discriminative network is trained at the same time to distinguish between the generated and real images . Specifically, we use an energy-based method to incorporate the adversarial loss. In this approach an autoencoder is used as the discriminative network, D\mathcal{D}. The adversarial loss for the generative network is defined as Eadv=D(I′)E_{\text{adv}}=D(I^{\prime}), where I′I^{\prime} is the reconstruction of the discriminator input IoI_{o}, hence D(.)D(.) is the L2L_{2} reconstruction loss of the discriminator D\mathcal{D}. We train D\mathcal{D} to minimize the margin-based reconstruction error proposed by .

Fully unsupervised training using only the reconstruction and adversarial loss on the output image will often result in semantically meaningless latent representations. The network architecture itself cannot prevent degenerate solutions, e.g. when AeA_{e} captures both albedo and shading information while SeS_{e} remains constant. Since each of the rendering elements has a specific physical meaning, and they are explicitly encoded as intermediate layers in the network, we introduce additional constraints through intermediate loss functions to guide the training.

First, we introduce N^\hat{N}, a “pseudo ground-truth” of the normal map NeN_{e}, to keep the normal map close to plausible face normals during the training process. We estimate N^\hat{N} by fitting coarse face geometry to every image in the training set using a 3D Morphable Model . We then introduce the following objective to NeN_{e}:

Similar to N^\hat{N}, we provide a L2L2 reconstruction loss w.r.t L^\hat{L}, on the lighting parameters ZLZ_{L}:

where L^\hat{L} is computed from N^\hat{N} and the input image using least square optimization and a constant albedo assumption .

Furthermore, following Retinex theory which assumes albedo to be piecewise constant and shading to be smooth, we introduce an L1L1 smoothness loss on the gradients of the albedo, AA:

in which ∇\nabla is the spatial image gradient operation. In addition, since the shading is assumed to vary smoothly, we introduce an L2L2 smoothness loss on the gradients of the shading, SeS_{e}:

For the implicit coordinate system (UVUV) variant (Fig. 2-(b)), we provide L2L2 supervisions to both UVUV and NiN_{i}:

UV^\hat{UV} and Ni^\hat{N_{i}} are obtained from the previously mentioned Morphable Model, in which vertex-wise correspondence on the 3D fit exists. We utilize the average shape of the Morphable Model Sˉ\bar{S} to construct a canonical coordinate map (UVUV) and surface normal map (NiN_{i}), and propagate it to each shape estimation via this correspondence. More details of this computation are presented in our supplemental document.

Due to ambiguity in the magnitude of lighting, and therefore the intensity of shading (Eq. 2), it is necessary to incorporate constraints on the shading magnitude to prevent the network from generating arbitrary bright/dark shading. Moreover, since the illumination is separated in individual colors Lr\mathbf{L}_{r}, Lg\mathbf{L}_{g} and Lb\mathbf{L}_{b}, we incorporate a constraint to prevent the shading from being too strong in one color channel vs. the others. To handle these ambiguities, we introduce a Batch-wise White Shading (BWS) constraint on SeS_{e}:

where sri(j)s_{r}^{i}(j) denotes the jj-th pixel of the ii-th example in the first (red) channel of SeS_{e}. sgs_{g} and sbs_{b} denote the second and the third channel of shading respectively. mm is the number of pixels in a training batch. In all experiments c=0.75c=0.75.

Since N^\hat{N} obtained by the Morphable Model comes with a region of interest only on the face surface, we use it as the mask under which we compute all foreground losses. In addition, this region of interest is also used as the mask pseudo ground truth at training time for learning the matte mask:

in which M^\hat{M} represents the Morphable Model mask.

Experiments

We use the CelebA dataset to train the network. For each image in the dataset, we detect landmarks , and fit a 3D Morphable Model to the face region to have a rough estimation of the rendering elements (N^,L^\hat{N},\hat{L}). These estimates are used to set-up the various losses detailed in the previous section. This data is subsequently used only for the training of the network as previously described.

For comparison, we train an auto-encoder B\mathcal{B} as a baseline. The encoder and decoder of B\mathcal{B} is identical to the encoder and decoder for albedo in our architecture. To make the comparison fair, the bottleneck layer of B\mathcal{B} is set to 265265 (=128+128+9=128+128+9) dimensions, which is more than twice as large as the bottleneck layer in our architecture (size 128128), yielding slightly more capacity for the baseline. Even though our architecture has a narrower bottleneck, the disentangling of the latent factors and the presence of physically based rendering layers, lead to reconstructions that are more robust to complex background, pose, illumination, occlusion, etc., (Fig. 3).

More importantly, given an input face images, our network provides explicit access to an estimation of the albedo, shading and normal map (Fig. 3) for the face. Notably, in the last row of Fig. 3, we compare the inferred normals from our network with the normals estimated from the input image using the 3D morphable model that we deployed to guide the training process. The data to construct the morphable model contains only 16 identities; this small subspace of identity variation leads to normals that are often inaccurate approximations of the true face shape (row 7 in Fig. 3). By using these estimates as weak supervision in combination with an appearance-based rendering loss, our network is able to generate normal maps (row 6 in Fig. 3) that extend beyond the morphable model subspace, better fit the shape of the input face, and exhibit more identity information. Please refer to our supplementary material for more comparisons.

2 Face Editing by Manifold Traversal

Our network enables manipulation of semantic face attributes, (e.g. expression, facial hair, age, makeup, and eye-wear) by traversing the manifold(s) of the disentangled latent spaces that are most appropriate for that edit.

For a given attribute, e.g., the smiling expression, we feed both positive data {xp}\{\mathbf{x}_{p}\} (smiling faces) and negative data {xn}\{\mathbf{x}_{n}\} (faces with other expressions) into our network to generate two sets of ZZ-codes {zp}\{\mathbf{z}_{p}\} and {zn}\{\mathbf{z}_{n}\}. These sets represent corresponding empirical distributions of the data on the low dimensional ZZ-space(s). Given an input face image IsourceI_{\text{source}} that is not smiling, we seek to make it smile by moving its ZZ-code(s) ZsourceZ_{\text{source}} towards the distribution {z}p\{\mathbf{z}\}_{p} to get a transformed code ZtransZ_{\text{trans}}. After that, we reconstruct the image corresponding to ZtransZ_{\text{trans}} with the decoders in our model.

In order to compute the distributions for each attribute, we sample a subset of 2000 images from the CelebA with the appropriate attribute label (e.g., smiling vs other expresssions). We use the manifold traversal method proposed by Gardner et al. independently on each appropriate variable. The extent of the traversal is parameterized by a regularization parameter, λ\lambda (see for details).

In Fig. 4, we compare the results using our network against the baseline auto-encoder. We traverse the albedo and normal variables to produce edits which make the faces smile and are able to capture changes in expression and the appearance of teeth, while preserving the other aspects of the image. In contrast, the results from traversing the baseline latent space are much poorer – in addition to not being able to reconstruct the pose and identity of the input properly, the traversal is not able to capture the smiling transformation as well as we do.

In Fig. 5 we demonstrate the utility of our implicit representation. While lips/mouth and teeth might map to the same region of the image space, they are in fact separated in the face UV-space. This allows the implicit variables to learn more targeted and accurate representations, hence traversing just the ZUVZ_{UV}, already results in a smiling face. Combining this with traversal along ZNiZ_{N_{i}} exaggerates the smile. In contrast, we do not expect smiling to be correlated with the implicit albedo space, and traversing along the ZAiZ_{A_{i}} leads to poorer results with an incorrect frontal pose.

In Fig. 6 we demonstrate more results for smiling and demonstrate that relaxing the traversal regularization parameter, λ\lambda, gradually leads to stronger smiling expressions.

We also address the editing task of aging via manifold traversal. For this experiment, we construct the latent space distributions using images and labels from the PubFig dataset corresponding to the most and least senior images. We expect aging to be correlated with both shape and texture, and show in Fig. 7 that traversing these manifolds leads to convincing age progression.

Note that all of these edits have been performed on the exact same network, indicating that our network architecture is general enough to represent the manifold of face appearance, and is able to disentangle the latent factors to support specific editing tasks. Refer to our supplementary material for more results, and comparisons.

Limitations. Our current face masks do not include hair. This results in less control over some edits, e.g. aging, that are inherently affecting the hair as well. However, this can trivially be addressed, if a mask that also includes the hair can be generated .

3 Relighting

A direct application of the albedo-normal-light decomposition in our network is that it allows us to manipulate the illumination of an input face via ZLZ_{L} while keeping the other latent variable fixed. We can directly “relight” the face by replacing its ZLtargetZ_{L}^{\text{target}} with some other ZLsourceZ_{L}^{\text{source}} (e.g. using the lighting variable of another face).

While our network is trained to reconstruct the input, due to its limited capacity (especially due to the bottleneck layer dimensionality), the reconstruction does not reproduce the input with all the details. For illumination editing, however, we can directly manipulate the shading, that is also available in our network. We pass the source IsourceI^{\text{source}} and target images ItargetI^{\text{target}} through our network to estimate their individual factors. We use the target shading StargetS^{\text{target}} with Eq. 2 to compute a “detailed” albedo AtargetA^{\text{target}}. Given the source light LsourceL^{\text{source}}, we render the shading of the target under this light with the target normals NtargetN^{\text{target}} (Eq. 3) to obtain the transferred shading StransferS^{\text{transfer}}. In the end, the lighting transferred image is rendered with AtargetA^{\text{target}} and StransferS^{\text{transfer}} using Eq. 2. This is demonstrated in Fig. 8 where we are able to successfully transfer the lighting from two sources with disparate identities, genders, and poses to a target while retaining all its details. We present more relighting results, as well as quantitative tests on illumination (i.e. spherical harmonics coefficients) prediction in the supplementary material.

Conclusions

We proposed a physically grounded rendering-based disentangling network specifically designed for faces. Such disentangling enables realistic face editing since it allows trivial constraints at manipulation time. We are the first to attempt in-network rendering for faces in the wild with real, arbitrary backgrounds. Comparisons with traditional auto-encoder approaches show significant improvements on final edits, and our intermediate outputs such as face normals show superior identity preservation compared to traditional approaches.

Acknowledgements

This work started when Zhixin Shu was an intern at Adobe Research. This work was supported by a gift from Adobe, NSF IIS-1161876, the Stony Brook SensorCAT and the Partner University Fund 4DVision project.

References

Appendix A Implementation: more details

In this section, we provide more details regarding the implementation of the rendering layers fshadingf_{\text{shading}} and fimage-formationf_{\text{image-formation}} as described in the paper.

The shading layer is rendered with a spherical harmonics illumination representation .

The forward process is described by equations (3),(4), and (5) in the main paper. We now provide the backward process, i.e., the partial derivatives ∂Sei∂nx\frac{\partial S_{e}^{i}}{\partial n_{x}}, ∂Sei∂ny\frac{\partial S_{e}^{i}}{\partial n_{y}},∂Sei∂nz\frac{\partial S_{e}^{i}}{\partial n_{z}} and ∂Sei∂Lj\frac{\partial S_{e}^{i}}{\partial L_{j}} as follows:

A.2 Image Formation Layer

The forward process of the image formation layer (for the foreground) is simply a per-element product (see equation (2) in the main paper), therefore the backward process (partial derivatives) is:

Appendix B Quantitative Experiments

In order to evaluate the illumination estimation of our network, we utilize the Multi-PIE dataset where controlled illumination is availableControlled illumination in this context means that the different subjects have been illuminated under the same lighting conditions. Hence, the correspondences across subjects for the same lighting is known, but the actual lighting setup, or conditions are not known.. We randomly sample 7,000 images with different identities and poses under 20 controlled light sources. We measure the variance of illumination coefficients LiL_{i} (i=1,...,9i=1,...,9) within an illumination condition. The average variance of a 3D Morphable Model with least square estimation is 0.360.36 while the average variance of our network is 0.160.16. In Fig. 9, we show the average (among 20 lighting conditions and RGB) variance of each lighting coefficient. Note that our model is only trained with CelebA dataset while the images from the Multi-PIE datasets are only used for quantitative evaluation.

We evaluate the quality of our normal reconstruction vs. a direct 3D Morphable Model (3DMM) fit. We create input images for five different individuals (two women, three men) using light stage data captured by Weyrich et al. . We fit a 3DMM to these images and compute normals. We also pass these image as inputs to our disentangling network to get a normal reconstruction. In Fig. 10, we compare the ground truth normals to our estimates and the 3DMM normals. Even though our network was trained on normals from morphable model fits, the additional reconstruction losses enable it to expand beyond this subspace. This is especially apparent for the reconstruction of the two women, whose faces look man-like in the 3DMM fits.

Appendix C Additional Results

We present more face editing results, as well as relighting results using our proposed method described in section 5.3 of the main paper.

Eye-glasses. Figure 12 shows results of wear eye-glasses. For this experiment, we sample 2000 images from the CelebA of faces wearing eye-glasses as {xp}\{\mathbf{x}_{p}\} and 2000 images of faces not wearing eye-glasses as {xn}\{\mathbf{x}_{n}\}. We compute the edits on the AiA_{i} manifold only, with λ=0.02\lambda=0.02.

Since eye-glasses only affect the reflectance in the image, no geometry or warping is associate with this editing, therefore the natural choice of manifold-to-be-edited is ZAiZ_{A_{i}}. We show in figure 13 a comparison of adding eye-glasses via different manifolds (using same the λ=0.02\lambda=0.02). We notice that (1) editing ZNiZ_{N_{i}} (Fig. 13-(c)) almost has no effect on the image; (2) editing ZUVZ_{UV} slightly aged the face (Fig. 13-(d)), mainly since senior people are more likely to wear eye-glasses; (3) manipulating all the manifolds leads to changes in geometry and appearance (beards, nostrils and shape of noses in Fig. 13-(e)); (4) editing through ZAiZ_{A_{i}} generates faithful results of wearing eye-glasses and has little effect on other attributes of the face.

Beards. In figures 14, 15, 16, we present additional results of grow beard. For this experiment, we sample 2000 images from the CelebA of male faces with beard as {xp}\{\mathbf{x}_{p}\} and 2000 images of male faces without beard as {xn}\{\mathbf{x}_{n}\}. We compute the traversal on the AiA_{i} manifold only, with λ=0.03\lambda=0.03, 0.020.02, and 0.010.01 respectively.

Aging. Figures 17, 18, show additional results of aging. We compute the traversal on manifolds ZAiZ_{A_{i}}, ZNiZ_{N_{i}}, and ZUVZ_{UV}, with λ=0.05\lambda=0.05, 0.030.03, and 0.020.02 respectively.

Smiling. Figures 19, 20 show additional results of smiling as described in Section 5.2 of the main paper. The data is as described in the main paper. We compute the traversal on manifolds ZNiZ_{N_{i}}, and ZUVZ_{UV}, with λ=0.07\lambda=0.07, 0.050.05, and 0.030.03 respectively.

Relighting. We present additional relighting results in Fig. 21. In addition, in Fig. 11, we provide comparisons to two previous techniques for re-lighting . Our results apply to the full face, capture the target lighting, and have fewer artifacts. Our general face editing technique is able to produce results that are qualitatively similar, or even better than other techniques specifically designed for this particular task.