GIF: Generative Interpretable Faces
Partha Ghosh, Pravir Singh Gupta, Roy Uziel, Anurag Ranjan, Michael Black, Timo Bolkart
Introduction
The ability to generate a person’s face has several uses in computer graphics and computer vision that include constructing a personalized avatar for multimedia applications, face recognition and face analysis. To be widely useful, a generative model must offer control over the generative factors such as expression, pose, shape, lighting, skin tone, etc. Early work focuses on learning a low dimensional representation of human faces using PCA (PCA) spaces or higher-order tensor generalizations . Although they provide some semantic control, these methods use linear transformations in the pixel-domain to model facial variation, resulting in blurry images. Further, effects of rotations are not well parameterized by linear transformations in 2D, resulting in poor image quality.
To overcome this, Blanz and Vetter introduced a statistical 3D morphable model, of facial shape and appearance. Such a statistical face model (e.g. ) is used to manipulate shape, expression, or pose of the facial 3D geometry. This is then combined with a texture map (e.g. ) and rendered to an image. Rendering of a 3D face model lacks photo-realism due to the difficulty in modeling hair, eyes, and the mouth cavity (i.e., teeth or tongue), along with the absence of facial details like wrinkles in the geometry of the facial model. Further difficulties arise from subsurface scattering of facial material. These factors affect the photo-realism of the final rendering.
On the other hand, generative adversarial networks (GANs) have recently shown great success in generating photo-realistic face images at high resolution . Methods like StyleGAN or StyleGAN2 even provide high-level control over factors like pose or identity when trained on face images. However, these controlling factors are often entangled, and they are unknown prior to training. The control provided by these models does not allow the independent change of attributes like facial appearance, shape (e.g. length, width, roundness, etc.) or facial expression (e.g. raise eyebrows, open mouth, etc.). Although these methods have made significant progress in image quality, the provided control is not sufficient for graphics applications.
In short, we close the control-quality gap. However, we find that a naive conditional version of StyleGAN2 yields unsatisfactory results. We overcome this problem by rendering the geometric and photo-metric details of the FLAME mesh to an image and by using these as conditions instead. This design combination results in a generative 2D face model called GIF (Generative Interpretable Faces) that produces photo-realistic images with explicit control over face shape, head and jaw pose, expression, appearance, and illumination (Figure 1).
Finally we remark that the community lacks an automated metric to effectively evaluate continuous conditional generative models. To this end we derive a quantitative score from a comparison-based perceptual study using an anonymous group of participants to facilitate future model comparisons.
In summary, our main contributions are 1) a generative 2D face model with FLAME control, 2) use of FLAME renderings as conditioning for better association, 3) use of a texture consistency constraint to improve disentanglement, 4) providing a quantitative comparison mechanism.
Related Work
Generative 3D face models: Representing and manipulating human faces in 3D have a long standing history dating back almost five decades to the parametric 3D face model of Parke . Blanz and Vetter propose a 3D morphable model (3DMM), the first generative 3D face model that uses linear subspaces to model shape and appearance variations. This has given rise to a variety of 3D face models to model facial shape , shape and expression , shape, expression and head pose , localized facial details and wrinkle details . However, renderings of these models do not reach photo-realism due to the lack of high-quality textures.
To overcome this, Saito et al. introduce high-quality texture maps, and Slossberg et al. and Gecer et al. train GANs to synthesize textures with high-frequency details. While these works enhance the realism when being rendered by covering more texture details, they only model the face region (i.e. ignore hair, teeth, tongue, eyelids, eyes, etc.) required for photo-realistic rendering. While separate part-based generative models of hair , eyes , eyelids , ears , teeth , or tongue exist, combining these into a complete realistic 3D face model remains an open problem.
Instead of explicitly modeling all face parts, Gecer et al. use image-to-image translation to enhance the realism of images rendered from a 3D face mesh. Nagano et al. generate dynamic textures that allow synthesizing expression dependent mouth interior and varying eye gaze. Despite significant progress of generative 3D face models , they still lack photo-realism.
Our approach, in contrast, combines the semantic control of generative 3D models with the image synthesis ability of generative 2D models. This allows us to generate photo-realistic face images, including hair, eyes, teeth, etc. with explicit 3D face model controls.
Generative 2D face models: Early parametric 2D face models like Eigenfaces , Fisherfaces , Active Shape Models , or Active Appearance Models parametrize facial shape and appearance in images with linear spaces. Tensor faces generalize these linear models to higher-order, generating face images with multiple independent factors like identity, pose, or expression. Although these models provided some semantic control, they produced blurry and unrealistic images.
StyleGAN , a member of broad category of GAN models, extends Progressive-GAN by incorporating a style vector to gain partial control over the image generation, which is broadly missing in such models. However, the semantics of these controls are interpreted only post-training. Hence, it is possible that desired controls might not be present at all. InterFaceGAN and StyleRig aim to gain control over a pre-trained StyleGAN . InterFaceGAN identifies hyper-planes that separate positive and negative semantic attributes in a GAN’s latent space. However, this requires categorical attributes for every kind of control, making it not suitable for a variety of aspects e.g. facial shape, expression etc. Further, many attributes might not be linearly separable. StyleRig learns mappings between 3DMM parameters and the parameter vectors of each StyleGAN layer. StyleRig learns to edit StyleGAN parameters and thereby controls the generated image, with 3DMM parameters. The setup is mainly tailored towards face editing or face reenactment tasks. GIF, in contrast, provides full generative control over the image generation process (similar to regular GANs) but with semantically meaningful control over shape, pose, expression, appearance, lighting, and style.
CONFIG leverages synthetic images to get ground truth control parameters and leverages real images to make the image generation look more realistic. However, generated images still lack photo-realism.
HoloGAN randomly applies rigid transformations to learnt features during training, which provide explicit control over 3D rotations in the trained model. While this is feasible for global transformations, it remains unclear how to extend this to transformations of local parts or parameters like facial shape or expression. Similarly, IterGAN also only models rigid rotation of generated objects using a GAN, but these rotations are restricted to 2D transformations.
Further works use variational autoencoders and flow-based methods to generate images. These provide controllability, but do not reach the image quality of GANs.
Facial animation: A large body of work focuses on face editing or facial animation, which can be grouped into 3D model-based approaches (e.g. ) or 3D model-free methods (e.g. )
Thies et al. build a subject specific 3DMM from a video, reenact this model with expression parameters from another sequence, and blend the rendered mesh with the target video. Follow-up work use similar 3DMM-based retargeting techniques but replace the traditional rendering pipeline, or parts of it, with learnable components . Ververas and Zafeiriou (Slider-GAN) and Geng et al. propose image-to-image translation models that, given an image of a particular subject, condition the face editing process on 3DMM parameters. Lombardi et al. learn a subject-specific autoencoder of facial shape and appearance from high-quality multi-view images that allow animation and photo-realistic rendering. Like GIF, all these methods use explicit control of a pre-trained 3DMM or learn a 3D face model to manipulate or animate faces in images, but in contrast to GIF, they are unable to generate new identities.
Zakharov et al. use image-to-image translation to animate a face image from 2D landmarks, ReenactGAN and Recycle-GAN transfer facial movement from a monocular video to a target person. Pumarola et al. learn a GAN conditioned on facial action units for control over facial expressions. None of these methods provide explicit control over a 3D face representation.
All methods discussed above are task specific, i.e., they are dedicated towards manipulating or animating faces, while GIF, in contrast, is a generative 2D model that is able to generate new identities, and also provides control of a 3D face model. Further, most of the methods are trained on video data , in contrast to GIF which is trained from static images only. Regardless, GIF can be used to generate facial animations.
Preliminaries
GANs and conditional GANs: GANs are a class of neural networks where a generator and a discriminator have opposing objectives. Namely, the discriminator estimates the probability of its input to be a generated sample, as opposed to a natural random sample from the training set, while the generator tries to make this task as hard as possible. This is extended in the case of conditional GANs . The objective function in such a setting is given as
where is the conditioning variable. Although sound in an ideal setting, this formulation suffers a major drawback. Specifically under incomplete data regime, independent conditions tend to influence each other. In Section 4.2, we discuss this phenomenon in detail.
StyleGAN2 , a revised version of StyleGAN , produces photo-realistic face images at resolution. Similar to StyleGAN, StyleGAN2, is controlled by a style vector . This vector is first transformed by a mapping network of fully connected layers to , which then transforms the activations of the progressively growing resolution blocks using adaptive instance normalization (AdaIN) layers. Although StyleGAN2 provides some high-level control, it still lacks explicit and semantically meaningful control. Our work addresses this shortcoming by distilling a conditional generative model out of StyleGAN2 and combining this with inputs from FLAME.
2 FLAME
Method
Here, the FLAME parameters control all aspects related to the geometry, appearance and lighting parameters control the face color (i.e. skin tone, lighting, etc.), while the style vector controls all factors that are not described by the FLAME geometry and appearance parameters, but are required to generate photo-realistic face images (e.g. hairstyle, background, etc.).
Our data set consists of Flickr images (FFHQ) introduced in StyleGAN and their corresponding FLAME parameters, appearance parameters, lighting parameters, and camera parameters. We obtain these using DECA , a publicly available monocular 3D face regressor. In total, we use about 65,500 FFHQ images, paired with the corresponding parameters. The obtained FFHQ training parameters are available for research purposes.
2 Condition representation
Condition cross-talk: The vanilla conditional GAN formulation as described in Section 3 does not encode any semantic factorization of the conditional probability distributions that might exist in the nature of the problem. Consider a situation where the true data depends upon two independent generating factors and , i.e. the true generation process of our data is where . Ideally, given complete data (often infinite) and a perfect modeling paradigm, this factorization should emerge automatically. However, in practice, neither of these can be assumed. Hence, the representation of the condition highly influences the way it gets associated with the output. We refer to this phenomenon of independent conditions influencing each other as – condition cross-talk. Inductive bias introduced by condition representation in the context of conditional cross-talk is empirically evaluated in Section 5.
Pixel-aligned conditioning: Learning the rules of graphical projection (orthographic or perspective) and the notion of occlusion as part of the generator is wasteful if explicit 3D geometry information is present, as it approximates classical rendering operations, which can be done learning-free and are already part of several software packages (e.g. ). Hence, we provide the generator with explicit knowledge of the 3D geometry by conditioning it with renderings from a classical renderer. This makes pixel-localized association between the FLAME conditioning signal and the generated image possible. We find that, although a vanilla conditional GAN achieves comparable image quality, GIF learns to better obey the given condition.
We condition GIF on two renderings, one provides pixel-wise color information (referred to as texture rendering), and the other provides information about geometry (referred to as normal rendering). Normal renderings are obtained by rendering the mesh with a color-coded map of the surface normals .
Both renderings use a scaled orthographic projection with the camera parameters provided with each training image. For the color rendering, we use the provided inferred lighting and appearance parameter. The texture and the normal renderings are concatenated along the color channel and used to condition the generator as shown in Figure 2. As demonstrated in Section 5, this conditioning mechanism helps reduce condition cross-talk.
3 GIF architecture
The model architecture of GIF is based on StyleGAN2 , with several key modifications, discussed as follows. An overview is shown in Figure 2.
Noise channel conditioning: StyleGAN and StyleGAN2 insert random noise images at each resolution level into the generator, which mainly contributes to local texture changes. We replace this random noise by the concatenated textured and normal renderings from the FLAME model, and insert scaled versions of these renderings at different resolutions into the generator. This is motivated by the observation that varying FLAME parameters, and therefore varying FLAME renderings, should have direct, pixel-aligned influence on the generated images.
4 Texture consistency
Since GIF’s generation is based on an underlying 3D model, we can further constrain the generator by introducing a texture consistency loss optimized during training. We first generate a set of new FLAME parameters by randomly interpolating between the parameters within a mini batch. Next, we generate the corresponding images with the same style embedding , appearance and lighting parameters by an additional forward pass through the model. Finally, the corresponding FLAME meshes are projected onto the generated images to get a partial texture map (also referred to as ‘texture stealing’). To enforce pixel-wise consistency, we apply an loss on the difference between pairs of texture maps, considering only pixels for which the corresponding 3D point on the mesh surface is visible in both the generated images. We find that this texture consistency loss improves the parameter association (see Section 5).
Experiments
Condition influence: As described in Section 4, GIF is parametrized by FLAME parameters , appearance parameters , lighting parameters , camera parameters and a style vector . Figure 3 shows the influence of each individual set of parameters by progressively exchanging one type of parameter in each row. The top and bottom rows show GIF generated images for two sets of parameters, randomly chosen from the training data.
Exchanging style (row 2) most noticeably changes hair style, clothing color, and the background. Shape (row 3) is strongly associated to the person’s identity among other factors. The expression parameters (row 4) control the facial expression, best visible around the mouth and cheeks. The change in pose parameters (row 5) affects the orientation of the head (i.e. head pose) and the extent of the mouth opening (jaw pose). Finally, appearance (row 6) and lighting (row 7) change the skin color and the lighting specularity.
Random sampling: To further evaluate GIF qualitatively, we sample FLAME parameters, appearance parameters, lighting parameters, and style embeddings and generate random images, shown in Figure 4. For shape, expression, and appearance parameters, we sample parameters of the first three principal components from a standard normal distribution and keep all other parameters at zero. For pose, we sample from a uniform distribution in (head pose) for rotation around the y-axis, and (jaw pose) around the x-axis. For lighting parameters and style embeddings, we choose random samples from the set of training parameters. Figure 4 shows that GIF produces photo-realistic images of different identities with a large variation in shape, pose, expression, skin color, and age. Figure 1 further shows rendered FLAME meshes for generated images, demonstrating that GIF generated images are well associated with the FLAME parameters.
Speech driven animation: As GIF uses FLAME’s parametric control, it can directly be combined with existing FLAME-based application methods such as VOCA , which animates a face template in FLAME mesh topology from speech. For this, we run VOCA for a speech sequence, fit FLAME to the resulting meshes, and use these parameters to drive GIF for different appearance embeddings (see Figure 5). For more qualitative results and the full animation sequence, see the supplementary video.
2 Quantitative evaluation
We conduct two AMT (AMT) studies to quantitatively evaluate i) the effects of ablating individual model parts, and ii) the disentanglement of geometry and style. We compare GIF in total with different ablated versions, namely vector conditioning, no texture interpolation, normal rendering, and texture rendering conditioning. For the vector conditioning model, we directly provide the FLAME parameters as a dimensional real valued vector. In the no texture interpolation model, we drop the texture consistency during training. The normal rendering and texture rendering conditioning models condition the generator and discriminator only on the normal rendering or texture rendering, respectively, while removing the other rendering. The supplementary video shows examples for both user studies.
Ablation experiment: Participants see three images, a reference image in the center, which shows a rendered FLAME mesh, and two generated images to the left and right in random order. Both images are generated from the same set of parameters, one using GIF, another an ablated GIF model. Participants then select the generated image that corresponds best with the reference image. Table 1 shows that with the texture consistency loss, normal rendering, and texture rendering conditioning, GIF performs slightly better than without each of them. Participants tend to select GIF generated results over a vanilla conditional StyleGAN2 (please refer to our supplementary material for details on the architecture). Furthermore, Figure 2 quantitatively evaluates the image quality with FID scores, indicating that all models produce similar high-quality results.
Geometry-style disentanglement: In this experiment, we study the entanglement between the style vector and FLAME parameters. We find qualitatively that the style vector mostly controls aspects of the image that are not influenced by FLAME parameters like background, hairstyle, etc. (see Figure 3). To evaluate this quantitatively, we randomly pair style vectors and FLAME parameters and conduct a perceptual study with AMT. Participants see a rendered FLAME image and GIF generated images with the same FLAME parameters but with a variety of different style vectors. Participants then rate the similarity of the generated image to shape, pose and expression of the FLAME rendering on a standard 5-Point Likert scale (i.e. 1: Strongly Disagree, 2. Disagree, 3. Neither agree nor disagree, 4. Agree, 5. Strongly Agree). We use randomized styles and random FLAME parameters totaling to images. We find that the majority of the participants agree that generated images and FLAME rendering are similar, irrespective of the style (see Figure 6).
Re-inference error: In the spirit of DiscoFaceGAN , we run DECA on the generated images and compute the Root Mean Square Error (RMSE) between the FLAME face vertices of the input parameters and the inferred parameters as reported in Table 3. We generate a population of images by randomly sampling one of the shape, pose and expression latent spaces while holding the rest of the generating factors to their neutral. Thus we find an association error for individual factors.
Discussion
GIF is trained on the FFHQ data-set and hence inherits some of its limitations, e.g. the images are roughly eye-centered. Although StyleGAN2 has to some extent addressed this issue (among others), but failed to do so completely. Hence, rotations of the head look like an eye-centered rotation, which involves a combination of 3D rotation and translation as opposed to a pure neck centered rotation. Adapting GIF to use an architecture other than StyleGAN2 or training it on a different data set to improve image quality is subject to future work.
As FLAME renderings for conditioning GIF must be similarly eye-centered as the FFHQ training data. We compute a suitable camera parameter from given FLAME parameters so that the eyes are located roughly at the centre of the image plane with a fixed distance between the eyes. However, for profile view poses, this can not be met without an extreme zoomed in view. This often causes severe artifacts. Please see the supplementary material for examples.
GIF requires a statistical 3D model, and a way to associate its parameters to a large data set of high-quality images. While ‘objects’ like human bodies or animals potentially fulfill these requirements, it remains unclear how to apply GIF to general object categories.
Faulty image to 3D model associations stemming from the parameter inference method potentially degrade GIF’s generation quality. One example is the ambiguity of lighting and appearance, which causes most color variations in the training data to be described by lighting variation rather than by appearance variation. GIF inherits these errors.
Finally, as GIF is solely trained from static images without multiple images per subject, generating images with varying FLAME parameters is not temporally consistent. As several unconditioned parts are only loosely correlated or uncorrelated to the condition (e.g. hair, mouth cavity, background, etc.), this results in jittery video sequences. Training or refining GIF on temporal data with additional temporal constraints during training remains subject to future work.
Conclusion
We present GIF, a generative 2D face model, with high realism and with explicit control from FLAME, a statistical 3D face model. Given a data set of approximately 65,500 high-quality face images with associated FLAME model parameters for shape, global pose, jaw pose, expression, and appearance parameters, GIF learns to generate realistic face images that associate with them. Our key insight is that conditioning a generator network on explicit information rendered from a 3D face model allows us to decouple shape, pose, and expression variations within the trained model. Given a set of FLAME parameters associated with an image, we render the corresponding FLAME mesh twice, once with color-coded normal details, once with an inferred texture, and insert these as condition to the generator. We further add a loss that enforces consistency in texture for the reconstruction of different FLAME parameters for the same appearance embedding. This encourages the network during training to disentangle appearance and FLAME parameters, and provides us with better temporal consistency when generating frames of FLAME sequences. Finally we devise a comparison-based perceptual study to evaluate continuous conditional generative models quantitatively.
Acknowledgements
We thank H. Feng for prepraring the training data, Y. Feng and S. Sanyal for support with the rendering and projection pipeline, and C. Köhler, A. Chandrasekaran, M. Keller, M. Landry, C. Huang, A. Osman and D. Tzionas for fruitful discussions, advice and proofreading. The work was partially supported by the International Max Planck Research School for Intelligent Systems (IMPRS-IS).
Disclosure
MJB has received research gift funds from Intel, Nvidia, Adobe, Facebook, and Amazon. While MJB is a part-time employee of Amazon, his research was performed solely at, and funded solely by, MPI. MJB has financial interests in Amazon and Meshcapade GmbH. PG has received funding from Amazon Web Services for using their Machine Learning Services and GPU Instances. AR’s research was performed solely at, and funded solely by, MPI.
References
Appendix A Vector conditioning architecture
Here we describe the model architecture of the vector condition model, used as one of the baseline models in Section 5.2 of the main paper. For this model, we pass the vector values conditioning parameters FLAME (), appearance () and lighting () as a dimensional vector through the dimensions of style vector of the original StyleGAN2 architecture as shown in Figure 7. We further input the same conditioning vector to the discriminator at the last fully connected layer of the discriminator by subtracting it from the last layer activation.
Appendix B Image centering for extreme rotations
As discussed in Section 6 of the main paper, GIF produces artifacts for extreme head poses close to profile view. This is due to the pixel alignment of the FLAME renderings and the generated images, which requires the images to be similarly eye-centered as the FFHQ training data. For profile views however it is unclear how the centering within the training data was achieved. The centering strategy used in GIF causes a zoom in for profile views, effectively cropping parts of the face, and hence the generator struggles to generate realistic images as shown in Figure 8.