ExprGAN: Facial Expression Editing with Controllable Expression Intensity

Hui Ding, Kumar Sricharan, Rama Chellappa

Introduction

Facial expression editing is the task that transforms the expression of a given face image to a target one without affecting the identity properties. It has applications in facial animation, human-computer interactions, entertainment, etc. The area has been attracting considerable attention both from academic and industrial research communities.

Existing methods that address expression editing can be divided into two categories. One category tries to manipulate images by reusing parts of existing ones (?; ?; ?) while the other resorts to synthesis techniques to generate a face image with the target expression (?; ?; ?). In the first category, traditional methods (?) often make use of the expression flow map to transfer an expression by image warping. Recently, ? (?) applies the idea to a variational autoencoder to learn the flow field. Although the generated face image has high resolution, paired data where one subject has different expressions are needed to train the model. In the second category, deep learning-based methods are mainly used. The early work by ? (?) uses a deep belief network to generate emotional faces, which can be controlled by the Facial Action Coding System (FACS) labels. In (?), a three-way gated Boltzmann machine is employed to model the relationships between the expression and identity. However, the synthesized image of these methods has low resolution (48 x 48), lacking fine details and tending to be blurry.

Moreover, existing works can only transform the expression to different classes, like Angry or Happy. However, in reality, the intensity of facial expression is often displayed over a range. For example, humans can express the Happy expression either with a huge grin or by a gentle smile. Thus it is appealing if both the type of the expression and its intensity can be controlled simultaneously. Motivated by this, in this paper, we present a new expression editing model, Expression Generative Adversarial Network (ExprGAN) which has the unique property that multiple diverse styles of the target expression can be synthesized where the intensity of the generated expression is able to be controlled continuously from weak to strong, without the need for training data with intensity values.

To achieve this goal, we specially design an expression controller module. Instead of feeding in a deterministic one-hot vector label like previous works, the expression code generated by the expression controller module is used. It is a real-valued vector conditioned on the label, thus more complex information such as expression intensity can be described. Moreover, to force each dimension of the expression code to capture a different factor of the intensity variations, the conditional mutual information between the generated image and the expression code is maximized by a regularizer network.

Our work is inspired by the recent success of the image generative model, where a generative adversarial network (?) learns to produce samples similar to a given data distribution through a two-player game between a generator and a discriminator. Our ExprGAN also adopts the generator and discriminator framework in addition to the expression controller module and regularizer network. However, to facilitate image editing, the generator is composed of an encoder and a decoder. The input of the encoder is a face image, the output of the decoder is a reconstructed one, and the learned identity and expression representations bridge the encoder and decoder. To preserve the most prominent facial structure, we adopt a multi-layer perceptual loss (?) in the feature space in addition to the pixel-wise L1L_{1} loss. Moreover, to make the synthesized image look more photo-realistic, two adversarial networks are imposed on the encoder and decoder, respectively. Because it is difficult to directly train our model on the small training set, a three-stage incremental learning algorithm is also developed.

The main contributions of our work are as follows:

We propose a novel model called ExprGAN that can change a face image to a target expression with multiple styles, where the expression intensity can also be controlled continuously.

We show that the synthesized face images have high perceptual quality, which can be used to improve the performance of an expression classifier.

Our identity and expression representations are explicitly disentangled which can be exploited for tasks such as expression transfer, image retrieval, etc.

We develop an incremental training strategy to train the model on a relative small dataset without the rigid requirement of paired samples.

Related Works

Deep generative models have achieved impressive success in recent years. There are two major approaches: generative adversarial network (GAN) (?) and variational autoencoder (VAE) (?). GAN is composed of a generator and a discriminator, where the training is carried out with a minimax two-player game. GAN has been used for image synthesis (?), image superresolution (?), etc. One interesting extension of GAN is Conditional GAN (CGAN) (?) where the generated image can be controlled by the condition variable. On the other hand, VAE is a probabilistic model with an encoder to map an image to a latent representation and a decoder to reconstruct the image. A reparametrization trick is proposed which enables the model to be trained by backpropogation (?). One variant of VAE is Adversarial Autoencoder (?), where an adversarial network is adopted to regularize the latent representation to conform to a prior distribution. Our ExprGAN also adopts an autoencoder structure, but there are two main differences: First, an expression controller module is specially designed, so a face with different types of expressions across a wide range of intensities can be generated. Second, to improve the generated image quality, a face identity preserving loss and two adversarial losses are incorporated.

Facial Expression Editing

Facial expression editing has been actively investigated in computer graphics. Traditional approaches include 3D model-based (?), 2D expression mapping-based (?) and flow-based (?). Recently, deep learning-based methods have been proposed. ? (?) studied a deep belief network to generate facial expression given high-level identity and facial action unit (AU) labels. In (?), a higher-order Boltzman machine with multiplicative interactions was proposed to model the distinct factors of variation. ? (?) proposed a decorrelating regularizer to disentangle the variations between identity and expression in an unsupervised manner. However, the generated image is low resolution with size of 48 x 48, which is not visually satisfying. Recently, ? (?) proposed to edit the facial expression by image warping with appearance flow. Although the model can generate high-resolution images, paired samples as well as the labeled query image are required.

The most similar work to ours is CFGAN (?), which uses a filter module to control the generated face attributes. However, there are two main differences: First, CFGAN adopts the CGAN architecture where an encoder needs to be trained separately for image editing. While for the proposed ExprGAN, the encoder and the decoder are constructed in a unified framework. Second, the attribute filter of CFGAN is mainly designed for a single class, while our expression controller module works for multiple categories. Most recently, ? (?) proposed a conditional AAE (CAAE) for face aging, which can also be applied for expression editing. Compared with these studies, our ExprGAN has two main differences: First, in addition to transforming a given face image to a new facial expression, our model can also control the expression intensity continuously without the intensity training labels; Second, photo-realistic face images with new identities can be generated for data augmentation, which is found to be useful to train an improved expression classifier.

Proposed Method

In this section, we describe our Expression Generative Adversarial Network (ExprGAN). We first describe the Conditional Generative Adversarial Network (CGAN) (?) and Adversarial Autoencoder (AAE) (?), which form the basis of ExprGAN. Then the formulation of ExprGAN is explained. The architectures of the three models are shown in Fig. 1.

CGAN is an extension of a GAN (?) for conditional image generation. It is composed of two networks: a generator network G and a discriminator network D that compete in a two-player minimax game. Network G is trained to produce a synthetic image x^=G(z,y)\hat{x}=G(z,y) to fool D to believe it is an actual photograph, where zz and yy are the random noise and condition variable, respectively. D tries to distinguish the real image xx and the generated one x^\hat{x}. Mathematically, the objective function for G and D can be written as follows:

Adversarial Autoencoder

AAE (?) is a probabilistic autoencoder which consists of an encoder GencG_{enc}, a decoder GdecG_{dec} and a discriminator DD. Apart from the reconstruction loss, the hidden code vector g(x)=Genc(x)g(x)=G_{enc}(x) is also regularized by an adversarial network to impose a prior distribution Pz(z)P_{z}(z). Network DD aims to discriminate g(x)g(x) from z∼Pz(z)z\sim P_{z}(z), while GencG_{enc} is trained to generate g(x)g(x) that could fool DD. Thus, the AAE objective function becomes:

where Lp(,)L_{p}(,) is the pthp_{th} norm: Lp(x′,x)=∣∣x′−x∣∣ppL_{p}(x^{\prime},x)=||x^{\prime}-x||_{p}^{p}

Expression Generative Adversarial Network

Given a face image xx with expression label yy, the objective of our learning problem is to edit the face to display a new type of expression at different intensities. Our approach is to train a ExprGAN conditional on the original image xx and the expression label yy with its architecture illustrated in Fig. 1 (c).

ExprGAN first applies an encoder GencG_{enc} to map the image xx to a latent representation g(x)g(x) that preserves identity. Then, an expression controller module FctrlF_{ctrl} is adopted to convert the one-hot expression label yy to a more expressive expression code cc. To further constrain the elements of cc to capture the various aspects of the represented expression, a regularizer QQ is exploited to maximize the conditional mutual information between cc and the generated image. Finally, the decoder GdecG_{dec} generates a reconstructed image x^\hat{x} combining the information from g(x)g(x) and cc. To further improve the generated image quality, a discriminator DimgD_{img} on the decoder GdecG_{dec} is used to refine the synthesized image x^\hat{x} to have photo-realistic textures. Moreover, to better capture the face manifold, a discriminator DzD_{z} on the encoder GencG_{enc} is applied to ensure the learned identity representation is filled and exhibits no “holes” (?).

In previous conditional image generation methods (?; ?), a binary one-hot vector is usually adopted as the condition variable. This is enough for generating images corresponding to different categories. However, for our problem, a stronger control over the synthesized facial expression is needed: we want to change the expression intensity in addition to generating different types of expressions. To achieve this goal, an expression controller module FctrlF_{ctrl} is designed to ensure the expression code cc can describe the property of the expression intensity except the category information. Furthermore, a regularizer network QQ is proposed to enforce the elements of cc to capture the multiple levels of expression intensity comprehensively.

Expression Controller Module FctrlF_{ctrl} To enhance the description capability, FctrlF_{ctrl} transforms the binary input yy to a continuous representation cc by the following operation:

where the inputs are the expression label y∈{0,1}Ky\in\{0,1\}^{K} and uniformly distributed zy∼U(−1,1)dz_{y}\sim U(-1,1)^{d}, while the output is the expression code c=[c1T,…,cKT]T∈RKdc=[c_{1}^{T},\dots,c_{K}^{T}]^{T}\in R^{Kd}, KK is the number of classes. If the ithi_{th} class expression is present, i.e., yi=1y_{i}=1, ci∈Rdc_{i}\in R^{d} is set to be a positive vector within 0 and 1, while cj,j≠ic_{j},j\neq i has negative values from -1 to 0. Thus, in testing, we can manipulate the elements of cc to generate the desired expression type. This flexibility greatly increases the controllability of cc over synthesizing diverse styles and intensities of facial expressions.

Regularizer on Expression Code QQ It is desirable if each dimension of cc could learn a different factor of the expression intensity variations. Then faces with a specific intensity level can be generated by manipulating the corresponding expression code. To enforce this constraint, we impose a regularization on cc by maximizing the conditional mutual information I(c;x^∣y)I(c;\hat{x}|y) between the generated image x^\hat{x} and the expression code cc. This ensures that the expression type and intensity encoded in cc is reflected in the image generated by the decoder. The direct computation of II is hard since it requires the posterior P(c∣x^,y)P(c|\hat{x},y), which is generally intractable. Thus, a lower bound is derived with variational inference which extends (?) to the conditional setting:

For simplicity, the distribution of cc is fixed, thus H(c∣y)H(c|y) is treated as a constant. Here the auxiliary distribution QQ is parameterized as a neural network, thus the final loss function is defined as follows:

Generator Network G𝐺G

The generator network G=(Genc,Gdec)G=(G_{enc},G_{dec}) adopts the autoencoder structure where the encoder GencG_{enc} first transforms the input image xx to a latent representation that preserves as much identity information as possible. After obtaining the identity code g(x)g(x) and the expression code cc, the decoder GdecG_{dec} then generates a synthetic image x^=Gdec(Genc(x),c)\hat{x}=G_{dec}(G_{enc}(x),c) which should be identical as xx. For this purpose, a pixel-wise image reconstruction loss is used:

To further preserve the face identity between xx and x^\hat{x}, a pre-trained discriminative deep face model is leveraged to enforce the similarity in the feature space:

where ϕl\phi_{l} are the lthl_{th} layer feature maps of a face recognition network, and βl\beta_{l} is the corresponding weight. We use the activations at the conv1_2conv1\_2, conv2_2conv2\_2, conv3_2conv3\_2, conv4_2conv4\_2 and conv5_2conv5\_2 layer of the VGG face model (?).

It is a well known fact that face images lie on a manifold (?; ?). To ensure that face images generated by interpolating between arbitrary identity representations do not deviate from the face manifold (?), we impose a uniform distribution on g(x)g(x), forcing it to populate the latent space evenly without “holes”. This is achieved through an adversarial training process where the training objective is:

Similar to existing methods (?; ?), an adversarial loss between the generated image x^\hat{x} and the real image xx is further adopted to improve the photorealism:

Overall Objective Function

The final training loss function is a weighted sum of all the losses defined above:

We also impose a total variation regularization LtvL_{tv} (?) on the reconstructed image to reduce spike artifacts.

Incremental Training

Empirically we find that jointly training all the subnetworks yields poor results as we have multiple loss functions. It is difficult for the model to learn all the functions at one time considering the small size of the dataset. Therefore, we propose an incremental training algorithm to train the proposed ExprGAN. Overall our incremental training strategy can be seen as a form of curriculum learning, and includes three stages: controller learning stage, image reconstruction stage and image refining stage. First, we teach the network to generate the image conditionally by training GdecG_{dec}, QQ and DimgD_{img} where the loss function only includes LQL_{Q} and LadvimgL_{adv}^{img}. g(x)g(x) is set to be random noise in this stage. After the training finishes, we then teach the network to learn the disentangled representations by reconstructing the input image with GencG_{enc} and GdecG_{dec}. To ensure that the network does not forget what is already learned, QQ is also trained but with a decreased weight. So the loss function has three parts: LpixelL_{pixel}, LidL_{id} and LQL_{Q}. Finally, we train the whole network to refine the image to be more photo-realistic by adding DimgD_{img} and DzD_{z} with the loss function defined in (10). We find in our experiments that stage-wise training is crucial to learn the desired model on the small dataset.

Experiments

We first describe the experimental setup then three main applications: expression editing with continuous control over intensity, facial expression transfer and conditional face image generation for data augmentation .

We evaluated the proposed ExprGAN on the widely used Oulu-CASIA (?) dataset. Oulu-CASIA has 480 image sequences taken under Dark, Strong, Weak illumination conditions. In this experiment, only videos with Strong condition captured by a VIS camera are used. There are 80 subjects and six expressions, i.e., Angry, Disgust, Fear, Happy, Sad and Surprise. The first frame is always neutral while the last frame has the peak expression. Only the last three frames are used, and the total number of images is 1440. Training and testing sets are divided based on identity, with 1296 for training and 144 for testing. We aligned the faces using the landmarks detected from (?), then cropped and resized the images to dimension of 128 x 128. Lastly, we normalized the pixel values into range of . To alleviate overfitting, we augmented the training data with random flipping.

Implementation Details

The ExprGAN mainly builds on multiple upsampling and downsampling blocks. The upsampling block consists of the nearest-neighbor upsampling followed by a 3 x 3 stride 1 convolution. The downsampling block consists of a 5 x 5 stride 2 convolution. Specifically, GencG_{enc} has 5 downsampling blocks where the numbers of channels are 64, 128, 256, 512, 1024 and one FC layer to get the identity representation g(x)g(x). For GdecG_{dec}, it has 7 upsampling blocks with 512, 256, 128, 64, 32, 16, 3 channels. DzD_{z} consists of 4 FC layers with 64, 32, 16, 1 channels. We model Q(c∣x^,y)Q(c|\hat{x},y) as a factored Gaussian, and share many parts of QQ with DimgD_{img} to reduce computation cost. The shared parts have 4 downsampling blocks with 16, 32, 64, 128 channels and one FC layer to output a 1024-dim representation. Then it is branched into two heads, one for DimgD_{img} and one for QQ. QQ has KK branches {Qi}i=1K\{Q_{i}\}_{i=1}^{K} where each QiQ_{i} has two individual FC layers with 64, dd channels to predict the expression code cic_{i}. Leaky ReLU (?) and batch normalization (?) are applied to DimgD_{img} and DzD_{z}, while ReLU (?) activation is used in GencG_{enc} and GdecG_{dec}. The random noise zz is uniformly distributed from -1 to 1. We fixed the dimensions of g(x)g(x) and cc to be 50 and 30, and found this configuration sufficient for representing the identity and expression variations.

We train the networks using the Adam optimizer (?), with learning rate of 0.0002, β1=0.5\beta_{1}=0.5, β2=0.999\beta_{2}=0.999 and mini-batch size of 48. In the image refining stage, we empirically set λ1=1\lambda_{1}=1, λ2=1\lambda_{2}=1, λ3=0.01\lambda_{3}=0.01, λ4=0.01\lambda_{4}=0.01, λ5=0.001\lambda_{5}=0.001. The model is implemented using Tensorflow (?).

Facial Expression Editing

In this part, we demonstrate our model’s ability to edit the expression of a given face image. To do this, we first input the image to GencG_{enc} to obtain an identity representation g(x)g(x). Then with the decoder GdecG_{dec}, a face image of the desired expression ii can be generated by setting cic_{i} to be positive and cj,j≠ic_{j},j\neq i to be negative. A positive (negative) value indicates the represented expression is present (absent). Here 1 and -1 are used. Some example results are shown in Fig. 2. The left column contains the original input images, while the middle row in the right column contains the synthesized faces corresponding to six different expressions. For comparison, the ground truth images and the results from the recent proposed CAAE (?) are also shown in the first and third row, respectively. We can see faces generated by our ExprGAN preserve the identities well. Even some subtle details like the transparent eyeglasses are also kept. Moreover, the synthesized expressions look natural. While CAAE failed to transform the input faces to new expressions with fine details, and the generated faces are blurry.

We now demonstrate that our model can transform a face image to new types of expressions with continuous intensity. This is achieved by exploiting the fact that each dimension of the expression code captures a specific level of expression intensity. In particular, to vary the intensity of the desired class ii, we set the individual element of the expression code cic_{i} to be 1, while the other dimensions of cic_{i} and all other cj,j≠ic_{j},j\neq i to be -1. The generated results are shown in Fig. 3. Take the Happy expression in the forth column as an example. The face in the first row which corresponds to the first element of cic_{i} being 1 displays a gentle smile with mouth closed, while a big smile with white teeth is synthesized in the last row that corresponds to the fifth element of cic_{i} being 1. Moreover, when we set all cic_{i} to be -1, a Neutral expression is able to be generated even though this expression class is not present in the training data. This validates that the expression code discovers the diverse spectrum of expression intensity in an unsupervised way, i.e., without the training data containing explicit labels for intensity levels.

Facial Expression Transfer

In this part, we demonstrate our model’s ability to transfer the expression of another face image xBx_{B} to a given face image xAx_{A}. To do this, we first input xAx_{A} to GencG_{enc} to get the identity representation g(xA)g(x_{A}). Then we train an expression classifier to predict the expression label yBy_{B} of xBx_{B}. With yBy_{B} and xBx_{B}, the expression code cBc_{B} can be obtained from QQ. Finally, we can get an image with identity A and expression B from Gdec(g(xA),cB)G_{dec}(g(x_{A}),c_{B}). The generated images are shown in Fig. 4. We observe that faces having the source identities and expressions similar to the target ones can be synthesized even for some very challenging cases. For example, when the expression Happy is transferred to an Angry face, the teeth region which does not exist in the source image is also able to be generated.

Face Image Generation for Data Augmentation

In this part, we first show our model’s ability to generate high-quality face images controlled by the expression label, then quantitatively demonstrate the usefulness of the synthesized images. To generate faces with new identities, we feed in random noise and expression code to GdecG_{dec}. The results are shown in Fig. 5. Each column shows the same subject displaying different expressions. We can see the synthesized face images look realistic. Moreover, because of the design of the expression controller module, the generated expressions for the same class are also diverse. For example, for the class Happy, there are big smile with teeth and slight smile with mouth closed.

We further demonstrate that the images synthesized by our model can be used for data augmentation to train a robust expression classifier. Specifically, for each expression category, we generate 0.5KK, 1KK, 5KK, and 10KK images, respectively. The classifier has the same network architecture as GencG_{enc} except one additional FC layer with six neurons is added. The results are shown in Table 1. We can see by only adding 3KK synthetic images, the improvement is marginal, with an accuracy of 78.47% vs. 77.78%. However, when the number is increased to 30KK, the recognition accuracy is improved significantly, reaching to 84.72% with a relative error reduction by 31.23%. The performance starts to saturate when more images (60KK) are utilized. This validates the synthetic face images have high perceptual quality.

Feature Visualization

In this part, we demonstrate that the identity g(x)g(x) and expression cc representations learned by our model are disentangled. To show this, we first use t-SNE (?) to visualize the 50-dim identity feature g(x)g(x) on a two dimensional space. The results are shown in Fig. 6. We can see that most of the subjects are well separated, which confirms the latent identity features g(x)g(x) learn to preserve the identity information.

To demonstrate that the expression code cc captures the high-level expression semantics, we perform image retrieval experiment based on cc in terms of Euclidean distance. For comparison, the results with expression label yy and image pixel space xx are also provided in Fig. 7. As expected, the pixel space xx sometimes fails to retrieve images from the same expression. While the images retrieved by yy do not always have the same style of expressions as the queries. For example, the query face in the second row shows a big smile with teeth, but the retrieved image by yy only has a mild smile with mouth closed. However, with the expression code cc, we observe that face images with similar expressions are always retrieved. This validates that the expression code learns a rich and diverse feature representation.

Conclusions

This paper presents ExprGAN for facial expression editing. To the best of our knowledge, it is the first GAN-based model that can transform the face image to a new expression where the expression intensity is allowed to be controlled continuously. The proposed model learns the disentangled identity and expression representations explicitly, allowing for a wide variety of applications, including expression editing, expression transfer, and data augmentation for training improved face expression recognition models. We further develop an incremental learning scheme to train the model on small datasets. Our future work will explore how to apply ExprGAN to a larger and more unconstrained facial expression dataset.

References