Geometry-Contrastive GAN for Facial Expression Transfer
Fengchun Qiao, Naiming Yao, Zirui Jiao, Zhihao Li, Hui Chen, Hongan Wang
Introduction
Facial expression transfer aims to transfer facial expressions from a source subject to a target subject. The newly-synthesized expressions of the target subject are supposed to be identity-preserving and exhibit similar emotions to the source subject. A wide range of applications such as virtual reality, affective interaction, and artificial agents could be facilitated by the advances of facial expression transfer. The past few decades has witnessed numerous techniques designed for facial expression transfer. Most representative methods are conducted in a sequence-to-sequence way where they assume the availability of a video of the target subject with variants in facial expressions and head poses, which limits their applications. Averbuch et al. proposed a warp-based method to animate a single facial image in a sequence-to-image way, capable of generating both photo-realistic and high-resolution facial expression videos, just like bringing a still portrait to life. However, there exists misalignment across different subjects without the assistance of neutral faces for alignment, leading to artifacts on the generated faces. Moreover, large manual adjustments or high computations are still required to generate realistic expressions. How to handle pixel-wise misalignments across different subjects with different emotions in an easy and automatic way is still an open problem.
Recently, Generative Adversarial Nets (GANs) have received extensive attentions to generate lavish and realistic faces. Conditional GANs (cGANs) have also been widely applied in many face-related tasks, such as face pose manipulation , face aging , and facial expression transfer . Existing works relating to GAN-based facial expression transfer mainly focus on generating facial expressions with discrete and limited emotion states. Emotion states (e.g., happy, sad) were usually encoded as conditions of one-hot codes to control the generation of expressions . Facial landmarks of action units (AU) in the Facial Action Coding System (FACS) were also adopted as conditions to guide the generated expressions . In , the target face and facial landmarks of the driving face are concatenated in image space, which brings extra artifacts when there exist big differences in facial shapes and expressions between target and driving faces. Since people express emotions in a continuous and vivid way, how to inject facial geometry into GANs to generate continuous emotions of the target subject is essential in facial expression transfer.
In order to transfer continuous emotions across different subjects, a Geometry-Contrastive Generative Adversarial Network (GC-GAN) is proposed in this paper. GC-GAN consists of a facial geometry embedding network, an image generator network, and an image discriminator network. Contrastive learning from geometry information is integrated in embedding network. Its bottleneck layer, representing a semantic manifold of facial expressions, is concatenated into the latent space in GAN. Therefore, continuous facial expressions are displayed within the latent space, and the pixel-to-pixel misalignments across different subjects are resolved via embedded geometry. Our main contributions are as follows: (1) We apply contrastive learning in GAN to embed geometry information onto a semantic manifold in the latent space. (2) We inject facial geometry to guide the facial expression transfer across different subjects with Lipschitz continuity. (3) Experimental results demonstrate that our proposed method can be applied in facial expression transfer even there exist big differences in facial shapes and expressions between different subjects.
Related work
GAN-based conditional image generation has been actively studied. Mirza et al. proposed conditional GAN (cGAN) to generate images controlled by labels or attributes. Larsen et al. proposed VAE-GAN combing variational autoencoder and GAN to learn an embedding representation which can be used to modify high-level abstract visual features. cGANs have also been applied in facial expression transfer . Zhou et al. proposed a conditional difference adversarial autoencoder (CDAAE) to generate faces conditioned on emotion states or AU labels. Choi et al. proposed StarGAN to perform image-to-image translations for multiple domains using only a single model, which can also be applied in facial expression transfer. Previous studies mainly focus on generating facial images conditioned on discrete emotion states. However, human emotion is expressed in a continuous way, thus discrete states are not sufficient to describe detailed characteristics of facial expressions.
Some researchers have attempted to incorporate continuous information such as geometry into cGANs . Ma et al. proposed a pose guided person generation network (PG2), which allows to generate images of individuals in arbitrary poses. The target pose is defined by a set of 18 joint locations and encoded as heatmaps, which are concatenated with the input image in PG2. In addition, they adopted a two-stage generation method to enhance the quality of generated images. Song et al. proposed a Geometry-Guided Generative Adversarial Network (G2GAN), which applies facial geometry information to guide the transfer of facial expressions. In G2GAN, facial landmarks are treated as an additional image channel to be concatenated with the input face directly. They use dual generators to perform the synthesis and the removal of facial expressions simultaneously. The neutral facial images are generated through the removal network and used for the subsequent facial expression transfer. This procedure brings additional artifacts and degrades the performance especially when the driving faces are collected from other subjects in different emotions.
Contrastive learning minimizes a discriminative loss function that drives the similarity metric to be small for pairs of the same class, and large for pairs from different classes. Although contrastive learning has been widely used in recognition tasks , potentials of the semantic manifold in latent space haven’t been fully explored, which means that inter-class transitions can be expressed in a continuous representation. In our method, we introduce contrastive learning in GANs to embed geometry information and image appearance in continuous latent space to smoothly minimize the pixel-wise misalignments between different subjects and to generate continuous emotion expressions.
Method
The overall framework of GC-GAN is shown in Figure 1. GC-GAN consists of three components: a facial geometry embedding network (, ), an image generator network (, ), and an image discriminator network . Input facial image and target landmarks are encoded by and into and respectively. Then and are concatenated into a single vector for . is reference facial landmarks used for contrastive learning against . Note that geometry expression features and image identity features are learned in a disentangled manner, thus we could modify expression and keep identity preserved.
2 Geometry-Contrastive GAN
The objective of original GANs is formulated as:
where and are the discriminator and generator respectively. In order to guide the generated samples, conditional information is input into generators as well as input prior in GANs. Classical and other cGANs mainly use a method of concatenation of prior variables and conditional variables, which is often represented as one-hot codes or embedded discrete variables .
Now we first introduce continuous conditions into generators to provide more precise guidance. The input of generator consists of two parts, and , i.e., . and obey different independent prior distributions and , so , and thus is also a prior. When we use the prior to drive the generation of images, it is usually accepted that the generated images are not abruptly changed when the input is gradually changed. Referring to the LS-GAN , this assumption can be represented as the Lipschitz continuity of for , i.e., , , , s.t. . When is fixed, i.e., , we get , which means is Lipschitz continuous for . Similarly, when is fixed, it is concluded that is Lipschitz continuous for :
We denote and , a certain continuous prior distribution. The Lipschitz of for indicates that the condition we introduce can continuously control the generation. It is obvious that if is a discrete distribution, the model is altered into an ordinary cGAN and the Lipschitz of for turns false, because does not exist in a metric space and makes no sense anymore.
In GC-GAN, we aim to solve the problem in Eq.(1), where and represent identity information and facial expression condition of generated faces respectively. And and here are independent since the known image and geometry vector don’t necessarily refer to the same subject. Compared to discrete distribution of facial expressions, if is some continuous distribution derived from facial landmarks, conditions can guide the generation of facial expressions more precisely.
However, it increases the risk of errors to introduce geometry variables into cGANs directly i.e., . Original facial landmark space is badly separable for facial expressions. So when changes as conditions, the changes of generated images will get out of control, i.e., the constant is too large or approximately infinity. It also has bad interpretability to concatenate landmark vector and prior in the input of generator. So an embedding method is necessary to transform the original geometry information into a more effective representation. Since the objective of contrastive learning is based on the calculation of metrics, we apply a contrastive learning network to the landmarks and define an embedding variable . Through contrastive learning, landmarks are encoded into a semantic space of facial expression, which reduces the dimensions, extracts facial expression features, and provides the different features with better separability and appropriate distance. Thus, when the conditions in semantic space are introduced and the prior fixed, the generated images are Lipschitz continuous over the facial expression conditions with the identity information preserved.
3 Learning strategy
Contrastive learning is conducted on facial geometry with reference facial landmarks, where the -coordinates of facial landmarks are arranged into a one-dimensional vector as input for the geometry embedding network. Landmark pairs are prepared for training, in which is the target expression while is the reference expression. represents a certain subject. After a transform function , the facial landmarks are mapped into the embedding space. Our goal is to measure the similarity between and according to their expression labels. The contrastive loss is formulated as below:
where , if the expression labels are the same, otherwise. is a margin which is always greater than 0. enables our embedding network to focus on facial expression itself regardless of different subjects and plays an essential part for pushing facial landmarks to reside in a semantic manifold. The semantic manifold is explained in detail in Section 4.3.
Adversarial learning
In our model, the adversarial losses for and are formulated as follows:
where and indicate the distribution of real facial images and real facial landmarks, respectively. represents the generated image computed by . The adversarial losses are optimized via WGAN-GP .
Reconstruction learning
Overall learning
First, network is pretrained on the training set. The loss function is . Then network and are trained together while the parameters of network are fixed. The loss function is . , , and are set to , , and , respectively. The detailed parameters of GC-GAN are provided in Appendix-A.
Experiments
To evaluate the proposed GC-GAN, experiments have been conducted on four popular facial datasets: Multi-PIE , CK+ , BU-4DFE and CelebA .
Multi-PIE consists of 754,200 images from 337 subjects with large variations in head pose, illumination, and facial expression. We select a subset of images with three head poses (), 20 illumination conditions, and all six expressions. CK+ is composed of 327 image sequences from 118 subjects with seven prototypical emotion labels. Each sequence starts with a neutral emotion and ends with the peak of a certain emotion. BU-4DFE contains 101 subjects, each one displaying six acted categorical facial expressions with moderate head pose variations. Since this dataset does not contain frame-wise facial expression, we manually select neutral and apex frames of each sequence. CelebA is a large-scale facial dataset, containing 202,599 face images of celebrities. All experiments are performed in a person-independent way, the training set and test set for each dataset are split based on subjects with proportions of 90% and 10%, respectively.
To prepare the training triplet , we create all triplets of facial expressions per identity. is taken from and is generated by sampling at random. Note that and facial expression labels are not required in test. In our experiments, 68 facial landmark points are obtained through dlib http://dlib.net/ which implements the method proposed in . Since faces in our experiments remain near-frontal, the accuracy of facial landmark detection is over 95% when dlib is used, which is sufficient for our work. The facial landmarks include points of two eyebrows, two eyes, the nose, the lips, and the jaw. The facial images are aligned according to inner eyes and bottom lip. Then, face regions are cropped and resized into . Pixel value of images and -coordinates of facial landmarks are both normalized into $$.
2 Facial expression generation
In order to evaluate whether the synthetic faces are generated by the guidance of the target facial expression, we qualitatively and quantitatively compare the generated faces with the ground truth. The results are shown in Figure 2, in which the generated faces are identity-preserving and are similar to the target facial expressions. For quantitative comparison, SSIM (structural similarity index measure) and PSNR (peak-signal-to-noise ratio) are used as two evaluation metrics. The detailed comparison results are shown in Table 1. We conduct ablation study to evaluate contributions of three losses , , and , respectively. According to Table 1, GC-GAN performs better when supervised by the three losses. We compare our method with CDAAE , which uses one-hot codes to represent facial expressions. As shown in Table 1, our method substantially outperforms CDAAE on both measures across all the datasets, indicating that the images generated using our method are more similar to the ground truth aided by the embedded geometry features. The experiments mentioned above demonstrate that our proposed GC-GAN effectively incorporates geometry information in conditioned facial expression generation.
3 Analysis of geometry-contrastive learning
First, we evaluate the Lipschitz continuity of generated faces over facial expression conditions . For interpolation of , we would like to seek a sequence of images with a smooth transition between two different facial expressions of the same person (e.g. and ). In our case, we conduct equal interval sampling on . Some generated facial expression sequences are shown in Figure 3(a). According to Eq.(3), we compute between two adjacent frames over the whole test set composed of 58,724 pairs from 33 subjects, and the violin plot is shown in Figure 3(b), where the mean at each time step is around 50 and the upper bound also exists. Therefore, the Lipschitz constant in Eq.(3) is approximately found.
In order to evaluate whether the manifold of is semantic-aware, we visualize the embedding manifold of facial landmarks by t-SNE . For the purpose of evaluating the effectiveness of contrastive learning, we train our model without contrastive loss . The corresponding manifold of facial landmarks is shown in Figure 4(a), in which the embedding features cannot be clustered properly without the guidance of contrastive learning. The manifold of GC-GAN is shown in Figure 4(b) where the embedding features are clustered according to their emotion states. It can be seen that the data scatters of disgust and squint are mixed up with each other due to the intrinsic geometry similarity between the two facial expressions. In addition, we present an embedding line consisting of a sequence of points corresponding to the transition of facial expressions from smile to scream of the same person obtained by equal interval sampling on facial landmarks, which is shown in Figure 4(c). It can be observed that the eyes are closing and the mouth is opening from smile to scream. So there could exist some points with open eyes and open mouth simultaneously which is exactly the characteristic of surprise, and could account for the reason why the embedding line comes across the region of surprise. The results indicate that the manifold of is both continuous and semantic-aware for GC-GAN.
4 Facial expression transfer
To further evaluate the generality of our proposed method, transferring a person’s emotion to different emotions of another person directly is also tested. Facial expression transfer between different persons are difficult due to the individual variances, especially in the absence of neutral faces for reference.
In our case, the target expressions are taken from other people of different datasets. Here, we randomly select 1000 facial images from CelebA. The facial landmarks of these images serve as guiding landmarks. Note that there are only six basic emotion states in Multi-PIE while the emotion states in CelebA are various and spontaneous. Figure 5 shows some representative samples from our experiment, in which the generated faces are identity-preserving and exhibit similar expressions with driving faces even there exist big differences in face shape between input faces and driving faces.
For comparison, a more general configuration SPG is adopted, which is similar to PG2 and G2GAN . In detail, we treat the guiding landmarks as an additional feature map and directly concatenate it with the input face, and train our model in this configuration. The corresponding results of SPG are shown in the lowest line of Figure 5.
Since there don’t exist ground truth faces corresponding to generated faces, SSIM and PSNR are not suitable for quantitative evaluation. In order to validate whether the generated faces are identity-preserving, a pretrained face identification model VGG-Face is adopted to conduct face verification on the generated faces. In detail, we extract the output of the last convolution layer for measuring the identity similarity between two faces. Two faces are considered as matched if their cosine similarity is no less than 0.65. In order to evaluate the similarity of facial expressions between generated and driving faces, the cosine similarity of of facial landmarks between generated and driving faces is adopted. The accuracy of face verification and the mean similarity of facial expressions are shown in Table 2. Compared with our method, SPG is not able to capture the target facial expression of driving faces and brings extra artifact, indicating that the model under this configuration cannot handle the misalignments between input faces and driving facial landmarks properly. Our model shows the promising capability of manipulating facial expressions while preserving identity information.
Conclusion
In this paper, we have proposed a Geometry-Contrastive Generative Adversarial Network (GC-GAN) for transferring facial expressions across different persons. Through contrastive learning, facial expressions are pushed to reside onto a continuous semantic-aware manifold. Benefited from this manifold, pixel-wise misalignments between different subjects are weakened and identity-preserving faces with more expression characteristics are generated. Experimental results demonstrate that there exists semantic consistency between the latent space of facial expressions and the image space of faces.
References
Appendix
Network architecture
The detailed architectures of GC-GAN are shown in Table 3. In detail, generator network has a similar architecture to U-Net, but we only add one skip connection between the output of the first convolution layer and the output of the penultimate deconvolution layer, which enables to reuse low-level facial features related with identity information. For discriminator network , the output is a feature map.
In training time, the margin in is set to 5. All the networks , , and are optimized by Adam optimizer with and . The batch size is 64 and the initial learning rate is set to . We implement GC-GAN using Tensorflow.
Reviews
The paper of expression transfer in face analyses is very relevant. Although reviewers agree that including geometry in transfer with GANs and the analyses with contrastive loss and Lipschitz continuous semantic latent space are interesting, there is a generalized opinion that most of the pieces of the model are standard techniques. Overall, the paper does not achieve the minimum score required to be published at NIPS.
1. Please provide an "overall score" for this submission.
6: Marginally above the acceptance threshold. I tend to vote for accepting this submission, but rejecting it would not be that bad.
2. Please provide a "confidence score" for your assessment of this submission.
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
3. Please provide detailed comments that explain your "overall score" and "confidence score" for this submission. You should summarize the main ideas of the submission and relate these ideas to previous work at NIPS and in other archival conferences and journals. You should then summarize the strengths and weaknesses of the submission, focusing on each of the following four criteria: quality, clarity, originality, and significance.
The paper describes a method for generating synthetic images of faces with different facial expressions. The authors propose a GAN-based architecture for the purpose, wherein, the facial identity and expression priors are disentangled and encoded in different latent spaces. The expression and identity facial latent vectors are concatenated and reconstructed by the generator to preserve the identity and change the facial expression. The facial expressions are represented by 68 facial fiducial points and encoded into a latent space. The authors use reconstruction losses for the input facial fiducial points and facial images along with adversarial training for the entire algorithm. Overall, neither the problem for synthesizing facial expressions, nor its proposed solution using GANs is novel. The paper is also specific to facial analysis and hence not directly applicable to many different problem domains. Hence the paper has decent novelty and some interesting findings, but is not earth-shattering or likely to be of broad impact.
In contrast to the prior work, the claimed superiority of this method lies in the fact that it allows to learn a Lipschitz continuous semantic latent space of facial expressions instead of a discrete one. The latter is what previous works have learnt via one-hot encoding of the facial expression classes. By comparing their approach to that of , the authors show empirically the value of learning a latent embedding of the facial fiducial points for continuous transitions of expressions versus using the facial fiducial points directly as inputs along with the facial image for the GAN-based approach. The latter they argue is not guaranteed to be Lipschitz continuous. This is an interesting finding.
One other finding of this work is that using the contrastive loss for encoding facial expressions helps to learn a more semantically meaningful expression latent space and hence helps to produce better expression synthesis results on moving linearly on the expression manifold. However, the authors’ use of the contrastive loss is at odds with their goal of trying to learn a *continuous* latent expression space. The contrastive loss tends to create discontinuous clusters for discrete facial expressions classes, which is exactly what the authors observe in the figure 4(b) on visualizing their learnt latent space. This is also something that the authors set out to avoid in the first place. How then do the authors explain the helpfulness of the constructive loss in learning a continuous latent expression space? This contradictory observation needs to explored further and clearly addressed and explained in the paper. I would have imagined that something like a variational auto encoder with a KL divergence loss term would have helped to create a more continuously varying latent expression space.
Other than that, the paper is well written, includes sufficient experimental evidence to demonstrate the superior performance of their method versus previous work, and includes sufficient details in the paper and the supplementary document of the network and training for someone to be able to replicate it.
4. How confident are you that this submission could be reproduced by others, assuming equal access to data and resources?
1. Please provide an "overall score" for this submission.
6: Marginally above the acceptance threshold. I tend to vote for accepting this submission, but rejecting it would not be that bad.
2. Please provide a "confidence score" for your assessment of this submission.
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
3. Please provide detailed comments that explain your "overall score" and "confidence score" for this submission. You should summarize the main ideas of the submission and relate these ideas to previous work at NIPS and in other archival conferences and journals. You should then summarize the strengths and weaknesses of the submission, focusing on each of the following four criteria: quality, clarity, originality, and significance.
The paper introduces Geometry-Contrastive Generative Adversarial Network (GC-GAN), a novel deep architecture for transferring facial emotions across different subjects. The proposed architecture is based on the idea of decoupling the identity information with the motion information in the latent representation of the main network. A novel contrastive loss is employed to learn the latent motion information from landmark images. The proposed approach is evaluated on three common face image benchmarks.
- This paper, together with , is one of the first attempts to inject geometry information into GAN models for facial expression transfer. This is an interesting and challenging problem.
- The idea of using contrastive learning to compute the latent representation from the landmark images is new, despite contrastive learning have been exploited in many other applications
- The experimental evaluation consider three different benchmarks where face have different characteristics. The proposed approach is shown to be more effective than previous method on the proposed task.
- The paper is well written and the proposed approach clearly explained.
- In my opinion the experimental evaluation can be improved. Specifically, when analyzing the dynamics of face sequences in section 4.3 a comparison with can be added. This will confirm the effectiveness of the proposed contrastive loss.
- In Table 2 it is not clear what SPG is. Which method was used exactly?
- The results from a user study could be added, as SSIM may be not the best metric for analyzing generated faces.
4. How confident are you that this submission could be reproduced by others, assuming equal access to data and resources?
1. Please provide an "overall score" for this submission.
4: An okay submission, but not good enough; a reject. I vote for rejecting this submission, although I would not be upset if it were accepted.
2. Please provide a "confidence score" for your assessment of this submission.
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
3. Please provide detailed comments that explain your "overall score" and "confidence score" for this submission. You should summarize the main ideas of the submission and relate these ideas to previous work at NIPS and in other archival conferences and journals. You should then summarize the strengths and weaknesses of the submission, focusing on each of the following four criteria: quality, clarity, originality, and significance.
This paper presented an framework - GC-GAN for facial expression transfer/generation. The proposed method consists of a facial geometry embedding network and a conditional GAN networks. The experimental on several datasets show the effectiveness of GC-GAN.
1. the integration of the facial geometry embedding network is somehow novel
1.the novelty contribution of the overall framework is minor. The generator/discriminator just follow general conditional GANs. The geometry-guided idea have been explored a lot in previous works (such as G2GAN,GAGAN).
2. the experimental part is not strong: a) the model has many hyper-parameters/pretrianed modules setting, making it difficult to reproduce the results. The details of those are not described clearly; b) just few baselines are compared in the results. There should be more about GANs with structured conditions.
3. The analysis of geometry-contrastive learning, i.e., the facial embedding network is not convincing. Section 4.3 can not give a convincing quantitative result. The measure in Section 4.4 only takes the overall measure but fails to show how the generated samples have better representation regarding to the continuity.
Some related and important references are missing about GANs:
a. Disentangled Representation Learning GAN for Pose-Invariant Face Recognition
4. How confident are you that this submission could be reproduced by others, assuming equal access to data and resources?
Rebuttal
Q: About the novelty and potential application to other problem domains.
A: Yes, the problem for synthesizing facial expressions is not new, but how to integrate continuity of facial expressions in GANs hasn’t been mentioned or studied carefully before, thus the detailed problem and our solution are novel to a certain extent. Facial expressions are intrinsically continuous but always annotated with discrete labels, so our method can also be applied in similar problems such as the pose of objects, the length of hair, the height of heels and so on.
Q: About the effect of contrastive learning for semantically continuous representation.
A: Contrastive learning has been widely used in face identification such as FaceNet, VGG-Face and so on. Face identification is a typical classification problem, however, these networks trained for it also generalize well on face verification where faces are unseen. Similarly, such learning mechanism enables the latent space to accommodate unseen emotion states. As seen in Figure 4(b), the data scatters of disgust and squint are not separated well with each other due to the intrinsic geometry similarity although they are different expressions. In Figure 4(c), we present an embedding line corresponding to the transition from smile to scream of the same person obtained by equal interval sampling on facial landmarks. It can be observed that the eyes are closing and the mouth is opening, which is exactly the embodiment of continuity. We try to give a general architecture of continuous latent space in cGANs and have embedded an autoencoder in current work, and we will test KL divergence loss term in the future. Thanks!
A: proposes a method to generate facial image sequences of diverse smiles, which is an excellent work. We’ll add this comparison in terms of analyzing the dynamics of face sequences in our future work.
Q: The explanation of SPG and user study.
A: SPG is the abbreviation of “similar to PG2 and G2GAN”, which means the method of concatenating geometry and images directly as in both PG2 and G2GAN. PG2 mainly focuses on the generation of persons of different poses and the input facial images of G2GAN should be neutral expressions. Besides, the ways of training GANs, neural network architectures and loss functions among PG2, G2GAN and GC-GAN are different. It’s improper to directly compare their methods with ours considering the reasons mentioned above, so we construct such configuration for fair comparison. We also took user study into consideration. However, user study is somewhat subjective and may cause some trouble if other people want to compare their methods with ours. We’ll conduct a user study if necessary. Thanks for your advice!
Q: The novelty contribution of the overall framework is minor.
A: Instead of following the general conditional GAN, we concentrate on how to establish semantically continuous conditions into GANs, and formal description of Lipschitz continuity was given in Section 3.2. Despite geometry-guided ideas have been explored directly in previous works, we are the first one to embed geometry information in contrastive learning to maintain continuity into cGANs.
A: a) The detailed architecture and the key parameters are provided in supplementary materials. This work can be easily reproduced and code will be available upon publication. b) Two types of baselines have been compared. First, a standard one-hot code based cGAN CDAAE and also an ablation study have been conducted in Section 4.2. Second, we compare our method with two state-of-the-art methods PG2 and G2GAN in Section 4.4. Three different benchmarks have been tested both qualitatively and quantitatively.
Q: The analysis of geometry-contrastive learning.
A: Firstly, we evaluate the Lipschitz continuity of generated faces over facial expression conditions in Section 4.3, and the quantitative result is given in Figure 3. Besides, the embedded manifold is visualized qualitatively in Figure 4 for intuitive interpretation. Secondly, in Section 4.4, we mainly demonstrate how geometry-contrastive learning handles the misalignments between different persons and the generalization on the spontaneous emotion states.
Q: Some related and important references are missing about GANs.
A: Paper (c) has been cited as in our paper. Paper (b) hasn’t been officially published before the submission deadline and doesn’t get much improvements compared with paper (c). Paper (a) concentrates on face frontalization and pose-invariant face recognition, which doesn’t have close relationship with our work.