Talking Face Generation by Adversarially Disentangled Audio-Visual Representation

Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, Xiaogang Wang

Introduction

Understanding talking faces visually is of great importance to machine perception and communication. Humans can not only guess the semantic meaning of words by observing lip movement but also imagine the scenario when a specific subject talks (i.e. face generation). Recent advances have focused on automatic lip reading, which surpasses human-level performance in certain domains. Here, we explore generating a video of arbitrary-subject speaking, which perfectly syncs with a specific speech where the speech information can be represented by either a clip of audio or video. We refer this problem as arbitrary-subject talking face generation, as shown in Fig. 1.

However, generating identity-preserving talking faces that clearly conveys certain speech information is a challenging task, since the continuous deformation of the face region relates to both intrinsic subject traits (?) and extrinsic speech vibrations. Previous efforts in this direction are mainly from computer graphics (?; ?; ?; ?; ?). Researchers construct specific 3D face model for a chosen subject and the talking faces are animated by manipulating 3D meshes of the face model. However, these approaches strongly rely on the 3D face model and are hard to scale up to arbitrary identities. More recent attempts (?) leverage the power of deep generative model and learn to generate talking faces from scratch. Though the resulting models can be applied to an arbitrary subject, the generated face sequences are sometimes blurry and not temporally meaningful. One important reason is that the subject-related and speech-related information are coupled together such that the talking faces are difficult to learn in a purely data-driven manner.

To address the aforementioned problems, we integrate the identity-related and speech-related information by learning disentangled audio-visual representation, as illustrated in Fig. 2. We aim to disentangle a talking face sequence into two complementary representations, one containing identity information while the other containing speech information. However, directly separating these two parts is not a trivial task because the variations of face deformation can be extremely large considering the diversity of potential subjects and speeches.

The key idea here is using audio-visual speech recognition (?; ?) (i.e. recognizing words from talking face sequence and audios, aka lip reading) as a probe task for associating audio-visual representations, and then employing adversarial learning to disentangle the subject-related and speech-related information inside them. Specifically, we first learn a joint audio-visual space where talking face sequence and its corresponding audio are embedded together. It is achieved by enforcing the lip reading result obtained from talking faces aligns with the speech recognition result obtained from audio. Next, we further utilize lip reading task to disentangle subject-related and speech-related information through adversarial learning (?). Notably, we enforce one of the representations extracted from talking faces to fool the lip reading system, in the sense that it only contains subject-related information, but not speech-related information. Overall, with the aid of associative-and-adversarial training, we can jointly embed audio-visual inputs and disentangle subject and speech-related information of talking faces.

The contributions of this work can be summarized as follows. (1) A joint audio-visual representation is learned through audio-visual speech discrimination by associating several supervisions. Experiments show that the joint-embedding improves the baseline of lip reading result on LRW dataset (?). (2) Thanks to the discriminative nature of our joint representation, we disentangle the person-identity and speech information through adversarial learning for better talking face generation. (3) By unifying audio-visual speech recognition and audio-visual synchronizing, we achieve arbitrary-identity talking face generation from either video or audio speech as inputs in an end-to-end framework, which synthesizes high-quality and temporally-accurate talking faces.

Related Work

Generating Talking Faces. The work of synthesizing lip motion from either audio (?; ?; ?; ?; ?; ?) or generating moving faces from videos (?; ?; ?) has long been a task of concern in both the community of computer vision and graphics. However, most synthesis works from audio require a large amount of video footage of the target person for training, modeling, or sampling. They could not transfer the speech information to an arbitrary photo in the wild.

(?) use a setting that is different from the traditional ones. They try to directly generate the whole face image with different lip motions in an image-to-image translation manner based on audios. But their method base on data-driven training using an autoencoder, which leads to blurry results and lacks continuity. More recently, (?) propose to use conditional RNN adversarial network, and (?) propose to use correlation loss and three-stream GAN. (?) use flow to generate high precision arbitrary-identity talking face based on videos and claim to be able to produce videos based on audios, but with no results shown. However, as a common problem, without specific disentangling face and lip motion information, they all cannot generate high-quality results.

Learning Audio-Visual Representation. The task of audio-visual speech recognition is a recognition problem uses either one or both video and audio as inputs. Using visual information only for recognition is also referred to as Lip Reading. A review of traditional methods for tackling this task has been made in (?) thoroughly. In recent years, this field develop quickly with the usage of convolutional neural networks (CNNs) and recurrent neural networks (RNNs) for end-to-end word-level (?; ?), sentence-level (?; ?), and multi-view (?) lip reading. In the meantime, the exploration of this topic has been greatly pushed forward by the build-up of large-scale word-level lip reading dataset (?), and the large sentence-level multi-view dataset (?).

For the correspondence between human faces and audio clips, a number of works have been proposed to solve the problem of the audio-video synchronization between mouth motion and speech (?; ?). Particularly, SyncNet (?; ?) used two stream CNNs to sync audio mfcc with 5 consecutive frames. In (?), they further fixed the sync image feature as the pretraining for lip reading, but the two tasks are still separate from each other. Recently, works from (?; ?) also attempt to learn the association between a human face and voice for identity recognition instead of semantic level synchronization.

Approach

We propose Disentangled Audio-Visual System (DAVS), an end-to-end trainable network for talking face generation by learning disentangled audio-visual representations, as shown in Fig. 3.

We leverage both talking video SvS^{v} and its corresponding audio SaS^{a} as training inputs. For learning the disentangled audio-visual representations between Person-ID space (pidpid) and the Word-ID space (widwid), there are three encoder networks involved:

Video to Word-ID space encoder (Ewv\text{E}^{v}_{w}): Ewv\text{E}^{v}_{w} learns to embed the video frame svs^{v} into a visual representation fwvf^{v}_{w} which only contains speech-related information. It is achieved by learning a joint embedding space which associates video and audio that correspond to the same word.

Audio to Word-ID space encoder (Ewa\text{E}^{a}_{w}): Ewa\text{E}^{a}_{w} learns to embed the speech sas^{a} into an audio representation fwaf^{a}_{w}, which resides in the shared space with fwvf^{v}_{w} as introduced above.

Video to Person-ID space encoder (Epv\text{E}^{v}_{p}): Epv\text{E}^{v}_{p} learns to embed the video frame svs^{v} into a representation fpvf^{v}_{p} which only contains subject-related information. It is achieved by the adversarial training process, forcing our target representation fpvf^{v}_{p} to fool the speech recognition system.

The whole idea of our pipeline is to first learn the discriminative audio-visual joint space widwid, then disentangle it from the pidpid space. Finally to combine features from the two spaces to get generation results. Specifically, for learning the widwid space, we employ three supervisions: the supervision of Word-ID labels with shared classifier Cw\text{C}_{w} for associating audio and visual signals with semantic meanings; contrastive loss LC\mathcal{L}_{C} for pulling paired video and audio samples closer; and an adversarial training supervision on audio and video features to make them indistinguishable. As for the pidpid space, Person-ID labels from extra labeled face data are used. For disentangling widwid and pidpid spaces, adversarial training is employed. As for generation, we introduce L1L_{1}-norm reconstruction loss LL1\mathcal{L}_{L_{1}} and temporal GAN loss LGAN\mathcal{L}_{GAN} for sharpness and continuity.

We learn a joint audio-visual space that associates representations from both sources. We constrain the extracted audio representation to be close to its corresponding visual representation, forcing the embedded features to share a same distribution and restricting fwa≃fwvf^{a}_{w}\simeq f^{v}_{w}, so that G(fpv,fwv)≃G(fpv,fwa)G(f^{v}_{p},f^{v}_{w})\simeq G(f^{v}_{p},f^{a}_{w}) can be achieved. While requiring information of person facial identity flows from the pidpid space, the other space of widwid would have to be person-ID invariant. The task of audio-visual speech recognition benefits us in achieving the shared latent space assumption and creating a discriminative space through mapping videos and audios to word labels. The implementation of learning the space is shown in Fig 4 (a). Then with the discriminative embedding, we can take the advantage of adversarial training for thoroughly information disentangling as described in Sec. 3.2

Sharing Classifier. After the embedded features are extracted from the widwid encoders Ewa\text{E}^{a}_{w}, Ewv\text{E}^{v}_{w} to get Fwv=[fwv(1),⋯ ,fwv(n)]F^{v}_{w}=[{f^{v}_{w}}_{(1)},\cdots,{f^{v}_{w}}_{(n)}] and Fwa=[fwa(1),⋯ ,fwa(n)]F^{a}_{w}=[{f^{a}_{w}}_{(1)},\cdots,{f^{a}_{w}}_{(n)}], normally they would be fed into different classifiers for visual and audio speech recognition. Here we share the classifier for both the modalities to enforce them to share their distributions. As a classifier’s weight wj{\bf{w}}_{j} tend to fall into the center of the clustering of the features belonging to the jj’th class, through sharing the weights, the features between both modalities are pulled towards the centroid of the class (?). The supervision is denoted as Lw\mathcal{L}_{w}.

Contrastive Loss. As the problem of mapping audio and visual together is very similar to feature mapping (?), retrieval and particularly the same as lip sync (?), we adopted the contrastive loss which aims at bringing closer paired data while dispelling unpaired as a baseline. During training, for a batch of NN audio-video samples, the mmth and nnth sample are drawn with labels lm=n=1l_{m=n}=1 while the others lm≠n=0l_{m\neq n}=0. The distance metric used to measure the distance between Fwa(m){F^{a}_{w}}_{(m)} and Fwv(n){F^{v}_{w}}_{(n)} here is the euclidean norm dmn=∥Fwv(m)−Fwa(n)∥2d_{mn}=\|{F^{v}_{w}}_{(m)}-{F^{a}_{w}}_{(n)}\|_{2}. The objective can be written as:

During our implementation, all features Fwv,FwaF^{v}_{w},F^{a}_{w} used in this loss are normalized first.

Domain Adversarial Training. To further push the face and audio features to be in the same distribution, we apply a domain adversarial training. An extra two-class domain classifier is appended for distinguishing the source of the feature. The audio and face encoders are then trained to prevent the classifier from success. This is mostly a simple version of the adversarial training described in section 3.2. We refer to the objective of this method as LadvD\mathcal{L}^{D}_{adv}.

2 Adversarial Training for Latent Space Disentangling

In this section, we describe how we disentangle the subject-related and speech-related information in the joint embedding space using adversarial training.

Specifically, we would like the Person-ID feature fpvf^{v}_{p} to be free of Word-ID information. The discriminator could be formed to be a classifier Cpw\text{C}^{w}_{p} to map the collection of Fpv=[fpv(1),⋯ ,fpv(n)]F^{v}_{p}=[{f^{v}_{p}}_{(1)},\cdots,{f^{v}_{p}}_{(n)}] to the NwN_{w} Word-ID classes. The objective function for training the classifier is the same as softmax cross-entropy loss. However, the parameter updating is only performed on Cpw\text{C}^{w}_{p}, where pwj{p_{w}}^{j} is the one-hot label of the identity classes:

Then we update the encoder while fixing the classifier. The way to ensure that the features have lost all information about speech information is that it produces the same prediction for all classes after being sent into Cpw\text{C}^{w}_{p}. One way to form this limitation is to assign the probabilities of each word-label to be 1Nw\frac{1}{N_{w}} in softmax cross-entropy loss. The problem of this loss is that it would still backward gradient for updating parameters even if it reaches the minimum, so we propose to implement the loss using Euclidean distance:

The dual feature fwvf^{v}_{w} should also be free of pidpid information accordingly, so the loss for encoding pidpid information from each fwvf^{v}_{w} using classifier Cwp\text{C}^{p}_{w} and loss for widwid encoder Ewv\text{E}^{v}_{w} to dispel pidpid information can be formed as follows:

NpN_{p} is the number of person identities in the training set for embedding pidpid space. We summarize the adversarial training procedure for classifier Cpw\text{C}^{w}_{p} and encoder Epv\text{E}^{v}_{p} as Fig. 5.

3 Inference: Arbitrary-Subject Talking Face Generation

In this section, we describe how we generate arbitrary-subject talking faces using the disentangled representations learned above. Combining pidpid feature fpvf^{v}_{p} with either of the video widwid feature fwvf^{v}_{w} or audio widwid feature fwaf^{a}_{w}, our system can generate a frame using the decoder G. The newly generated frame can be expressed as G(fpv,fwv)\text{G}({f^{v}_{p}},{f^{v}_{w}}), G(fpv,fwa)\text{G}({f^{v}_{p}},{f^{a}_{w}}).

Here we take synthesizing talking faces from audio widwid information as example. The generation results can be expressed as G(fpv(k),Fwa)={G(fpv(k),fwa(1)),⋯ ,G(fpv(k),fwa(n))}\text{G}({f^{v}_{p}}_{(k)},F^{a}_{w})=\{\text{G}({f^{v}_{p}}_{(k)},{f^{a}_{w}}_{(1)}),\cdots,\text{G}({f^{v}_{p}}_{(k)},{f^{a}_{w}}_{(n)})\}, where fpv(k){f^{v}_{p}}_{(k)} is the pidpid feature of the random kkth frame, which acts as identity guidance. Our overall loss function consists of a L1{L}_{1} reconstruction loss and a temporal GAN loss, where a discriminator Dseq\text{D}_{seq} takes the generated sequence G(fpv(k),Fwa)\text{G}({f^{v}_{p}}_{(k)},F^{a}_{w}) as input. These two terms can be formulated as follows:

The overall reconstruction loss can be written as LRe{\mathcal{L}}_{Re}, α\alpha is a hyper-parameter that leverages the two losses.

The same procedure can be applied to generation from video information by substituting FwaF^{a}_{w} with FwvF^{v}_{w}. As the reconstruction from audio and video can perform at the same time during training, we use LRe{\mathcal{L}}_{Re} to denote the overall reconstruction loss function.

Experiments

Datasets. Our model is trained and evaluated on the LRW dataset (?), which is currently the largest word-level lip reading dataset with 11-of-500500 diverse word labels. For each class, there are more than 800 training samples and 50 validation/test samples. Each sample is a one-second video with the target word spoken. Besides, the identity-preserving module of the network is trained on a subset of the MS-Celeb-1M dataset (?). All the talking faces in the videos are detected and aligned using RSA algorithm (?), and then resized to 256×256256\times 256. For the audio stream, we follow the implementation in (?) to extract the mfccmfcc features at the sampling rate of 100Hz. Then we match each image with a mfccmfcc audio input with the size of 12∗2012*20.

Network Architecture. We adopted a modified VGG-M (?) as the backbone for encoder Epv\text{E}^{v}_{p}, and for encoder Ewv\text{E}^{v}_{w}, we modified a simple version of FAN (?). The encoder Ewa\text{E}^{a}_{w} has a similar structure as that used in (?). Meanwhile, our decoder contains 10 convolution layers with 6 bilinear upsampling layers to obtain a full-resolution output image. All the latent representations are set to be 256-dimensional.

Implementation Details. We implemented DAVS using Pytorch. The batch size is set to be 18 with 1e-4 learning rate and trained on 6 Titan X GPUs. It takes about 4 epochs for the audio-visual speech recognition and person-identity recognition to converge and another 5 epochs for further tuning the generator. The whole training process takes about a week. Due to the alignment of the training set, the directly generated results may suffer from a scale changing problem, so we apply the subspace video stabilization (?) for smoothness.

At test time, the input identity guidance spvs^{v}_{p} to Epv\text{E}^{v}_{p} is any person’s face image and only one of the source for speech information SwvS^{v}_{w}, SwaS^{a}_{w} is needed to generate a sequence of images.

Quantitative Results. To verify the effectiveness of our GAN loss for improving image quality, we evaluate the PSNR and SSIM (?) score on the test set of LRW based on reconstruction. We compare the results with and without the GAN loss in Table 1. We can see that both the scores are improved by changing LL1\mathcal{L}_{L_{1}} to LRe\mathcal{L}_{Re}.

Qualitative Results. Video results are shown in supplementary materials. Here we show image results in Fig 6. The input guidance photos are celebrities chosen randomly from the Internet. Our model is capable of generating talking faces based on both audios or videos. The focus of our work is to improve audio guided generation results by using joint audio-visual embedding, so we compare our work with (?) at Fig 7. It can be clearly seen that our results outperform theirs from both the perspective of identity preserving and image quality.

User Study. We also conduct user study to investigate the visual quality of our generated results comparing with a fair reproduction of (?) with our network structure. They are evaluated w.r.t two different criteria: whether participants could regard the generated talking faces as realistic (true or false), and how much percent of the time steps the generated talking faces temporally sync with the corresponding audio. We generate videos with the identity guidance to be 10 different celebrity photos. As for speech content information, we use clips from the test set of LRW dataset and selections from the Voxceleb dataset (?), which is not used for training. There are overall 1010 participants involved, and the results are average over persons and video time steps. The ground-truth is not included in the user study. Different subjects may behave different lip motion given the same audio clip and it is not desirable for the ground-truth to interfere with the participants’ perception. When conducting the user study for lip sync evaluation, we asked the participants to only focus on whether the lip motion and given audio are temporally synchronized. Their ratings indicate that our generation results outperform the baseline by synchronizing rate and the extent of realistic, according to Table 2.

2 Effectiveness of Audio-Visual Representation

In order to inspect the quality of our embedded audio-visual representation, we evaluate the discriminative power and the closeness of our co-embedded features.

Word-level Audio-Visual Speech Recognition. We report audio-visual speech recognition accuracy on the test set of LRW dataset. Containing the task of visual recognition (lip reading) and audio recognition (speech recognition).

Our model structure for lip reading is similar to the Multiple-Towers method which reaches the highest lip reading results in (?), so we consider it as a baseline. The difference is that the concatenation of features is performed at the spacial size of 1×11\times 1 in our setting. This would not be a reasonable choice for this task alone for the spatial information in images would be lost across time. However, as shown in Table 3, our results adding the contrastive loss alone outperforms the baseline. With the help of sharing classifier and domain adversarial training, the results improve a large margin.

Audio-Video Retrieval. To evaluate the closeness between the audio and face features, we borrow protocols used in the retrieval community. The retrieval experiments are conducted on the test set of LRW with 25000 samples, which means that given a test target video (audio), we try to find the closest audio (video) based on the distance of widwid features FwvF^{v}_{w}, FwaF^{a}_{w} among all the test samples. Here we report the R@1R@1, R@10R@10 and MedMed RR measurements which is the same as (?). As we can see in Table 3, with all supervisions, the highest results can be achieved.

Qualitative Results. Figure 8 shows the sequence generation quality from audio with different supervisions provided above. We can observe from the figure that given the same clip of audio, the duration of the mouth opening and to what extent it is opened is affected by different supervisions. Sharing the classifier apparently lengthens the time and strength of the mouth opening to make the image closer to the ground truth. Combining with the adversarial training makes the image quality improves. Note that it is not a one-to-one mapping between audio and lip motion; different subjects may behave different lip motion given the same audio clip so the final results may not perform the same as the ground truth.

3 Identity-Speech Disentanglement

To validate our adversarial training is able to disentangle speech information from person-ID branch, we use person-ID encoder on every frame of a video and concatenate them to get Fpv={fp(1)v,...,fp(nv)}F^{v}_{p}=\{f^{v}_{p(1)},...,f^{v}_{p(n})\}. Then we train an SVM to map training samples to their widwid labels and test the results, which implies that we attempt to find the widwid information left in the pidpid encoder. The whole procedure is repeated before and after the feature disentanglement. Before the disentanglement, 27.8% of the test set can be assigned to the right class, but only 9.7% left after, indicating that considerable speech content information within the encoder Epv\text{E}^{v}_{p} is gone.

We then highlight the merits of adversarial disentanglement from two aspects, identity preserving and lip sync quality. For identity preserving, we use OpenFace’s squared L2 similarity score as an indicator and compare the identity distance between the generated faces and the original ones (lower indicates more similar). For lip sync quality, we detect 20 landmarks using dlib library (?) around the lips to characterize its deviation from ground truth, measured by the averaged L2-norm (lower is better). Then we conduct retrieval experiments between all generated results and source videos based on extracted FwidvF^{v}_{wid} features. Experiments are also conducted on a direct replication of every video clip, to prove that the retrieval results are affected by lip motion rather than appearance features. From Table 4, we can observe that adversarial disentanglement indeed helps improves lip sync quality.

Conclusion

In this paper, we propose a novel framework called Disentangled Audio-Visual System (DAVS), which generates high quality talking face videos using disentangled audio-visual representation. Specifically, we first learn a joint audio-visual embedding space widwid with discriminative speech information by leveraging the word-ID labels. Then we disentangled the widwid space from the person-ID pidpid space through adversarial learning. Compared to prior works, DAVS has several appealing properties: (1) A joint audio-visual representation is learned through audio-visual speech discrimination by associating several supervisions. The disentangled audio-visual representation significantly improves lip reading performance; (2) Audio-visual speech recognition and audio-visual synchronizing are unified in an end-to-end framework; (3) Most importantly, arbitrary-subject talking face generation with high-quality and temporal accuracy can be achieved by our framework; both audio and video speech information can be employed as input guidance.

Acknowledgements

We thank Yu Xiong for helpful discussions and his assistance with our video. This work is supported by SenseTime Group Limited, the General Research Fund sponsored by the Research Grants Council of Hong Kong and the Hong Kong Innovation and Technology Support Program (No.ITS/121/15FX).

References