One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

Ting-Chun Wang, Arun Mallya, Ming-Yu Liu

Introduction

We study the task of generating a realistic talking-head video of a person using one source image of that person and a driving video, possibly derived from another person. The source image encodes the target person’s appearance, and the driving video dictates motions in the output video.

We propose a pure neural rendering approach, where we render a talking-head video using a deep network in the one-shot setting without using a graphics model of the 3D human head. Compared to 3D graphics-based models, 2D-based methods enjoy several advantages. First, it avoids 3D model acquisition, which is often laborious and expensive. Second, 2D-based methods can better handle the synthesis of hair, beard, \etc, while acquiring detailed 3D geometries of these regions is challenging. Finally, they can directly synthesize accessories present in the source image, including eyeglasses, hats, and scarves, without their 3D models.

However, existing 2D-based one-shot talking-head methods come with their own set of limitations. Due to the absence of 3D graphics models, they can only synthesize the talking-head from the original viewpoint. They cannot render the talking-head from a novel view.

Our approach addresses the fixed viewpoint limitation and achieves local free-view synthesis. One can freely change the viewpoint of the talking-head within a large neighborhood of the original viewpoint, as shown in Fig. 1(c). Our model achieves this capability by representing a video using a novel 3D keypoint representation, where person-specific and motion-related information is decomposed. Both the keypoints and their decomposition are learned unsupervisedly. Using the decomposition, we can apply 3D transformations to the person-specific representation to simulate head pose changes such as rotating the talking-head in the output video. Figure 2 gives an overview of our approach.

We conduct extensive experimental validation with comparisons to state-of-the-art methods. We evaluate our method on several talking-head synthesis tasks, including video reconstruction, motion transfer, and face redirection. We also show how our approach can be used to reduce the bandwidth of video conferencing, which has become an important platform for social networking and remote collaborations. By sending only the keypoint representation and reconstructing the source video on the receiver side, we can achieve a 10x bandwidth reduction as compared to the commercial H.264 standard without compromising the visual quality.

Contribution 1. A novel one-shot neural talking-head synthesis approach, which achieves better visual quality than state-of-the-art methods on the benchmark datasets.

Contribution 2. Local free-view control of the output video, without the need for a 3D graphics model. Our model allows changing the viewpoint of the talking-head during synthesis.

Contribution 3. Reduction in bandwidth for video streaming. We compare our approach to the commercial H.264 standard on a benchmark talking-head dataset and show that our approach can achieve 10×\times bandwidth reduction.

Related Works

GANs. Since its introduction by Goodfellow et al. , GANs have shown promising results in various areas , such as unconditional image synthesis , image translation , text-to-image translation , image processing , and video synthesis . We focus on using GANs to synthesize talking-head videos in this work.

3D model-based talking-head synthesis. Works on transferring the facial motion of one person to another—face reenactment—can be divided into subject-dependent and subject-agnostic models. Traditional 3D-based methods usually build a subject-dependent model, which can only synthesize one subject. Moreover, they focus on transferring the expressions without the head movement . This line of works starts by collecting footage of the target person to be synthesized using an RGB or RGBD sensor . Then a 3D model of the target person is built for the face region . At test time, the new expressions are used to drive the 3D model to generate the desired motions.

More recent 3D model-based methods are able to perform subject-agnostic face synthesis . While they can do an excellent job synthesizing the inner face region, they have a hard time generating realistic hair, teeth, accessories, etc. Due to the limitations, most modern face reenactment frameworks adopt the 2D approach. Another line of works focuses on controllable face generation, providing explicit control over the generated face from a pretrained StyleGAN . However, it is not clear how they can be adapted to modifying real images since the inverse mapping from images to latent codes is nontrivial.

2D-based talking-head synthesis. Again, 2D approaches can be classified into subject-dependent and subject-agnostic models. Subject-dependent models can only work on specific persons since the model is only trained on the target person. On the other hand, subject-agnostic models only need a single image of the target person, who is not seen during training, to synthesize arbitrary motions. Siarohin et al. warp extracted features from the input image, using motion fields estimated from sparse keypoints. On the other hand, Zakharov et al. demonstrate that it is possible to achieve promising results using direct synthesis methods without any warping. Few-shot vid2vid injects the information into their generator by dynamically determining the the parameters in the SPADE modules. Zakharov et al. decompose the low and high frequency components of the image and greatly accelerate the inference speed of the network. While demonstrating excellent result qualities, these methods can only synthesize fixed viewpoint videos, which produce less immersive experiences.

Video compression. A number of recent works propose using a deep network to compress arbitrary videos. The general idea is to treat the problem of video compression as one of interpolating between two neighboring keyframes. Through the use of deep networks to replace various parts of the traditional pipeline, as well as techniques such as hierarchical interpolation and joint encoding of residuals and optical flows, these prior works reduce the required bit-rate. Other works focus on restoring the quality of low bit-rate videos using deep networks. Most related to our work is DAVD-Net , which restores talking-head videos using information from the audio stream. Our proposed method is different from these works in a number of aspects, in both the goal as well as the method used to achieve compression. We specifically focus on videos of talking faces. People’s faces have an inherent structure—from the shape to the relative arrangement of different parts such as eyes, nose, mouth, \etc. This allows us to use keypoints and associated metadata for efficient compression, an order of magnitude better than traditional codecs. Our method does not guarantee pixel-aligned output videos; however, it faithfully models facial movements and emotions. It is also better suited for video streaming as it does not use bi-directional or B-frames.

Method

Let ss be an image of a person, referred to as the source image. Let {d1,d2,...,dN}\{d_{1},d_{2},...,d_{N}\} be a talking-head video, called the driving video, where did_{i}’s are the individual frames, and NN is the total number of frames. Our goal is to generate an output video {y1,y2,...,yN}\{y_{1},y_{2},...,y_{N}\}, where the identity in yiy_{i}’s is inherited from ss and the motions are derived from did_{i}’s. Several talking-head synthesis tasks fall in the above setup. When ss is a frame of the driving video (\eg, the first frame: s≡d1s\equiv d_{1}.), we have a video reconstruction task. When ss is not from the driving video, we have a motion transfer task.

We propose a pure neural synthesis approach that does not use any 3D graphics models, such as the well-known 3D morphable model (3DMM) . Our approach contains three major steps: 1) source image feature extraction, 2) driving video feature extraction, and 3) video generation. In Fig. 3, we illustrate 1) and 2), while Fig. 5 shows 3). Our key ingredient is an unsupervised approach for learning a set of 3D keypoints and their decomposition. We decompose the keypoints into two parts, one that models the facial expressions and the other that models the geometric signature of a person. These two parts are combined with the target head pose to generate the image-specific keypoints. After the keypoints are estimated, they are then used to learn a mapping function between two images. We implement these steps using a set of networks and train them jointly. In the following, we discuss the three steps in detail.

Synthesizing a talking-head requires knowing the appearance of the person, such as the skin and eye colors. As shown in Fig. 3(a), we first apply a 3D appearance feature extraction network FF to map the source image ss to a 3D appearance feature volume fsf_{s}. Unlike a 2D feature map, fsf_{s} has three spatial dimensions: width, height, and depth. Mapping to a 3D feature volume is a crucial step in our approach. It allows us to operate the keypoints in the 3D space for rotating and translating the talking-head during synthesis.

The final keypoints are image-specific and contain person-signature, pose, and expression information. Figure 5 visualizes the keypoint computation pipeline.

The 3D keypoint decomposition in (1) is of paramount importance to our approach. It commits to a prior decomposition of keypoints: geometry-signatures, head poses, and expressions. It helps learn manipulable representations and differs our approach from prior 2D keypoint-based neural talking-head synthesis approaches . Also note that unlike FOMM , our model does not estimate Jacobians. The Jacobian represents how a local patch around the keypoint can be transformed into the corresponding patch in another image via an affine transformation. Instead of explicitly estimating them, our model assumes the head is mostly rigid and the local patch transformation can be directly derived from the head rotation via Js=RsJ_{s}=R_{s}. Avoiding estimating Jacobians allows us to further reduce the transmission bandwidth for the video conferencing application, as detailed in Sec. 5.

2 Driving video feature extraction

We use dd to denote a frame in {d1,d2,...,dN}\{d_{1},d_{2},...,d_{N}\} as individual frames are processed in the same way. To extract motion-related information, we apply the head pose estimator HH to get RdR_{d} and tdt_{d} and apply the expression deformation estimator Δ\Delta to obtain δd,k\delta_{d,k}’s, as shown in Fig. 3(b).

Now, instead of extracting canonical 3D keypoints from the driving image dd using LL, we reuse xc,kx_{c,k}, which were extracted from the source image ss. This is because the face in the output image must have the same identity as the one in the source image ss. There is no need to compute them again. Finally, the identity-specific information and the motion-related information are combined to compute the driving keypoints for the driving image dd in the same way we obtained source keypoints:

We apply this processing to each frame in the driving video, and each frame can be compactly represented by RdR_{d}, tdt_{d}, and δd,k\delta_{d,k}’s. This compact representation is very useful for low-bandwidth video conferencing. In Sec. 5, we will introduce an entropy coding scheme to further compress these quantities to reduce the bandwidth utilization.

Our approach allows manual changes to the 3D head pose during synthesis. Let RuR_{u} and tut_{u} be user-specified rotation and translation, respectively. The final head pose in the output image is given by Rd←RuRdR_{d}\leftarrow R_{u}R_{d} and td←tu+tdt_{d}\leftarrow t_{u}+t_{d}. In video conferencing, we can change a person’s head pose in the video stream freely despite the original view angle.

3 Video generation

As shown in Fig. 5, we synthesize an output image by warping the source feature volume and then feeding the result to the image generator GG to produce the output image yy. The warping approximates the nonlinear transformation from ss to dd. It re-positions the source features for the synthesis task.

To obtain the required warping function ww, we take a bottom-up approach. We first compute the warping flow wkw_{k} induced by the kk-th keypoint using the first order approximation , which is reliable only around the neighborhood of the keypoint. After obtaining all KK warping flows, we apply each of them to warp the source feature volume. The KK warped features are aggregated to estimate a flow composition mask mm using the motion field estimation network MM. This mask indicates which of the KK flows to use at each spatial 3D location. We use this mask to combine the KK flows to produce the final flow ww. Details of the operation are given in Appendix A.1.

4 Training

We train our model using a dataset of talking-head videos where each video contains a single person. For each video, we sample two frames: one as the source image ss and the other as the driving image dd. We train the networks FF, Δ\Delta, HH, LL, MM, and GG by minimizing the following loss:

In short, the first two terms ensure the output image looks similar to the ground truth. The next two terms enforce the predicted keypoints to be consistent and satisfy some prior knowledge about the keypoints. The last two terms constrain the estimated head pose and keypoint perturbations. We briefly discuss these losses below and leave the implementation details in Appendix A.2.

Perceptual loss LP\mathcal{L}_{P}. We minimize the perceptual loss between the output and the driving image, which is helpful in producing sharp-looking outputs.

GAN loss LG\mathcal{L}_{G}. We use a multi-resolution patch GAN where the discriminator predicts at the patch-level. We also minimize the discriminator feature matching loss .

Equivariance loss LE\mathcal{L}_{E}. This loss ensures the consistency of image-specific keypoints xd,k{x}_{d,k}. For a valid keypoint, when applying a 2D transformation to the image, the predicted keypoints should change according to the applied transformation . Since we predict 3D instead of 2D keypoints, We use an orthographic projection to project the keypoints to the image plane before computing the loss.

Keypoint prior loss LL\mathcal{L}_{L}. We use a keypoint coverage loss to encourage the estimated image-specific keypoints xd,k{x}_{d,k}’s to spread out across the face region, instead of crowding around a small neighborhood. We compute the distance between each pair of the keypoints and penalize the model if the distance falls below a preset threshold. We also use a keypoint depth prior loss that encourages the mean depth of the keypoints to be around a preset value.

Head pose loss LH\mathcal{L}_{H}. We penalize the prediction error of the head rotation RdR_{d} compared to the ground truth Rˉd\bar{R}_{d}. Since acquiring the ground truth head pose for a large-scale video dataset is expensive, we use a pre-trained pose estimation network to approximate Rˉd\bar{R}_{d}.

Deformation prior loss LΔ\mathcal{L}_{\Delta}. The loss penalizes the magnitude of the deformations δd,k\delta_{d,k}’s. As the deformations model the deviation from the canonical keypoints due to expression changes, their magnitudes should be small.

Experiments

Implementation details. The network architecture and training hyper-parameters are available in Appendix A.3.

Datasets. Our evaluation is based on VoxCeleb2 and TalkingHead-1KH, a newly collected large-scale talking-head video dataset. It contains 180180K videos, which are often with higher quality and larger resolution than those in VoxCeleb2. Details are available in Appendix B.1.

Baselines. We compare our neural talking-head model with three state-of-the-art methods: FOMM , few-shot vid2vid (fs-vid2vid) , and bi-layer neural avatars (bi-layer) . We use the released pre-trained model on VoxCeleb2 for bi-layer , and retrain from scratch for others on the corresponding datasets. Since bi-layer does not predict the background, we subtract the background when doing quantitative analyses.

Metrics. We evaluate a synthesis model on 1) reconstruction faithfulness using L1L_{1}, PSNR, SSIM/MS-SSIM, 2) output visual quality using FID, and 3) semantic consistency using average keypoint distance (AKD). Please consult Appendix B.2 for details of the performance metrics.

Same-identity reconstruction. We first compare the face synthesis results where the source and driving images are of the same person. The quantitative evaluation is shown in Table 1. It can be seen that our method outperforms other competing methods on all metrics for both datasets. To verify that our superior performance does not come from more parameters, we train another large FOMM model with doubled filter size (FOMM-L), which is larger than our model. We can see that enlarging the model actually hurts the performance, proving that simply making the model larger does not help. Figures 7 and 7 show the qualitative comparisons. Our method can more faithfully reproduce the driving motions.

Cross-identity motion transfer. Next, we compare results where the source and driving images are from different persons (cross-identity). Table 3 shows that our method achieves the best results compared to other methods. Figure 9 compares results from different approaches. It can be seen that our method generates more realistic images while still preserving the original identity. For cross-identity motion transfer, it is sometimes useful to use relative motion , where only motion differences between two neighboring frames in the driving video are transferred. We report comparisons using relative motion in Appendix B.3.

Ablation study. We benchmark the performance gains from the proposed keypoint decomposition scheme, the mask estimation network, and pose supervision in Appendix B.4.

Failure cases. Our model fails when large occlusions and image degradation occur, as visualized in Appendix B.5.

Face recognition. Since the canonical keypoints are independent of poses and expressions, they can also be applied to face recognition. In Appendix B.6, we show that this achieves 5x accuracy than using facial landmarks.

2 Face redirection.

Baselines. We benchmark our talking-head model’s face redirection capability using latest face frontalization methods: pixel2style2pixel (pSp) and Rotate-and-Render (RaR) . pSp projects the original image into a latent code and then uses a pre-trained StyleGAN to synthesize the frontalized image. RaR adopts a 3D face model to rotate the input image and re-renders it in a different pose.

Metrics. The results are evaluated by two metrics: identity preservation and head pose angles. We use a pre-trained face recognition network to extract high-level features, and compute the distance between the rotated face and the original one. We use a pre-trained head pose estimator to obtain head angles of the rotated face. For a rotated image, if its identity distance to the original image is within some threshold, and/or its head angle is within some tolerance to the desired angle, we consider it as a “good” image.

We report the ratio of “good” images using our metric for each method in Table 3. Example comparisons can be found in Fig. 9. It can be seen that while pSp can always frontalize the face, the identity is usually lost. RaR generates more visually appealing results since it adopts 3D face models, but has problems outside the inner face regions. Besides, both methods have issues regarding the temporal stability. Only our method can realistically frontalize the inputs.

Neural Talking-Head Video Conferencing

Our talking-head synthesis model distills motions in a driving image using a compact representation, as discussed in Sec. 3. Due to this advantage, our model can help reduce the bandwidth consumed by video conferencing applications. We can view the process of video conferencing as the receiver watching an animated version of the sender’s face.

Figure 11 shows a video conferencing system built using our neural talking-head model. For a driving image dd, we use the driving image encoder, consisting of Δ\Delta and HH, to extract the expression deformations δd,k\delta_{d,k} and the head pose Rd,tdR_{d},t_{d}. By representing a rotation matrix using Euler angles, we have a compact representation of dd using 3K+63K+6 numbers: 3 for the rotation, 3 for the translation, and 3K for the deformations. We further compress these values using an entropy encoder . Details are in Appendix C.1.

The receiver receives the entropy-encoded representation and uses the entropy decoder to recover δd,k\delta_{d,k} and Rd,tdR_{d},t_{d}. They are then fed into our talking-head synthesis framework to reconstruct the original image dd. We assume that the source image ss is sent to the receiver at the beginning of the video conferencing session or re-used from a previous session. Hence, it does not consume additional bandwidth. We note that the source image is different from the I-frame in traditional video codecs. While I-frames are sent frequently in video conferencing, our source image only needs to be sent once at the beginning. Moreover, the source image can be an image of the same person captured on a different day, a different person, or even a face portrait painting.

Adaptive number of keypoints. Our basic model uses a fixed number of keypoints during training and inference. However, since the transmitted bits are proportional to the number of keypoints, it is advantageous to change this number to accommodate varying bandwidth requirements dynamically. Using keypoint dropouts at training time, we derive a model that can dynamically use a smaller number of keypoints for reconstruction. This allows us to transmit even fewer bits without compromising visual quality. On average, with the adaptive scheme, the number of sent keypoints is reduced from K=20K=20 to 11.5211.52 (Appendix C.2).

Benchmark dataset. We manually select a dataset of 222222 high-quality talking-head videos. Each video’s resolution is 512×512512\times 512 and the length is up to 10241024 frames for evaluation.

Baselines. We compare our video streaming method with the popular H.264 codec. In order to conform to real-time video streaming cases, we disable the use of bidirectional B-frames, as this uses information from the future. By varying the constant rate factor (CRF) while encoding the ground truth input videos, we can obtain a set of videos of varying qualities and sizes suitable for a range of bandwidth availability. We also compare with FOMM and fs-vid2vid , which also use keypoints or facial landmarks. For a fair comparison, we also compress their keypoints and Jacobians using our entropy coding scheme.

Metrics. We compare the compression effectiveness using the average number of bits required per pixel (bpp) for each output frame. We measure the compression quality using both automatic and human evaluations. Unlike traditional compression methods, our method does not reproduce the input image in a pixel-aligned manner but can faithfully reproduce facial motions and gestures. Metrics based on exact pixel alignments are ill-suited for measuring the quality of our output videos. We hence use the LPIPS perceptual similarity metric for measuring compression quality .

As shown on the left side of Fig. 16, compared to the other neural talking-head synthesis methods, ours obtains better quality while requiring much lower bandwidth. This is because other methods send the full keypoints and Jacobians , while ours only sends the head pose and keypoint deformations. Compared with H.264 videos of the same quality, ours requires significantly lower bandwidth. For human evaluation, we show MTurk workers two videos side-by-side, one produced by H.264 and the other produced by our method’s adaptive version. We then ask the workers to choose the video that they feel is of better quality. The preference scores are visualized on the right side of Fig. 16. Based on these scores, our compression method is comparable to the H.264 codec at a CRF value of 3636, which means our adaptive and 2020 keypoint scheme obtains 10.37×10.37\times and 6.5×6.5\times reduction in bandwidth compared to the H.264 codec, respectively. To handle challenging corner cases for our video conferencing system and out-of-distribution videos, we further develop a binary latent encoding network that can efficiently encode the residual at the expense of additional bandwidth, the details of which are in Appendix C.3.

Conclusion

In this work, we present a novel framework for neural talking-head video synthesis and compression. We show that by using our unsupervised 3D keypoints, we are able to decompose the representation into person-specific canonical keypoints and motion-related transformations. This decomposition has several benefits: By modifying the keypoint transformation only, we are able to generate free-view videos. By transmitting just the keypoint transformations, we can achieve much better compression ratios than existing methods. These features provide users a great tool for streaming live videos. By dramatically reducing the bandwidth and ensuring a more immersive experience, we believe this is an important step towards the future of video conferencing.

Acknowledgements. We thank Jan Kautz for his valuable comments throughout the development of the work. We thank Timo Aila, Koki Nagano, Sameh Khamis, Jaewoo Seo, and Xihui Liu for providing very useful feedback to shape our draft. We thank Henry Lin, Rochelle Pereira, Santanu Dutta, Simon Yuan, Brad Nemire, Margaret Albrecht, Siddharth Sharma, and Eric Ladenburg for their helpful suggestions in presenting our visualization results.

References

Appendix A Additional Network and Training Details

Here, we present the architecture of our neural talking-head model. We also discuss the training details.

The implementation details of the networks in our model are shown in Fig. 12 and described below.

Appearance feature extractor FF. The network extracts 3D appearance features from the source image. It consists of a number of downsampling blocks, followed by a convolution layer that projects the input 2D features to 3D features. We then apply a number of 3D residual blocks to compute the final 3D features fsf_{s}.

Canonical keypoint detector LL. The network takes the source image and applies a U-Net style encoder-decoder to extract canonical keypoints. Since we need to extract 3D keypoints, we project the encoded features to 3D through a 1×11\times 1 convolution. The output of the 1×11\times 1 convolution is the bottleneck of the U-Net. The decoder part of the U-Net consists of 3D convolution and upsampling layers.

Head pose estimator HH and expression deformation estimator Δ\Delta. We adopt the same architecture as in Ruiz et al. . It consists of a series of ResNet bottleneck blocks, followed by a global pooling to remove the spatial dimension. Different linear layers are then used to estimate the rotation angles, the translation vector, and the expression deformations. The full angle range is divided into 6666 bins for rotation angles, and the network predicts which bin the target angle is in. The estimated head pose and deformations are used to transform the canonical keypoints to obtain the source or driving keypoints.

Motion field estimator MM. After the keypoints are predicted, they are used to estimate warping flow maps. We generate a warping flow map wkw_{k} based on the kk-th keypoint using the first-order approximation . Let pdp_{d} be a 3D coordinate in the feature volume of the driving image dd. The kk-th flow field maps pdp_{d} to a 3D coordinate in the 3D feature volume of the source image ss, denoted by psp_{s}, by:

This builds a correspondence between the source and driving.

Using the flow field wkw_{k} obtained from the kk-th keypoint pair, we can warp the source feature fsf_{s} to construct a candidate warped volume, wk(fs)w_{k}(f_{s}). After we obtain the warped source features wk(fs)w_{k}(f_{s}) using all KK flows, they are concatenated together and fed to a 3D U-Net to extract features. Then a softmax function is employed to obtain the flow composition mask mm, which consists of KK 3D masks, {m1,m2,...,mK}\{m_{1},m_{2},...,m_{K}\}. These maps satisfy the constraints that ∑kmk(pd)=1\sum_{k}m_{k}(p_{d})=1 and 0≤mk(pd)≤10\leq m_{k}(p_{d})\leq 1 for all pdp_{d}. These KK masks are then linearly combined with the KK warping flow maps, wkw_{k}’s, to construct the final warping map ww by ∑k=1Kmk(pd)wk(pd)\sum_{k=1}^{K}m_{k}(p_{d})w_{k}(p_{d}). To handle occlusions caused by the warping, we also predict a 2D occlusion mask oo, which will be inputted to the generator GG.

Generator GG. The generator takes the warped 3D appearance features w(fs)w(f_{s}) and projects them back to 2D. Then, the features are multiplied with the occlusion mask oo obtained from the motion field estimator MM. Finally, we apply a series of 2D residual blocks and upsamplings layers to obtain the final image.

A.2 Losses

We present details of the loss terms in the following.

Perceptual loss LP\mathcal{L}_{P}. We use the multi-scale implementation introduced by Siarohin et al. . In particular, a pre-trained VGG network is used to extract features from both the ground truth and the output image, and the L1L_{1} distance between the features is computed. Then both images are downsampled, and the same VGG network is used to extract features and compute the L1L_{1} distance again. This process is repeated 33 times to compute losses at multiple image resolutions. We use layers relu_1_1, relu_2_1, relu_3_1, relu_4_1, relu_5_1 of the VGG19 network with weights 0.03125,0.0625,0.125,0.25,1.00.03125,0.0625,0.125,0.25,1.0, respectively. Moreover, since we are synthesizing face images, we also compute a single-scale perceptual loss using a pre-trained face VGG network . These losses are then summed together to give the final perceptual loss.

GAN loss LG\mathcal{L}_{G}. We adopt the same patch GAN implementation as in , and use the hinge loss. Feature matching loss is also adopted to stabilize training. We use single-scale discriminators for training 256×256256\times 256 images, and two-scale discriminators for 512×512512\times 512 images.

Equivariance loss LE\mathcal{L}_{E}. This loss ensures the consistency of estimated keypoints . In particular, let the original image be dd and its detected keypoints be xdx_{d}. When a known spatial transformation T\mathbf{T} is applied on image dd, the detected keypoints xT(d)x_{\mathbf{T}(d)} on this transformed image T(d)\mathbf{T}(d) should be transformed in the same way. Based on this observation, we minimize the L1L_{1} distance ∥xd−T−1(xT(d))∥1\|x_{d}-\mathbf{T}^{-1}(x_{\mathbf{T}(d)})\|_{1}. Affine transformations and randomly sampled thin plate splines are used to perform the transformation. Since all these are 2D transformations, we project our 3D keypoints to 2D by simply dropping the zz values before computing the losses.

Keypoint prior loss LL\mathcal{L}_{L}. As described in the main paper, we penalize the keypoints if the distance between any pair of them is below some threshold DtD_{t}, or if the mean depth value deviates from a preset target value ztz_{t}. In other words,

where Z(⋅)Z(\cdot) extracts the mean depth value of the keypoints. This ensures the keypoints are more spread out and used more effectively. We set DtD_{t} to 0.10.1 and ztz_{t} to 0.330.33 in our experiments.

Head pose loss LH\mathcal{L}_{H}. We compute the L1L_{1} distance between the estimated head pose RdR_{d} and the one predicted by a pre-trained pose estimator Rˉd\bar{R}_{d}, which we treat as ground truth. In other words, LH=∥Rd−Rˉd∥1\mathcal{L}_{H}=\|R_{d}-\bar{R}_{d}\|_{1}, where the distance is computed as the sum of differences of the Euler angles.

Deformation prior loss LΔ\mathcal{L}_{\Delta}. Since the expression deformation Δ\Delta is the deviation from the canonical keypoints, their magnitude should not be too large. To ensure this, we put a loss on their L1\mathcal{L}_{1} norm: LΔ=∥δd,k∥1\mathcal{L}_{\Delta}=\|\delta_{d,k}\|_{1}.

where λ\lambda’s are the weights and are set to 10,1,20,10,20,510,1,20,10,20,5 respectively in our implementation.

A.3 Optimization

We adopt the ADAM optimizer with β1=0.5\beta_{1}=0.5 and β2=0.999\beta_{2}=0.999. The learning rate is set to 0.00020.0002. We apply Spectral Norm to all the layers in both the generator and the discriminator. We use synchronized BatchNorm for the generator. Training is conducted on an NVIDIA DGX1 with 8 32GB V100 GPUs.

We adopt a coarse-to-fine approach for training. We first train our model on 256×256256\times 256 images for 100100 epochs. We then finetune on 512×512512\times 512 images for another 1010 epochs.

Appendix B Additional Experiment Details

We use the following datasets in our evaluations.

VoxCeleb2 . The dataset contains about 1M talking-head videos of different celebrities. We follow the training and test split proposed in the original paper. where we use 280K videos with high bit-rates to train our model. We report our results on the validation set, which contains about 36k videos.

TalkingHead-1KH. We compose a dataset containing about 10001000 hours of videos from various sources. A large portion of them is from the YouTube website with the creative common license. We also use videos from the Ryerson audio-visual dataset as well as a set of videos that we recorded with the permission from the subject ourselves. We only use videos whose resolution and bit-rate are both high. We call this dataset TalkingHead-1KH. The videos in the TalkingHead-1KH are in general with higher resolutions and better image quality than those in the VoxCeleb2.

B.2 Metrics

We use a set of metrics to evaluate a talking-head synthesis method. We use L1L_{1}, PSNR, SSIM, and MS-SSIM for quantifying the faithfulness of the recreated videos. We use FID to measure how close is the distribution of the recreated videos to that of the original videos. We use AKD to measure how close the facial landmarks extracted by an off-the-shelf landmark detector from the recreated video are to those in the original video. In the following, we discuss the implementation details of these metrics.

L1L_{1}. We compute the average L1L_{1} distance between generated and real images.

PSNR measures the image reconstruction quality by computing the mean squared error (MSE) between the ground truth and the reconstructed image.

SSIM/MS-SSIM. SSIM measures the structural similarity between patches of the input images. Therefore, it is more robust to global illumination changes than PSNR, which is based absolute errors. MS-SSIM is a multi-scale variant of SSIM that works on multiple scales of the images and has been shown to correlate better with human perception.

FID measures the distance between the distributions of synthesized and real images. We use the pre-trained InceptionV3 network to extract features from both sets of images and estimate the distance between them.

Average keypoint distance (AKD). We use a facial landmark detector to detect landmarks of real and synthesized images and then compute the average distance between the corresponding landmarks in these two images.

B.3 Relative motion transfer

For cross-identity motion transfer results in our experiment section, we transfer absolute motions in the driving video. For completeness, we also report quantitative comparisons using relative motion proposed in in Table 4. As can be seen, our method still performs the best.

B.4 Ablation study

We perform the following ablation studies to verify the effectiveness of our several important design choices.

Two-step vs. direct keypoint prediction. We estimate the keypoints in an image by first predicting the canonical keypoints and then applying the transformation and the deformations. To compare this approach with direct keypoint location prediction, we train another network that directly predicts the final source and driving keypoints in the image. In particular, the keypoint detector LL directly predicts the final source and driving keypoints in the image instead of the canonical ones, and there is no pose estimator HH and deformation estimator Δ\Delta. Since there is no pose estimator, we do not need any pose supervision (i.e., head pose loss) for this direct prediction network. Note that while this model is only slightly inferior to our final (two-step) model quantitatively, it has no pose control for the output video since the head pose is no longer estimated, so a major feature of our method would be lost.

3D vs. 2D warping. We generate 3D flow fields from the estimated keypoints to warp 3D features. Another option is to project the keypoints to 2D, estimate a 2D flow field, and extract 2D features from the source image. The estimated 2D flow filed is then used to warp 2D image features.

Number of keypoints. We show that our approach’s output quality is positively correlated with the number of keypoints.

As can be seen in Table 5, our model works better than all the other alternatives on all of the performance metrics.

B.5 Failure cases

While our model is in general robust to different situations, it cannot handle large occlusions well. For example, when the face is occluded by the person’s hands or other objects, the synthesis quality will degrade, as shown in Fig. 14

B.6 Canonical keypoints for face recognition

Our canonical keypoints are formulated to be independent of the pose and expression change. They should only contain a person’s geometry signature, such as the shapes of face, nose, and eyes. To verify this, we conduct an experiment using the canonical keypoints for face recognition.

We extract canonical keypoints from 384384 identities in the VoxCeleb2 dataset to form a training set. For each identity, we also pick a different video of the same identity to form the test set. The training and test videos of the same subject have different head poses and expressions. A face recognition algorithm would fail if it could not filter out pose and expression information. To prove our canonical keypoints are independent to poses and expressions, we apply a simple nearest neighbor classifier using our canonical keypoints for the face recognition task.

Overall, our canonical keypoints reaches an accuracy of 0.0700.070, while a random guess has an accuracy of 0.00260.0026 (Ours is 27×27\times better than the random guess.). On the other hand, a classifier using the off-the-shelf dlib landmark detector only achieves an accuracy of 0.0130.013, which means our keypoints are 5×5\times more effective for face recognition.

Appendix C Additional Video Conferencing Details

We represent each rotation angle, translation, and deformation value as an fp16 floating-point number. Each number consumes two bytes. Naively transmitting the 3K+63K+6 floating numbers will result in transmitting 6K+126K+12 bytes. We adopt arithmetic coding to encode the 3K+63K+6 numbers. Arithmetic coding is one kind of entropy coding. It assigns different codeword lengths to different symbols based on their frequencies. The symbol that appears more often will have a shorter code.

We first apply the driving image encoder to a validation set of 127127 videos that are not included in the test set. Each frame will give us 6K+126K+12 bytes. We treat each of the bytes separately and build a frequency table for each byte. This gives us 6K+126K+12 frequency tables. When encoding the test set, we encode each byte using the associated frequency table learned from the validation set. This results in a varying-length representation that is much smaller than 6K+126K+12 bytes on average.

Table 6 shows the sizes of the per-frame metadata in bytes that needs to be transmitted for various talking-head methods before and after performing the arithmetic compression for an image size of 512×\times512. Our adaptive scheme requires a per-frame metadata size of 53.0353.03 B, which corresponds to (53.03×8/5122)=0.001618(53.03\times 8/512^{2})=0.001618 bits per pixel.

C.2 Adaptive number of keypoints

Our basic model uses a fixed number of keypoints during training and inference. However, on a video call, it is advantageous to adaptively change the number of keypoints used to accommodate varying bandwidth and internet connectivity. We devise a scheme where our synthesis model can dynamically use a smaller number of keypoints for reconstruction. This is based on the intuition that not all of the images are of the same complexity. Some just require fewer keypoints. Using fewer keypoints, we can reduce the bandwidth required for video conferencing because we just need to send a subset of δd,k\delta_{d,k}’s. To train a model that supports a varying number of keypoints, we randomly choose an index into the array of ordered keypoints, and dropout all values from that index till the end of the array. This dropout percentage ranges from 0% to 75%. This scheme is also helpful when the available bandwidth suddenly drops.

C.3 Binary encoding of the residuals

When the contents of the video being streamed change drastically, \egwhen new objects are introduced into the video or the person changes, it becomes necessary to update the source frame being used to perform the talking-head synthesis. This can be done by sending a new image to the receiver. We also devise a more efficient scheme to encode and send only the residual between the ground truth frame and the reconstructed frame, instead of an entirely new source image. To encode a residual image of size 512×512512\times 512, we use a 3-layer network with convolutions of kernel size 3, stride 2, and 32 channels, similar to the network proposed by Tsai et al. . We compute the sign of the latent code of size 32×64×6432\times 64\times 64 to obtain binary latent codes. The decoder also consists of 3 convolutional layers of 128 filters and uses the pixel shuffle layer to perform upsampling. After arithmetic coding, the binary latent code requires 13.40 KB on average. Note that we do not need to send the encoded binary residual every frame. We just need to send it when the current source image is not good enough to reconstruct the current driving image. In the receiver side, we will use the encoded residual to improve the quality of the reconstructed image. The reconstructed image will become the new source image for decoding future frames using the encoded rotation, translation, and deformations. Example improvements after adding the residual are shown in Fig. 16.

C.4 Dataset

For testing, we collect a set of high-resolution talking-head videos from the web. We ensure that the head is of size at least 512×\times512 pixels and manually check each video to ensure its quality. This results in a total of 222 videos, with a mean of 608 frames, a median of 661 frames, and a min and max of 20 and 1024 frames, respectively.

C.5 Additional experiment results

In Fig. 16(a), we show the achieved LPIPS score by our approach under the adaptive setting (red circle), our approach under the 20 keypoint setting (red triangle), FOMM (green square), fs-vid2vid (orange diamond), H.264, and H.265 using different bpp rates. We observe that our method requires much lower bandwidth than the competing methods.

User study. Here, we describe the details of our user study. We use the Amazon Mechanical Turk (MTurk) platform for the user preference score. A worker needs to have a lift-time approval rate greater than 98 to be qualified for our study. This means that the requesters approve 98% of his/her task assignments. For comparing two competing methods, we generate 222 videos from each method. We show the corresponding pair of videos from two competing methods to three different MTurk workers and ask them to select which one has better visual quality. This gives 666 preference scores for each comparison. We report the average preference score achieved by our method. We compare our adaptive approach to both H.264 and H.265. The user preference scores of our approach when compared to H.264 and H.265 are shown in Fig. 16(b) and (c), respectively. We found that our approach renders comparable performance to H.264 with CRF value 36. For H.265, our approach is comparable to CRF value 37. Our approach was able to achieve the same visual quality using a much lower bit-rate.