Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation

Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, Bolei Zhou

Introduction

Generating high-fidelity video portraits based on speech audio is of great importance to various applications like digital human, film-making and video dubbing. Many researchers tackle the task of audio-driven talking face or video portrait generation by using deep generative models. Several works rely solely on learning-based image reconstruction, which typically synthesize static results of low-resolution . Other methods utilize explicit structural intermediate representations such as 2D landmarks or 3D facial models . Though some of them can generate high-fidelity images , the errors in structured representation prediction (e.g., expression parameters of a 3D Morphable Model (3DMM) ) lead to inaccurate face deformation .

Recently, the implicit 3D scene representation of Neural Radiance Fields (NeRF) provides a new perspective for realistic generation. It enables free-view control with higher image quality compared to explicit methods, which is suitable for the video portrait generation task. Gafni et al. first involve NeRF in the dynamic human head modeling from single-view data in a video-driven manner. However, an accurate explicit 3D model is still required in their settings. Moreover, they model torso consistently with the head, which leads to unstable results. Guo et al. further propose AD-NeRF for audio-driven talking head synthesis. In particular, they build two individual sets of NeRF for head and torso modeling conditioned on audio input. Such a straightforward pipeline suffers from head-torso separation during the render stage, making generated results unnatural.

Based on previous studies, we identify two key challenges for incorporating NeRF into video portrait generation: 1) Each facial part’s appearance and moving patterns are intrinsically connected but substantially different, especially when associated with audios. Thus weighing all rendering areas equally without semantic guidance would lead to blurry details and difficulties in training. 2) While it is easy to bind head pose with camera pose, the global movements of the head and torso are in significant divergence. As the human head and torso are non-rigidly connected, modeling them with one set of NeRF is an ill-posed problem.

In this work, we develop a method called Semantic-aware Speaking Portrait NeRF (SSP-NeRF), which generates stable audio-driven video portraits of high-fidelity. We show that semantic awareness is the key to handle both local facial dynamics and global head-torso relationship. Our intuition lies in the fact that different parts of a speaking portrait have different associations with speech audio. While other organs like ears move along with the head, the high-frequency mouth motion that is strongly correlated with audios requires additional attention. To this end, we devise an Semantic-Aware Dynamic Ray Sampling module, which consists of an Implicit Portrait Parsing branch and a Dynamic Sampling Strategy. Specifically, the parsing branch supervises the modeling with facial semantics in 2D plane. Then the number of rays sampled at each semantic region could be adjusted dynamically according to the parsing difficulty. Thus more attention can be paid to the small but important areas like lip and teeth for better lip-synced results. Besides, we also enhance the semantic information by anchoring a set of latent codes to the vertices of a roughly predicted 3DMM without expression parameters.

On the other hand, since the head and torso motions are rigidly bound together in the current NeRF, a correctly positioned torso cannot be rendered even with the portrait parsing results. We further observe the relationship between head and torso: while they share the same translational movements, the orientation of torso seldom changes with head pose under the speaking portrait setting. Thus we model non-rigid deformation through a Torso Deformation module. Concretely, for each point (x,y,z)(x,y,z) in the 3D scene, we predict a displacement (Δx,Δy,Δz)(\Delta x,\Delta y,\Delta z) based on the head-canonical view information and time flows. Interestingly, although there are local deformations on the face, the deformation module implicitly learns to focus on the global parts. This design facilitates portrait stabilization in one unified set of NeRF. Experiments demonstrate that our method generates high-fidelity video portraits with better lip-synchronization and better image quality efficiently.

To summarize, our work has three main contributions: (1) We propose the Semantic-Aware Dynamic Ray Sampling module to grasp the detailed appearance and local dynamics of each portrait part without using accurate structural information. (2) We propose the Torso Deformation module that implicitly learns the global torso motion to prevent unnatural head-torso separated results. (3) Extensive experiments show that the proposed SSP-NeRF renders high-fidelity audio-driven video portraits with one unified NeRF in an efficient manner, which outperforms state-of-the-art methods on both objective evaluations and human studies.

Related Work

Audio-Driven Talking Head Synthesis. Driving talking head with speech audio has a bunch of applications, which is of great research interest to computer vision and graphics. Conventional works mostly resort to stitching techniques , where a predefined set of phoneme-mouth correspondence rules is used to modify mouth shapes. With the rapid growth of deep neural networks, end-to-end frameworks are proposed. One category of methods, namely image reconstruction-based methods, generate talking face by latent feature learning and image reconstruction . For example, Chung et al. propose the first end-to-end method with an encoder-decoder pipeline. Zhou et al. explicitly disentangle identity and word information for better feature extraction. Prajwal et al. achieve synchronous lip movements with a pretrained lip-sync expert. However, these methods can only generate fix-sized images with low resolution. Another line of approaches named model-based methods utilize structural intermediate representations like 2D facial landmarks or 3D representations to bridge the mappings from audio to complicated facial images . Typically, Chen et al. and Das et al. first predict 2D landmarks then generate faces. Thies et al. and Song et al. infer facial expression parameters from audio in the first stage, then generate 3D mesh for final image synthesis. But errors in intermediate prediction often hinder accurate results. In contrast to these two lines of works, our method can render more realistic speaking portraits of high-fidelity without any accurate structural information.

Implicit Representation Methods. Recent works leverage implicit functions for learning scene representations , where multi-layer perceptron (MLP) weights are used to represent the mapping from spatial coordinates to a signal in continuous space like occupancy , signed distance function , color and volume density , semantic label and neural feature map . A recent popular work named Neural Radiance Fields (NeRF) optimizes an underlying continuous volumetric scene mapping from 5D coordinate of spatial location and view direction to implicit fields of color and density for photo-realistic view results. Naturally, naive NeRF is confined to static scenes, which triggers a branch of studies to extend NeRF for dynamic scenes . However, few works focus on complicated dynamic scenes like speaking portraits. The main difficulty lies in the learning of cross-modal associations between different portrait parts and speech audio. One relevant work synthesizes talking head with two individual sets of NeRF for head and torso, making generated results fall apart. In this work, we take semantics as guidance to grasp each portrait part’s local dynamics and appearances for fine-grained results efficiently. A deformation module further enables us to synthesize stable video portraits using one unified set of NeRF.

Our Approach

We present Semantic-aware Speaking Portrait NeRF (SSP-NeRF) that generates delicate audio-driven portraits with one unified set of NeRF. The whole pipeline is depicted in Fig. 1. In this section, we first review the preliminaries and the problem setting of video portrait synthesis with neural radiance fields (Sec. 3.1). We then introduce the Semantic-Aware Dynamic Ray Sampling module, which facilitates fine-grained appearance and dynamics modeling for each portrait part with semantic information (Sec. 3.2). Furthermore, we elaborate the Torso Deformation module that handles non-rigid torso motion by learning location displacements (Sec. 3.3). Finally, the volume rendering process and network training details are described (Sec. 3.4).

Given images with calibrated camera intrinsics and extrinsics, NeRF represents a scene using a continuous volumetric radiance field FF. Specifically, FF is modeled by an MLP, which takes 3D spatial coordinates x=(x,y,z)\mathbf{x}=(x,y,z) and 2D view directions d=(θ,ϕ)\mathbf{d}=(\theta,\phi) as input, then outputs the implicit fields of color c=(r,g,b)\mathbf{c}=(r,g,b) and density σ\sigma. In this way, the MLP weights store scene information by the mapping of F:(x,d)→(c,σ)F:(\mathbf{x},\mathbf{d})\rightarrow(\mathbf{c},\sigma). To compute the color of a single pixel, NeRF approximates the volume rendering integral using numerical quadrature . Consider the ray r(v)=o+vd\mathbf{r}(v)=\mathbf{o}+v\mathbf{d} from camera center o\mathbf{o}, its expected color C^(r)\hat{C}(\mathbf{r}) with near and far bounds vnv_{n} and vfv_{f} is calculated as:

where T(v)=exp⁡(−∫vnvσ(r(u))du)T(v)=\exp(-\int_{v_{n}}^{v}\sigma(\mathbf{r}(u))du) is the accumulated transmittance along the ray from vnv_{n} to vv. With the hierarchical volume sampling, both coarse and fine MLPs are optimized by minimizing the photometric discrepancy.

Guo et al. use an off-the-shelf parsing method to divide training images into head and torso for individual NeRF modeling. Following their settings, we assume that the semantic parsing maps are also available in our method.

2 Semantic-Aware Dynamic Ray Sampling

To avoid the unnatural head-torso separation problem described in Sec.1, we render the whole portrait with one unified set of NeRF. However, two problems remain: 1) The associations between different portrait parts and audio are different. For example, audio is more related to lip movements than torso motions. How to grasp the fine-grained appearance and dynamics of each portrait part remains unsolved; 2) Since the rays are uniformly sampled over the whole image, how to make the model pay more attention to small but important regions like mouth is challenging.

Implicit Portrait Parsing Branch. Our solution to the first problem is to add a parsing branch. Since the portrait parts of the same semantic category share similar motion patterns and texture information, it will be beneficial for the appearance and geometry learning in NeRF, which is also proven in recent implicit representation studies . As shown in Fig. 1, we extend the original NeRF with an additional parsing branch that predicts the semantic information. Note that since a certain 3D coordinate’s semantic label is view-invariant, the parsing branch does not condition on view direction d\mathbf{d}. Specifically, suppose there are totally KK semantic categories, the parsing branch maps the 3D spatial coordinate x\mathbf{x} to semantic logits s(x)\mathbf{s(x)} over KK classes, which is further conditioned on audio a\mathbf{a}. Hence the expected semantic logits S^(r)\hat{S}(\mathbf{r}) along the ray r(v)\mathbf{r}(v) with near and far bounds vnv_{n} and vfv_{f} can be calculated as:

Such semantic awareness can naturally distinguish each part over the whole image, thus figuring out different associations between audio and different portrait regions.

Dynamic Ray Sampling Strategy. To generate delicate facial images with lip-synced results, we have to care for each portrait part, especially those small but crucial regions. Original NeRF uniformly samples rays on the image plane . Such an unconstrained ray sampling process focuses on big regions (e.g., background and cheek) yet ignores small regions (e.g., lip and teeth) that are important for fine-grained results. Therefore, we use semantic information to guide the ray sampling process dynamically. In particular, we denote all the points that are sampled on the image as Ω=⋃i=1KΩi\Omega=\bigcup_{i=1}^{K}\Omega_{i}, where KK is the total number of semantic categories in parsing map and Ωi\Omega_{i} is the set of points that are sampled on the ii-th semantic class. During the training stage, we calculate the average loss of each category Li\mathcal{L}_{i} for the previous epoch (the sum of semantic loss and RGB loss, which will be introduced in Sec. 3.4), and then dynamically sample rays across KK categories by:

where NΩiN_{\Omega_{i}} denotes the number of rays distributed to the ii-th category and NsN_{s} is the total number of sampled rays. We identify two benefits for such design: 1) The average loss of a semantic category is area-agnostic. Thus the learning process will equally sample those small-area regions; 2) Some image parts are comparatively easier to learn. For example, the texture of eye is more complicated than that of background. This leads to lower loss of background category and dynamically drives the implicit function to pay more attention to hard-to-learn regions. Our experiment further shows that this design can accelerate training as well.

3 Torso Deformation Module

As mentioned in Sec. 3.1, the estimated head pose serves as camera pose. However, such straightforward treatment ignores the fact that head and torso motions are inconsistent. To tackle this problem, we design a Torso Deformation module to stabilize the large-scale non-rigid torso motions.

Notably, although such deformation is added to the whole image, we empirically find that only the torso part tends to be deformed, while the facial dynamics are naturally modeled by semantic-aware implicit function in Eq. 6. Such disentanglement will be further analyzed in Sec. 4.5.

Overall Implicit Function. Combine the semantic-aware implicit function with our proposed Torso Deformation module, we can model the overall implicit function as:

4 Volume Rendering and Network Training

With the deformed 3D coordinate x′(v,t)\mathbf{x}^{\prime}(v,t) along the modified ray path r′(v,t)\mathbf{r}^{\prime}(v,t), we can calculate the expected color C^(r′(v),t)\hat{C}(\mathbf{r}^{\prime}(v),t) and semantic logits S^(r′(v),t)\hat{S}(\mathbf{r}^{\prime}(v),t) with near and far bounds vnv_{n} and vfv_{f} under semantic-aware setting as:

where T′(v,t)T^{\prime}(v,t) is the accumulated transmittance along the ray path r′(v,t)\mathbf{r}^{\prime}(v,t) from vnv_{n} to vv. Note that the estimated semantic logits S^(r′)\hat{S}(\mathbf{r}^{\prime}) are subsequently transformed into multi-class distribution p(r′)p(\mathbf{r}^{\prime}) through softmax operation.

where R′\mathcal{R}^{\prime} is the set of deformed camera rays passing through image pixels; C(r′)C(\mathbf{r}^{\prime}), C^c(r′)\hat{C}_{c}(\mathbf{r}^{\prime}) and C^f(r′)\hat{C}_{f}(\mathbf{r}^{\prime}) denote the ground-truth, coarse volume predicted and fine volume predicted pixel color for the deformed ray r′\mathbf{r}^{\prime}, respectively; and pk(r′)p^{k}(\mathbf{r}^{\prime}), p^ck(r′)\hat{p}_{c}^{k}(\mathbf{r}^{\prime}) and p^fk(r′)\hat{p}_{f}^{k}(\mathbf{r}^{\prime}) denote the ground-truth, coarse volume predicted and fine volume predicted multi-class semantic distribution for the deformed ray r′\mathbf{r}^{\prime}, respectively. The overall learning objective for the framework is:

where λ\lambda is the weight balancing coefficient. At the training stage, the network parameters Θ\Theta and Φ\Phi of the implicit functions in Eq. 3.3 are updated based on above loss function.

Experiments

Dataset Collection. Our method targets to synthesize audio-driven facial images. Hence a certain person’s speaking portrait video with audio track is needed. Unlike previous studies that demand large-corpus data or hours-long videos, we can achieve high-fidelity results with short videos of merely a few minutes. In particular, we extend the publicly-released video set of Guo et al. and obtain videos of average length 6,750 frames in 25 fps.2

2 Experimental Settings

Comparison Baselines. We compare our method with recent representative works: (1) ATVG , which uses 2D landmark to guide facial image synthesis; (2) Wav2Lip that achieves state-of-the-art lip-sync performance by pretraining a lip-sync expert; (3) MakeitTalk , a representative 3D landmark-based approach; (4) PC-AVS which generates pose-controllable talking face by modularized audio-visual representation; (5) NVP that first infers expression parameters from audio, then generates images with a neural renderer; (6) SynObama which learns mouth shape changes for facial image warping; (7) AD-NeRF , which is the first work that uses implicit representation of NeRF to achieve arbitrary-size talking head synthesis. In particular, we also show the evaluations directly on the Ground Truth for a clearer comparison.

Implementation Details.Please refer to supplementary material for more details. The FΘsemanticF^{\text{semantic}}_{\Theta} and FΦdeformF^{\text{deform}}_{\Phi} together with their associated fine models all consist of simple 8-layers MLPs with hidden size of 128 and ReLU activations. Following NeRF , positional encoding is applied to each 3D coordinate x\mathbf{x}, view direction d\mathbf{d} and time instant tt to map the input into higher dimensional space for better learning. The positional encoder is formulated as: γ(q)=<(sin⁡(2lπq),cos⁡(2lπq))>0L\gamma(q)=<(\sin(2^{l}\pi q),\cos(2^{l}\pi q))>^{L}_{0}, where we use L=10L=10 for x\mathbf{x}, and L=4L=4 for d\mathbf{d} and tt. For the parsing maps, we use K=11K=11 categories for semantic guidance, including cheek, eye, eyebrow, ear, nose, teeth, lip, neck, torso, hair and background. The structured 3D feature extractor is borrowed from that processes feature volume with 3D sparse convolutions and outputs latent code with 2×2\times, 4×4\times, 8×8\times, 16×16\times downsampled sizes. The semantic weight λ\lambda is empirically set to 0.040.04. The model is trained with 450×450450\times 450 images during 400k400k iterations with a batch size of Ns=1024N_{s}=1024 rays. The framework is implemented in PyTorch and trained with Adam optimizer of learning rate 5e−45e-4 on a single Tesla V100 GPU for 36 hours.

3 Quantitative Evaluation

Evaluation Metrics. We employ evaluation metrics that have been previously used in talking face generation. We adopt PSNR and SSIM to evaluate the image quality of generated results; Landmark Distance (LMD) and SyncNet Confidence to account for the accuracy of mouth shapes and lip sync. Note that the landmarks are detected from synthesized images for the computation of LMD metric. Other metrics such as CSIM for measuring identity preserving and CPBD for measuring result sharpness are shown in supplementary material.

Comparison Settings. The reconstruction/model-based methods require large-corpus training data or long videos, hence we directly inference with their publicly-released best models. Note that all baseline methods except for fail to generate the whole portrait with full resolution, we divide our comparisons into two settings: 1) The cropped setting in Table 1, where we crop the generated facial image with same region and resize into same size for fair evaluation metric comparison. 2) The full resolution setting in Table 2, where we compare with AD-NeRF that could also synthesize the whole portrait with full resolution of 450×450450\times 450.

In the first setting, since NVP and SynObama do not provide pretrained models, we conduct comparisons on three datasets: (1) Testset A, the collected dataset mentioned in Sec. 4.1; (2) Testset B, where we extract speech audio from the demo of NVP to drive other baselines; (3) Testset C, where the audio from SynObama’s demo is used for animation. Note that the metrics for measuring image quality (PSNR and SSIM) are not evaluated on Testset B and C due to the low image quality of original videos. In the second setting, the experiment is only conducted on Testset A for high-resolution comparison. We further compare the number of model parameters against AD-NeRF to show the efficiency of our proposed approach.

Evaluation Results. The results of the cropped setting and full resolution setting are shown in Table 1 and Table 2, respectively. It can be seen that the proposed SSP-NeRF achieves the best evaluation results in most metrics: (1) In the cropped setting, we synthesize fine-grained facial images with detailed local appearance and dynamics of each portrait part. Note that Wav2Lip uses SyncNet for pretraining, which makes their results on SyncNet Confidence even better than the ground truth. Our performance on the LMD metric is the best, and the SyncNet Confidence of our model is close to the ground truth on all three datasets, showing that we can generate accurate lip-sync video portraits. (2) In the full resolution setting, the human face as well as torso part is evaluated. Different from AD-NeRF’s separated rendering pipeline, our design of Torso Deformation module facilitates steady results. The statistics on both model’s parameter number are shown in Table 2. Notably, our method is trained with 400k×1024400k\times 1024 sampled rays, while AD-NeRF uses 400k×2048400k\times 2048 rays for each model. Hence we generate portraits of better image quality and better lip-synchronization in a more compact model with fewer iterations, proving the effectiveness and efficiency of SSP-NeRF.2

4 Qualitative Evaluation

To compare the generated results of each method, we show the key frames of two clips in Fig. 2. The figure shows that our method synthesizes more lip-synced video portraits of higher image quality. In particular, ATVG and MakeitTalk rely on precise facial landmarks, which leads to inaccurate mouth shapes (green arrow); Wav2Lip creates static talking heads; PC-AVS fails to preserve the speaker’s identity, making generated results unrealistic. Moreover, all the image reconstruction-based methods or model-based methods fail to synthesize the whole portrait of high-fidelity simultaneously. Although AD-NeRF manages to create full-resolution results, the separated rendering pipeline with uniform ray sampling leads to head-torso separation (as highlighted by blue arrows) and blurry results (orange arrow).

User Study.2 Since subjective evaluation can reflect the quality of audio-driven portrait, a user study is further conducted. Specifically, we sample 30 audio clips from Testset A, B and C for all methods to generate results, and then involve 18 participants for user study. The Mean Opinion Scores rating protocol is adopted for evaluation, which requires the participants to rate three aspects of generated speaking portraits: (1) Lip-sync Accuracy; (2) Video Realness; (3) Image Quality. The rating is based on a scale of 1 to 5, with 5 being the maximum and 1 being the minimum.

The results are shown in Table 3. Since NeRF enables full-resolution whole portrait generation with expressive pose, both AD-NeRF and our method score comparatively high on Image Quality and Video Realness. Besides, the users prefer our generated speaking portraits to AD-NeRF’s due to the fine-grained local rendering and stable torso motions provided by our framework design. Although PC-AVS also creates pose-controllable talking faces, the inaccuracy of implicit pose code extraction weakens their realness. Note that Wav2Lip , NVP and SynObama achieve competitive scores on Lip-sync Accuracy. However, they rely on large corpus or long training videos, while we merely take a short video as input, showing the efficacy of our method. To further measure the disagreement on scoring among the participants, the Fleiss’s-Kappahttps://en.wikipedia.org/wiki/Fleiss%27_kappa statistic is calculated on 18 participants’ ratings. The Fleiss-Kappa value is 0.8160.816, which can be interpreted as “almost perfect agreement”.

5 Ablation Study

In this section, we present ablation study on the Testset A in terms of two key modules proposed in our framework.

Torso Deformation Module. We conduct ablation experiments under two settings: (1) w/o FΦdeformF^{\text{deform}}_{\Phi}, where we directly synthesize the whole portrait without deforming 3D coordinates. The results are shown in Table 4 (line1), where the ill-posed rendering leads to blurry torso with low image quality. To further investigate the efficacy of Torso Deformation module, we visualize the heatmap of learned displacements in Fig. 3. Since audio feature is not input to the deformation implicit function, it tends to warp the weakly audio-related torso part, while the strongly audio-related mouth movements are mostly modeled by FΘsemanticF^{\text{semantic}}_{\Theta}. The marginal drop in lip-sync metrics also suggests that the deformation module majorly takes effect on the torso part.

Broader Impact

Ethical Consideration. Animating realistic talking portrait has extensive applications like digital human and film-making. On the other hand, it could be misused for malicious purposes such as identity theft, deepfake generation, and media manipulation. Recent studies have shown promising results in detecting deepfakes . However, the lack of realistic data limits their performance. As part of our responsibility, we feel obliged to share our generated results with the deepfake detection community to improve the model’s robustness. We believe that the proper use of this technique will enhance the healthy development of both machine learning research and digital entertainment.

Limitation and Future Work. Our proposed SSP-NeRF achieves audio-driven video portrait generation of high-fidelity. However, the method still has limitations. The speed of synthesizing images is slow due to the heavy computation of rendering high-quality images. We also observe that the language gap between training and driven audio makes the synthesized mouth look unnatural occasionally . We will address these issues in future work.

Conclusion

In this paper, we propose a novel framework Semantic-aware Speaking Portrait NeRF (SSP-NeRF) for audio-driven portrait generation. We introduce Semantic-Aware Dynamic Ray Sampling module to grasp the detailed appearance and the local dynamics of each portrait part without using accurate structural information. We then propose a Torso Deformation module to learn global torso motion and prevent head-torso separated results. Extensive experiments show that our approach can synthesize more realistic video portraits compared to the previous methods.

References