FaceVerse: a Fine-grained and Detail-controllable 3D Face Morphable Model from a Hybrid Dataset
Lizhen Wang, Zhiyuan Chen, Tao Yu, Chenguang Ma, Liang Li, Yebin Liu
Introduction
3D human face modeling has been a hot topic in computer vision and computer graphics, which enables a wide range of applications such as film, video games, mixed reality, etc. Since 3D Morphable Model (3DMM) was proposed in 1999, it has been one of the most powerful tools in face-related researches due to its effective control of facial shape, expression and texture. However, recent researches pose more challenges to 3DMM in terms of accuracy, photo-realistic details and editability. On one hand, the performance of 3DMMs is limited due to the difficulty of data acquisition. On the other hand, given a coarse face model, detailed facial geometry and texture are still not changeable in the previous methods , which limits the detailed adjustment of facial features. To overcome the above issues, we propose a hybrid dataset and design a coarse-to-fine structure to combine high generalization ability and fidelity. Furthermore, facial geometry and texture details, like small changes of facial features, can also be parameter-changeable.
At one end of the spectrum, existing 3D face datasets are usually limited in either scale or fidelity. The capturing system can be divided into two categories: sparse or dense camera arrays and consumer depth sensors . The former system requires elaborated setup and the data collection process is quite time-consuming, which limits the scale of captured dataset to a few hundreds. The latter system is off-the-shelf and takes less time in data acquisition, which allows collecting RGB-D data from a large number of identities. However, the captured RGB-D data usually suffers from low resolution and low precision. The insufficiency of scale or fidelity limits the performance of previous works in either generalization or fidelity. Therefore, we propose to build a hybrid dataset.
At the other end, the formulation of previous 3DMMs can not represent parameter-changeable facial details. PCA-based methods can describe shape and expression changes in an effective way. Multi-linear methods present a larger parameter space to cover more information of the corresponding datasets. Non-linear methods use neural networks to achieve better flexibility. However, all the above methods can not represent the facial details, like the detailed shape of facial features. Recent methods show strong capability in 3D fine-grained face model reconstruction, but they still rely on pre-trained super-resolution or displacement prediction networks, which means the facial details are not parameter changeable. To conclude, a 3DMM representation with changeable facial details has not been proposed yet.
To overcome the limitations above, in this paper, we propose FaceVerse, which achieves high generalization ability and fidelity using a hybrid dataset and can generate parameter-changeable facial details. Firstly, we collect a hybrid dataset of East Asians consisting of a large-scale dataset captured by consumer depth sensors and a high-fidelity dataset captured by a multi-camera system. Secondly, we propose a coarse-to-fine structure to scheme our parametric model. The base model is first built from the large-scale dataset, which guarantees strong generalization ability and basic fidelity of the base model. Then, input with the UV maps unwrapped from the base model, we build our detailed model using a novel conditional StyleGAN architecture, which can generate changeable facial details input with additional latent code and noise while preserving the basic facial attributes provided by the input base model. Different from the original StyleGAN , our generator takes advantage of multi-scale features encoded from the input maps to constrain the output maps and we use an additional normal discriminator to further enrich the geometry details. Note that two conditional StyleGAN networks are used in two phases: detail generation and expression refinement. Finally, we propose a single-image fitting pipeline based on differentiable rendering, which also follows the coarse-to-fine idea. Benefiting from the hybrid dataset, the coarse-to-fine scheme and the novel conditional StyleGAN architecture, the proposed FaceVerse shows better performance than previous 3DMM methods both qualitatively and quantitatively.
Our contributions are summarized as follows:
We build a hybrid dataset and propose a coarse-to-fine scheme to make better use of the dataset: the large-scale RGB-D dataset guarantees high generalization ability of our base model and the high-fidelity scan dataset helps to enrich the geometry and texture details of our detailed model.
We propose a conditional StyleGAN architecture with normal discriminators, which allows changing facial details while preserving basic facial attributes.
The proposed FaceVerse provides a powerful tool for face modeling of East Asians and we have released our pre-trained models and the detailed dataset to public for research purposehttps://github.com/LizhenWangT/FaceVerse.
Related Work
The 3D face morphable model (3DMM) has been a long-standing research topic in computer vision since first proposed by Blanz et al. in 1999. 3DMM was first formulated as a linear model by the PCA algorithm, which can represent the shape and texture of 3D face model. The following researches improved the performance using larger 3D face datasets. Moreover, new representations including multi-linear and non-linear models for 3DMM were also proposed in .
Recent 3D face datasets show higher diversity in both identities and expressions. LSFM was built from a large 3D face dataset containing 10,000 face scans and shows better generalization in facial shape fitting. In the meanwhile, 3D face datasets with rich expressions were also collected to incorporate the facial expression bases into 3DMM . Furthermore, with the development of elaborated capturing system like dense camera arrays, recent 3DMM methods exhibited even higher accuracy in 3D face modeling.
Besides the improvement in 3D face datasets, novel modeling mechanisms were also presented for better performance and flexibility. Vlasic et al. first proposed a multi-linear model to jointly estimate the variations in identity and expression, Cao et al. and Yang et al. built comprehensive bilinear models which decompose the face meshes in both identity and expression dimensions. Recently, non-linear models were also proposed to enable adaptive and high-level facial deformations. Neumann et al. decomposed the captured face mesh sequences into the sparse and localized deformation components. With the development of neural networks, generative adversarial networks (GAN) were also used to build the non-linear 3DMMs , the face representations of which can be controlled by high-level semantics.
Monocular Face Reconstruction Based on 3DMM.
Monocular 3D face reconstruction based on 3DMM plays an important role in many applications like face alignment and face view synthesis . With the assistance of 3DMM, the 3D face reconstruction task can be simplified as a model fitting problem. Early methods mainly tried to regress the parameters of 3DMM using the facial landmarks or some other facial features. Then the convolutional neural networks were used to directly predict the parameters from an input face image . Recently, self-supervised methods based on differentiable rendering were presented and show great performance in fitting 3D face models from a single face image.
The above methods based on model parameters prediction are limited in representing facial details, and thus multi-layer refinement structures are proposed to reconstruct detailed face models. Recent works firstly generated a rough face model through the model parameters prediction, and then refined the facial details by adjusting the rendered depth or predicting a displacement map. Lin et al. generated high-fidelity models by the optimization of the albedo and normal maps. However, the detailed facial features are still not parameter-changeable in these researches, which limits the adjustment of facial details in 3D face models.
Compared with the state-of-the-art 3DMM and monocular facial reconstruction methods, our approach is superior in the following aspects: (a) our model is built from a hybrid dataset, which contains a large-scale coarse dataset and a high-fidelity detailed dataset; (b) we propose a coarse-to-fine model which consists of a PCA-based model and a novel conditional styleGAN-based non-linear model; (c) our coarse-fine model-fitting pipeline based on differentiable rendering can not only reconstruct high-fidelity 3D face models from in-the-wild face images, but also generate facial details which can be adjusted by our detailed parameters.
Hybrid Dataset
We chose the structured-light depth sensors to collect coarse 3D face data from volunteers, which show better performance than ToF-based devices in distance below 1 meter. Compared with dense camera array, the structured-light depth sensors are cost-friendly and more convenient for parallel setup, which allows collecting RBG-D data from a large number of identities. In practice, as shown in Fig. 2.a, we collect about 5 RGB-D frames for each volunteer and the frames are fused by ICP registration to generate a smooth facial point cloud. The whole capturing process for each volunteer only costs 5 to 10 seconds. With the assistance of several data acquisition companies and parallel capturing, we finally get 60K textured facial point clouds of East Asians after data cleaning. Volunteers are required to keep neutral expression during the capturing to ensure the consistence of data distribution in expression.
In order to generate a topologically uniformed parametric model, we use a pre-designed 3D facial template mesh to fit the point clouds. We firstly detect facial landmarks using OpenSeeFacehttps://github.com/emilianavt/OpenSeeFace from the captured RGB images and project them to the fused point clouds. Then we roughly align the point clouds to our template mesh by 3D landmarks. Finally, a Non-rigid ICP algorithm is utilized to deform the template mesh to the aligned point clouds. The distribution of age and gender is presented in Fig. 2.b.
2 Detailed Dataset
Our camera system for 3D scan model collection consists of 128 DSLR cameras, which equip with 85 mm lenses and are placed about 2.5 meters away from the volunteer, as shown in Fig. 3. The cameras are arranged in cylinder facing towards the center by 16 pillars with 8 cameras on each, which is similar to high quality full body scan system in . During data collection, 128 images with resolution will be synchronously collected from different view points. We follow the Data acquisition process of FaceWarehouse , where the volunteers are required to perform 21 specific expressions including neutral expression. We finally collect 2,310 scan models (110 identities in 21 expressions) for training and 378 scan models (18 identities in 21 expressions) for testing, which has been released to public for research purpose.
After the data collection, the 3D scans are fitted to our topologically uniformed template. Firstly, 3D landmarks are marked for rigid-ICP alignment by projecting 2D landmarks onto the 3D scans. Our base model generated from the coarse dataset (Sec. 4.1) is used to fit the scans with corresponding 3D landmarks. Then, the resulting fitted models are up-sampled in the UV space (from to ) for the subsequent registration. Finally, we conduct the detailed deformation on the fitted models using Non-rigid ICP .
FaceVerse Model
A coarse-to-fine scheme is proposed to generate the proposed model, FaceVerse, from the hybrid dataset: our base model is built from the large-scale coarse dataset by PCA and the detailed model is built from the high-fidelity detailed dataset by our conditional StyleGAN networks. In addition, we also present a single-image fitting framework based on differentiable rendering.
Benefiting from the large-scale coarse dataset, our base model shows strong performance in fitting faces of different ages and genders quantitatively. However, our base model can not preserve the facial geometry and texture details, which will be generated by the following detailed model.
2 Detailed Model Generation
As shown in Fig. 4, to incorporate more detailed facial geometry and texture, we propose a neural representation for our detailed model, which can take better advantage of the detailed dataset. The base model is first unwrapped into the UV space and up-sampled to to facilitate subsequent processing. The whole refinement work is divided into a shape&texture refinement part and a expression refinement part.
In the shape&texture refinement part, we use a conditional StyleGAN to generate facial geometry and texture details. Firstly, the input base model in the neutral expression is unwrapped into a geometry UV map and a texture UV map . Note that we believe geometry details and texture details should have a strong correlation, so we concatenate the geometry and texture UV maps into a 6-channel input . Due to the combined training of geometry and texture channels, the output geometry and texture are influenced by each other, which further facilitates the subsequent detailed geometry fitting (Sec. 4.3). The concatenated 12-channel input and output UV map of is entered into the discriminator and a 3-channel UV normal map is entered into the normal discriminator . As discussed in Sec. 5.3, the output detailed model shows fine-grained facial geometry and texture details, which can be controlled by and the injected noise. Moreover, the basic shape and texture provided the input base model are still retained. The loss terms used in the training of can be formulated as
where represents the adversarial loss term and the path length regularization term of StyleGAN provided by and . Note that our training process is under incomplete supervision and thus the used training data contains not only the data pairs from the detailed dataset but also the conditional UV maps generated from our coarse dataset, which further guarantees the effective interpolation capability of our detail generator .
In the expression refinement part, detailed expression-related geometry changes like a smiling mouth will be further refined by another conditional StyleGAN network . Given the detailed geometry UV map in neutral expression and the base expression formulated by the UV offset map , will refine the detailed geometry while preserving the basic shape and expression. Specifically, the 6-channel conditional input consists of a basic geometry which is the sum of and and an additional expression offset , where the basic geometry input is used to constrain the similarity of the input and output geometry and the additional expression input provides priors of the facial expression. The concatenated 9-channel input and output UV map of is entered into the discriminator and the concatenated 6-channel UV map consisting of a normal map and an expression offset map is entered into the normal discriminator . After training of , the 3-channel output geometry can represent more detailed expression changes, as discussed in Sec. 5.3, and the generation is also controlled by the latent code and injected noise. The training process of utilizes the paired data generated from our detailed dataset and the training loss terms can be formulated as
3 Coarse-to-Fine Single-Image Fitting
We further propose a single-image fitting pipeline, which adopts an optimization algorithm based on differentiable rendering, in this subsection, as shown in Fig. 6. The fitting process is divided into three phases: base model fitting, detailed model fitting and expression refinement.
where denotes the mean square loss of the detected 2D facial landmarks and the projected landmarks from the 3D model, denotes the mean square loss of the rendered image and the input image and denotes the L2 regular terms of , and . The resulting shape, texture and expression are unwrapped into UV maps for subsequent phases.
In the detailed model fitting phase, we recover the identity-related facial geometry and texture details through the optimization of our pre-trained detail generator . The expression offset UV map, and generated from the previous phase are fixed in this and the next phase. Input with the UV maps of shape and texture generated from the previous phase, our detail latent code and the injected noise are first randomly sampled and then optimized by differentiable rendering with the similar loss terms, where is changed to the L2 regular terms of the injected noise. Note that the detailed geometry can also be generated through the associations between geometry and texture established by . The latent code mainly controls the generation of medium-grained face details like detailed shape of facial features, while the small-grained facial details like freckles is controlled by the injected noise. As shown in Fig. 6, the facial details can be generated after the optimization.
In the expression refinement phase, expression-related geometry changes will be further refined through the optimization of our pre-trained expression refinement generator, . Input with a conditional image composed of the expression offset UV map generated from the base model fitting phase and the output detailed geometry generated from the detail model fitting phase, the expression latent code and the injected noise are also first randomly sampled and then optimized by differentiable rendering with the same loss terms with the detailed model fitting phase. After the final optimization, more detailed geometry like a smiling mouth are further refined, as presented in Fig. 6.
Experiments
We firstly evaluate the performance of our coarse-to-fine 3D model fitting framework in predicting 3D face model from a single face image. As shown in Fig. 7, based on FaceVerse, our method shows both high-generalization and high-fidelity in predicting 3D face models from various input East Asian face images among different ages, genders or skin colors. On one hand, the base model built on large-scale base face dataset provide more prior knowledge in rough face fitting which makes the method more robust to various face images. On the other hand, the conditional styleGAN-based detailed model trained on the detailed dataset shows powerful ability in facial geometry and texture details generation. The texture details and geometry details in Fig. 7 proves that even the facial details in pupils and eyebrows can still be described by our details generator.
In addition, both the base shape and the detailed shape of facial features can be adjusted by parameters in our method. To demonstrate the changeability of our model, we conduct a detail transfer experiment over the single-images fitting results, as shown in Fig. 8. Using the base parameters fitted from images in the left column and the detail parameters fitted from images in the top row, our method can generate new face models which has the basic shape of the source face and the details of the target face (e.g. bigger eyes, thinner lips or a broader nose). Some images are sampled from the FFHQ dataset.
2 Comparisons to Prior Works
We compare our monocular fitting results with the state-of-art monocular facial reconstruction methods, including FaceScape and Hifi3DFace which are also proposed for East Asian facial reconstruction, as well as DECA and 3DDFAv2 which is based on BFM and FLAME respectively. As shown in Fig. 9, gaining from the large-scale base model and GAN-based detail generator, our method shows better qualitative performance in both fitting face rough shape and generating face details compared with other methods. We also conduct a quantitative comparison using a single image and the corresponding detailed 3D models sampled from our testing set. As shown in Fig. 10, the generated models of different methods are fitted to the ground-truth models by a rigid-ICP algorithm. The calculated MAE error is displayed below the models and our method shows the best quantitative performance.
To further demonstrate the effectiveness of our parametric base model, we conduct a quantitative comparison with the state-of-the-art asian facial parametric models proposed by FaceScape and Hifi3DFace , as well as the BFM , on 3D scans from our testing set, which contains 357 models from 17 people and the models are fixed in the length of 200mm. We fit the parametric models to 3D scans by an optimization algorithm, which is based on the back-propagation through ICP (the algorithm is explained in detail in our supplementary pdf file). Note that our detailed model needs additional texture input, so we only use our base model to make a fair comparison. As shown in Fig. 12, benefiting from the large-scale dataset, our base model shows the best quantitative performance in 3D model fitting. The visualized results are also presented in Fig. 11.
3 Ablation Study
In order to demonstrate the effectiveness of the modules used in our approach, we compare the fitting results of our base model, detailed results generated by the detail generator , refined results of the expression refinement generator and results generated by the detailed model trained without our normal discriminator. As shown in Fig. 13, given a base shape and texture, our detail generator can add reasonable details but still lack description power of expressions. The geometric changes caused by expressions can be further refined by , as indicated by the blue rectangles in Fig. 13. Besides, the detailed model trained without our normal discriminator shows messy geometry, which demonstrates the effectiveness of our normal discriminator. Furthermore, the ablation study of the effects of the injected noise and latent code of our detailed model is presented in our supplementary materials. Please watch our supplementary video for more results.
To further prove the superiority of introducing the coarse dataset into our base model, we generate an additional base model only using our detailed dataset, which contains 50 shape principal components and using the same expression principal components with our full base model. The 3D fitting results are also presented in Fig. 11 and Fig. 12 (labelled as “Ours w/o coarse dataset”). The quantitative results prove that the fitting ability is significantly improved after introducing our coarse dataset.
Discussion and Conclusion
Limitations. On one hand, our dataset only contains faces of East Asians, and thus our performance declines when fitting the faces from other regions. On the other hand, our detailed model still suffers from a lack of detailed 3D face scans of the old people. As a result, as shown in Fig. 14, our method can not generate extreme textures like a thick beard and can not generate deep wrinkles of the old people.
Potential Social Impact. Our method enables 3D face reconstruction from a single image. Therefore, it can be used to generate a 3D fake model of a person, which needs to be addressed carefully before deploying the technology.
Conclusion. In this paper, we have presented FaceVerse, a fine-grained and detail-changeable 3D face morphable model from a hybrid dataset. We have collected a large-scale coarse dataset and a high-fidelity detailed dataset and proposed a coarse-to-fine scheme to build our model, which guarantees the high generalization ablity and high fidelity of our model. The proposed conditional StyleGAN is able to generate and control the facial geometry and texture details while perserving the basic facial atrributes from the base model. Experiments have demonstated the superiority of our method in 3D face model fitting and monocular face reconstruction compared with the state-of-the-art methods. We believe FaceVerse can be a powerful tool for face-related researches and our pipeline will inspire the following research of 3DMM and monocular 3D facial reconstruction.
Acknowledgement. This work is supported by Ant Group through Ant Research Program and is sponsed by NSFC No. 62125107 and No. 62171255.