Morphable Face Models - An Open Framework
Thomas Gerig, Andreas Morel-Forster, Clemens Blumer, Bernhard Egger, Marcel Lüthi, Sandro Schönborn, Thomas Vetter
I Introduction
In this paper, we derive a method for face registration based on Gaussian Process Morphable Models, where face specific domain knowledge is modeled with a Gaussian process. This approach has the following advantages:
The method is conceptually simple because problem specific adaptions are decoupled from the registration algorithm.
Domain knowledge is modeled intuitively using building blocks in terms of kernels.
Extending the model does not require changing the registration algorithm.
As the deformation prior is generative, random samples can be drawn to check modeling assumptions visually.
We show how to build a prior model for faces incrementally:
1) The geometric variability of the face can be decoupled into multiple levels of detail. We model this variability with multi-scale B-spline kernels and propose an adaption scheme to damp the predefined regions on different deformation scales spatially.
2) Facial shapes are nearly mirror symmetric. We model this by modeling symmetry with a mirror symmetric kernel.
3) We propose a simple statistical shape model kernel built from facial expression prototypes to model the opening and closing of the mouth.
4) To build a shape and texture model from the registered data, we propose a model-building method, which also handles regions with missing data.
A further primary purpose of this work is full reproducibility on publicly available data. We release the full face registration and model-building pipeline together with experiments on the model adaption of a single 2D image. By releasing the complete pipeline, tested on publicly available data, we provide full reproducibility for all the individual steps and the end-result of the pipeline.
We also release a new Basel Face Model (BFM-2017). The model contains facial expressions, is based on an improved age distribution compared to the model published by and is built with training samples that have been recorded in a well-controlled environment.
The paper is organized as follows: Section II describes work related to this topic. Gaussian processes and their usage for modeling deformation priors are described in Section III-A and III-B. In Section III-C we propose a new kernel for face registration. The registration pipeline itself is explained in Section III-D. Section IV explains the different datasets that are used for this work. In Section V the quantitative and qualitative results of the BU3D-FE registration and a model adaption application of single 2D images are shown. At last, our conclusions are drawn in Section VI.
II Related Work
The iterative closest point algorithm (ICP) and its non-rigid extension (NICP), introduced by and are the most popular algorithms used for establishing the correspondence of 3D face shapes (,,,,). Extending the non-rigid ICP to a specific problem domain or data-set requires changes in point search heuristics and stiffness weights, which makes the method complicated to adapt in practice. Additional extensions the NICP algorithm have also been proposed: introduced independent local components for the NICP algorithm to handle the difficulty of facial expressions. In , local statistical models, trained from registered data, are embedded as constraints in the NICP algorithm. In addition to NICP, alternative approaches for face registration have been proposed: In , a registration algorithm with a B-spline based deformation model is shown. In , the authors propose an algorithm based on thin-plate-splines, which handles different levels of detail and mirror symmetry. propose to model facial expressions with mouth opening as isometric deformations on the face surface. handle the expression problem by fitting an expression model of blendshapes before the shape registration step. In case of model-building, the BFM is the most used Morphable Model in literature, and it was built on 200 neutral faces using NICP. Recently a large scale Morphable Model built from 10’000 faces has been proposed using NICP for registration ). Both those models lack facial expressions. The Surrey face model contains facial expressions, which are built from 6 blendshapes and provides multiple resolutions of their shape model . A statistical shape model (no color) was built on the BU-3DFE face database using a multilinear expression model . After registration and model-building, we demonstrate the applicability of the model with an inverse rendering application of 2D face images. Unlike the approach by , which is used in this work, most methods only recover shape but ignore color and illumination. An overview over current inverse rendering techniques is contained in . A recent publication presents an end-to-end learning of rendering and model adaptation incorporating a 3DMM .
III Method
To turn this conceptual problem into a practical one, we need to fix the likelihood function and find a strategy to optimize the problem. For the likelihood function we define the distance between a point and the target surface as
with as a loss function and as the closest point on surface to x:
Assuming independence of the errors at every vertex, we obtain the likelihood function:
This is a parametric optimization problem, which can be approached using standard optimization algorithms. For this work, an implementation of LBGFS was used.
III-B Combining Kernels
III-C A Shape Prior tailored for Face Registration
In this section we show how to build a deformation prior for face registration. As the reference surface we have chosen the mean shape of the Basel Face Model. It is therefore a good assumption to choose the mean deformation to be the zero function,
As the basis of the model, we chose the multi-scale B-spline kernel, introduced in . Given a univariate third order B-spline and the function , the kernel reads
with evaluated in the support of the B-spline. The multiple scales are defined as
with level from coarse to fine and multiplied by the identity matrix to get a matrix valued kernel. The value is the deformation scale per level and thus defines how far the correlating points can deform. It is chosen such that coarse scale levels are able to deform more than finer levels. This kernel defines smoothly varying function on multiple scale-levels. The individual scales can be decoupled as a superposition of different levels as shown Figure 1. All scale layers combined to the fully detailed registration result are shown next to the target shape in blue. The increasing level of scale from Level 1 until Level 4 is shown for comparison. Level 1 is defined as with and Level 4 as . While in Level 1 and Level 2 coarse details are globally adapted, skin details and eye shape are deformed with small scale deformations (Level 4).
III-C2 Spatially Varying Scales
Typical face shapes contain small scale variability around the eyes and mouth, but are rather smooth around the cheeks. Therefore we have divided the face into smooth regions and combine this information with the multi-scale B-spline kernel. This leads to a model with small scale deformations around the eyes and mouth region, while the cheeks are still restricted to smooth, large scale deformations (see Figure 2: Regions).
where are smooth indicator functions that determine if the kernel is active (i.e. ) at location for level and is a single scale B-spline kernel defined as in (11). In Figure 2, random samples of the described kernel are shown in comparison to samples from the standard non-varying kernel.
III-C3 Symmetry
where is the identity matrix and
Intuitively, this construction takes a definition of how the function values at two points on the surface are correlated. Then the correlations of the three components of the resulting deformation field are constructed by multiplying with the identity matrix . To achieve mirror-symmetry, the minus sign is introduced in the first component to ensure that correlation between two points, which are on the opposite side of the symmetry plane lead to the inverse correlations. The symmetry is integrated in the face model by combining symmetric and asymmetric deformations to make the face samples look more realistic. In Figure 3 a comparison between a normal and a symmetrized kernel is visualized.
III-C4 Core Expression Model
In facial expressions, the opening and closing of the mouth cannot be modelled simply with a smooth kernel, such as a B-spline or radial basis function. Since the points on the upper and lower lip are close, they correlate strongly, which hinders an opening deformation. One approach is the usage of a new reference with an open mouth. However, the registration with multiple templates is inconvenient in practice. The second row in Figure 4 visualizes the registration using an open mouth reference. Although it gives perfect results for open mouth registrations, the mouth does not close properly for neutral faces. To build a model that can cope with both situations, we combine a simple statistical shape model with the previously described prior model. To build this facial expression kernel we make use of the facial expression reference shapes (anger, disgust, fear, happy, sad, surprise) to compute the mean
The facial expression kernel can be combined with another kernel according to the rules in (9), which again results in a valid kernel function. For the face registration we use the kernel formulated in the previous sections in combination with the core model. The final deformation model is defined by symmetrizing the spatially-varying kernel
and augment the function with the facial expression kernel:
In Figure 4, bottom row, the results of a registration using are shown. All the test-cases, the closed, as well as the open mouth samples, have been accurately registered.
III-D The registration algorithm
So far we have described how the registration algorithm works in principle: We formulate a Gaussian process model as a prior and minimize (8). The steps are summarized in Algorithm 1. In the first step we make use of the provided landmark points in the registration. Gaussian process morphable models make it possible to include those landmarks directly into the prior by considering the deformation between a landmark pair , as a noisy observation of the true deformation , i.e.
and applying Gaussian regression to it, as described in . The resulting posterior distribution assigns a low probability to any deformation that does not match the specified landmarks (up to the specified uncertainty ). The posterior model is again a Gaussian process, and thus can be used instead of the original prior, without changing the algorithm. The registration problem (8) is optimized in different steps with decreasing regularization weights. In each step, all the points of the model for which the current fit is further away from the target surface than some predefined threshold or whose closest point is a boundary point (indicating a hole in the target surface) are eliminated from the optimization.
To describe the distance metric that has been used, we denote as defined in (4) and calculate the distance (3) with as the Huber loss function defined by
III-E Building the Morphable Model
To build a color model, the closest corresponding color value of the target mesh is extracted at all points on of the registered mesh. Since the target scans are often incomplete and contain holes, not every point in the registration can be assigned a color value. To address this issue we introduce a binary indicator variable to specify whether a reliable color at a point exists or not. We then compute the color mean using only the available colors
When estimating the covariance function for the color model we use an additional kernel to express our prior similar to the smoothness assumption for the shape surface.In practice, we use a single level square exponential kernel with a scaling of and a correlation of , where the units are millimeters. Based on and we use either the empirical covariance or the the covariance specified by the prior kernel. The full color covariance function handling missing data is then
with as the same term used in the empirical estimate:
III-E2 Expression Model
We extend the original face model to a multi-linear model to handle expressions as described in . The multi-linear statistical model consists of two independent models for face shape and face color as well as an additional model for the deformations by facial expression. Facial expression is modeled as a difference from the neutral face shape.
IV Data
The Binghamton University 3D Facial Expression Database (BU-3DFE) has neutral and expression scans of 100 individuals. Per individual it contains 6 facial expressions with 4 levels of strength. For the registration pipeline we use the raw data without cropping and a single expression strength (Level 4).
For the F3D data 83 detected landmarks are given. The RAW data has 5 landmarks. We used an ICP alignment to transfer the F3D landmarks onto the RAW data. Additionally, we clicked 23 landmarks for all neutral scans and expression scans of level 4 for the registration and used the F3D points for correspondence evaluation.
All scans have a texture file which is constructed from two pictures ( degrees), see also the bottom row of Figure 4. However, no ambient illumination was ensured, and many illumination effects (e.g., shadows on both sides of the nose or strong specular highlights) are visible. Additional disadvantages like make-up, facial hair or hair falling into the facial area do occur. We demonstrate the advantages of controlled data compared to the BU-3DFE data in Section V.
IV-B Basel Scans
We used the face scans introduced in which are scanned under a strictly controlled environment. For further details about the data we refer the reader to the original publication. In contrast to , we used an improved age distribution which includes more people over 40 years. The advantages of using the Basel scans for building a high quality morphable face model are:
Strict setting for scanning: No make-up, beards or hair in the facial area.
Number of scans: more individuals than in BU-3DFE.
Texture quality: Ambient illumination and the texture in high resolution and good quality.
Expressions: 6 types (anger, disgust, fear, happy, sad, surprise), controlled conditions.
In case of the Basel scans, also contour lines are available. We have included them in the registration by calculating a posterior model to the closest points on the line before every registration step, as mentioned in Algorithm 1.
For the new Basel Face Model (BFM-2017), a more representative data distribution compared to the original BFM (see also ) has been selected. Also, a facial expression model has been included, built from 160 examples. The data has been chosen the following:
160 expression examples, equally distributed on expression types.
In Figure 5 it is visible that the selected data is closer to the real age distribution (e.g., from the European Union). Moreover, the improved age distribution reflects the importance of people older than 40 years. This group has facial attributes like sacking and wrinkles, which young people (below 30) are mostly lacking off.
V Results
To provide a measure of the registration accuracy with the BU-3DFE database, we compare our registrations to the landmarks, which are provided with the BU-3DFE database. To evaluate an average distance error, the landmarks of the registrations are compared to the positions that are provided with the BU-3DFE dataset. In Table I, the average distance error per region is shown. We sorted the BU-3DFE landmarks as described in , to match their evaluation scheme. The proposed registration shows a similar correspondence as annotated in the BU-3DFE database and are on par with the evaluation in . The high standard deviation is due to the fact that all expressions are evaluated together. Expressions, such as anger and fear heavily affect the shape of the eye and eyebrows, which in turn has impacts on the standard deviation. To enable a better comparison in future work, we provide our manually clicked landmarks on the reference mesh together with the source code.
V-B Inverse Rendering
The original application of 3D Morphable Face Models proposed in is an inverse rendering task. Inverse rendering aims to estimate all necessary parameters of an image formation process to generate a given target image. The full model consists of the statistical shape, color and expression model, a pinhole camera model as well as spherical harmonics for illumination modeling (). To complete the framework, we include an implementation of a recent technique to estimate the parameters from a single still image. The framework we are implementing is a fully probabilistic model adaptation framework based on Markov chain Monte Carlo sampling. The face model is integrated as a prior on facial shape, expressions and color appearance into this model adaptation framework. Such a strong prior is necessary to be able to reconstruct the 3D shape from a single 2D image. We extended the model adaptation framework to handle facial expressions by including expression proposals like the ones for the shape and color coefficients as described in . This is the first publicly available implementation of a 3D Morphable Model adaptation framework in an Analysis-by-Synthesis setting including facial expressions.
We present results of our face model adaptation method on the Multi-PIE database . For our experiments we used the neutral and smiling photographs of 249 individuals in the first session in four poses (0∘ camera 051, 15∘ camera 140, 30∘ camera 130, 45∘ camera 080) under frontal illumination (illumination 16). We show the different poses and expressions together with there fitting results in Figure 6. We perform an unconstrained face recognition experiment over pose and expressions, see Table II. The face recognition results are competitive compared to state of the art inverse rendering techniques. Additionally we present qualitative results in a more realistic setting on the Labeled Faces in the Wild (LFW) database in Figure 7. For all fitting experiments we initialized the pose with 9 manually annotated landmarks.
VI Conclusion
As a first central contribution, we presented a non-rigid registration method for facial shapes based on Gaussian process registration. The framework cleanly separates domain-specific knowledge as modeled by a Gaussian process from the actual registration algorithm. We specifically demonstrated how to build a prior model for face registration by combining multiple deformation scales, symmetry and mouth opening for facial expressions using kernel modeling techniques. The pipeline has been made available open source together with the publicly available face database to reach full reproducibility based on open data. Part of the framework is a model-building pipeline, which enables the construction of a morphable model from registered data, and an inverse rendering software, which applies the built model to 2D images of faces. Furthermore, we release a new BFM-2017 based on high-quality shape and color data with facial expressions and an improved age distribution. With qualitative and quantitative evaluations, we compared the model performance for inverse rendering and face recognition. We showed that the new Basel face model outperforms the model built on the BU3D-FE dataset and also its predecessor from 2009 . With this work on face registration and the release of an open pipeline for registration, model-building and model fitting, we enable the community to reproduce and compare the results of neutral and facial expression registration, model-building and model-fitting. The pipeline code has been released on Githubhttps://github.com/unibas-gravis/basel-face-pipeline and the new BFM-2017 with expressions is available for download on our websitehttp://gravis.dmi.unibas.ch/pmm/.