Pose-Invariant Face Alignment with a Single CNN

Amin Jourabloo, Mao Ye, Xiaoming Liu, Liu Ren

Introduction

Face alignment, also known as face landmark detection, is an essential process for many facial analysis tasks, such as face recognition , expression estimation and 33D face reconstruction . During the last decade, face alignment technologies have been substantially improved . One recent advancement in this area is to tackle challenging cases with large face poses, e.g., frontal to profile views with ±90∘\pm 90^{\circ} yaw angles .

The dominant technology for large-pose face alignment (LPFA) utilizes a cascade of regressors which combines different types of regression designs with feature extraction methods . At each stage of this procedure, the target parameters, e.g., 22D landmarks or the head pose and 33D face shape, are refined by regressing an update of these parameters. Due to the proven power of Convolutional Neural Network (CNN) in vision tasks, it is also adopted as the regressor in this framework and has achieved the state-of-the-art performance on face alignment .

Despite the recent success, the cascade of CNNs, when applied to LPFA, suffers from the following drawbacks.

Lack of end-to-end training: It is a consensus that end-to-end training is desired for CNN . However, one CNN regressor is trained independently at each cascade stage. Sometimes even multiple CNNs are applied independently at each stage. E.g., locations of different landmark sets are estimated by various CNNs and combined by a separate fusing module . Therefore, these CNNs can not be optimized jointly and might lead to a sub-optimal solution.

Hand-crafted feature extraction: Since the CNNs are trained independently, feature extraction is required to utilize the result of previous CNN and provide input to the current CNN. Simple feature extraction methods are used, e.g., extracting patches based on 22D or 33D face shapes without considering other factors including pose and expression. Normally, the cascade of CNNs is a collection of shallow CNNs where each one has less than five layers. Hence, this framework can not extract deep features by building upon the extracted features of early-stage CNNs.

Slow training speed: Training a cascade of CNNs is usually time-consuming for two reasons. Firstly, the CNNs are trained sequentially, one after another. Secondly, feature extraction is required between two consecutive CNNs.

To address these issues, as shown in Fig. 1, we introduce a novel layer, named the visualization layer, into a CNN architecture, for the LPFA problem. Our CNN architecture consists of several blocks, which are called visualization blocks. This architecture can be considered as a cascade of shallow CNNs. The new layer visualizes the alignment result of the previous visualization block and utilizes it in the current block. It is designed based on several guidelines. Firstly, it is derived from the surface normals of the underlying 33D face model and encodes the relative pose between the face and camera. The use of surface normals is partially inspired by the success of adopting surface normals for 33D face recognition . Secondly, the visualization layer is differentiable, which allows the gradient to be computed analytically, enabling end-to-end training. Lastly, a mask is utilized to differentiate between pixels in the middle and contour parts of a face, and to also make the pixel values of the visualized images similar across various poses.

Benefiting from the design of the visualization layer, our method has the following advantages and contributions:

⋄\diamond The proposed method allows a block in the CNN to utilize the extracted features from previous blocks and extract deeper features. Therefore, extraction of hand-crafted features is no longer necessary.

⋄\diamond The visualization layer is differentiable, allowing for backpropagation of an error from a later block to an earlier one. To the best of our knowledge, this is the first method for large-pose face alignment, that utilizes only one single CNN and allows end-to-end training.

⋄\diamond The proposed method converges faster during the training phase compared to the cascade of CNNs. Therefore, the training time is dramatically reduced.

The source code of the proposed method with the trained model are released at here.

Prior Work

This section reviews the relevant prior work in three topics: cascade of regressors for face alignment, convolutional recurrent neural network and visualization in deep learning.

Cascade of Regressors for Face Alignment Cascade of Regressors is a classic approach in not only conventional face alignment , but also the large-pose face alignment . To handle large poses, many approaches go beyond 22D landmarks and also estimate 33D landmarks and 33D face shapes . Zhu et al. use a set of local regressors to estimate the 22D shape update, and fuse their results with another regressor. The occlusion-invariant approach of RCPR is applicable to large poses since self-occlusion is one type of occlusions. An iterative probabilistic method is utilized in for registering 33D shape to the pre-computed 22D landmarks. Tulyakov et al. also use a cascade of regressors to estimate 33D landmark updates directly from a single image. Some even use two regressors at each cascade stage. Wu et al. use one regressor to estimate the 22D shape update and the other to estimate the visibility of each landmark. Similarly, Liu et al. employ one regressor for 22D shape update and the other uses the 22D shape to estimate the 33D face shape.

Among methods with cascade of regressors, CNN is a popular choice of regressors due to its strong learning ability. These methods typically extract hand-crafted features between consecutive regressors. TCDCN use one CNN to estimate five landmarks, with yaw angles within ±60∘\pm 60^{\circ}. A cascade of stacked autoencoder (SAE) progressively estimates 22D landmark updates from extracted patches . Similarly, cascades of CNNs with global or local patches are combined at each stage, and their results are fused via averaging . The methods in combine cascade of CNNs with 33D feature extraction to estimate the dense 33D face shape. All aforementioned methods lack the ability to end-to-end train the network, which is our novel contribution to large-pose face alignment.

Convolutional Recurrent Neural Network (CRNN) The face alignment methods based on the CRNNs are the first attempts to combine cascade of regressors with joint optimization, for aligning mostly frontal faces. Their convolutional part extracts features from the whole image or from the patches at the landmark locations . The recurrent part facilitates the joint optimization by sharing information among all regressors. The main differences between the proposed method and CRNNs are: 1) existing CRNN methods are designed for near-frontal face alignment, while ours is for LPFA; 2) the CRNN methods share the same CNN at all stages, while our CNN of each block is different which might be more suitable for estimating the coarse-to-fine mappings during the course of alignment; 3) due to our new differentiable visualization layer, our method has one additional flow of the gradient back-propagation (note the two blue arrows between consectutive blocks in Fig. 2).

Visualization in Deep Learning Visualization techniques have been used in deep learning to assist in making a relative comparison among the input data and focusing on the region of interest. These methods can be categorized in two groups. The first exploits the deconvolutional and upsampling layers to either expand response maps or represent estimated parameters . Alternatively, various types of feature maps, e.g., heatmaps and Z-Buffering, can represent the current estimation of landmarks and parameters. In , 22D landmark heatmaps represent the landmarks’ locations. proposes a two step large pose alignment based on heatmaps to make more precise estimations. The heatmaps suffer from three drawbacks: 1) lack of the capability to represent objects in details; 2) requirement of one heatmap per landmark due to its weak representation power. 3) they cannot estimate visibility of landmarks. The Z-Buffer rendered using the estimated 33D face is fed to the CNNs to convey the results of a previous CNN to the next one. However, the Z-Buffer representation is not differentiable, and hence does not allow end-to-end training. In contrast, our visualization layer is differentiable and encodes the face geometry details via surface normals. It guides the CNN to focus on the face area that incorporates both the pose and expression information.

Proposed Method

Given a single face image with an arbitrary pose, our goal is to estimate the 22D landmarks with their visibility labels by fitting a 33D face model. Towards this end, we propose a CNN architecture with end-to-end training for model fitting, as shown in Fig. 2. In this section, we will first describe the underlying 33D face model used in this work, followed by our CNN architecture and the visualization layer.

We use the 33D Morphable Model (33DMM) for representing the 33D shape of a face. 33DMM represents a 33D face Sp\textbf{S}_{p} as a linear combination of mean shape S0\textbf{S}_{0}, identity bases SI\textbf{S}^{I} and expression bases SE\textbf{S}^{E} as follows:

We use vector p=[pI,pE]\textbf{p}=[\textbf{p}^{I},\textbf{p}^{E}] to indicate the 33D shape parameters, where pI=[p0I,⋯ ,pNII]\textbf{p}^{I}=[p_{0}^{I},\cdots,p_{N_{I}}^{I}] are the identity parameters and pE=[p0E,⋯ ,pNEE]\textbf{p}^{E}=[p_{0}^{E},\cdots,p_{N_{E}}^{E}] are the expression parameters. We use the Basel 33D face model , which has 199199 bases, as our identity bases and the face wearhouse model with 2929 bases as our expression bases. Each 33D face shape consists of a set of QQ 33D vertexes:

The 22D face shapes are the projection of 33D shapes. In this work, we use the weak perspective projection model with 66 degrees of freedoms, i.e., one for scale, three for rotation angles and two for translations, which projects the 33D face shape Sp\textbf{S}_{p} onto 22D images to obtain the 22D shape U:

Here U collects a set of NN 22D landmarks, M is the camera projection matrix, with misuse of notation P={M,p}\textbf{P}=\{\textbf{M},\textbf{p}\}, and the NN-dim vector b includes 33D vertex indexes which are semantically corresponding to 22D landmarks. We denote m1=[m1  m2  m3]\textbf{m}_{1}=[m_{1}\;m_{2}\;m_{3}] and m2=[m5  m6  m7]\textbf{m}_{2}=[m_{5}\;m_{6}\;m_{7}] as the first two rows of the scaled rotation component, while m4m_{4} and m8m_{8} are the translations.

Eqn. 3 establishs the relationship, or equivalency, between 22D landmarks U and P, i.e., 33D shape parameters p and the camera projection matrix M. Given that almost all the training images for face alignment have only 22D labels, i.e., U, we preform a data augmentation step similar to to compute their corresponding P. Given an input image, our goal is to estimate the parameter P, based on which the 22D landmarks and their visibilities can be naturally derived.

2 Proposed CNN Architecture

Our CNN architecture resembles the cascade of CNNs, while each “shallow CNN” is defined as a visualization block. Inside each block, a visualization layer based on the latest parameter estimation serves as a bridge between consecutive blocks. This design enables us to address the drawbacks of typical cascade of regressors in Sec. 1. We now describe the visualization block and CNN architecture, and dive into the details of the visualization layer in Sec. 3.3.

Visualization Block Fig. 3 shows the structure of our visualization block. The visualization layer generates a feature map based on the current estimated, or input, parameter P, and will be described in Sect. 3.3. Each convolutional layer is followed by a batch normalization (BN) layer and a ReLU layer, it extracts deeper features based on the input features provided by the previous visualization block and visualization layer output. Between the two fully connected layers, the first one is followed by a ReLU layer and a dropout layer, while the second one simultaneously estimates the update of M and p, ΔP\Delta\textbf{P}. The outputs of the visualization block are deeper features and the new estimation of the parameters, when adding ΔP\Delta\textbf{P} to the input P. As in Fig. 3, basically the top part of the visualization block focuses on learning deeper features, while the bottom part utilizes such features to estimate the parameters in a ResNet-like structure . During the backward pass of the training phase, the visualization block backpropagates the loss through both of its inputs to adjust the convolutional and fully connected layers in the previous blocks. This allows the block to extract better features that are suitable for the next block and improve the overall parameter estimation.

CNN Architecture The proposed CNN architecture consists of several connected visualization blocks as shown in Fig. 2. The inputs include the image and an initial estimation of the parameter P0\textbf{P}^{0}; and the output is the final estimation of the parameters. Compared to the typical cascade of CNNs, due to the joint optimization of all visualization blocks with backpropagation of the loss functions, the proposed architecture is able to converge in substantially fewer epochs during training.

Loss Functions Two types of loss functions are employed in our CNN architecture. The first one is an Euclidean loss between the estimation and the target of the parameter update, with each parameter weighted separately:

where EPi\textbf{E}_{P}^{i} is the loss, ΔPi\Delta\textbf{P}^{i} is the estimation and ΔPˉi\Delta\bar{\textbf{P}}^{i} is the target (or ground truth) at the ii-th visualization block. The diagonal matrix W\mathbf{W} contains the weights. For each element of the shape parameter p, its weight is the inverse of the standard deviation that was obtained from the data used in 33DMM training. To compensate the relative scale among the parameters of M, we compute the ratio rr between the average of scaled rotation parameters and average of translation parameters in the training data. We set the weights of the scaled rotation parameters of M to 1r\frac{1}{r} and the weights of the translation of M to 11. The second type of loss function is the Euclidean loss on the resultant 22D landmarks:

where Uˉ\bar{\textbf{U}} is the ground truth 22D landmarks, and Pi\textbf{P}^{i} is the input parameter to the ii-th block, i.e., the output of the i−1i-1-th block. f(⋅)f(\cdotp) computes 22D landmark locations using the currently updated parameters via Eqn. 3. For backpropagation of this loss function to the parameter ΔP\Delta\textbf{P}, we use the chain rule to compute the gradient (see supplemental material for the detailed derivation).

For the first three visualization blocks, the Euclidean loss on the parameter updates (Eqn. 6) is used, while the Euclidean loss on 22D landmarks (Eqn. 7) is applied to the last three blocks. The first three blocks estimate parameters to align 33D shape to the face image roughly and the last three blocks leverage the good initialization to estimate the parameters and the 22D landmark locations more precisely.

3 Visualization Layer

Several visualization techniques have been explored for facial analysis. In particular, Z-Buffering, which is widely used in prior works , is a simple and fast 22D representation for the 33D shape. However, this representation is not differentiable. In contrast, our visualization is based on surface normals of the 33D face, which describes surface’s orientation in a local neighbourhoods. It has been successfully utilized for different facial analysis tasks, e.g., 33D face reconstruction and 33D face recognition .

In this work, we use the zz coordinate of surface normals of each vertex, transformed with the pose. It is an indicator of “frontability” of a vertex, i.e., the amount that the surface normal is pointing towards the camera. This quantity is used to assign an intensity value at its projected 22D location to construct the visualization image. The frontability measure g, a QQ-dim vector, can be computated as,

where ×\times is the cross product, and ∥.∥\|.\| denotes the L2L_{2} norm. The 3×Q3\times Q matrix N0{\textbf{N}}_{0} is the surface normal vectors of a 33D face shape. To avoid the high computational cost of computing the surface normals after each shape update, we approximate N0{\textbf{N}}_{0} as the surface normals of the mean 33D face. Note that both the face shape and pose are still continuously updated across various visualization blocks, and are used to determine the projected 22D location. Hence, this approximation would only slightly affect the intensity value. To transform the surface normal based on the pose, we apply the estimation of the scaled rotation matrix (m1\textbf{m}_{1} and m2\textbf{m}_{2}) to the surface normals computed from the mean face. The value is then truncated with the lower bound of (Eqn. 8).

The pixel intensity of a visualized image V(u,v)\textbf{V}(u,v) is computed as the weighted average of the frontability measures within a local neighbourhood:

a is a QQ-dim mask vector with positive values for vertexes in the middle area of the face and negative values for vertexes around the contour area of the face:

where (xn,yn,zn)(x^{n},y^{n},z^{n}) is the vertex coordinate of the nose tip. a is pre-computed and normalized for zero-mean and unit standard deviation. The mask is utilized to discriminate between the central and boundary areas of the face, as well as to increase similarity across visualization of different faces. A visualization of the mask is provided in Fig. 4.

Since the human face is a 33D object, visualizing it at an arbitrary view angle requires the estimation of the visibility of each 33D vertex. To avoid the computationally expensive visibility test via rendering, we adopt two strategies for approximation. Firstly, we prune the vertexes whose frontability measures g equal , i.e., the vertexes pointing against the camera. Secondly, if multiple vertexes projects to a same image pixel, we keep only the one with the smallest depth values. An example is illustrated in Fig. 5.

Backpropagation To allow backpropagation of the loss functions through the visualization layer, we compute the derivative of V with respect to the elements of the parameters M and p. Firstly, we compute the partial derivatives, ∂g∂mk\frac{\partial\textbf{g}}{\partial m_{k}}, ∂w(u,v,xit,yit)∂mk\frac{\partial w(u,v,x_{i}^{t},y_{i}^{t})}{\partial m_{k}} and ∂w(u,v,xit,yit)∂pj\frac{\partial w(u,v,x_{i}^{t},y_{i}^{t})}{\partial p_{j}}, then the derivatives of ∂V∂mk\frac{\partial\textbf{V}}{\partial m_{k}} and ∂V∂pj\frac{\partial\textbf{V}}{\partial p_{j}} can be computed based on Eqn. 9 (the details are provided in the supplemental material).

Experimental Results

We evaluate our proposed method on two challenging LPFA datasets, namely AFLW and AFW, both qualitatively and quantitatively, as well as the near-frontal face dataset of 300300W. Further, we conduct experiments on different CNN architectures to validate our visualization layer design.

Implementation details Our implementation is built upon the Caffe toolbox . In all of the experiments, we use six visualization blocks (NvN_{v}) with two convolutional layers (NcN_{c}) and fully connected layers in each block (Fig. 3). Details of the network structure are provided in Tab. 1.

Instead of using the sequentially pretrain strategy , we perform the joint end-to-end training from scratch. To better estimate the parameter update in each block and to increase the effectiveness of using visualization block, we set the weight of the loss function in the first visualization block to 11, and linearly increase the weights by one for each block, i.e., the loss weight of the last block is 66. This strategy helps the CNN to pay more attention to the landmark loss used in later blocks. On the one hand, backpropagation of loss functions in the last blocks has more impact in the first block, and on the other hand the last block can adopt itself more quickly to the changes in the first block.

The AFLW dataset is a very challenging dataset with large-pose face images (±90∘\pm 90^{\circ} yaw). We use the subset of this dataset released by , which includes 3,9013,901 training images and 1,2991,299 testing images. All face images in this subset are labeled with 3434 landmarks and a bounding box. The AFW dataset contains 205205 images with 468468 faces. Each face image is labeled with at most 66 landmarks with visibility labels, as well as a bounding box. AFW is used only for testing in our experiments. The bounding boxes in both datasets are used as initilization for our algorithm, as well as the baselines. We crop the face image inside the bounding box and normalize it to 114×114114\times 114. Due to the memory constraint of GPUs, we have a pooling layer in the first visualization block after the first convolutional layer to decrease the size of feature maps to half, and the input to the subsequent visualization blocks is of 57×5757\times 57. To augment the training data, we generate 2020 different variations for each training image by adding noise to the location, width and height of the provided bounding boxes.

For quantitative evaluations, we use two conventional metrics. The first one is Mean Average Pixel Error (MAPE) , which is the average of the pixel errors for the visible landmarks. The other one is Normalized Mean Error (NME), i.e., the average of the normalized estimation error of visible landmarks. The normalization factor is the square root of the face bounding box size , instead of the eye-to-eye distance in the frontal-view face alignment.

We compare our method with several state-of-the-art methods in LPFA. For AFLW, we compare with LPFA , PIFA and RCPR with the NME metric. Tab. 2 shows that the proposed method achieved a higher accuracy than the baseline methods. Also, CALE , a heatmap-based 22D face alignment method, reports NME of 2.96%2.96\% on the AFLW. We discuss about the advantage of the proposed method over heatmap-based methods in section 2. To demonstrate the capabilities of each visualization block, the NME computed using the estimated P\mathbf{P} after each block is shown in Tab. 3. If a higher alignment speed is desirable, it is possible to skip the last two visualization blocks with a reasonable NME.

On the AFW dataset, the comparisons are conducted with LPFA , PIFA , CDM and TSPM with the MAPE metric. The evaluations are provided in Tab. 4, which also shows the superiority of the proposed method.

Some examples of alignment results of the proposed method on AFLW and AFW datasets are shown in Fig. 9. Three examples of visualization layer output at each visualization block are shown in Fig. 10.

2 Evaluation on 300W dataset

While our main goal is LPFA, we further evaluate on the most widely used near frontal 300300W dataset . 300300W containes 3,1483,148 training and 689689 testing images, which are divide into common and challenging sets with 554554 and 135135 images, respectively. Tab. 5 shows the NME (normalized by the interocular distance) of the proposed and state-of-the-art methods. The most related method to ours is 33DDFA , which also estimates M and p. Our method outperforms it on both common and challenging sets. Other near frontal alignment methods do not employ shape constraints e.g., 33DMM which is an advantage for them. Because the span of the 33D shape bases cannot cover all possible locations of landmarks. To comapre with the MDM , we compute the failure rate with threshold of 0.080.08. The failure rates of our method are 16.83%16.83\% (6.80%6.80\% for MDM) and 8.99%8.99\% (4.20%4.20\% for MDM) with 6868 and 5151 landmarks.

3 Analysis of the Visualization Layer

We perform four sets of experiments to study the properties of the visualization layer and network architectures.

Influence of visualization layers To analyze the influence of the visualization layer in the testing phase, we add 5%5\% noise to the fully connected layer parameters of each visualization block, and compute the alignment error on the AFLW test set. The NMEs are [4.464.46, 4.534.53, 4.604.60, 4.664.66, 4.804.80, 5.165.16] when each block is modified seperately. This analysis shows that visualized image has more influence on the later blocks, since imprecise parameters of early blocks could be compensated by later blocks. To evaluate the influence of the visualization layer in the training phase, we train the network without any visualization layer. The final NME on AFLW is 7.18%7.18\% which shows the importance of visualization layers for guiding the network training.

Advantage of deeper features We train three CNN architectures (Fig. 6) on AFLW. The inputs of the visualization block in the first architecture are the input images I, feature maps F and the visualization image V. The inputs of the second and the third architectures are {F,V}\{\textbf{F},\textbf{V}\} and {I,V}\{\textbf{I},\textbf{V}\}, respectively. The NME of each architecture is shown in Tab. 6. While the first one performs the best, the substantial lower performance of the third one demonstrates the importance of deeper features learned across blocks.

At the first convolutional layer of each visualization block, we compute the average of the filter weights, across both the kernel size and number of maps. The averages for three types of input features are shown in Fig. 7. As can be observed, from the first to the sixth block, the weights continue to decrease, making a more precise estimation of small-scale parameter updates. Considering the number of filters in Tab. 1, the total impact of feature maps are higher than the other two inputs in all blocks. This again shows the importance of deeper features in guiding the network to estimate parameters. Furthermore, the average of the visualization filter is higher than that of the input image filter, which validates the stronger influence of the proposed visualization during training.

Advantage of using masks To show the advantage of using the mask in the visualization layer, we conduct an experiment with different masks. Specifically, we define another mask for comparison, which is shown in Fig. 8. It has five positive areas, i.e., the eyes, nose tip and two lip corners. The values are normalized to zero-mean and unit standard deviation. Compared to the original mask in Fig. 4, this mask is more complicated and conveys more information about the informative facial areas to the network. Moreover, to show the necessity of using the mask, we also test using visualization layers without any mask. The NMEs of the trained networks with different masks are shown in Tab. 7. Comparing the first and third columns shows the advantage of using the mask in the network. The mask makes the pixel value of visualized images to be similar for faces with different poses and discriminate between the middle-area and contour-area of the face. By comparing the first and second columns, we can see that utilizing more complicated mask does not further improve the result, meaning the original mask provides sufficient information for its purpose.

Different numbers of blocks and layers Given the total number of 1212 convolutional layers in our network, we can partition them to visualization blocks in various sizes. To compare their performance, we train two additional CNNs, one with 44 visualization blocks and each with 33 convolutional layers; and the other with 33 block and 44 convolutional layers per block, where all three architectures have 1212 total convolutional layers. The NME of these architectures are shown in Tab. 8. It shows the same conclusion as in that the number of regressors is important for face alignment and we can potentially achieve a higher accuracy by increasing the number of visualization blocks.

4 Time complexity

Compared to the cascade of CNNs, one of the main advantages of end-to-end training a single CNN is the reduced training time. The training of the proposed method needs 3333 epochs and takes around 2.52.5 days. The state of the art , that uses the same train and test sets as ours, trains six CNNs and each needs 7070 epochs. The total time of is around 77 days. Similarly, the method in needs around 1212 days to train three CNNs each one with 2020 epochs, despite using different training data. Compared to , the proposed method reduces the training time by more than half. The testing speed of proposed method is 4.34.3 FPS on a Titan X GPU. It is much faster than the 0.60.6 FPS speed of and is simalar to 44 FPS speed of .

Conclusions

We propose a large-pose face alignment method with end-to-end training in a single CNN. We present a differentiable visualization layer, which is integrated to the network and enables joint optimization by backpropagating the error from a later visualization blocks to early ones. It allows the visualization block to utilize the extracted features from previous blocks and extract deeper features, without extracting hand-crafted features. Also, the proposed method converges faster during the training phase compare to the cascade of CNNs. Finally, we demonstrate the superior results of the proposed method over the state-of-the-art methods.

References