When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition

Guosheng Hu, Yongxin Yang, Dong Yi, Josef Kittler, William Christmas, Stan Z. Li, Timothy Hospedales

Introduction

The conventional face recognition pipeline consists of four stages: face detection, face alignment, feature extraction (or face representation) and classification. Perhaps the single most important stage is feature extraction. In constrained environments, the hand-crafted features such as Local Binary Patterns (LBP) and Local Phase Quantisation (LPQ) have achieved respectable face recognition performance. However, the performance using these features degrades dramatically in unconstrained environments where face images cover complex and large intra-personal variations such as pose, illumination, expression and occlusion. It remains an open problem to find an ideal facial feature which is robust for face recognition in unconstrained environments (FRUE). In the last three years, convolutional neural network (CNN) rebranded as ‘deep learning’ has achieved very impressive results on FRUE. Unlike the traditional hand-crafted features, the CNN learning-based features are more robust to complex intra-personal variations. More notably, the top three face recognition rates reported on the FRUE benchmark database LFW (Labeled Faces in the Wild) have been achieved by CNN methods . The success of the latest CNNs on FRUE and more general object recognition task stems from the following facts: (1) much larger labeled training sets are available; (2) GPU implementations greatly reduce the time of training a large CNN; (3) CNNs greatly improve the model generation capacity by introducing effective regularisation strategies, such as dropout .

Despite the promising performance achieved by CNNs, it remains unclear how to design a ‘good’ CNN architecture to adapt to a specific classification task due to the lack of theoretical guidance. However, some insights into CNN design can be gained by experimental comparisons of different CNN architectures. The work made such comparisons and comprehensive analysis for the task of object recognition. However, face recognition is very different from object recognition. Specifically, faces are aligned via 2D similarity transformation or 3D pose correction to a fixed reference position in images before feature extraction while object recognition usually does not conduct such alignment, and therefore objects appear in arbitrary positions. As a result, the CNN architectures used for face recognition are rather different from those for object recognition . For the task of face recognition, it is important to make a systematic evaluation of the effect of different CNN design and implementation choices. In addition, those published CNNs are trained in different face databases, most of which are not publicly available. The difference of training sets might result in unfair comparisons of CNN architectures. To avoid this unfairness, the comparison of different CNNs should be conducted on a common ground.

To clarify the contributions of different components of CNN-based face recognition systems, in this paper, a systematic evaluation is conducted. To make our work reproducible, all the networks evaluated are trained on the publicly available LFW database. Specifically, our contributions are as follows:

Different CNN architectures including number of filters and layers are compared. In addition, we evaluate the impact of multiple network fusion introduced by .

Various implementation choices, such as data augmentation, pixel value type (colour or grey) and similarity, are evaluated.

We quantitatively analyse how downstream metric learning methods such as joint Bayesian can boost the effectiveness of the CNN-learned features.

Finally, source code for our CNN architectures and trained networks will be made publicly available (the training data is already public). This provides an extremely competitive baseline for face recognition to the community. To our knowledge, we are the first to publish fully reproducible CNNs for face recognition.

Related Work

CNN methods have drawn considerable attention in the field of face recognition in recent years. In particular, CNNs have achieved impressive results on FRUE. In this section, we briefly review these CNNs.

The researchers in Facebook AI group trained an 8-layer CNN named DeepFace . The first three layers are conventional convolution-pooling-convolution layers. The subsequent three layers are locally connected, followed by 2 fully connected layers. Pooling layers make learned features robust to local transformations but result in missing local texture details. Pooling layers are important for object recognition since the objects in images are not well aligned. However, face images are well aligned before training a CNN. It is claimed in that one pooling layer is a good balance between local transformation robustness and preserving texture details. DeepFace is trained on the largest face database to-date which contains four million facial images of 4,000 subjects. Another contribution of is the 3D face alignment. Traditionally, face images are aligned using 2D similarity transformation before they are fed into CNNs. However, this 2D alignment cannot handle out-of-plane rotations. To overcome this limitation, proposes a 3D alignment method using an affine camera model.

In , a CNN-based face representation, referred to as Deep hidden IDentity feature (DeepID), is proposed. Unlike DeepFace whose features are learned by one single big CNN, DeepID is learned by training a collection of small CNNs (network fusion). The input of one single CNN is the crops/patches of facial images and the features learned by all CNNs are concatenated to form a powerful feature. Both RGB and grey crops extracted around facial points are used to train the DeepID. The length of DeepID is 2 (RGB and Grey images) ×\times 60 (crops) ×\times 160 (feature length of one network) = 19,200. One small network consists of 4 convolutional layers, 3 max pooling layers and 2 fully connected layers shown in Table 1. DeepID uses identification information only to supervise the CNN training. In comparison, DeepID2 , an extension of DeepID, uses both identification and verification information to train a CNN, aiming to maximise the inter-class difference but minimise the intra-class variations. To further improve the performance of DeepID and DeepID2, DeepID2+ is proposed. DeepID2+ adds the supervision information to all the convolutional layers rather than the topmost layers like DeepID and DeepID2. In addition, DeepID2+ improves the number of filters of each layer and uses a much bigger training set than DeepID and DeepID2 . In , it is also discovered that DeepID2+ has three interesting properties: being sparse, selective and robust.

The work proposes another face recognition pipeline, refereed to as WebFace, which also learns the face representation using a CNN. WebFace collects a database which contains around 10,000 subjects and 500,000 images and makes this database publicly available. Motivated by very deep architectures of , WebFace trains a much deeper CNN than those used for face recognition as shown in Table 1. Specifically, WebFace trains a 17-layer CNN which includes 10 convolutional layers, 5 pooling layers and 2 fully connected layers detailed in Table 1. Note that the use of very small convolutional filters (3×\times3), which avoids too much texture information decrease along a very deep architecture, is crucial to learn a powerful feature. In addition, WebFace stacks two 3×\times3 convolutional layers (without pooling in between) which is as effective as a 5×\times5 convolutional layer but with fewer parameters.

Table 1 compares three typical CNNs (DeepFace , DeepID , WebFace ). It is clear that their architectures and implementation choices are rather different, which motivates our work. In this study, we make systematic evaluations to clarify the contributions of different components on a common ground.

Methodology

LFW is the de facto benchmark database for FRUE. Most exisiting CNNs train their networks on private databases and test the trained models on LFW. In comparison, we train our CNNs only using LFW data to make our work easily reproducible. In this way, we cannot directly use the reported CNN architectures since our training data is much less extensive. We introduce three architectures adapting to our training set in subsection 3.1. To further improve the discrimination of CNN-learned features, metric learning method is usually used. One metric learning method, Joint Bayesian model , is detailed in subsection 3.2.

How to design a ‘good’ CNN architecture remains an open problem. Generally, the architecture depends on the size of training data. Less data should drive a smaller network (fewer layers and filters) to avoid overfitting. In this study, the size of our training data is much smaller than that used by the state of the art methods ; therefore, smaller architectures are designed.

We propose three CNN architectures adapting to the size of training data in LFW. These architectures are of three different sizes: small (CNN-S), medium (CNN-M), and large (CNN-L). CNN-S and CNN-M have 3 convolutional layers and two fully connected layers, while CNN-M has more filters than CNN-S. Compared with CNN-S and CNN-M, CNN-L has 4 convolutional layers. The activation function we used is REctification Linear Unit (RELU) . In our experiments, dropout does not improve the preformance of our CNNs, therefore, it is not applied to our networks. Following , softmax function is used in the last layer for predicting a single class of K (the number of subjects in the context of face recognition) mutually exclusive classes. During training, the learning rate is set to 0.001 for three networks, and the batch size is fixed to 100. Table 2 details these three architectures.

2 Metric Learning

Metric Learning (MeL), which aims to find a new metric to make two classes more separable, is often used for face verification. MeL is independent of the feature extraction process and any feature (hand-crafted and learning-based) can be fed into a MeL method. Joint Bayesian (JB) model is a well-known MeL method and it is the most widely used MeL method which is applied to the features learned by CNNs .

JB models the face verification task as a Bayesian decision problem. Let HIH_{I} and HEH_{E} represent intra-personal (matched) and extra-personal (unmatched) hypotheses, respectively. Based on the MAP (Maximum a Posteriori) rule, the decision is made by:

where x1x_{1} and x2x_{2} are features of one face pair. It is assumed that P(x1,x2∣HI)P(x_{1},x_{2}\mid H_{I}) and P(x1,x2∣HE)P(x_{1},x_{2}\mid H_{E}) have Gaussian distributions N(0,SI)N(0,S_{I}) and N(0,SE)N(0,S_{E}), respectively.

Before discussing the way of computing SIS_{I} and SES_{E}, we first explain the distribution of a face feature. A face xx is modelled by the sum of two independent Gaussian variables (identity μ\mu and intra-personal variations ε\varepsilon):

μ\mu and ε\varepsilon follow two Gaussian distributions N(0,Sμ)N(0,S_{\mu}) and N(0,Sε)N(0,S_{\varepsilon}), respectively. SμS_{\mu} and SεS_{\varepsilon} are two unknown covariance matrices and they are regarded as face prior. For the case of two faces, the joint distribution of {x1,x2}\{x_{1},x_{2}\} is also assumed as a Gaussian with zero mean. Based on Eq. (2), the covariance of two faces is:

Then SIS_{I} and SES_{E} can be derived as:

Clearly, r(x1,x2)r(x_{1},x_{2}) in Eq. (1) only depends on SμS_{\mu} and SεS_{\varepsilon}, which are learned from data using an EM algorithm .

Evaluation

LFW contains 5,749 subjects and 13,233 images and the training and test sets are defined in . For evaluation, LFW is divided into 10 predefined splits for 10-fold cross validation. Each time nine of them are used for model training and the other one (600 image pairs) for testing. LFW defines three standard protocols (unsupervised, restricted and unrestricted) to evaluate face recognition performance. ‘Unrestricted’ protocol is applied here because the information of both subject identities and matched/unmatched labels is used in our system. The face recognition rate is evaluated by mean classification accuracy and standard error of the mean.

The images we used are aligned by deep funneling . Each image is cropped to 58×\times58 based on the coordiates of two eye centers. Some sample crops are visualised in Fig. 1. It is commonly believed that data augmentation can boost the generalisation capacity of a neural network; therefore, each image is horizontally flipped. The mean of the images is subtracted before network training. The open source implementation MatConvNet is used to train our CNNs. In this section, different components of our CNN-based face recognition system are evaluated and analysed.

Choosing a ‘good’ architecture is crucial for CNN training. Overlarge or extremely small networks relative to the training data can lead to overfitting or underfitting, in the case of which the network does not converge at all during training. In this comparison, the RGB colour images are fed into CNNs and feature distance is measured by cosine distance. The performances of the three architectures are compared in Table 3. CNN-M achieves the best face recognition performance, indicating that the CNN-M generalises best among these three architectures using only LFW data. From this point, all the other evaluations are conducted using CNN-M. The face recognition rate 0.7882 of CNN-M is considered as the baseline, and all the remaining investigations will be compared with it.

Feature Distance

The exisiting research offers little discussion about the distance measurement for CNN-learned features. In particular, it is interesting to know what is the best distance measure for face recognition. Table 4 compares the impact of six distance measures on face recognition accuracy. Cosine and correlation achieve the best recognition rates, however, the standard deviation of cosine is smaller than that of correlation. Therefore, cosine distance is the best among these distances.

Grey vs Colour

In and , CNNs are trained using grey-level and RGB colour images, respectively. In comparison, both grey and colour images are used in . We quantitatively compare the impact of these two images types on face recognition. Their comparative evaluation yields face recognition accuracies using grey and colour images of 0.7830±\pm0.0077 and 0.7882±\pm0.0118, respectively. The performances using grey and colour images are very close to each other. Although colour images contain more information, they do not deliver a significant improvement.

Data Augmentation

Flip, mirroring images horizontally producing two samples from each, is a commonly used data augmentation technique for face recognition. Both original and mirrored images are used for training in all our evaluations. However, little discussion in the existing work was made to analyse the impact of image flipping during testing. Naturally, the test images can also be mirrored. A pair of test images can produce 2 new mirrored ones. These 4 images can generate 4 pairs instead of one original pair. To combine these 4 images/pairs, two fusion strategies (feature and score fusion) are implemented in this work. For feature fusion, the learned features of a test image and its mirrored one are concatenated to one feature, which is then used for score computing. For score fusion, 4 scores generated from 4 pairs are averaged to one score. Table 5 compares the three scenarios: no flip during the test, feature and score fusions. As is shown in Table 5, mirroring images does improve the face recognition performance. In addition, feature fusion works slightly better than score fusion, however, the improvements are not statistically significant.

Learned Feature Analysis

It is interesting to investigate the properties of CNN-learned face representations. First, we discuss feature normalisation, which standardises the range of features and is generally performed during the data preprocessing step. For example, to implement eigenface , the features (pixel values) are usually normalised via Eq. (6) before training a PCA space.

where x∈R\mathbf{x}\in R and x^∈R\hat{\mathbf{x}}\in R are original and normalised feature vectors, respectively. μx\mu_{\mathbf{x}} and σx\sigma_{\mathbf{x}} are the mean and standard deviation of x\mathbf{x}. Motivated by this, our CNN features are normalised by Eq. (6) before computing cosine distance. The accuracies with and without normalisation are 0.7927±\pm0.0126 and 0.7882±\pm0.0118, respectively. Thus normalisation is effective to improve recognition rate.

Second, we perform dimensionality reduction on the learned 160D features using PCA. As shown in Figure 2, only 16 dimensions of the PCA feature space can achieve comparable face recognition rates to those of the original space. It is a very interesting property of CNN-learned features because low dimensionality can significantly reduce storage space and computation, which is crucial for large scale applications or mobile devices such as smartphone.

Network Fusion

The work DeepID and its variants apply the fusion of multiple networks. Specifically, the images of different facial regions and scales are separately fed into the networks that have the same architecture. The features learned from different networks are concatenated to a powerful face representation, which implicitly captures the spatial information of facial parts. The size of these images can be different as shown in Table 1. In , 120 networks are trained separately for this fusion. However, it is not very clear how greatly this fusion improves the face recognition performance. To clarify this issue, we implement the network fusion.

We extract d×dd\times d crops from four corners and center and then upsample them to the original image size 58×5858\times 58. The crops have 6 different scales: d=floor(58×{0.3,0.4,0.5,0.6,0.7,0.8})d=floor(58\times\{0.3,0.4,0.5,0.6,0.7,0.8\}), where floorfloor is the operator to get the integer part. Therefore we obtain 30 local patches with size of 58×5858\times 58 from one original image. Figure 3 shows these 30 crops. To evaluate the performance of network fusion, we separately train 30 different networks using these crops. Then one face image can be represented by concatenating the features learned from different networks. Table 6 compares the performance of single network and network fusion. Note that we choose 16 best networks of 30 ones for the fusion. It is clear that network fusion works much better than a single network. Specifically, the fusion of 16 best networks improves the face recognition accuracy of single network by 4.51%. Clearly, the face representation of network fusion is actually the fusion of features of different facial componets and scales. Similar ideas have widely been used to improve the facial representation capacity of hand-crafted features such as multi-scale local binary pattern , multi-scale local phase quantisation and high-dimensional local features .

Metric Learning

For metric learning, the features of the fusion of best 16 networks are used. The feature dimensionality (2560=160×\times16) is reduced to 320 via PCA before they are fed into JB. Figure 4 compares the face recognition accuracies with and without JB in each split of LFW database. JB consistently and significantly improves the face recognition rates, showing the importance of metric learning.

Table 7 compares our method with non-commercial state-of-the-art methods. The performance of our method is slightly better than but worse than . However, the feature dimensionality of is much higher than ours. In , a large number of new pairs are generated in addition to those provided by LFW to train the model, while we do not generate new pairs.

Conclusions

Recently, convolutional neural networks have attracted a lot of attention in the field of face recognition. In this work, we present a rigorous empirical evaluation of CNN-based face recognition systems. Specifically, we quantitatively evaluate the impact of different architectures and implementation choices of CNNs on face recognition performances on common ground. We have shown that network fusion can significantly improve the face recognition performance because different networks capture the information from different regions and scales to form a powerful face representation. In addition, metric learning such as Joint Bayesian method can improve the face recognition greatly.

Since network fusion and metric learning are the two most important factors affecting CNN performance, they will be the subject of future investigation.

This work is supported by the European Union’s Horizon 2020 research and innovation program under grant agreement No 640891, EPSRC/dstl project ‘Signal processing in a networked battlespace’ under contract EP/K014307/1, EPSRC Programme Grant ‘S3A: Future Spatial Audio for Immersive Listener Experiences at Home’ under contract EP/L000539, and the European Union project BEAT. We also gratefully acknowledge the support of NVIDIA Corporation for the donation of the GPUs used for this research.

References