AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations
Xiao Zhang, Rui Zhao, Yu Qiao, Xiaogang Wang, Hongsheng Li
Introduction
Recent years witnessed the breakthrough of deep Convolutional Neural Networks (CNNs) on significantly improving the performance of one-to-one face verification and one-to-many face identification tasks. The successes of deep face CNNs can be mainly credited to three factors: enormous training data , deep neural network architectures and effective loss functions . Modern face datasets, such as LFW , CASIA-WebFace , MS1M and MegaFace , contain huge number of identities which enable the training of deep networks. A number of recent studies, such as DeepFace , DeepID2 , DeepID3 , VGGFace and FaceNet , demonstrated that properly designed network architectures also lead to improved performance.
Apart from the large-scale training data and deep structures, training losses also play key roles in learning accurate face recognition models . Unlike image classification tasks, face recognition is essentially an open set recognition problem, where the testing categories (identities) are generally different from those used in training. To handle this challenge, most deep learning based face recognition approaches utilize CNNs to extract feature representations from facial images, and adopt a metric (usually the cosine distance) to estimate the similarities between pairs of faces during inference.
However, such inference evaluation metric is not well considered in the methods with softmax cross-entropy loss functionWe denote it as “softmax loss” for short in the remaining sections., which train the networks with the softmax loss but perform inference using cosine-similarities. To mitigate the gap between training and testing, recent works directly optimized cosine-based softmax losses. Moreover, angular margin-based terms are usually integrated into cosine-based losses to maximize the angular margins between different identities. These methods improve the face recognition performance in the open-set setup. In spite of their successes, the training processes of cosine-based losses (and their variants introducing margins) are usually tricky and unstable. The convergence and performance highly depend on the hyperparameter settings of loss, which are determined empirically through large amount of trials. In addition, subtle changes of these hyperparameters may fail the entire training process.
In this paper, we investigate state-of-the-art cosine-based softmax losses , especially those aiming at maximizing angular margins, to understand how they provide supervisions for training deep neural networks. Each of the functions generally includes several hyperprameters, which have substantial impact on the final performance and are usually difficult to tune. One has to repeat training with different settings for multiple times to achieve optimal performance. Our analysis shows that different hyperparameters in those cosine-based losses actually have similar effects on controlling the samples’ predicted class probabilities. Improper hyperparameter settings cause the loss functions to provide insufficient supervisions for optimizing networks.
Based on the above observation, we propose an adaptive cosine-based loss function, AdaCos, which automatically tunes hyperparameters and generates more effective supervisions during training. The proposed AdaCos dynamically scales the cosine similarities between training samples and corresponding class center vectors (the fully-connection vector before softmax), making their predicted class probability meets the semantic meaning of these cosine similarities. Furthermore, AdaCos can be easily implemented using built-in functions from prevailing deep learning libraries . The proposed AdaCos loss leads to faster and more stable convergence for training without introducing additional computational overhead.
To demonstrate the effectiveness of the proposed AdaCos loss function, we evaluated it on several face benchmarks, including LFW face verification , MegaFace one-million identification and IJB-C . Our method outperforms state-of-the-art cosine-based losses on all these benchmarks.
Related Works
Cosine similarities for inference. For learning deep face representations, feature-normalized losses are commonly adopted to enhance the recognition accuracy. Coco loss and NormFace studied the effect of normalization and proposed two strategies by reformulating softmax loss and metric learning. Similarly, Ranjan et al. in also discussed this problem and applied normalization on learned feature vectors to restrict them lying on a hypersphere. Movrever, compared with these hard normalization, ring loss came up with a soft feature normalization approach with convex formulations.
Margin-based softmax loss. Earlier, most face recognition approaches utilized metric-targeted loss functions, such as triplet and contrastive loss , which utilize Euclidean distances to measure similarities between features. Taking advantages of these works, center loss and range loss were proposed to reduce intra-class variations via minimizing distances within each class . Following this, researchers found that constraining margin in Euclidean space is insufficient to achieve optimal generalization. Then angular-margin based loss functions were proposed to tackle the problem. Angular constraints were integrated into the softmax loss function to improve the learned face representation by L-softmax and A-softmax . CosFace , AM-softmax and ArcFace directly maximized angular margins and employed simpler and more intuitive loss functions compared with aforementioned methods.
Automatic hyperparameter tuning. The performance of an algorithm highly depends on hyperparameter settings. Grid and random search are the most widely used strategies. For more automatic tuning, sequential model-based global optimization is the mainstream choice. Typically, it performs inference with several hyperparameters settings, and chooses setting for the next round of testing based on the inference results. Bayesian optimization and tree-structured parzen estimator approach are two famous sequential model-based methods. However, these algorithms essentially run multiple trials to predict the optimized hyperparameter settings.
Investigation of hyperparameters in cosine-based softmax losses
In recent years, state-of-the-art cosine-based softmax losses, including L2-softmax , CosFace , ArcFace , significantly improve the performance of deep face recognition. However, the final performances of those losses are substantially affected by their hyperparameters settings, which are generally difficult to tune and require multiple trials in practice. We analyze two most important hyperparameters, the scaling parameter and the margin parameter , in cosine-based losses. Specially, we deeply study their effects on the prediction probabilities after softmax, which serves as supervision signals for updating entire neural network.
Let denote the deep representation (feature) of the -th face image of the current mini-batch with size , and be the corresponding label. The predicted classification probability of all samples in the mini-batch can be estimated by the softmax function as
where is logit used as the input of softmax, represents its softmax-normalized probability of assigning to class , and is the number of classes. The cross-entropy loss associated with current mini-batch is
Conventional softmax loss and state-of-the-art cosine-based softmax losses calculate the logits in different ways. In conventional softmax loss, logits are obtained as the inner product between feature and the -th class weights as . In the cosine-based softmax losses , cosine similarity is calculated by . The logits are calculated as , where is a scale hyperparameter. To enforce angular margin on the representations, ArcFace modified the loss to the form
Intuitively, on one hand, the parameter scales up the narrow range of cosine distances, making the logits more discriminative. On the other hand, the parameter enlarges the margin between different classes to enhance classification ability. These hyperparameters eventually affect . Empirically, an ideal hyperparameter setting should help to satisfy the following two properties: (1) Predicted probabilities of each class (identity) should span to the range $P_{i,y_{i}}1P_{i,y_{i}}\theta_{i,y_{i}}$ to make training effective.
The scale parameter can significantly affect . Intuitively, should gradually increase from to as the angle decreases from to Mathematically, can be any value in . We empirically found, however, the maximum is always around . See the red curve in Fig. 1 for examples., i.e., the smaller the angle between and its corresponding class weight is, the larger the probability should be. Both improper probability range and probability curves w.r.t. would negatively affect the training process and thus the recognition performance.
We first study the range of classification probability . Given scale parameter , the range of probabilities in all cosine-based softmax losses is
where the lower boundary is achieved when and for all in Eq. (1). Similarly, the upper bound is achieved when and for all . The range of approaches 1 when , i.e.,
which means that the requirement of the range spanning $s{0.94}C=10s=5.0$. But an oversized scale would lead to poor probability distribution, as will be discussed in the following paragraphs.
We investigate the influences of parameter by taking as a function of and angle where denotes the label of . Formally, we have
where are the logits summation of all non-corresponding classes for feature . We observe that the values of are almost unchanged during the training process. This is because the angles for non-corresponding classes always stay around during training (see red curve in Fig. 1).
Therefore, we can assume is constant, i.e., . We then plot curves of probabilities w.r.t. under different setting of parameter in Fig. 2(a). It is obvious that when is too small (e.g., for class/identity number and ), the maximal value of could not reach . This is undesirable because even when the network is very confident on a sample ’s corresponding class label , e.g. , the loss function would still penalize the classification results and update the network.
On the other hand, when is too large (e.g., ), the probability curve w.r.t. is also problematic. It would output a very high probability even when is close to , which means that the loss function with large may fail to penalize mis-classified samples and cannot effectively update the networks to correct mistakes.
In summary, the scaling parameter has substantial influences to the range as well as the curves of the probabilities , which are crucial for effectively training the deep network.
2 Effects of the margin parameter m𝑚m
In this section, we investigate the effect of margin parameters in cosine-based softmax losses (Eqs. (3) & (4)), and their effects on feature ’s predicted class probability . For simplicity, we here study the margin parameter for ArcFace (Eq. 3); while the similar conclusions also apply to the parameter in CosFace (Eq. (4)).
We first re-write classification probability following Eq. (7) as
To study the influence of parameter on the probability , we assume both and are fixed. Following the discussion in Section 3.1, we set , and fix . The probability curves w.r.t. under different are shown in Fig. 2(b).
According to Fig. 2(b), increasing the margin parameter shifts probability curves to the left. Thus, with the same , larger margin parameters lead to lower probabilities and thus larger loss even with small angles . In other words, the angles between the feature and its corresponding class’s weights have to be very small for sample being correctly classified. This is the reason why margin-based losses provide stronger supervisions for the same than conventional cosine-based losses. Proper margin settings have shown to boost the final recognition performance in .
Although larger margin provides stronger supervisions, it should not be too large either. When is oversized (e.g., ), the probabilities becomes unreliable. It would output probabilities around even is very small. This lead to large loss for almost all samples even with very small sample-to-class angles, which makes the training difficult to converge. In previous methods, the margin parameter selection is an ad-hoc procedure and has no theoretical guidance for most cases.
3 Summary of the hyparameter study
According to our analysis, we can draw the following conclusions:
(1) Hyperparameters scale and margin can substantially influence the prediction probability of feature with ground-truth identity/category . For the scale parameter , too small would limit the maximal value of . On the other hand, too large would make most predicted probabilities to be , which makes the training loss insensitive to the correctness of . For the margin parameter , a too small margin is not strong enough to regularize the final angular margin, while an oversized margin makes the training difficult to converge.
(2) The effect of scale and margin can be unified to modulate the mapping from cosine distances to the prediction probability . As shown in Fig. 2(a) and Fig. 2(b), both small scales and large margins have similar effect on for strengthening the supervisions, while both large scales and small margins weaken the supervisions. Therefore it is feasible and promising to control the probability using one single hyperparameter, either or . Considering the fact that is more related to the range of that required to span $s$ in the reminder of this paper.
The cosine-based softmax loss with adaptive scaling
Based on our previous studies on the hyperparameters of the cosine-based softmax loss functions, in this section, we propose a novel loss with a self-adaptive scaling scheme, namely AdaCos, which does not require the ad-hoc and time-consuming manual parameter tuning. Training with the proposed loss does not only facilitate convergence but also results in higher recognition accuracy.
Our previous studies on Fig. 1 show that during the training process, the angles for between the feature and its non-corresponding weights are almost always close to , In other words, we could safely assume that in Eq. (7). Obviously, it is the probability of feature belonging to its corresponding class that has the most influence on supervision for network training. Therefore, we focus on designing an adaptive scale parameter for controling the probabilities .
From the curves of w.r.t. (Fig. 2(a)), we observe that the scale parameter does not only simply affect ’s boundary of of determining correct/incorrect but also squeezes/stretches the curvature; In contrast to scale , margin parameter only shifts the curve in phase. We therefore propose to automatically tune the scale parameter and eliminate the margin parameter from our loss function, which makes our proposed AdaCos loss different from state-of-the-art softmax loss variants with angular margin. With softmax function, the predicted probability can be defined by
where . Combining Eqs. (7) and (10), we obtain an transcendental equation. Considering that is close to , the relation between the scale parameter and the point can be approximated as
Since is in the center of , it is natural to regard as the point, i.e. setting for figuring out an effective mapping from angle to the probability . Then the supervisions determined by would be back-propagated to update and further to update network parameters. According to Eq. (11), we can estimate the corresponding scale parameter as
2 Dynamically adaptive scale parameter
As Fig. 1 shows, the angles between features and their ground-truth class weights gradually decrease as the training iterations increase; while the angles between features and non-corresponding classes become stabilize around , as shown in Fig. 1.
At the begin of the training process, the median angle of each mini-batch might be too large to impose enough supervisions for training. We therefore force the central angle to be less than . Our dynamic scale parameter for the -th iteration could then be formulated as
Experiments
We examine the proposed AdaCos loss function on several public face recognition benchmarks and compare it with state-of-the-art cosine-based softmax losses. The compared losses include -softmax , CosFace , and ArcFace . We present evaluation results on LFW , MegaFace 1-million Challenge , and IJB-C data. We also present results on some exploratory experiments to show the convergence speed and robustness against low-resolution images.
Preprocessing. We use two public training datasets, CASIA-WebFace and MS1M , to train CNN models with our proposed loss functions. We carefully clean the noisy and low-quality images from the datasets. The cleaned WebFace and MS1M contain about M and M facial images, respectively. All models are trained based on these training data and directly tested on the test splits of the three datasets. RSA is applied to the images to extract facial areas. Then, according to detected facial landmarks, the faces are aligned through similarity transformation and resized to the size . All image pixel values are subtracted with the mean and dividing by .
The LFW dataset collected thousands of identities from the inertnet. Its testing protocol contains about images for about identities with a total of ground-truth matches. Half of the matches are positive while the other half are negative ones. LFW’s primary difficulties lie in face pose variations, color jittering, illumination variations and aging of persons. Note portion of the pose variations can be eliminated by the RSA facial landmark detection and alignment algorithm, but there still exist some non-frontal facial images which can not be aligned by RSA and then aligned manually.
For all experiments on LFW , we train ResNet-50 models with batch size of on the cleaned WebFace dataset. The input size of facial image is and the feature dimension input into the loss function is . Different loss functions are compared with our proposed AdaCos losses.
Results in Table 1 show the recognition accuracies of models trained with different softmax loss functions. Our proposed AdaCos losses with fixed and dynamic scale parameters (denoted as Fixed AdaCos and Dyna. AdaCos) surpass the state-of-the-art cosine-based softmax losses under the same training configuration. For the hyperparameter settings of the compared losses, the scaling parameter is set as for -softmax , CosFace and ArcFace ; the margin parameters are set as and for CosFace , and ArcFace , respectively. Since LFW is a relatively easy evaluation set, we train and test all losses for three times. The average accuracy of our proposed dynamic AdaCos is higher than state-of-the-art ArcFace and than -softmax .
1.2 Exploratory Experiments
Convergence rates. Convergence rate is an important indicator of efficiency of loss functions. We examine the convergence rates of several cosine-based losses at different training iterations. The training configurations are same as Table 1. Results in Table 2 reveal that the convergence rates when training with the AdaCos losses are much higher.
2 Results on MegaFace
We then evaluate the performance of proposed AdaCos on the MegaFace Challenge , which is a publicly available identification benchmark, widely used to test the performance of facial recognition algorithms. The gallery set of MegaFace incorporates over million images from K identities collected from Flickr photos . We follow ArcFace ’s testing protocol, which cleaned the dataset to make the results more reliable. We train the same Inception-ResNet models with CASIA-WebFace and MS1M training data, where overlapped subjects are removed.
Table 3 and Fig. 5 summarize the results of models trained on both WebFace and MS1M datasets and tested on the cleaned MegaFace dataset. The proposed AdaCos and state-of-the-art softmax losses are compared, where the dynamic AdaCos loss outperforms all compared losses on the MegaFace.
3 Results on IJB-C 1:1 verification protocol
The IJB-C dataset contains about identities with a total of still facial images and unconstrained video frames. In the 1:1 verification, there are positive matches and negative matches, which allow us to evaluate TARs at various FARs (e.g., ).
We compare the softmax loss functoins, including the proposed AdaCos, -softmax , CosFace , and ArcFace with the same training data (WebFace and MS1M ) and network architecture (Inception-ResNet ). We also report the results of FaceNet , VGGFace listed in Crystal loss . Table 4 and Fig. 6 exhibit their performances on the IJB-C 1:1 verification. Our proposed dynamic AdaCos achieves the best performance.
Conclusions
Acknowledgements. This work is supported in part by SenseTime Group Limited, in part by the General Research Fund through the Research Grants Council of Hong Kong under Grants CUHK14202217, CUHK14203118, CUHK14205615, CUHK14207814, CUHK14213616, CUHK14208417, CUHK14239816, in part by CUHK Direct Grant, and in part by National Natural Science Foundation of China (61472410) and the Joint Lab of CAS-HK.