AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations

Xiao Zhang, Rui Zhao, Yu Qiao, Xiaogang Wang, Hongsheng Li

Introduction

Recent years witnessed the breakthrough of deep Convolutional Neural Networks (CNNs) on significantly improving the performance of one-to-one (1:1)(1:1) face verification and one-to-many (1:N)(1:N) face identification tasks. The successes of deep face CNNs can be mainly credited to three factors: enormous training data , deep neural network architectures and effective loss functions . Modern face datasets, such as LFW , CASIA-WebFace , MS1M and MegaFace , contain huge number of identities which enable the training of deep networks. A number of recent studies, such as DeepFace , DeepID2 , DeepID3 , VGGFace and FaceNet , demonstrated that properly designed network architectures also lead to improved performance.

Apart from the large-scale training data and deep structures, training losses also play key roles in learning accurate face recognition models . Unlike image classification tasks, face recognition is essentially an open set recognition problem, where the testing categories (identities) are generally different from those used in training. To handle this challenge, most deep learning based face recognition approaches utilize CNNs to extract feature representations from facial images, and adopt a metric (usually the cosine distance) to estimate the similarities between pairs of faces during inference.

However, such inference evaluation metric is not well considered in the methods with softmax cross-entropy loss functionWe denote it as “softmax loss” for short in the remaining sections., which train the networks with the softmax loss but perform inference using cosine-similarities. To mitigate the gap between training and testing, recent works directly optimized cosine-based softmax losses. Moreover, angular margin-based terms are usually integrated into cosine-based losses to maximize the angular margins between different identities. These methods improve the face recognition performance in the open-set setup. In spite of their successes, the training processes of cosine-based losses (and their variants introducing margins) are usually tricky and unstable. The convergence and performance highly depend on the hyperparameter settings of loss, which are determined empirically through large amount of trials. In addition, subtle changes of these hyperparameters may fail the entire training process.

In this paper, we investigate state-of-the-art cosine-based softmax losses , especially those aiming at maximizing angular margins, to understand how they provide supervisions for training deep neural networks. Each of the functions generally includes several hyperprameters, which have substantial impact on the final performance and are usually difficult to tune. One has to repeat training with different settings for multiple times to achieve optimal performance. Our analysis shows that different hyperparameters in those cosine-based losses actually have similar effects on controlling the samples’ predicted class probabilities. Improper hyperparameter settings cause the loss functions to provide insufficient supervisions for optimizing networks.

Based on the above observation, we propose an adaptive cosine-based loss function, AdaCos, which automatically tunes hyperparameters and generates more effective supervisions during training. The proposed AdaCos dynamically scales the cosine similarities between training samples and corresponding class center vectors (the fully-connection vector before softmax), making their predicted class probability meets the semantic meaning of these cosine similarities. Furthermore, AdaCos can be easily implemented using built-in functions from prevailing deep learning libraries . The proposed AdaCos loss leads to faster and more stable convergence for training without introducing additional computational overhead.

To demonstrate the effectiveness of the proposed AdaCos loss function, we evaluated it on several face benchmarks, including LFW face verification , MegaFace one-million identification and IJB-C . Our method outperforms state-of-the-art cosine-based losses on all these benchmarks.

Related Works

Cosine similarities for inference. For learning deep face representations, feature-normalized losses are commonly adopted to enhance the recognition accuracy. Coco loss and NormFace studied the effect of normalization and proposed two strategies by reformulating softmax loss and metric learning. Similarly, Ranjan et al. in also discussed this problem and applied normalization on learned feature vectors to restrict them lying on a hypersphere. Movrever, compared with these hard normalization, ring loss came up with a soft feature normalization approach with convex formulations.

Margin-based softmax loss. Earlier, most face recognition approaches utilized metric-targeted loss functions, such as triplet and contrastive loss , which utilize Euclidean distances to measure similarities between features. Taking advantages of these works, center loss and range loss were proposed to reduce intra-class variations via minimizing distances within each class . Following this, researchers found that constraining margin in Euclidean space is insufficient to achieve optimal generalization. Then angular-margin based loss functions were proposed to tackle the problem. Angular constraints were integrated into the softmax loss function to improve the learned face representation by L-softmax and A-softmax . CosFace , AM-softmax and ArcFace directly maximized angular margins and employed simpler and more intuitive loss functions compared with aforementioned methods.

Automatic hyperparameter tuning. The performance of an algorithm highly depends on hyperparameter settings. Grid and random search are the most widely used strategies. For more automatic tuning, sequential model-based global optimization is the mainstream choice. Typically, it performs inference with several hyperparameters settings, and chooses setting for the next round of testing based on the inference results. Bayesian optimization and tree-structured parzen estimator approach are two famous sequential model-based methods. However, these algorithms essentially run multiple trials to predict the optimized hyperparameter settings.

Investigation of hyperparameters in cosine-based softmax losses

In recent years, state-of-the-art cosine-based softmax losses, including L2-softmax , CosFace , ArcFace , significantly improve the performance of deep face recognition. However, the final performances of those losses are substantially affected by their hyperparameters settings, which are generally difficult to tune and require multiple trials in practice. We analyze two most important hyperparameters, the scaling parameter ss and the margin parameter mm, in cosine-based losses. Specially, we deeply study their effects on the prediction probabilities after softmax, which serves as supervision signals for updating entire neural network.

Let x⃗i\vec{x}_{i} denote the deep representation (feature) of the ii-th face image of the current mini-batch with size NN, and yiy_{i} be the corresponding label. The predicted classification probability Pi,jP_{i,j} of all NN samples in the mini-batch can be estimated by the softmax function as

where fi,jf_{i,j} is logit used as the input of softmax, Pi,jP_{i,j} represents its softmax-normalized probability of assigning x⃗i\vec{x}_{i} to class jj, and CC is the number of classes. The cross-entropy loss associated with current mini-batch is

Conventional softmax loss and state-of-the-art cosine-based softmax losses calculate the logits fi,jf_{i,j} in different ways. In conventional softmax loss, logits fi,jf_{i,j} are obtained as the inner product between feature x⃗i\vec{x}_{i} and the jj-th class weights W⃗j\vec{W}_{j} as fi,j=W⃗jTx⃗if_{i,j}={\vec{W}_{j}^{\text{T}}{\vec{x}_{i}}}. In the cosine-based softmax losses , cosine similarity is calculated by cos⁡θi,j=⟨x⃗i,W⃗j⟩/∥x⃗i∥∥W⃗j∥\cos\theta_{i,j}=\langle\vec{x}_{i},\vec{W}_{j}\rangle/\|\vec{x}_{i}\|\|\vec{W}_{j}\|. The logits fi,jf_{i,j} are calculated as fi,j=s⋅cos⁡θi,jf_{i,j}=s\cdot\cos{\theta_{i,j}}, where ss is a scale hyperparameter. To enforce angular margin on the representations, ArcFace modified the loss to the form

Intuitively, on one hand, the parameter ss scales up the narrow range of cosine distances, making the logits more discriminative. On the other hand, the parameter mm enlarges the margin between different classes to enhance classification ability. These hyperparameters eventually affect Pi,yiP_{i,y_{i}}. Empirically, an ideal hyperparameter setting should help Pi,jP_{i,j} to satisfy the following two properties: (1) Predicted probabilities Pi,yiP_{i,y_{i}} of each class (identity) should span to the range $:thelowerboundaryof: the lower boundary ofP_{i,y_{i}}shouldbenearwhiletheupperboundarynearshould be near while the upper boundary near1;(2)Changingcurveof; (2) Changing curve ofP_{i,y_{i}}shouldhavelargeabsolutegradientsaroundshould have large absolute gradients around\theta_{i,y_{i}}$ to make training effective.

The scale parameter ss can significantly affect Pi,yiP_{i,y_{i}}. Intuitively, Pi,yiP_{i,y_{i}} should gradually increase from to 11 as the angle θi,yi\theta_{i,y_{i}} decreases from π2\frac{\pi}{2} to Mathematically, θ\theta can be any value in [0,π][0,\pi]. We empirically found, however, the maximum θ\theta is always around π2\frac{\pi}{2}. See the red curve in Fig. 1 for examples., i.e., the smaller the angle between x⃗i\vec{x}_{i} and its corresponding class weight W⃗yi\vec{W}_{y_{i}} is, the larger the probability should be. Both improper probability range and probability curves w.r.t. θi,yi\theta_{i,y_{i}} would negatively affect the training process and thus the recognition performance.

We first study the range of classification probability Pi,jP_{i,j}. Given scale parameter ss, the range of probabilities in all cosine-based softmax losses is

where the lower boundary is achieved when fi,j=s⋅0=0f_{i,j}=s\cdot 0=0 and fi,k=s⋅1=sf_{i,k}=s\cdot 1=s for all k≠jk\neq j in Eq. (1). Similarly, the upper bound is achieved when fi,j=sf_{i,j}=s and fi,k=0f_{i,k}=0 for all k≠jk\neq j. The range of Pi,jP_{i,j} approaches 1 when s→∞s\rightarrow\infty, i.e.,

which means that the requirement of the range spanning $couldbesatisfiedwithalargecould be satisfied with a larges.Howeveritdoesnotmeanthatthelargerthescaleparameter,thebettertheselectionis.Infacttheprobabilityrangecaneasilyapproachahighvalue,suchas. However it does not mean that the larger the scale parameter, the better the selection is. In fact the probability range can easily approach a high value, such as{0.94}whenclassnumberwhen class numberC=10andscaleparameterand scale parameters=5.0$. But an oversized scale would lead to poor probability distribution, as will be discussed in the following paragraphs.

We investigate the influences of parameter ss by taking Pi,yiP_{i,y_{i}} as a function of ss and angle θi,yi\theta_{i,y_{i}} where yiy_{i} denotes the label of x⃗i\vec{x}_{i}. Formally, we have

where Bi=∑k≠yiefi,k=∑k≠yies⋅cos⁡θi,kB_{i}=\sum_{k\neq y_{i}}e^{f_{i,k}}=\sum_{k\neq y_{i}}e^{s\cdot\cos\theta_{i,k}} are the logits summation of all non-corresponding classes for feature x⃗i\vec{x}_{i}. We observe that the values of BiB_{i} are almost unchanged during the training process. This is because the angles θi,k\theta_{i,k} for non-corresponding classes k≠yik\neq y_{i} always stay around π2\frac{\pi}{2} during training (see red curve in Fig. 1).

Therefore, we can assume BiB_{i} is constant, i.e., Bi≈∑k≠yies⋅cos⁡(π/2)B_{i}\approx\sum_{k\neq y_{i}}e^{s\cdot\cos(\pi/2)} =C−1=C-1. We then plot curves of probabilities Pi,yiP_{i,y_{i}} w.r.t. θi,yi\theta_{i,y_{i}} under different setting of parameter ss in Fig. 2(a). It is obvious that when ss is too small (e.g., s=10s=10 for class/identity number C=2,000C=2,000 and C=20,000C=20,000), the maximal value of Pi,yiP_{i,y_{i}} could not reach 11. This is undesirable because even when the network is very confident on a sample xi⃗\vec{x_{i}}’s corresponding class label yiy_{i}, e.g. θi,yi=0\theta_{i,y_{i}}=0, the loss function would still penalize the classification results and update the network.

On the other hand, when ss is too large (e.g., s=64s=64), the probability curve Pi,yiP_{i,y_{i}} w.r.t. θi,yi\theta_{i,y_{i}} is also problematic. It would output a very high probability even when θi,yi\theta_{i,y_{i}} is close to π/2\pi/2, which means that the loss function with large ss may fail to penalize mis-classified samples and cannot effectively update the networks to correct mistakes.

In summary, the scaling parameter ss has substantial influences to the range as well as the curves of the probabilities Pi,yiP_{i,y_{i}}, which are crucial for effectively training the deep network.

2 Effects of the margin parameter m𝑚m

In this section, we investigate the effect of margin parameters mm in cosine-based softmax losses (Eqs. (3) & (4)), and their effects on feature x⃗i\vec{x}_{i}’s predicted class probability Pi,yiP_{i,y_{i}}. For simplicity, we here study the margin parameter mm for ArcFace (Eq. 3); while the similar conclusions also apply to the parameter mm in CosFace (Eq. (4)).

We first re-write classification probability Pi,yiP_{i,y_{i}} following Eq. (7) as

To study the influence of parameter mm on the probability Pi,yiP_{i,y_{i}}, we assume both ss and BiB_{i} are fixed. Following the discussion in Section 3.1, we set Bi≈C−1B_{i}\approx{C-1}, and fix s=30s=30. The probability curves Pi,yiP_{i,y_{i}} w.r.t. θi,yi\theta_{i,y_{i}} under different mm are shown in Fig. 2(b).

According to Fig. 2(b), increasing the margin parameter shifts probability Pi,yiP_{i,y_{i}} curves to the left. Thus, with the same θi,yi\theta_{i,y_{i}}, larger margin parameters lead to lower probabilities Pi,yiP_{i,y_{i}} and thus larger loss even with small angles θi,yi\theta_{i,y_{i}}. In other words, the angles θi,yi\theta_{i,y_{i}} between the feature x⃗i\vec{x}_{i} and its corresponding class’s weights W⃗yi\vec{W}_{y_{i}} have to be very small for sample ii being correctly classified. This is the reason why margin-based losses provide stronger supervisions for the same θi,yi\theta_{i,y_{i}} than conventional cosine-based losses. Proper margin settings have shown to boost the final recognition performance in .

Although larger margin mm provides stronger supervisions, it should not be too large either. When mm is oversized (e.g., m=1.0m=1.0), the probabilities Pi,yiP_{i,y_{i}} becomes unreliable. It would output probabilities around even θi,yi\theta_{i,y_{i}} is very small. This lead to large loss for almost all samples even with very small sample-to-class angles, which makes the training difficult to converge. In previous methods, the margin parameter selection is an ad-hoc procedure and has no theoretical guidance for most cases.

3 Summary of the hyparameter study

According to our analysis, we can draw the following conclusions:

(1) Hyperparameters scale ss and margin mm can substantially influence the prediction probability Pi,yiP_{i,y_{i}} of feature x⃗i\vec{x}_{i} with ground-truth identity/category yiy_{i}. For the scale parameter ss, too small ss would limit the maximal value of Pi,yiP_{i,y_{i}}. On the other hand, too large ss would make most predicted probabilities Pi,yiP_{i,y_{i}} to be 11, which makes the training loss insensitive to the correctness of θi,yi\theta_{i,y_{i}}. For the margin parameter mm, a too small margin is not strong enough to regularize the final angular margin, while an oversized margin makes the training difficult to converge.

(2) The effect of scale ss and margin mm can be unified to modulate the mapping from cosine distances cos⁡θi,yi\cos{\theta_{i,y_{i}}} to the prediction probability Pi,yiP_{i,y_{i}}. As shown in Fig. 2(a) and Fig. 2(b), both small scales and large margins have similar effect on θi,yi\theta_{i,y_{i}} for strengthening the supervisions, while both large scales and small margins weaken the supervisions. Therefore it is feasible and promising to control the probability Pi,yiP_{i,y_{i}} using one single hyperparameter, either ss or mm. Considering the fact that ss is more related to the range of Pi,yiP_{i,y_{i}} that required to span $,wewillfocusonautomaticallytuningthescaleparameter, we will focus on automatically tuning the scale parameters$ in the reminder of this paper.

The cosine-based softmax loss with adaptive scaling

Based on our previous studies on the hyperparameters of the cosine-based softmax loss functions, in this section, we propose a novel loss with a self-adaptive scaling scheme, namely AdaCos, which does not require the ad-hoc and time-consuming manual parameter tuning. Training with the proposed loss does not only facilitate convergence but also results in higher recognition accuracy.

Our previous studies on Fig. 1 show that during the training process, the angles θi,k\theta_{i,k} for k≠yik\neq y_{i} between the feature x⃗i\vec{x}_{i} and its non-corresponding weights W⃗k≠yi\vec{W}_{k\neq y_{i}} are almost always close to π2\frac{\pi}{2}, In other words, we could safely assume that Bi≈∑k≠yies⋅cos⁡(π/2)=C−1B_{i}\approx\sum_{k\neq y_{i}}e^{s\cdot\cos(\pi/2)}=C-1 in Eq. (7). Obviously, it is the probability Pi,yiP_{i,y_{i}} of feature xix_{i} belonging to its corresponding class yiy_{i} that has the most influence on supervision for network training. Therefore, we focus on designing an adaptive scale parameter for controling the probabilities Pi,yiP_{i,y_{i}}.

From the curves of Pi,yiP_{i,y_{i}} w.r.t. θi,yi\theta_{i,y_{i}} (Fig. 2(a)), we observe that the scale parameter ss does not only simply affect Pi,yiP_{i,y_{i}}’s boundary of of determining correct/incorrect but also squeezes/stretches the Pi,yiP_{i,y_{i}} curvature; In contrast to scale ss, margin parameter mm only shifts the curve in phase. We therefore propose to automatically tune the scale parameter ss and eliminate the margin parameter mm from our loss function, which makes our proposed AdaCos loss different from state-of-the-art softmax loss variants with angular margin. With softmax function, the predicted probability can be defined by

where θ0∈[0,π2]\theta_{0}\in[0,\frac{\pi}{2}]. Combining Eqs. (7) and (10), we obtain an transcendental equation. Considering that P(θ0)P(\theta_{0}) is close to 12\frac{1}{2}, the relation between the scale parameter ss and the point (θ0,P(θ0))(\theta_{0},P(\theta_{0})) can be approximated as

Since π4\frac{\pi}{4} is in the center of [0,π2][0,\frac{\pi}{2}], it is natural to regard π/4\pi/4 as the point, i.e. setting θ0=π/4\theta_{0}=\pi/4 for figuring out an effective mapping from angle θi,yi\theta_{i,y_{i}} to the probability Pi,yiP_{i,y_{i}}. Then the supervisions determined by Pi,yiP_{i,y_{i}} would be back-propagated to update θi,yi\theta_{i,y_{i}} and further to update network parameters. According to Eq. (11), we can estimate the corresponding scale parameter sfs_{f} as

2 Dynamically adaptive scale parameter

As Fig. 1 shows, the angles θi,yi\theta_{i,y_{i}} between features x⃗i\vec{x}_{i} and their ground-truth class weights W⃗yi\vec{W}_{y_{i}} gradually decrease as the training iterations increase; while the angles between features x⃗i\vec{x}_{i} and non-corresponding classes W⃗j≠yi\vec{W}_{j\neq y_{i}} become stabilize around π2\frac{\pi}{2}, as shown in Fig. 1.

At the begin of the training process, the median angle θmed(t)\theta_{\text{med}}^{(t)} of each mini-batch might be too large to impose enough supervisions for training. We therefore force the central angle θmed(t)\theta_{\text{med}}^{(t)} to be less than π4\frac{\pi}{4}. Our dynamic scale parameter for the tt-th iteration could then be formulated as

Experiments

We examine the proposed AdaCos loss function on several public face recognition benchmarks and compare it with state-of-the-art cosine-based softmax losses. The compared losses include l2l2-softmax , CosFace , and ArcFace . We present evaluation results on LFW , MegaFace 1-million Challenge , and IJB-C data. We also present results on some exploratory experiments to show the convergence speed and robustness against low-resolution images.

Preprocessing. We use two public training datasets, CASIA-WebFace and MS1M , to train CNN models with our proposed loss functions. We carefully clean the noisy and low-quality images from the datasets. The cleaned WebFace and MS1M contain about 0.450.45M and 2.352.35M facial images, respectively. All models are trained based on these training data and directly tested on the test splits of the three datasets. RSA is applied to the images to extract facial areas. Then, according to detected facial landmarks, the faces are aligned through similarity transformation and resized to the size 144×144144\times 144. All image pixel values are subtracted with the mean 127.5127.5 and dividing by 128128.

The LFW dataset collected thousands of identities from the inertnet. Its testing protocol contains about 13,00013,000 images for about 1,6801,680 identities with a total of 6,0006,000 ground-truth matches. Half of the matches are positive while the other half are negative ones. LFW’s primary difficulties lie in face pose variations, color jittering, illumination variations and aging of persons. Note portion of the pose variations can be eliminated by the RSA facial landmark detection and alignment algorithm, but there still exist some non-frontal facial images which can not be aligned by RSA and then aligned manually.

For all experiments on LFW , we train ResNet-50 models with batch size of 512512 on the cleaned WebFace dataset. The input size of facial image is 144×144144\times 144 and the feature dimension input into the loss function is 512512. Different loss functions are compared with our proposed AdaCos losses.

Results in Table 1 show the recognition accuracies of models trained with different softmax loss functions. Our proposed AdaCos losses with fixed and dynamic scale parameters (denoted as Fixed AdaCos and Dyna. AdaCos) surpass the state-of-the-art cosine-based softmax losses under the same training configuration. For the hyperparameter settings of the compared losses, the scaling parameter is set as 3030 for l2l2-softmax , CosFace and ArcFace ; the margin parameters are set as 0.250.25 and 0.50.5 for CosFace , and ArcFace , respectively. Since LFW is a relatively easy evaluation set, we train and test all losses for three times. The average accuracy of our proposed dynamic AdaCos is 0.26%0.26\% higher than state-of-the-art ArcFace and 1.52%1.52\% than l2l2-softmax .

1.2 Exploratory Experiments

Convergence rates. Convergence rate is an important indicator of efficiency of loss functions. We examine the convergence rates of several cosine-based losses at different training iterations. The training configurations are same as Table 1. Results in Table 2 reveal that the convergence rates when training with the AdaCos losses are much higher.

2 Results on MegaFace

We then evaluate the performance of proposed AdaCos on the MegaFace Challenge , which is a publicly available identification benchmark, widely used to test the performance of facial recognition algorithms. The gallery set of MegaFace incorporates over 11 million images from 690690K identities collected from Flickr photos . We follow ArcFace ’s testing protocol, which cleaned the dataset to make the results more reliable. We train the same Inception-ResNet models with CASIA-WebFace and MS1M training data, where overlapped subjects are removed.

Table 3 and Fig. 5 summarize the results of models trained on both WebFace and MS1M datasets and tested on the cleaned MegaFace dataset. The proposed AdaCos and state-of-the-art softmax losses are compared, where the dynamic AdaCos loss outperforms all compared losses on the MegaFace.

3 Results on IJB-C 1:1 verification protocol

The IJB-C dataset contains about 3,5003,500 identities with a total of 31,33431,334 still facial images and 117,542117,542 unconstrained video frames. In the 1:1 verification, there are 19,55719,557 positive matches and 15,638,93215,638,932 negative matches, which allow us to evaluate TARs at various FARs (e.g., 10−710^{-7}).

We compare the softmax loss functoins, including the proposed AdaCos, l2l2-softmax , CosFace , and ArcFace with the same training data (WebFace and MS1M ) and network architecture (Inception-ResNet ). We also report the results of FaceNet , VGGFace listed in Crystal loss . Table 4 and Fig. 6 exhibit their performances on the IJB-C 1:1 verification. Our proposed dynamic AdaCos achieves the best performance.

Conclusions

Acknowledgements. This work is supported in part by SenseTime Group Limited, in part by the General Research Fund through the Research Grants Council of Hong Kong under Grants CUHK14202217, CUHK14203118, CUHK14205615, CUHK14207814, CUHK14213616, CUHK14208417, CUHK14239816, in part by CUHK Direct Grant, and in part by National Natural Science Foundation of China (61472410) and the Joint Lab of CAS-HK.

References