Large-Margin Softmax Loss for Convolutional Neural Networks

Weiyang Liu, Yandong Wen, Zhiding Yu, Meng Yang

Introduction

Over the past several years, convolutional neural networks (CNNs) have significantly boosted the state-of-the-art performance in many visual classification tasks such as object recognition, (Krizhevsky et al., 2012; Sermanet et al., 2014; He et al., 2015b, a), face verification (Taigman et al., 2014; Sun et al., 2014, 2015) and hand-written digit recognition (Wan et al., 2013). The layered learning architecture, together with convolution and pooling which carefully extract features from local to global, renders the strong visual representation ability of CNNs as well as their current significant positions in large-scale visual recognition tasks. Facing the increasingly more complex data, CNNs have continuously been improved with deeper structures (Simonyan & Zisserman, 2014; Szegedy et al., 2015), smaller strides (Simonyan & Zisserman, 2014) and new non-linear activations (Goodfellow et al., 2013; Nair & Hinton, 2010; He et al., 2015b). While benefiting from the strong learning ability, CNNs also have to face the crucial issue of overfilling. Considerable effort such as large-scale training data (Russakovsky et al., 2014), dropout (Krizhevsky et al., 2012), data augmentation (Krizhevsky et al., 2012; Szegedy et al., 2015), regularization (Hinton et al., 2012; Srivastava et al., 2014; Wan et al., 2013; Goodfellow et al., 2013) and stochastic pooling (Zeiler & Fergus, 2013) has been put to address the issue.†Authors contributed equally. Code is available at https://github.com/wy1iu/LargeMargin_Softmax_Loss

A recent trend towards learning with even stronger features is to reinforce CNNs with more discriminative information. Intuitively, the learned features are good if their intra-class compactness and inter-class separability are simultaneously maximized. While this may not be easy due to the inherent large intra-class variations in many tasks, the strong representation ability of CNNs make it possible to learn invariant features towards this direction. Inspired by such idea, the contrastive loss (Hadsell et al., 2006) and triplet loss (Schroff et al., 2015) were proposed to enforce extra intra-class compactness and inter-class separability. A consequent problem, however, is that the number of training pairs and triplets can theoretically go up to O(N2)\mathcal{O}(N^{2}) where NN is the total number of training samples. Considering that CNNs often handle large-scale training sets, a subset of training samples need to be carefully selected for these losses. The softmax function is widely adopted by many CNNs (Krizhevsky et al., 2012; He et al., 2015a, b) due to its simplicity and probabilistic interpretation. Together with the cross-entropy loss, they form arguably one of the most commonly used components in CNN architectures. In this paper, we define the softmax loss as the combination of a cross-entropy loss, a softmax function and the last fully connected layer (see Fig. 1). Under such definition, many prevailing CNN models can be viewed as the combination of a convolutional feature learning component and a softmax loss component, as shown in Fig. 1. Despite its popularity, current softmax loss does not explicitly encourage intra-class compactness and inter-class-separability. Our key intuition is that the separability between sample and parameter can be factorized into amplitude ones and angular ones with cosine similarity: Wcx=∥Wc∥2∥x∥2cos⁡(θc)\bm{W}_{c}\bm{x}=\|\bm{W}_{c}\|_{2}\|\bm{x}\|_{2}\cos(\theta_{c}), where cc is the class index, and the corresponding parameters Wc\bm{W}_{c} of the last fully connected layer can be regarded as the linear classifier of class cc. Under softmax loss, the label prediction decision rule is largely determined by the angular similarity to each class since softmax loss uses cosine distance as classification score. The purpose of this paper, therefore, is to generalize the softmax loss to a more general large-margin softmax (L-Softmax) loss in terms of angular similarity, leading to potentially larger angular separability between learned features. This is done by incorporating a preset constant mm multiplying with the angle between sample and the classifier of ground truth class. mm determines the strength of getting closer to the ground truth class, producing an angular margin. One shall see, the conventional softmax loss becomes a special case of the L-Softmax loss under our proposed framework. Our idea is verified by Fig. 2 where the learned features by L-Softmax become much more compact and well separated.

The L-Softmax loss is a flexible learning objective with adjustable inter-class angular margin constraint. It presents a learning task of adjustable difficulty where the difficulty gradually increases as the required margin becomes larger. The L-Softmax loss has several desirable advantages. First, it encourages angular decision margin between classes, generating more discriminative features. Its geometric interpretation is very clear and intuitive, as elaborated in Section 3.2. Second, it partially avoids overfitting by defining a more difficult learning target, casting a different viewpoint to the overfitting problem. Third, L-Softmax benefits not only classification problems, but also verification problems where ideally learned features should have the minimum inter-class distance being greater than the maximum intra-class distance. In this case, learning well separated features can significantly improve the performance.

Our experiments validate that L-Softmax can effectively boost the performance in both classification and verification tasks. More intuitively, the visualizations of the learned features in Fig. 2 and Fig. 5 show great discriminativeness of the L-Softmax loss. As a straightforward generalization of softmax loss, L-Softmax loss not only inherits all merits from softmax loss but also learns features with large angular margin between different classes. Besides that, the L-Softmax loss is also well motivated with clear geometric interpretation as elaborated in Section 3.3.

Related Work and Preliminaries

Current widely used data loss functions in CNNs include Euclidean loss, (square) hinge loss, information gain loss, contrastive loss, triplet loss, Softmax loss, etc. To enhance the intra-class compactness and inter-class separability, (Sun et al., 2014) trains the CNN with the combination of softmax loss and contrastive loss. The contrastive loss inputs the CNNs with pairs of training samples. If the input pair belongs to the same class, the contrastive loss will require their features are as similar as possible. Otherwise, the contrastive loss will require their distance larger than a margin. (Schroff et al., 2015) uses the triplet loss to encourage a distance constraint similar to the contrastive loss. Differently, the triplet loss requires 3 (or a multiple of 3) training samples as input at a time. The triplet loss minimizes the distance between an anchor sample and a positive sample (of the same identity), and maximizes the distance between the anchor sample and a negative sample (of different identity). Both triplet loss and contrastive loss require a carefully designed pair selection procedure. Both (Sun et al., 2014) and (Schroff et al., 2015) suggest that enforcing such a distance constraint that encourages intra-class compactness and inter-class separability can greatly boost the feature discriminativeness, which motivates us to employ a margin constraint in the original softmax loss.

Unlike any previous work, our work cast a novel view on generalizing the original softmax loss. We define the ii-th input feature xi\bm{x}_{i} with the label yiy_{i}. Then the original softmax loss can be written as

where fjf_{j} denotes the jj-th element (j∈[1,K]j\in[1,K], K is the number of classes) of the vector of class scores f\bm{f}, and NN is the number of training data. In the softmax loss, f\bm{f} is usually the activations of a fully connected layer W\bm{W}, so fyif_{y_{i}} can be written as fyi=WyiTxif_{y_{i}}=\bm{W}_{y_{i}}^{T}\bm{x}_{i} in which Wyi\bm{W}_{y_{i}} is the yi{y_{i}}-th column of W\bm{W}. Note that, we omit the constant bb in fj,∀jf_{j},\forall j here to simplify analysis, but our L-Softmax loss can still be easily modified to work with bb. (In fact, the performance is nearly of no difference, so we do not make it complicated here.) Because fjf_{j} is the inner product between Wj\bm{W}_{j} and xi\bm{x}_{i}, it can be also formulated as fj=∥Wj∥∥xi∥cos⁡(θj)f_{j}=\|\bm{W}_{j}\|\|\bm{x}_{i}\|\cos(\theta_{j}) where θj\theta_{j} (0≤θj≤π0\leq\theta_{j}\leq\pi) is the angle between the vector Wj\bm{W}_{j} and xi\bm{x}_{i}. Thus the loss becomes

Large-Margin Softmax Loss

We give a simple example to describe our intuition. Consider the binary classification and we have a sample x\bm{x} from class 1. The original softmax is to force W1Tx>W2Tx\bm{W}_{1}^{T}\bm{x}>\bm{W}_{2}^{T}\bm{x} (i.e. ∥W1∥∥x∥cos⁡(θ1)>∥W2∥∥x∥cos⁡(θ2)\|\bm{W}_{1}\|\|\bm{x}\|\cos(\theta_{1})>\|\bm{W}_{2}\|\|\bm{x}\|\cos(\theta_{2})) in order to classify x\bm{x} correctly. However, we want to make the classification more rigorous in order to produce a decision margin. So we instead require ∥W1∥∥x∥cos⁡(mθ1)>∥W2∥∥x∥cos⁡(θ2)\|\bm{W}_{1}\|\|\bm{x}\|\cos(m\theta_{1})>\|\bm{W}_{2}\|\|\bm{x}\|\cos(\theta_{2}) (0≤θ1≤πm0\leq\theta_{1}\leq\frac{\pi}{m}) where mm is a positive integer. Because the following inequality holds:

Therefore, ∥W1∥∥x∥cos⁡(θ1)>∥W2∥∥x∥cos⁡(θ2)\|\bm{W}_{1}\|\|\bm{x}\|\cos(\theta_{1})>\|\bm{W}_{2}\|\|\bm{x}\|\cos(\theta_{2}) has to hold. So the new classification criteria is a stronger requirement to correctly classify x\bm{x}, producing a more rigorous decision boundary for class 1.

2 Definition

Following the notation in the preliminaries, the L-Softmax loss is defined as

where mm is a integer that is closely related to the classification margin. With larger mm, the classification margin becomes larger and the learning objective also becomes harder. Meanwhile, D(θ)\mathcal{D}(\theta) is required to be a monotonically decreasing function and D(πm)\mathcal{D}(\frac{\pi}{m}) should equal cos⁡(πm)\cos(\frac{\pi}{m}).

To simplify the forward and backward propagation, we construct a specific ψ(θi)\psi(\theta_{i}) in this paper:

where k∈[0,m−1]k\in[0,m-1] and kk is an integer. Combining Eq. (1), Eq. (4) and Eq. (6), we have the L-Softmax loss that is used throughout the paper. For forward and backward propagation, we need to replace cos⁡(θj)\cos(\theta_{j}) with WjTxi∥Wj∥∥xi∥\frac{\bm{W}_{j}^{T}\bm{x}_{i}}{\|\bm{W}_{j}\|\|\bm{x}_{i}\|}, and replace cos⁡(mθyi)\cos(m\theta_{y_{i}}) with

where nn is an integer and 2n≤m2n\leq m. After getting rid of θ\theta, we could perform derivation with respect to x\bm{x} and W\bm{W}. It is also trivial to perform derivation with mini-batch input.

3 Geometric Interpretation

We aim to encourage aa angle margin between classes via the L-Softmax loss. To simplify the geometric interpretation, we analyze the binary classification case where there are only W1\bm{W}_{1} and W2\bm{W}_{2}.

First, we consider the ∥W1∥=∥W2∥\|\bm{W}_{1}\|=\|\bm{W}_{2}\| scenario as shown in Fig. 4. With ∥W1∥=∥W2∥\|\bm{W}_{1}\|=\|\bm{W}_{2}\|, the classification result depends entirely on the angles between x\bm{x} and W1\bm{W}_{1}(W2\bm{W}_{2}). In the training stage, the original softmax loss requires θ1<θ2\theta_{1}<\theta_{2} to classify the sample x\bm{x} as class 1, while the L-Softmax loss requires mθ1<θ2m\theta_{1}<\theta_{2} to make the same decision. We can see the L-Softmax loss is more rigor about the classification criteria, which leads to a classification margin between class 1 and class 2. If we assume both softmax loss and L-Softmax loss are optimized to the same value and all training features can be perfectly classified, then the angle margin between class 1 and class 2 is given by m−1m+1θ1,2\frac{m-1}{m+1}\theta_{1,2} where θ1,2\theta_{1,2} is the angle between classifier vector W1\bm{W}_{1} and W2\bm{W}_{2}. The L-Softmax loss also makes the decision boundaries for class 1 and class 2 different as shown in Fig 4, while originally the decision boundaries are the same. From another viewpoint, we let θ1′=mθ1\theta^{\prime}_{1}=m\theta_{1} and assume that both the original softmax loss and the L-Softmax loss can be optimized to the same value. Then we can know θ1′\theta^{\prime}_{1} in the original softmax loss is m−1m-1 times larger than θ1\theta_{1} in the L-Softmax loss. As a result, the angle between the learned feature and W1\bm{W}_{1} will become smaller. For every class, the same conclusion holds. In essence, the L-Softmax loss narrows the feasible angleFeasible angle of the ii-th class refers to the possible angle between x\bm{x} and Wi\bm{W}_{i} that is learned by CNNs. for every class and produces an angle margin between these classes.

For both the ∥W1∥>∥W2∥\|\bm{W}_{1}\|>\|\bm{W}_{2}\| and ∥W1∥<∥W2∥\|\bm{W}_{1}\|<\|\bm{W}_{2}\| scenarios, the geometric interpretation is a bit more complicated. Because the length of W1\bm{W}_{1} and W2\bm{W}_{2} is different, the feasible angles of class 1 and class 2 are also different (see the decision boundary of original softmax loss in Fig. 4). Normally, the larger Wj\bm{W}_{j} is, the larger the feasible angle of its corresponding class is. As a result, the L-Softmax loss also produces different feasible angles for different classes. Similar to the analysis of the ∥W1∥=∥W2∥\|\bm{W}_{1}\|=\|\bm{W}_{2}\| scenario, the proposed loss will also generate a decision margin between class 1 and class 2.

4 Discussion

The L-Softmax loss utilizes a simple modification over the original softmax loss, achieving a classification angle margin between classes. By assigning different values for mm, we define a flexible learning task with adjustable difficulty for CNNs. The L-Softmax loss is endowed with some nice properties such as

The L-Softmax loss has a clear geometric interpretation. mm controls the margin among classes. With bigger mm (under the same training loss), the ideal margin between classes becomes larger and the learning difficulty is also increased. With m=1m=1, the L-Softmax loss becomes identical to the original softmax loss.

The L-Softmax loss defines a relatively difficult learning objective with adjustable margin (difficulty). A difficult learning objective can effectively avoid over-fitting and take full advantage of the strong learning ability from deep and wide architectures.

The L-Softmax loss can be easily used as a drop-in replacement for standard loss, as well as used in tandem with other performance-boosting approaches and modules, including learning activation functions, data augmentation, pooling functions or other modified network architectures.

Optimization

It is easy to compute the forward and backward propagation for the L-Softmax loss, so it is also trivial to optimize the L-Softmax loss using typical stochastic gradient descent. For LiL_{i}, the only difference between the original softmax loss and the L-Softmax loss lies in fyif_{y_{i}}. Thus we only need to compute fyif_{y_{i}} in forward and backward propagation while fj,j≠yif_{j},j\neq{y_{i}} is the same as the original softmax loss. Putting in Eq. (6) and Eq. (7), fyif_{y_{i}} is written as

where WyiTx∥Wyi∥∥x∥∈[cos⁡(kπm),cos⁡((k+1)πm)]\frac{\bm{W}_{y_{i}}^{T}\bm{x}}{\|\bm{W}_{y_{i}}\|\|\bm{x}\|}\in[\cos(\frac{k\pi}{m}),\cos(\frac{(k+1)\pi}{m})] and kk is an integer that belongs to [0,m−1][0,m-1]. For the backward propagation, we use the chain rule to compute the partial derivative: ∂Li∂xi=∑j∂Li∂fj∂fj∂xi\frac{\partial L_{i}}{\partial\bm{x}_{i}}=\sum_{j}\frac{\partial L_{i}}{\partial f_{j}}\frac{\partial f_{j}}{\partial\bm{x}_{i}} and ∂Li∂Wyi=∑j∂Li∂fj∂fj∂Wyi\frac{\partial L_{i}}{\partial\bm{W}_{y_{i}}}=\sum_{j}\frac{\partial L_{i}}{\partial f_{j}}\frac{\partial f_{j}}{\partial\bm{W}_{y_{i}}}. Because ∂Li∂fj\frac{\partial L_{i}}{\partial f_{j}} and ∂fj∂xi,∂fj∂Wyi,∀j≠yi\frac{\partial f_{j}}{\partial\bm{x}_{i}},\frac{\partial f_{j}}{\partial\bm{W}_{y_{i}}},\forall j\neq y_{i} are the same for both original softmax loss and L-Softmax loss, we leave it out for simplicity. ∂fyi∂xi\frac{\partial f_{y_{i}}}{\partial\bm{x}_{i}} and ∂fyi∂Wyi\frac{\partial f_{y_{i}}}{\partial\bm{W}_{y_{i}}} can be computed via

In implementation, kk can be efficiently computed by constructing a look-up table for WyiTxi∥Wyi∥∥xi∥\frac{\bm{W}_{y_{i}}^{T}\bm{x}_{i}}{\|\bm{W}_{y_{i}}\|\|\bm{x}_{i}\|} (i.e. cos⁡(θyi)\cos(\theta_{y_{i}})). To be specific, we give an example of the forward and backward propagation when m=2m=2. Thus fif_{i} is written as

where, \small k=\left\{{\begin{array}[]{*{20}{l}}{1,\ \ \ \ \frac{\bm{W}_{y_{i}}^{T}\bm{x}_{i}}{\|\bm{W}_{y_{i}}\|\|\bm{x}_{i}\|}\leq\cos(\frac{\pi}{2})}\\ {0,\ \ \ \ \frac{\bm{W}_{y_{i}}^{T}\bm{x}_{i}}{\|\bm{W}_{y_{i}}\|\|\bm{x}_{i}\|}>\cos(\frac{\pi}{2})}\end{array}}\right..

In the backward propagation, ∂fyi∂xi\frac{\partial f_{y_{i}}}{\partial\bm{x}_{i}} and ∂fyi∂Wyi\frac{\partial f_{y_{i}}}{\partial\bm{W}_{y_{i}}} can be computed with

While m≥3m\geq 3, we can still use Eq. (8), Eq. (9) and Eq. (10) to compute the forward and backward propagation.

Experiments and Results

We evaluate the generalized softmax loss in two typical vision applications: visual classification and face verification. In visual classification, we use three standard benchmark datasets: MNIST (LeCun et al., 1998), CIFAR10 (Krizhevsky, 2009), and CIFAR100 (Krizhevsky, 2009). In face verification, we evaluate our method on the widely used LFW dataset (Huang et al., 2007). We only use a single model in all baseline CNNs to compare our performance. For convenience, we use L-Softmax to denote the L-Softmax loss. Both Softmax and L-Softmax in the experiments use the same CNN shown in Table 1.

General Settings: We follow the design philosophy of VGG-net (Simonyan & Zisserman, 2014) in two aspects: (1) for convolution layers, the kernel size is 3×\times3 and 1 padding (if not specified) to keep the feature map unchanged. (2) for pooling layers, if the feature map size is halved, the number of filters is doubled in order to preserve the time complexity per layer. Our CNN architectures are described in Table 1. In convolution layers, the stride is set to 1 if not specified. We implement the CNNs using the Caffe library (Jia et al., 2014) with our modifications. For all experiments, we adopt the PReLU (He et al., 2015b) as the activation functions, and the batch size is 256. We use a weight decay of 0.0005 and momentum of 0.9. The weight initialization in (He et al., 2015b) and batch normalization (Ioffe & Szegedy, 2015) are used in our networks but without dropout. Note that we only perform the mean substraction preprocessing for training and testing data. For optimization, normally the stochastic gradient descent will work well. However, when training data has too many subjects (such as CASIA-WebFace dataset), the convergence of L-Softmax will be more difficult than softmax loss. For those cases that L-Softmax has difficulty converging, we use a learning strategy by letting fyi=λ∥Wyi∥∥xi∥cos⁡(θyi)+∥Wyi∥∥xi∥ψ(θyi)1+λf_{y_{i}}=\frac{\lambda\|\bm{W}_{y_{i}}\|\|\bm{x}_{i}\|\cos(\theta_{y_{i}})+\|\bm{W}_{y_{i}}\|\|\bm{x}_{i}\|\psi({\theta_{y_{i}}})}{1+\lambda} and start the gradient descent with a very large λ\lambda (it is similar to optimize the original softmax). Then we gradually reduce λ\lambda during iteration. Ideally λ\lambda can be gradually reduced to zero, but in practice, a small value will usually suffice.

MNIST, CIFAR10, CIFAR100: We start with a learning rate of 0.1, divide it by 10 at 12k and 15k iterations, and eventually terminate training at 18k iterations, which is determined on a 45k/5k train/val split.

Face Verification: The learning rate is set to 0.1, 0.01, 0.001 and is switched when the training loss plateaus. The total number of epochs is about is about 30 for our models.

Testing: we use the softmax to classify the testing samples in MNIST, CIFAR10 and CIFAR100 dataset. In LFW dataset, we use the simple cosine distance and the nearest neighbor rule for face verification.

2 Visual Classification

MNIST: Our network architecture is shown in Table 1. Table 2 shows the previous best results and those for our proposed L-Softmax loss. From the results, the L-Softmax loss not only outperforms the original softmax loss using the same network but also achieves the state-of-the-art performance compared to the other deep CNN architectures. In Fig. 2, we also visualize the learned features by the L-Softmax loss and compare them to the original softmax loss. Fig. 2 validates the effectiveness of the large margin constraint within L-Softmax loss. With larger mm, we indeed obtain a larger angular decision margin.

CIFAR10: We use two commonly used comparison protocols in CIFAR10 dataset. We first compare our L-Softmax loss under no data augmentation setup. For the data augmentation experiment, we follow the standard data augmentation in (Lee et al., 2015) for training: 4 pixels are padded on each side, and a 32×\times32 crop is randomly sampled from the padded image or its horizontal flip. In testing, we only evaluate the single view of the original 32×\times32 image. The results are shown in Table 3. One can observe that our L-Softmax loss greatly boosts the accuracy, achieving 1%-2% improvement over the original softmax loss and the other state-of-the-art CNNs.

CIFAR100: We also evaluate the generalize softmax loss on the CIFAR100 dataset. The CNN architecture refers to Table 1. One can notice that the L-Softmax loss outperform the CNN with softmax loss and all the other competitive methods. The L-Softmax loss improves more than 2.5% accuracy over the CNN and more than 1% over the current state-of-the-art CNN.

Confusion Matrix Visualization: We also give the confusion matrix comparison between the softmax baseline and the L-Softmax loss (m=4) in Fig. 5. Specifically we normalize the learned features and then calculate the cosine distance between these features. From Fig. 5, one can see that the intra-class compactness is greatly enhanced while the inter-class separability is also enlarged.

Error Rate vs. Iteration: Fig. 6 illustrates the relation between the error rate and the iteration number with different mm in the L-Softmax loss. We use the same CNN (same as the CIFAR10 network) to optimize the L-Softmax loss with m=1,2,3,4m=1,2,3,4, and then plot their training and testing error rate. One can observe that the original softmax suffers from severe overfitting problem (training loss is very low but testing loss is higher), while the L-Softmax loss can greatly avoid such problem. Fig. 7 shows the relation between the error rate and the iteration number with different number of filters in the L-Softmax loss (m=4). We use four different CNN architecture to optimize the L-Softmax loss with m=4m=4, and then plot their training and testing error rate. These four CNN architectures have the same structure and only differ in the number of filters (e.g. 32/32/64/128 denotes that there are 32, 32, 64 and 128 filters in every convolution layer of Conv0.x, Conv1.x Conv2.x and Conv3.x, respectively). On both the training set and testing set, the L-Softmax loss with larger number of filters performs better than those with smaller number of filters, indicating L-Softmax loss does not easily suffer from overfitting. The results also show that our L-Softmax loss can be optimized easily. Therefore, one can learn that the L-Softmax loss can make full use of the stronger learning ability of CNNs, since stronger learning ability leads to performance gain.

3 Face Verification

To further evaluate the learned features, we conduct an experiment on the famous LFW dataset (Huang et al., 2007). The dataset collects 13,233 face images from 5749 persons from uncontrolled conditions. Following the unrestricted with labeled outside data protocol (Huang et al., 2007), we train on the publicly available CASIA-WebFace (Yi et al., 2014) outside dataset (490k labeled face images belonging to over 10,000 individuals) and test on the 6,000 face pairs on LFW. People overlapping between the outside training data and the LFW testing data are excluded. As preprocessing, we use IntraFace (Asthana et al., 2014) to align the face images and then crop them based on 5 points. Then we train a single network for feature extraction, so we only compare the single model performance of current state-of-the-art CNNs. Finally PCA is used to form a compact feature vector. The results are given in Table 5. The generalize softmax loss achieves the current best results while only trained with the CASIA-WebFace outside data, and is also comparable to the current state-of-the-art CNNs with private outside data. Experimental results well validate the conclusion that the L-Softmax loss encourages the intra-class compactness and inter-class separability.

Concluding Remarks

We proposed the Large-Margin Softmax loss for the convolutional neural networks. The large-margin softmax loss defines a flexible learning task with adjustable margin. We can set the parameter mm to control the margin. With larger mm, the decision margin between classes also becomes larger. More appealingly, the Large-Margin Softmax loss has very clear intuition and geometric interpretation. The extensive experimental results on several benchmark datasets show clear advantages over current state-of-the-art CNNs and all the compared baselines.

Acknowledgement

The authors would like to thank Prof. Le Song (Georgia Tech) for constructive suggestions. This work is partially supported by the National Natural Science Foundation for Young Scientists of China (Grant no.61402289) and National Science Foundation of Guangdong Province (Grant no. 2014A030313558).

References