DisturbLabel: Regularizing CNN on the Loss Layer

Lingxi Xie, Jingdong Wang, Zhen Wei, Meng Wang, Qi Tian

Introduction

Deep Convolutional Neural Networks (CNNs) have shown significant performance gains in image recognition . The large image repository, ImageNet , and the high-performance computational resources such as GPUs played very important roles in the resurgence of CNNs. Meanwhile, a number of research attempts on various aspects have been made to learn the deep hierarchical structure better and faster . CNN also provides efficient visual features for other tasks .

In this paper, we propose DisturbLabel which imposes the regularization within the loss layer. In each training iteration, DisturbLabel randomly selects a small subset of samples (from those in the current mini-batch) and randomly sets their ground-truth labels to be incorrect, which results in a noisy loss function and, consequently, noisy gradient back-propagation. To the best of our knowledge, this is the first attempt to regularize the CNN on the loss layer. We show that DisturbLabel is an alternative approach to combining a large number of models that are trained with different noisy data. Experimental results show that DisturbLabel achieves comparable performance with Dropout and that it, in conjunction with Dropout, obtains better performance on several image classification benchmarks.

The rest of this paper is organized as follows. Section 2 briefly introduces related work. The DisturbLabel algorithm is presented in Section 3. The discussions and the cooperation with Dropout are presented in Sections 4 and 5, respectively. Experimental results are shown in Section 6, and we conclude our work in Section 7.

Related Work

The recent great success of CNNs in image recognition has benefitted from and inspired a wide range of research efforts, such as designing deeper network structures , exploring or learning non-linear activation functions , developing new pooling operations , introducing better optimization techniques , regularization techniques preventing the network from over-fitting , etc.

Recently, various regularization methods have been introduced. Data augmentation generates more training data as the input of the CNN by randomly cropping, rotating and flipping the input images , and adding noises to the image pixels . Dropout randomly discards a part of neuron response (the hidden neurons) during the training and only updates the remaining weights in each mini-batch iteration. DropConnect instead only updates a randomly-selected subset of weights. Stochastic Pooling changes the deterministic pooling operation and randomly samples one input as the pooling result in probability during training. Probabilistic Maxout instead turns the Maxout operation to stochastic. In contrast, our approach (DisturbLabel) imposes the regularization at the loss layer. A comparison of different regularization methods is summarized in Table 1.

There are other research works that are related to noisy labeling. explores the performance of discriminatively-trained CNNs on the noisy data, where there are some freely available labels for each image which may or may not be accurate. In contrast, our approach (DisturbLabel) assumes the labeling is correct and randomly changes the labels in a small probability in each mini-batch iteration, which means that the label of a training sample is correct in most iterations. Different from other works adding noise to the input unit , our work is a way of regularizing a neural network by adding noise to its loss layer (related to the output unit).

The DisturbLabel Algorithm

θ\boldsymbol{\theta} is often initialized as a set of white noises θ0\boldsymbol{\theta}_{0}, then updated using stochastic gradient descent (SGD) . The tt-th iteration of SGD updates the current parameters θt\boldsymbol{\theta}_{t} as:

Here, l ⁣(x,y)l\!\left(\mathbf{x},\mathbf{y}\right) is a loss function, e.g., softmax or square loss. ∇θt ⁣[l ⁣(x,y)]\nabla_{\boldsymbol{\theta}_{t}}\!\left[l\!\left(\mathbf{x},\mathbf{y}\right)\right] is computed using gradient back-propagation. Dt\mathcal{D}_{t} is a mini-batch randomly drawn from the training dataset D\mathcal{D}, and γt\gamma_{t} is the learning rate.

DisturbLabel works on each mini-batch independently. It performs an extra sampling process, in which a disturbed label vector y~=[y~1,y~2,…,y~C]⊤{\widetilde{\mathbf{y}}}={\left[\widetilde{y}_{1},\widetilde{y}_{2},\ldots,\widetilde{y}_{C}\right]^{\top}} is randomly generated for each data (x,y)\left(\mathbf{x},\mathbf{y}\right) from a Multinoulli (generalized Bernoulli) distribution P ⁣(α)\mathcal{P}\!\left(\alpha\right):

The Multinoulli distribution P ⁣(α)\mathcal{P}\!\left(\alpha\right) is defined as pc=1−C−1C⋅α{p_{c}}={1-\frac{C-1}{C}\cdot\alpha} and pi=1C⋅α{p_{i}}={\frac{1}{C}\cdot\alpha} for i≠c{i}\neq{c}. α\alpha is the noise rate, and cc is the ground-truth label (i.e., in the true label vector y\mathbf{y}, yc=1{y_{c}}={1}). In other words, we disturb each training sample with the probability α\alpha. For each disturbed sample, the label is randomly drawn from a uniform distribution over {1,2,…,C}\left\{1,2,\ldots,C\right\}, regardless of the true label.

The pseudo codes of DisturbLabel are listed above. An illustration of DisturbLabel is shown in Figure 1.

The noise rate α\alpha determines the expected fraction (C−1C⋅α\frac{C-1}{C}\cdot\alpha) of training data in a mini-batch which are assigned incorrect labels. When α=0%{\alpha}={0\%}, there are no noises involved, and the algorithm degenerates to the ordinary case. When α→100%{\alpha}\rightarrow{100\%}, we are actually discarding most of the labels and the training process becomes nearly unsupervised (as the probability assigned to any class is nearly 1C\frac{1}{C}). It is often necessary to set a relatively small α\alpha, although it is possible to obtain an efficient network with a rather large α\alpha (e.g., training the LeNet with α=90%{\alpha}={90\%} achieves <2%<2\% testing error rate on MNIST).

We evaluate the recognition accuracy on both MNIST and CIFAR10, by training the LeNet with different noise rates α\alpha, We summarize the results in Figures 3 and 3, respectively. It is observed that DisturbLabel with a relatively small α\alpha (e.g., 10%10\% or 20%20\%) achieves higher recognition accuracy than the model without regularization (α=0%{\alpha}={0\%}). This verifies that DisturbLabel does improve the generalization ability of the trained CNN model (Section 3.2 provides an empirical verification that the improvement comes from preventing over-fitting). When the noise rate α\alpha goes up to 50%50\%, DisturbLabel significantly causes the network to converge slower, meanwhile produces lower recognition accuracy compared to smaller noises, which is reasonable as the labels of the training data are not reliable enough.

2 DisturbLabel as a Regularizer

We empirically show that DisturbLabel is able to prevent the network training from over-fitting. Figures 5 and 5 show the results on the MNIST and CIFAR10 datasets using the normal training without regularization and the DisturbLabel algorithm over the same CNN structure. In addition, we also report the results when training the CNN with Dropout which is an alternative approach of CNN regularization.

We can observe that without regularization, the training error quickly drops to quite a low level, e.g., almost 0%0\% on MNIST and close to 3%3\% on CIFAR10, but the testing error stops decreasing at a high level. In contrast, the training error with DisturbLabel drops slower and is consistently larger than that without regularization. However, the testing error becomes lower, verifying that the improvement comes from preventing over-fitting. Similar results are also obtained in the case with the Dropout regularization. This reveals that training a CNN with regularization (either DisturbLabel or Dropout) yields stronger generalization ability.

Discussions

Consider the soft labeling problem, where each data point x\mathbf{x} is assigned to a label with probability. Denote the label vector by y′\mathbf{y}^{\prime}, which is a CC-dimensional vector, with the cc-th dimension being 1−C−1C⋅α1-\frac{C-1}{C}\cdot\alpha and all others being αC\frac{\alpha}{C}, and α\alpha is the noise level. Then we can normally train the CNN over the same data x\mathbf{x} and the soft label y′\mathbf{y}^{\prime}, using which the loss function is changed. Here we consider the standard cross-entropy loss function, and the derivation of other choices (e.g., logistic loss, square loss, etc.) is similar.

The gradient for a training sample (x,y′)\left(\mathbf{x},\mathbf{y}^{\prime}\right) is:

However in DisturbLabel, the gradient is computed as:

2 Interpretation as Model Ensemble

We show that DisturbLabel can be interpreted as an implicit model ensemble. Consider a normal noisy dataset D~={(xn,y~n)n=1N}{\widetilde{\mathcal{D}}}={\left\{\left(\mathbf{x}_{n},\widetilde{\mathbf{y}}_{n}\right)_{n=1}^{N}\right\}}, which is generated by assigning an incorrect label y~n≠yn{\widetilde{\mathbf{y}}_{n}}\neq{\mathbf{y}_{n}} to xn\mathbf{x}_{n} with a probability α\alpha for each data point in D\mathcal{D}. Combining neural network models that are trained on different noisy sets is usually helpful . However, separately training nets is prohibitively expensive as there are exponentially many noisy datasets. Even if we have already trained many different networks, combining them at the testing stage is very costly and often infeasible.

Each iteration in the DisturbLabel training process is like an iteration when the network is trained over a different noisy dataset D~\widetilde{\mathcal{D}} where a mini-batch of samples D~t\widetilde{\mathcal{D}}_{t} are drawn. Thus, training a neural network with DisturbLabel can be regarded as training many networks with massive weight sharing but over different training data, where each training sample is used very rarely.

It is interesting that Dropout can be interpreted as a way of approximately combining exponentially many different neural network architectures trained on the same data efficiently, while DisturbLabel can be regarded as a way of approximately combining exponentially many neural networks with the same architecture but trained on different noisy data efficiently. In Section 5, we will show that DisturbLabel cooperates with Dropout to produce better results than individual models.

3 Interpretation as Data Augmentation

We analyze DisturbLabel from a data augmentation perspective. Considering a data point (x,y)\left(\mathbf{x},\mathbf{y}\right) and its incorrect label y~\widetilde{\mathbf{y}}, the contribution to the loss is ∣f ⁣(x)−y~∣1\left|f\!\left(\mathbf{x}\right)-\widetilde{\mathbf{y}}\right|_{1}. This contribution can be rewritten as ∣(f ⁣(x)−y~+y)−y∣1\left|\left(f\!\left(\mathbf{x}\right)-\widetilde{\mathbf{y}}+\mathbf{y}\right)-\mathbf{y}\right|_{1}, where f~ ⁣(x)≐f ⁣(x)−y~+y{\widetilde{f}\!\left(\mathbf{x}\right)}\doteq{f\!\left(\mathbf{x}\right)-\widetilde{\mathbf{y}}+\mathbf{y}} can be viewed as a noisy output. Inspired by , we can project the noisy output f~ ⁣(x)\widetilde{f}\!\left(\mathbf{x}\right) back into the input space by minimizing the squared error ∥f~ ⁣(x)−f ⁣(x~)∥22\left\|\widetilde{f}\!\left(\mathbf{x}\right)-f\!\left(\widetilde{\mathbf{x}}\right)\right\|_{2}^{2}, where x~\widetilde{\mathbf{x}} is the augmented sample. In summary, the data point with a disturbed label (x,y~)\left(\mathbf{x},\widetilde{\mathbf{y}}\right) can be transformed to an augmented data point (x~,y)\left(\widetilde{\mathbf{x}},\mathbf{y}\right).

To verify that DisturbLabel indeed acts as data augmentation, we evaluate the algorithm on the MNIST dataset with only 1%1\% (600600) and 10%10\% (60006000) training samples, meanwhile keep the total number of iterations unchanged, i.e., each training sample is used 100×100\times and 10×10\times times as it is used in the original training process. With the LeNet , we obtain 10.92%10.92\% and 2.83%2.83\% error rates on the original testing set, respectively, which are dramatic compared to 0.86%0.86\% when the network is trained on the complete set. Meanwhile, in both cases, the training error rates quickly decrease to , implying that the limited training data cause over-fitting. DisturbLabel significantly decreases the error rates to 6.38%6.38\% and 1.89%1.89\%, respectively. As a reference, the error rate on the complete training set is further decreased to 0.66%0.66\% by DisturbLabel. This indicates that DisturbLabel improves the quality of network training with implicit data augmentation, thus it serves as an effective algorithm especially in the case that the amount training data is limited.

4 Relationship to Other CNN Training Methods

We briefly discuss the relationship between DisturbLabel and other network training algorithms.

Other regularization methods. There exist other network regularization methods, including DropConnect , Stochastic Pooling , Probabilistic Maxout , etc. Like Dropout which regularizes CNNs on the hidden neurons, these methods add regularization on other places such as neuron connections and pooling operations. DisturbLabel regularizes CNNs on the loss layer, which, to the best of our knowledge, has never been studied before. As we will show in the next section, DisturbLabel cooperates well with Dropout to obtain superior results to individual modules. We believe that DisturbLabel can also provide complementary information to other network regularization methods.

Other methods dealing with noises. Some previous works aim at training CNNs with noisy labels. We emphasize that DisturbLabel is intrinsically different with these methods since the problem settings are completely different. In these problems, training data suffer from noises (incorrect annotations), and researchers discuss the possibility of overcoming the noises towards an accurate training process. DisturbLabel, on the other hand, assumes that all the ground-truth labels are correct, and intentionally generates incorrect labels on a small fraction of data to prevent the network from over-fitting. In summary, inaccurate annotations may be harmful, but factitiously introducing noises is helpful to training a robust network.

Other network structures. We are also interested in other sophisticated network structures, such as the Maxout Networks , the Deeply-Supervised Nets , the Network-in-Network and the Recurrent Neural Networks . Thanks to the generality, DisturbLabel can be adopted on these networks to improve their generalization ability.

Cooperation with Dropout

We have shown that DisturbLabel regularizes the CNN on the loss layer. This is different from Dropout , which regularizes the CNN on hidden layers. DisturbLabel is an approximate ensemble of many CNN models with the same structure trained over different noisy datasets, while Dropout is an approximate ensemble of many CNN models with different structures trained over the same data. These two regularization strategies are complementary. We empirically discuss the combination of DisturbLabel with Dropout, which leads to an ensemble of many CNN models with different structures trained over different noisy data.

We report the results with various noise levels α\alpha for DisturbLabel and a fixed drop rate for Dropout over the MNIST and CIFAR10 datasets in Figures 9 and 9, respectively. In general, the proper combination of Dropout with DisturbLabel benefits the recognition accuracy improvement. In the MNIST dataset, the best result is obtained when α=10%{\alpha}={10\%}, meanwhile α=20%{\alpha}={20\%} performs much worse. We note that in Figure 3, without Dropout, α=10%{\alpha}={10\%} and α=20%{\alpha}={20\%} produce comparable results. In the CIFAR10 dataset, α=10%{\alpha}={10\%} does not help to improve the accuracy (close to baseline), while in Figure 3, we get comparable better results using α=10%{\alpha}={10\%} and α=20%{\alpha}={20\%}. The above experiments show that both DisturbLabel and Dropout add regularization to network training. If both strategies are adopted, we need to reduce the regularization power properly to prevent “under-fitting”.

In the later experiments, if both Dropout and DisturbLabel are used, we will reduce the noise level α\alpha by half to prevent the regularization on network training from being too strong. In the case that DisturbLabel provides strong regularization, e.g., in the case of ImageNet training where the wrong label is distributed over all the 10001000 categories, we will slightly decrease the drop rate of Dropout for the same purpose.

Experiments

We evaluate DisturbLabel on five popular datasets, i.e., MNIST and SVHN for digit recognition, CIFAR10/CIFAR100 for natural image recognition, and ImageNet for large-scale visual recognition.

MNIST is one of the most popular datasets for handwritten digit recognition. This dataset consists of 6000060000 training images and 1000010000 testing images, uniformly distributed over 1010 classes (0–9). All the samples are 28×2828\times 28 grayscale images.

We use a modified version of the LeNet as the baseline. The input image is passed through two units consisting of convolution, ReLU and max-pooling operations. In which, the convolutional kernels are of the scale 5×55\times 5, the spatial stride 11, and max-pooling operators are of the scale 2×22\times 2 and the spatial stride 22. The number of convolutional kernels are 3232 and 6464, respectively. After the second max-pooling operation, a fully-connect layer with 512512 filters is added, followed by ReLU and Dropout. The final layer is a 1010-way classifier with the softmax loss function. We use a set of abbreviation to represent the above network configuration as: [C55(S11P)@3232-MP22(S22)]-[C55(S11P)@6464-MP22(S22)]-FC512512-D0.50.5-FC1010.

To obtain higher recognition accuracy, we also train a more complicated BigNet. The cross-map normalization is adopted after each pooling layer, and the parameter KK for normalization is proportional to the logarithm of the number of kernels. The network configuration is abbreviated as: [C55(S11P22)@128128-MP33(S22)]-[C33(S11P11)@128128-D0.70.7-C33(S11P11)@256256-MP33(S22)]-D0.60.6- [C33(S11P11)@512512]-D0.50.5-[C33(S11P11)@10241024-MPSS(S11)]-D0.40.4-FC1010. Here, the number SS is the map size before the final (global) max-pooling, before which the down-sampling rate is 44. Therefore, if the input image size is W×WW\times W, S=⌊W/4⌋{S}={\left\lfloor W/4\right\rfloor}. The BigNet is feasible for data augmentation based on image cropping as the input image size is variable.

For data augmentation, we randomly crop input images into 24×2424\times 24 pixels. We apply (40,20,20,20)\left(40,20,20,20\right) training epochs for the LeNet-based configurations with learning rates (10−3,10−4,10−5,10−6)\left(10^{-3},10^{-4},10^{-5},10^{-6}\right). For the BigNet-based configurations, the numbers are (200,100,100,100)\left(200,100,100,100\right) and (10−2,10−3,10−4,10−5)\left(10^{-2},10^{-3},10^{-4},10^{-5}\right), respectively.

We evaluate DisturbLabel with the noise level α=20%{\alpha}={20\%}. According to the results shown in Table 2, DisturbLabel produces consistent accuracy gain over models without regularization, and also cooperates with Dropout to further improve the recognition performance. Train the BigNet using both Dropout and DisturbLabel achieves a 0.33%0.33\% error rate without data augmentation, which outperforms several recently reported results . In comparison with which applies more complicated data augmentation (e.g., image rotation), we only use randomly image cropping and obtain a comparable error rate (0.28%0.28\% vs. 0.23%0.23\%).

2 The SVHN Dataset

The SVHN dataset is a larger collection of 32×3232\times 32 RGB images, i.e., 7325773257 training samples, 2603226032 testing samples, and 531131531131 extra training samples. We preprocess the data as in the previous methods , i.e., selecting 400400 samples per category from the training set as well as 200200 samples per category from the extra set, using these 60006000 images for validation, and the remaining 598388598388 images as training samples. We also use Local Contrast Normalization (LCN) for data preprocessing .

We use another version of the LeNet. A 32×32×332\times 32\times 3 image is passed through three units consisting of convolution, ReLU and max-pooling operations. Using abbreviation, the network configuration can be written as: [C55(S11P22)@3232-MP33(S22)]-[C55(S11P22)@3232-MP33(S22)]-[C55(S11P22)@6464-MP33(S22)]-FC6464-D0.50.5-FC1010. Padding of 22 pixels wide is added in each convolution operation to preserve the width and height of the data. The BigNet is also used to achieve higher accuracy. The training epochs, learning rates and data augmentation settings remain the same as in the MNIST experiments.

We evaluate DisturbLabel with the noise level α=20%{\alpha}={20\%}, and summarize the results in Table 2. We can observe that DisturbLabel improves the recognition accuracy, either with or without using Dropout. With data augmentation and both regularization methods, we achieve a competitive 2.02%2.02\% error rate.

3 The CIFAR Datasets

The CIFAR10 and CIFAR100 datasets are both subsets drawn from the 8080-million tiny image database . There are 5000050000 images for training, and 1000010000 images for testing, all of them are 32×3232\times 32 RGB images. CIFAR10 contains 1010 basic categories, and CIFAR100 divides each of them into a finer level. In both datasets, training and testing images are uniformly distributed over all the categories. We use exactly the same network configuration as in the SVHN experiments, and add left-right image flipping into data augmentation with the probability 50%50\%.

We evaluate DisturbLabel with the noise level α=10%{\alpha}={10\%}. In CIFAR-100, we slightly modify DisturbLabel by only allowing disturbing the label among 1010 finer-level categories. We compare our results with the state-of-the-arts in Table 3. Once again, DisturbLabel produces consistent accuracy gain in every single case, either with or without Dropout. On CIFAR10, the BigNet with Dropout produces an excellent baseline (7.08%7.08\% error rate), and DisturbLabel further improves the performance (6.98%6.98\% error rate) with a complementary regularization function to Dropout. The success on the CIFAR datasets verifies that DisturbLabel is generalized: it not only works well in relatively simple digit recognition, but also helps natural image recognition tasks.

4 The ImageNet Database

The top-11 and top-55 error rates produced by the original AlexNet are 43.1%43.1\% and 19.9%19.9\%, respectively. When DisturbLabel is adopted, the error rates are reduced to 42.8%42.8\% and 19.7%19.7\%, respectively. We emphasize that the accuracy gain is not so small as it seems (e.g., the VGGNet combines two individually trained nets to get a 0.1%0.1\% gain), which, once again, verifies that DisturbLabel and Dropout cooperate well to provide regularization in different aspects. Figure 10 shows the error rate curve on the validation set. After about 2020 epochs, the model with DisturbLabel produces higher recognition accuracy at each testing phase.

While we only evaluate DisturbLabel on the AlexNet, we believe that it can also cooperate with other network architectures, such as the GoogLeNet and the VGGNet , since regularization is a common requirement of deep neural networks.

Conclusions

In this paper, we present DisturbLabel, a novel algorithm which regularizes CNNs on the loss layer. DisturbLabel is surprisingly simple, which works by randomly choosing a small subset of training data, and intentionally setting their ground-truth labels to be incorrect. We show that DisturbLabel consistently improves the network training process by preventing it from over-fitting, and that DisturbLabel can be explained as an alternative solution of implicit model ensemble and data augmentation. Meanwhile, DisturbLabel cooperates well with Dropout, which regularizes CNNs on the hidden neurons. Experiments verify that DisturbLabel achieves competitive performance on several popular image classification benchmarks.

References