Random Erasing Data Augmentation

Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, Yi Yang

Introduction

The ability to generalize is a research focus for the convolutional neural network (CNN). When a model is excessively complex, such as having too many parameters compared to the number of training samples, over-fitting might happen and weaken its generalization ability. A learned model may describe random error or noise instead of the underlying data distribution . In bad cases, the CNN model may exhibit good performance on the training data, but fail drastically when predicting new data. To improve the generalization ability of CNNs, many data augmentation and regularization approaches have been proposed, such as random cropping , flipping , dropout , and batch normalization .

Occlusion is a critical influencing factor on the generalization ability of CNNs. It is desirable that invariance to various levels of occlusion is achieved. When some parts of an object are occluded, a strong classification model should be able to recognize its category from the overall object structure. However, the collected training samples usually exhibit limited variance in occlusion. In an extreme case when all the training objects are clearly visible, i.e., no occlusion happens, the learned CNN will probably work well on testing images without occlusion, but, due to the limited generalization ability of the CNN model, may fail to recognize objects which are partially occluded. While we can manually add occluded natural images to the training set, it is costly and the levels of occlusion might be limited.

To address the occlusion problem and improve the generalization ability of CNNs, this paper introduces a new data augmentation approach, Random Erasing. It can be easily implemented in most existing CNN models. In the training phase, an image within a mini-batch randomly undergoes either of the two operations: 1) kept unchanged; 2) we randomly choose a rectangle region of an arbitrary size, and assign the pixels within the selected region with random values (or the ImageNet mean pixel value)In Section 5.1.2, we show erasing with random values achieves approximately equal performance to the ImageNet mean pixel value.. During Operation 2), an image is partially occluded in a random position with a random-sized mask. In this manner, augmented images with various occlusion levels can be generated. Examples of Random Erasing are shown in Fig. 1.

Two commonly used data augmentation approaches, i.e., random flipping and random cropping, also work on the image level and are closely related to Random Erasing. Both techniques have demonstrated the ability to improve the image recognition accuracy. In comparison with Random Erasing, random flipping does not incur information loss during augmentation. Different from random cropping, in Random Erasing, 1) only part of the object is occluded and the overall object structure is preserved, 2) pixels of the erased region are re-assigned with random values, which can be viewed as adding block noise to the image.

Working primarily on the fully connected (FC) layer, Dropout is also related to our method. It prevents over-fitting by discarding (both hidden and visible) units of the CNN with a probability pp. Random Erasing is somewhat similar to performing Dropout on the image level. The difference is that in Random Erasing, 1) we operate on a continuous rectangular region, 2) no pixels (units) are discarded, and 3) we focus on making the model more robust to noise and occlusion. The recent A-Fast-RCNN proposes an occlusion invariant object detector by training an adversarial network that generates examples with occlusion. Comparison with A-Fast-RCNN, Random Erasing does not require any parameter learning, can be easily applied to other CNN-based recognition tasks and still yields competitive accuracy with A-Fast-RCNN in object detection. To summarize, Random Erasing has the following advantages:

A lightweight method that does not require any extra parameter learning or memory consumption. It can be integrated with various CNN models without changing the learning strategy.

A complementary method to existing data augmentation and regularization approaches. When combined, Random Erasing further improves the recognition performance.

Consistently improving the performance of recent state-of-the-art deep models on image classification, object detection, and person re-ID.

Improving the robustness of CNNs to partially occluded samples. When we randomly adding occlusion to the CIFAR-10 testing dataset, Random Erasing significantly outperforms the baseline model.

Related Work

Regularization is a key component in preventing over-fitting in the training of CNN models. Various regularization methods have been proposed . Dropout randomly discards (setting to zero) the output of each hidden neuron with a probability during the training and only considers the contribution of the remaining weights in forward pass and back-propagation. Latter, Wan et al. propose a generalization of dropout approach, DropConect, which instead randomly selects weights to zero during training. In addition, Adaptive dropout is proposed where the dropout probability for each hidden neuron is estimated through a binary belief network. Stochastic Pooling randomly selects activation from a multinomial distribution during training, which is parameter free and can be applied with other regularization techniques. Recently, a regularization method named “DisturbLabel” is introduced by adding noise at the loss layer. DisturbLabel randomly changes the labels of small part of samples to incorrect values during each training iteration. PatchShuffle randomly shuffles the pixels within each local patch while maintaining nearly the same global structures with the original ones, it yields rich local variations for training of CNN.

Data augmentation is an explicit form of regularization that is also widely used in the training of deep CNN . It aims at artificially enlarging the training dataset from existing data using various translations, such as, translation, rotation, flipping, cropping, adding noises, etc. The two most popular and effective data augmentation methods in training of deep CNN are random flipping and random cropping. Random flipping randomly flips the input image horizontally, while random cropping extracts random sub-patch from the input image. As an analogous choice, Random Erasing may discard some parts of the object. For random cropping, it may crop off the corners of the object, while Random Erasing may occlude some parts of the object. Random Erasing maintains the global structure of object. Moreover, it can be viewed as adding noise to the image. The combination of random cropping and Random Erasing can produce more various training data. Recently, Wang et al. learn an adversary with Fast-RCNN detection to create hard examples on the fly by blocking some feature maps spatially. Instead of generating occlusion examples in feature space, Random Erasing generates images from the original images with very little computation which is in effect, computationally free and does not require any extra parameters learning.

Datasets

For image classification, we evaluate on three image classification datasets, including two well-known datasets, CIFAR-10 and CIFAR-100 , and a new dataset Fashion-MNIST . CIFAR-10 and CIFAR-100 contain 50,000 training and 10,000 testing 32×\times32 color images drawn from 10 and 100 classes, respectively. Fashion-MNIST consists of 60,000 training and 10,000 testing 28x28 gray-scale images. Each image is associated with a label from 10 classes. We evaluate top-1 error rates in the format “mean ±\pm std” based on 5 runs.

For object detection, we use the PASCAL VOC 2007 dataset which contains 9,963 images of 24,640 annotated objects in training/validation and testing sets. We use the “trainval” set for training and “test” set for testing.

For person re-identification (re-ID), the Market-1501 dataset contains 32,668 labeled bounding boxes of 1,501 identities captured from 6 different cameras. The dataset is split into two parts: 12,936 images with 751 identities for training and 19,732 images with 750 identities for testing. In testing, 3,368 hand-drawn images with 750 identities are used as probe set to identify the correct identities on the testing set. DukeMTMC-reID contains 36,411 images of 1,812 identities shot by 8 high-resolution cameras. Similar to Market-1501, it contains 16,522 training images of 702 identities, 2,228 query images of the other 702 identities and 17,661 gallery images. CUHK03 contains 14,096 images of 1,467 identities. We use the new training/testing protocol proposed in to evaluate the multi-shot re-ID performance. There are 767 identities in the training set and 700 identities in the testing set. We conduct experiment on both “detected” and “labeled” sets. We evaluate using rank-1 accuracy and mean average precision (mAP) on these three datasets.

Our Approach

This section presents the Random Erasing data augmentation method for training the convolutional neural network (CNN). We first describe the detailed procedure of Random Erasing. Next, the implementation of Random Erasing in different tasks is introduced. Finally, we analyze the differences between Random Erasing and random cropping.

In training, Random Erasing is conducted with a certain probability. For an image II in a mini-batch, the probability of it undergoing Random Erasing is pp, and the probability of it being kept unchanged is 1−p1-p. In this process, training images with various levels of occlusion are generated.

Random Erasing randomly selects a rectangle region IeI_{e} in an image, and erases its pixels with random values. Assume that the size of the training image is W×HW\times H. The area of the image is S=W×HS=W\times H. We randomly initialize the area of erasing rectangle region to SeS_{e}, where SeS\frac{S_{e}}{S} is in range specified by minimum sls_{l} and maximum shs_{h}. The aspect ratio of erasing rectangle region is randomly initialized between r1r_{1} and r2r_{2}, we set it to rer_{e}. The size of IeI_{e} is He=Se×reH_{e}=\sqrt{S_{e}\times r_{e}} and We=SereW_{e}=\sqrt{\frac{S_{e}}{r_{e}}}. Then, we randomly initialize a point P=(xe,ye)\mathcal{P}=(x_{e},y_{e}) in II. If xe+We≤Wx_{e}+W_{e}\leq W and ye+He≤Hy_{e}+H_{e}\leq H, we set the region, Ie=(xe,ye,xe+We,ye+He)I_{e}=(x_{e},y_{e},x_{e}+W_{e},y_{e}+H_{e}), as the selected rectangle region. Otherwise repeat the above process until an appropriate IeI_{e} is selected. With the selected erasing region IeI_{e}, each pixel in IeI_{e} is assigned to a random value in , respectively. The procedure of selecting the rectangle area and erasing this area is shown in Alg. 1.

2 Random Erasing for Image Classification and Person Re-identification

In image classification, an image is classified according to its visual content. In general, training data does not provide the location of the object, so we could not know where the object is. In this case, we perform Random Erasing on the whole image according to Alg. 1.

Recently, the person re-ID model is usually trained in a classification network for embedding learning . In this task, since pedestrians are confined with detected bounding boxes, persons are roughly in the same position and take up the most area of the image. In this scenario, we adopt the same strategy as image classification, as in practice, the pedestrian can be occluded in any position. We randomly select rectangle regions on the whole pedestrian image and erase it. Examples of Random Erasing for image classification and person re-ID are shown in Fig. 1.

3 Random Erasing for Object Detection

Object detection aims at detecting instances of semantic objects of a certain class in images. Since the location of each object in the training image is known, we implement Random Erasing with three schemes:1) Image-aware Random Erasing (IRE): selecting erasing regions on the whole image, the same as image classification and person re-identification; 2) Object-aware Random Erasing (ORE): selecting erasing regions in the bounding box of each object. In the latter, if there are multiple objects in the image, Random Erasing is applied on each object separately. 3) Image and object-aware Random Erasing (I+ORE): selecting erasing regions in both the whole image and each object bounding box. Examples of Random Erasing for object detection with the three schemes are shown in Fig. 2.

4 Comparison with Random Cropping

Random cropping is an effective data augmentation approach, it reduces the contribution of the background in the CNN decision, and can base learning models on the presence of parts of the object instead of focusing on the whole object. In comparison to random cropping, Random Erasing retains the overall structure of the object, only occluding some parts of object. In addition, the pixels of erased region are re-assigned with random values, which can be viewed as adding noise to the image. In our experiment (Section 5.1.2), we show that these two methods are complementary to each other for data augmentation. The examples of Random Erasing, random cropping, and the combination of them are shown in Fig. 3.

Experiment

In all of our experiment, we compare the CNN models trained with or without Random Erasing. For the same deep architecture, all the models are trained from the same weight initialization. Note that some popular regularization techniques (e.g., weight decay, batch normalization and dropout) and various data augmentations (e.g., flipping, padding and cropping) are employed. The compared CNN architectures are summarized as below:

Architectures. Four architectures are adopted on CIFAR-10, CIFAR-100 and Fashion-MNIST: ResNet , pre-activation ResNet , ResNeXt , and Wide Residual Networks . We use the 20, 32, 44, 56, 110-layer network for ResNet and pre-activation ResNet. The 18-layer network is also adopted for pre-activation ResNet. We use ResNeXt-29-8×\times64 and WRN-28-10 in the same way as and , respectively. The training procedure follows . Specially, the learning rate starts from 0.1 and is divided by 10 after the 150th and 225th epoch. We stop training by the 300th epoch. If not specified, all models are trained with data augmentation: randomly performs horizontal flips, and takes a random crop with 32×\times32 for CIFAR-10 and CIFAR-100 (28×\times28 for Fashion-MNIST) from images padded by 4 pixels on each side.

1.2 Classification Evaluation

Classification accuracy on different datasets. The results of applying Random Erasing on CIFAR-10 ,CIFAR-100 and Fashion-MNIST with different architectures are shown in Table 1. We set p=0.5p=0.5, sl=0.02s_{l}=0.02, sh=0.4s_{h}=0.4, and r1=1r2=0.3r_{1}=\frac{1}{r_{2}}=0.3. Results indicate that models trained with Random Erasing have significant improvement, demonstrating that our method is applicable to various CNN architectures. For CIFAR-10, our method improves the accuracy by 0.49% and 0.33% using ResNet-110 and ResNet-110-PreAct, respectively. In particular, our approach obtains 3.08% error rate using WRN-28-10, which improves the accuracy by 0.72% and achieves new state of the art. For CIFAR-100, our method obtains 17.73% error rate which gains 0.76% than the WRN-28-10 baseline. Our method also works well for gray-scale images: Random erasing improves WRN-28-10 from 4.01% to 3.65% in top-1 error on Fashion-MNIST.

The impact of hyper-parameters. When implementing Random Erasing on CNN training, we have three hyper-parameters to evaluate, i.e., the erasing probability pp, the area ratio range of erasing region sls_{l} and shs_{h}, and the aspect ratio range of erasing region r1r_{1} and r2r_{2}. To demonstrate the impact of these hyper-parameters on the model performance, we conduct experiment on CIFAR-10 based on ResNet18 (pre-act) under varying hyper-parameter settings. To simplify experiment, we fix sls_{l} to 0.02, r1=1r2r_{1}=\frac{1}{r_{2}} and evaluate pp, shs_{h}, and r1r_{1}. We set p=0.5p=0.5, sh=0.4s_{h}=0.4 and r1=0.3r_{1}=0.3 as the base setting. When evaluating one of the parameters, we fixed the other two parameters. Results are shown in Fig. 4.

Notably, Random Erasing consistently outperforms the ResNet18 (pre-act) baseline under all parameter settings. For example, when p∈[0.2,0.8]p\in[0.2,0.8] and sh∈[0.2,0.8]s_{h}\in[0.2,0.8], the average classification error rate is 4.48%4.48\%, outperforming the baseline method (5.17%5.17\%) by a large margin. Random Erasing is also robust to the aspect ratios of the erasing region. Specifically, our best result (when r1=0.3r_{1}=0.3, error rate = 4.31%4.31\%) reduces the classification error rate by 0.86% compared with the baseline. In the following experiment for image classification, we set p=0.5p=0.5, sl=0.02s_{l}=0.02, sh=0.4s_{h}=0.4, and r1=1r2=0.3r_{1}=\frac{1}{r_{2}}=0.3, if not specified.

Four types of random values for erasing. We evaluate Random Erasing when pixels in the selected region are erased in four ways: 1) each pixel is assigned with a random value ranging in , denoted as RE-R; 2) all pixels are assign with the mean ImageNet pixel value i.e., , denoted as RE-M; 3) all pixels are assigned with 0, denoted as RE-0; 4) all pixels are assigned with 255, denoted as RE-255. Table 2 presents the result with different erasing values on CIFAR10 using ResNet18 (pre-act). We observe that, 1) all erasing schemes outperform the baseline, 2) RE-R achieves approximately equal performance to RE-M, and 3) both RE-R and RE-M are superior to RE-0 and RE-255. If not specified, we use RE-R in the following experiment.

Comparison with Dropout and random noise. We compare Random Erasing with two variant methods applied on image layer. 1) Dropout: we apply dropout on image layer with probability λ1\lambda_{1}. 2) Random noise: we add different levels of noise on the input image by changing the pixel to a random value in with probability λ2\lambda_{2}. The probability of whether an image undergoes dropout or random noise is set to 0.5 as Random Erasing. Results are presented in Table 3. It is clear that applying dropout or adding random noise at the image layer fails to improve the accuracy. As the probability λ1\lambda_{1} and λ2\lambda_{2} increase, performance drops quickly. When λ2=0.4\lambda_{2}=0.4, the number of noise pixels for random noise is approximately equal to the number of erasing pixels for Random Erasing, the error rate of random noise increases from 5.17% to 6.52%, while Random Erasing reduces the error rate to 4.31%.

Comparing with data augmentation methods. We compare our method with random flipping and random cropping in Table 4. When applied alone, random cropping (6.33%) outperforms the other two methods. Importantly, Random Erasing and the two competing techniques are complementary. Particularly, combining these three methods achieves 4.31% error rate, a 7% improvement over the baseline without any augmentation.

Robustness to occlusion. Last, we show the robustness of Random Erasing against occlusion. In this experiment, we add different levels of occlusion to the CIFAR-10 dataset in testing. We randomly select a region of area and fill it with random values. The aspect ratio of the region is randomly chosen from the range of [0.3, 3.33]. Results as shown in Fig. 5. Obviously, the baseline performance drops quickly when increasing the occlusion level ll. In comparison, the performance of the model training with Random Erasing decreases slowly. Our approach achieves 56.36% error rate when the occluded area is half of the image (l=0.5l=0.5), while the baseline rapidly drops to 75.04%. It demonstrates that Random Erasing improves the robustness of CNNs against occlusion.

2 Object Detection

Experiment is conducted based on the Fast-RCNN detector. The model is initialized by the ImageNet classification models, and then fine-tuned on the object detection data. We experiment with VGG16 architecture. We follow A-Fast-RCNN for training. We apply SGD for 80K to train all models. The training rate starts with 0.001 and decreases to 0.0001 after 60K iterations. With this training procedure, the baseline mAP is slightly better than the report mAP in . We use the selective search proposals during training. For Random Erasing, we set p=0.5p=0.5, sl=0.02s_{l}=0.02, sh=0.2s_{h}=0.2, and r1=1r2=0.3r_{1}=\frac{1}{r_{2}}=0.3.

2.2 Detection Evaluation

We report results with using IRE, ORE and I+ORE during training Fast-RCNN in Table 5. The detector is trained with two training set, VOC07 trainval and union of VOC07 and VOC12 trainval. When training with VOC07 trainval, the baseline is 69.1% mAP. The detector learned with IRE scheme achieves an improvement to 70.5% mAP and the ORE scheme obtains 71.0% mAP. The ORE performs slightly better than IRE. When implementing Random Erasing on overall image and objects, the detector training with I+ORE obtains further improved in performance with 71.5% mAP. Our approach (I+ORE) outperforms A-Fast-RCNN by 0.5% in mAP. Moreover, our method does not require any parameter learning and is easy to implement. When using the enlarged 07+12 training set, the baseline is 74.8% which is much better than only using 07 training set. The IRE and ORE schemes give similar results, in which the mAP of IRE is improved by 0.8% and ORE is improved by 1.0%. When applying I+ORE during training, the mAP of Fast-RCNN increases to 76.2%, surpassing the baseline by 1.4%.

3 Person Re-identification

Three baselines are used in person re-ID, i.e., the ID-discriminative Embedding (IDE) , TriNet , and SVDNet . IDE and SVDNet are trained with the Softmax loss, while TriNet is trained with the triplet loss. The input images are resized to 256 ×\times 128.

For IDE, we basically follow the training strategy in . We further add a fully connected layer with 128 units after the Pool5 layer, followed by batch normalization, ReLU and Dropout. The Dropout probability is set to 0.5. We use SGD to train IDE. The learning rate starts with 0.01 and is divided by 10 after each 40 epochs. We train 100 epochs in total. In testing, we extract the output of Pool5 as feature for Market-1501 and DukeMTMC-reID datasets, and the fully connected layer with 128 units as feature for CUHK03.

For TriNet and SVDNet, we use the same model as proposed in and , respectively, and follow the same training strategy. In testing, we extract the last fully connected layer with 128 units as feature for TriNet and extract the output of Pool5 for SVDNet. Note that, we use 256 ×\times 128 as the input size to train SVDNet which achieves higher performance than the original paper using size 224 ×\times 224.

We use the ResNet-18, ResNet-34, and ResNet-50 architectures for IDE and TriNet, and ResNet-50 for SVDNet. We fine-tune them on the model pre-trained on ImageNet . We also perform random cropping and random horizontal flipping during training. For Random Erasing, we set p=0.5p=0.5, sl=0.02s_{l}=0.02, sh=0.2s_{h}=0.2, and r1=1r2=0.3r_{1}=\frac{1}{r_{2}}=0.3.

3.2 Person Re-identification Performance

Baseline Evaluation. The results of Random Erasing on Market-1501, DukeMTMC-reID, and CUHK03 with different baselines and architectures are shown in Table 6. For Market-1501 and DukeMTMC-reID, the IDE and SVDNet baselines outperform the TriNet baseline . Since there exists plenty of samples in each ID, the models with using the Softmax loss can learn better features. Specially, the IDE achieves 83.14% and 71.99% rank-1 accuracy on Market-1501 and DukeMTMC-reID with using ResNet-50, respectively. SVDNet gives rank-1 accuracy of 84.41% and 76.82% on Market-1501 and DukeMTMC-reID with ResNet-50, respectively. This is 1.81% higher for Market-1501 and 4.38% higher for DukeMTMC-reID than the TriNet with ResNet-50. However, on CUHK03, the performance of TriNet is higher than IDE and SVDNet, since the lack of training samples compromises the Softmax loss. TriNet obtains 49.86% rank-1 accuracy and 46.74% mAP on CUHK03 for the labeled setting with ResNet-50.

Random Erasing improves different baseline models. When implementing Random Erasing in these baseline models, we can observe that, Random Erasing consistently improves the rank-1 accuracy and mAP. Specifically, for Market-1501, Random Erasing improves the rank-1 by 3.10% and 2.67% for IDE and SVDNet with using ResNet-50. For DukeMTMC-reID, Random Erasing increases the rank-1 accuracy from 71.99% to 74.24% for IDE (ResNet-50) and from 76.82% to 79.31% for SVDNet (ResNet-50). For CUHK03, TriNet gains 8.28% and 5.0% in rank-1 accuracy when applying Random Erasing on the labeled and detected settings with ResNet-50, respectively. We note that, due to lack of adequate training data, over-fitting tend to occur on CUHK03. For example, a deeper architecture, such as ResNet-50, achieves lower performance than ResNet-34 when using the IDE mode on the detected subset. However, with our method, IDE (ResNet-50) outperforms IDE (ResNet-34). This indicates that our method can reduce the risk of over-fitting and improves the re-ID performance.

Comparison with the state-of-the-art methods. We compare our method with the state-of-the-art methods on Market-1501, DukeMTMC-reID, and CUHK03 in Table 7, Table 8, and Table 9, respectively. Our method achieves competitive results with the state of the art. Specifically, based on SVDNet, our method obtains rank-1 accuracy = 87.08% for Market-1501, and rank-1 accuracy = 79.31% for DukeMTMC-reID. On CUHK03, based on TriNet, our method achieves rank-1 accuracy = 58.14% for the labeled setting, and rank-1 accuracy = 55.50% for the detected setting.

When we further combine our system with re-ranking , the final rank-1 performance arrives at 89.13% for Market-1501, 84.02% for DukeMTMC-reID, and 64.43% for CUHK03 under the detected setting.

Conclusion

In this paper, we propose a new data augmentation approach named “Random Erasing” for training the convolutional neural network (CNN). It is easy to implemented: Random Erasing randomly occludes an arbitrary region of the input image during each training iteration. Experiment conducted on CIFAR10, CIFAR100, and Fashion-MNIST with various architectures validate the effectiveness of our method. Moreover, we obtain reasonable improvement on object detection and person re-identification, demonstrating that our method has good performance on various recognition tasks. In the future work, we will apply our approach to other CNN recognition tasks, such as, image retrieval and face recognition.

References