Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
Sangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han, Junsuk Choe, Sanghyuk Chun
Introduction
The ImageNet dataset has been at the center of modern advances in computer vision. Since the introduction of ImageNet, image recognition models based on convolutional neural networks have made quantum jumps in performances . Improving the model performance on ImageNet is seen as a litmus test for the general applicability of the model and the transfer learning performances on downstream tasks .
ImageNet, however, turns out to be noisier than one would expect. Recent studies have shed light on an overlooked problem with ImageNet that a significant portion of the dataset is composed of images with multiple possible labels. This contradicts the underlying assumption that there is only a single object class per image: the evaluation metrics penalize any prediction beyond the single ground-truth class. Thus, researchers have refined the ImageNet validation samples with multi-labeling policy using human annotators , and proposed new multi-label evaluation metrics. Under these new evaluation schemes, recent state-of-the-art models that seem to have surpassed the human level of recognition have been found to fall short of the human performance level.
The mismatch between the multiplicity of object classes per image and the assignment of single labels results in problems not only for evaluation, but also for training: the supervision becomes noisy. The widespread adoption of random crop augmentation aggravates the problem. A random crop of an image may contain an entirely different object from the original single label, introducing potentially wrong supervision signals during training, as in Figure 1.
The random crop augmentation makes supervision noisy not only for images with multiple classes. Even for images with a single class, the random crop often contains no foreground object. It is estimated that, under the standard training setupA random crop is sampled from % to of the entire image area., of the random crops have no overlap with the ground truths. Only of the random crops have the intersection-over-union (IoU) measure greater than with the ground truth boxes (see Figure 2). Training a model on ImageNet inevitably involves a lot of noisy supervision.
Ideally, for each training image, we want a human annotation telling the model (1) the full set of classes present (multi-label) and (2) where each object is located (localized label). One such format would be a dense pixel labeling where is the number of classes, as done for semantic segmentation ground truths. However, it is hardly scalable to collect even just the multi-label annotations for the 1.28 million ImageNet training samples. It took more than three months for five human experts (authors of ) to label mere 2,000 images.
We present an extensive set of evaluations for various model architectures trained with ReLabel on multiple datasets and tasks. On ImageNet classification, training ResNet-50 with ImageNet ReLabel has achieved a top-1 accuracy of 78.9%, a +1.4 pp gain over the baseline model trained with the original labels. The accuracy of ResNet-50 reaches 80.2% by employing the CutMix regularization on top, a new state-of-the-art performance on ImageNet to the best of our knowledge. Models trained with ReLabel have also consistently improved accuracies on ImageNet multi-label evaluation metrics proposed by . ReLabel and LabelPooling result in consistent improvements for transfer learning experiments, including the object detection and instance segmentation tasks on COCO and fine-grained classifications tasks. We further test LabelPooling on the multi-label classification task on COCO. Finally, we show that models trained with ReLabel are more resilient to test-time perturbations, as will be verified through experiments on several robustness benchmarks.
Related Works
We start this section by introducing prior works discussing the issues with ImageNet labels. We then discuss a few other research areas that share similarities with our approach. We describe the key differences from our approach.
Labeling issues in ImageNet. ImageNet has effectively served as the standard benchmark for the image classifiers: “methods live or die by their performance on this benchmark”, as argued by Shankar et al. . The reliability of the benchmark itself has thus come to be the subject of careful research and analysis. As with many other datasets, ImageNet contains much label noise . One of the most persistent and systematic types of label error on ImageNet is the erroneous single labels , referring to the cases where only one out of multiple present categories is annotated. Such errors are prevalent, as ImageNet contains many images with multiple classes. Shankar et al. and Beyer et al. have identified three subcategories for the erroneous single labels: (1) an image has multiple object classes, (2) there exist multiple labels that are synonymous or hierarchically including the other, and (3) inherent ambiguity in an image makes multiple labels plausible. Those studies have refined the validation set labels into multi-labels to establish an truthful and fair evaluation of models on effectively multi-label images. The focus of , however, has been only the validation, not training. has introduced a clean-up scheme to remove training samples with potentially erroneous labels by validating them with predictions from a strong classifier. Our work focuses on the clean-up strategy for the ImageNet training labels. Like , we utilize strong classifiers to clean up the training labels. Unlike , we correct the wrong labels, not remove. Our labels are also given per region. In our experiments, our method shows improved results compared to .
Knowledge distillation. Knowledge distillation (KD) also utilizes machine supervisions generated by the “teacher” network. Studies on KD have enriched and diversified the options for the teacher, such as feature map distillation , relation-based distillation , ensemble distillation , or iterative self-distillation . While those studies pursue stronger forms of supervision, none of them have considered a strong, state-of-the-art network as a teacher because it makes the KD supervision far heavier and impractical. With the random crop augmentation in place, every training iteration would involve a forward pass through the strong yet heavy teacher. Ours is similar in that the model is trained with machine supervision, but is more efficientFrom a more general view of KD that utilizes teacher and student in any form, our method can be seen as a new and efficient type of KD.. LabelPooling supervises a network with pre-computed label maps, rather than generating the label on the fly through the teacher for every random crop during training. We present the key advantages of ours against KD in Table 1.
Training tricks for ImageNet. Data augmentation is a simple yet powerful strategy for ImageNet training. The standard augmentation setting includes random cropping, flipping, and color jittering, as used in . In particular, the random crop augmentation, which crops random coordinates in an image and resize to a fixed size, is indispensable for a reasonable performance on ImageNet. Our work considers localized labels that make the supervision provided for each random crop region more sensible. There are additional training tricks for training classifiers that are orthogonal to our re-labeled training data. We show that those tricks can be combined with our re-labeling for improved performances.
Method
We propose a re-labeling strategy ReLabel to obtain pixel-level ground truth labels on the ImageNet training set. The label maps have two characteristics: (1) multi-class labels and (2) localized labels. The labels maps are obtained from a machine annotator: a strong image classifier trained on an extra data. We describe how to obtain the label maps and present a novel training framework, LabelPooling, to train image classifier using such localized multi-labels.
We obtain dense ground truth labels from a machine annotator, a state-of-the-art classifier that has been pre-trained on a super-ImageNet scale (\egJFT-300M or InstagramNet-1B ) and fine-tuned on ImageNet to predict ImageNet classes. Predictions from such a model are arguably close to human predictions . Since training the machine annotators requires an access to proprietary training data and hundreds of GPU or TPU days, we have adopted the open-source trained weights as the machine annotators. We show the comparison of different available machine annotators later in Section 3.3.
We remark that while the machine annotators are trained with single-label supervision on ImageNet, they still tend to make multi-label predictions for images with multiple categories. As an illustration, consider an image with two correct categories and . Assume that the model is fed with both and equal number of times during training, with those noisy labels. Then, the cross-entropy loss is given by where is the one-hot vector with at index and is the prediction vector for . Note that the minimal value for the function with respect to is taken at . Thus, in this example, the model minimizes the loss by predicting . Thus, if there exist much label noise in the dataset, a model trained with the single-label cross-entropy loss tends to predict multi-label outputs.
2 Training a Classifier with Dense Multi-labels
3 Discussion
So far we have introduced our labeling strategy and the supervision scheme using the label maps. We study the space and time consumption for our approach and examine design choices.
Time consumption. ReLabel requires a one-time cost for forward passing the ImageNet training images through the machine annotator. This procedure takes about 10 GPU-hours, which is only 3.3% of the entire train time for ResNet-50 (328 GPU-hours300 epochs on four NVIDIA V100 GPUs.). For each training iteration, LabelPooling performs the label map loading and regional pooling operations on top of the standard ImageNet supervision, which leads to only 0.5% additional training time. Note that ReLabel is much more computationally efficient than knowledge distillation which requires a forward pass through the teacher at every iteration. For example, KD with EfficientNet-B7 teacher takes more than four times the original training time.
Which machine annotator should we select? Ideally, we want the machine annotator to provide precise labels on training images. For this we consider ReLabel generated by a few state-of-the-art classifiers EfficientNet-{B1,B3,B5,B7,B8} , EfficientNet-L2 trained with JFT-300M , and ResNeXT-101_32x{32d,48d} trained with InstagramNet-1B . We train ResNet-50 with the above label maps from diverse classifiers. Note that ResNet-50 achieves the top-1 validation accuracy of when trained on vanilla single labels. We show the results in Figure 4. The performance of the target model overall follows the performance of the machine annotator. When the machine supervision is not sufficiently strong (\eg, EfficientNet-B1), the trained model shows a severe performance drop (76.1%). We choose EfficientNet-L2 as the machine annotator that has led to the best performance for ResNet-50 () in the rest of the experiments.
Factor analysis of ReLabel. ReLabel is both multi-label and pixel-wise. To examine the necessity of the two properties, we conduct an experiment by ablating each of them. We consider the localized single labels by taking argmax operation instead of softmax after the RoIAlign regional pooling, resulting in . For global multi-labels, we take the global average pooling, instead of the RoIAlign, over the label map, resulting in the label . Finally, by first performing the global average pooling and then performing argmax, we obtain the global single-labels, . Note that labels have the same format as the original ImageNet labels, but are machine-generated.
The results for those four variants are in Table 2. We observe that from the ReLabel performance of , the removal of multi-labels and localized labels results in -0.5 pp and -0.4 pp drops, respectively. When both are missing, there is a significant -1.4 pp drop. We thus argue that both ingredients are indispensable for a good performance. Note also that the global, single labels generated by a machine do not bring about any gain compared to the original ImageNet labels. This further signifies the importance of the aforementioned properties to benefit maximally from the machine annotations.
Confidence of ReLabel supervision. We study the confidence of ReLabel supervisions at different simulated levels of overlap between the random crop and the ground-truth bounding box. We draw M random crop samples as done for Figure 2. We measure the confidence for the ReLabel’s supervision in terms of the maximum class probability of the pooled label (\ie, where ). The results are shown in Figure 5. The averaged degree of supervision of ReLabel overall follows the degree of object existence, in particular, with small overlaps with object region (IoU ). For example, when IoU is zero (\ie, random crops are outside the object region), the label confidence is below , providing some uncertainty signals for the trained model.
Experiments
We present various experiments where we apply our labeling and training schemes for localized multi-label training. We first show the effectiveness of ReLabel on ImageNet classification with various network architectures and evaluation metrics, including the recently proposed multi-label evaluation metrics and robustness benchmarks (Section 4.1). Next, we show the transfer-learning performances for models trained with ReLabel when they are fine-tuned for object detection, instance segmentation, and fine-grained classification tasks (Section 4.2). We show that ReLabel improves the performances also for models on COCO multi-label classification tasks (Section 4.3). The re-labeled ImageNet training set, pre-trained weights, and the source code are availalble at https://github.com/naver-ai/relabel_imagenet.
We evaluate ReLabel strategy on the ImageNet-1K containing 1.28 million training images and 50,000 validation images of 1,000 object categories. We use standard data augmentation such as random cropping, flipping, color jittering, as in for all the models considered. We have trained the models with SGD for epochs with the initial learning rate and the cosine learning rate scheduling without restarts . The batch size and weight decay are set to and , respectively.
Comparison against other label manipulations. We compare ReLabel against prior methods that directly adjust the ImageNet labels. Label smoothing assigns a slightly weaker weight on the foreground class () and distributes the remaining weight uniformly across background classes. Label cleaning by Beyer et al. prunes out all training samples where the ground truth annotation does not agree with the prediction of a strong teacher classifier, namely BiT-L . For this, we use the list of clean sampled provided by the authors with our own training setting. We conducted the above label manipulation methods and ReLabel on ResNet-50. Results are given in Table 3. We measure the single-label accuracies on ImageNet validation and ImageNetV2 (Top-Images Results on ImageNetV2 “MatchedFrequency” and “ Threshold 0.7” are in Appendix C.). We show multi-label accuracies on two versions: ReaL and Shankar et al. . The metrics are identical: , where is the indicator function and is the top-1 prediction for a model . The ground-truth multi-label for image is given as a set . The difference between the metrics lies in the ground-truth multi-label annotation. We observe that ReLabel consistently achieves the best performance over all the metrics. We obtain validation accuracy with pp gain from the original labels, while the label smoothing and label cleaning boost only pp and pp, respectively. On ImageNetV2, ReaL, and Shankar et al.metrics, ReLabel achieves , , and accuracies, where the gains are pp, pp, and pp, respectively. It is notable that only ReLabel achieves remarkable boosts on the multi-label benchmarks. Label smoothing and cleaning shows only marginal gains or even worse multi-label accuracies (\eglabel cleaning results in a 0.1 pp worse result on Shankar et al.). We confirm that ReLabel improves the performances of image classifiers and that it helps models truly learn to make better multi-label predictions.
Results on various network architectures. We have trained various architectures with ReLabel to show that ReLabel is applicable to a wide range of networks with different training recipes. We consider ResNet-18, ResNet-101, EfficientNet-{B0,B1,B2,B3} , and ReXNet . Training details to make the best performance out of EfficientNet models are different from our base setting; we describe them in Appendix D.2. We follow the original paper’s training details for ReXNet . Results are shown in Table 4. ReLabel consistently enhances the performance of various network architectures. The 81.7% accuracy of EfficientNet-B3 is further improved to 82.5% with ReLabel.
State-of-the-art performance. ReLabel is complementary to many other training tricks used for achieving the best model performances. For example, we combine a strong regularizer CutMix with ReLabel. CutMix mixes two training images via cut-and-paste manner and likewise mixes the labels. To use it with ReLabel, we perform CutMix on the randomly cropped images. The pooled labels are then mixed according to the CutMix algorithm. We set the hyper-parameter of CutMix to . We show the results in Table 5. ReLabel with CutMix achieves the state-of-the-art ImageNet top-1 accuracies of 80.2% and 81.6% for the ResNet-50 and ResNet-101 backbones. On top of this, we further consider using the extra training data based on the ImageNet-21K dataset : M images with K categories. Unlike the previous work utilizing the ImageNet-21K with their original single-class labels over 21K categories, we perform ReLabel on them to generate multi-labels over the K classes. We then sub-sample M training data from the entire M training images by balancing the top-1 class distributions, as done in . Training with this extra data and CutMix on top of ReLabel boosts the accuracy of ResNet-50 to 81.2%. In summary, ReLabel is a practical addition to existing training tricks that consistently improves the backbone performances.
Comparison against knowledge distillation. We compare ReLabel against knowledge distillation (KD) in terms of the performance and training time costs. We train ResNet-50 with EfficientNet teachers: EfficientNet-{B1,B3,B5,B7}; we have not considered performing KD with EfficientNet-L2 as it would take 160 GPU days, beyond our computational capacity. Training details for KD are in Appendix D.3. Figure 6 shows the results. We plot the target model’s performance versus the required number of GPU days. KD with smaller teacher variants (EfficientNet-{B1,B3}) shows worse top-1 accuracies than ReLabel at higher training costs. For larger teachers (EfficientNet-{B5,B7}), KD achieves comparable performances with ReLabel (\eg for KD with B7 and for ReLabel). However, they require and GPU days to train, compared to mere using ReLabel. ReLabel training is almost as fast as the original training.
Storage-performance trade off. We study the trade off between the storage space for the label maps and the model performance. ReLabel only saves top- prediction maps for the interest of efficient storage, where the default value is 5. We explore . We also study the impact of quantization levels for label maps: -bit and -bit floating point, instead of the default -bit floating point. The results are in Table 6. ReLabel achieves a good performance-efficiency trade-off when . Quantizing label maps tend to yield only small performance drops ( pp to pp). When storage space is a crucial constraint, we advise users to adopt labels maps of coarser formats.
Combination with original labels. When we combine ReLabel’s annotation with the original label as , the performance degrades from 78.9% to 78.3% accuracy on ImageNet. They do not seem to make a good combination.
Robustness. We evaluate the robustness of ReLabel-trained models against test-time perturbations. We consider adversarial and natural perturbations: FGSM , ImageNet-A , ImageNet-C , and background challenge (BGC) . FGSM introduces one-step adversarial perturbations on images, while ImageNet-A samples consistent of common failure cases for modern image classifiers. ImageNet-C consists of 15 different types of natural perturbations. BGC evaluates the robustness against backgrounds by selecting background images adversarially from the dataset. Results are in Table 7. We observe that ReLabel consistently improves the resilience of models on adversarial and natural perturbations. Especially, ReLabel shows remarkable improvements in the background robustness (+8.7%) owing the localized supervision. Furthermore, combining ReLabel with other training strategies, \eg, CutMix and extra training data, significantly boosts the performances in the all robustness benchmarks.
ReLabel examples on ImageNet. We present examples generated by ReLabel during ImageNet training in Appendix E. As shown in the examples, ReLabel can generate location-specific multi-labels with more precise supervision than the original ImageNet labels.
2 Transfer Learning
Apart from serving as the standard benchmark, ImageNet has contributed to the computer vision research and engineering with its suite of pre-trained models. When the target task has only a small number of annotated data, transfer learning from the ImageNet pre-training usually helps . We examine here whether the ReLabel-induced improvements on the ImageNet performances transfer to various downstream tasks. We present the results of 5 fine-grained classification tasks and the object detection and instance segmentation tasks on COCO with models pre-trained on ImageNet with ReLabel.
Fine-grained classification tasks. We evaluate ReLabel-pretrained ResNet-50 on five fine-grained classification tasks: Food-101 , Stanford Cars , DTD , FGVC Aircraft , and Oxford Pets . We use the standard data augmentation as in Section 4.1. Models are fine-tuned with SGD for 5,000 iterations, following the convention for fine-tuning tasks . To find the best learning rate and weight decay values for each task, we perform a grid search per task and report the best performance. Table 8 shows the results. Note that the ReLabel-trained model results in a consistent improvement over the vanilla pre-trained model. For example, on FGVC Aircraft, ReLabel pre-training improves the downstream task performance by pp.
Object detection and instance segmentation. We used Faster-RCNN and Mask-RCNN with feature pyramid network (FPN ) as the base models for object detection and instance segmentation tasks, respectively. The backbone networks of Faster-RCNN and Mask-RCNN are initialized with ReLabel-pretrained ResNet-50 model, and then fine-tuned on COCO dataset by the original training strategy with the image size of . Table 9 shows the results. Pre-training with ReLabel improves the bbox AP of Faster-RCNN by pp and the mask AP of Mask-RCNN by pp. Pre-training a model with cleaner supervision like ReLabel leads to better feature representations and boosts the object detection and instance segmentation performances.
3 Multi-label Classification
Conclusion
We have proposed a re-labeling strategy, ReLabel, for the 1.28 million training images on ImageNet. ReLabel transforms the single-class labels assigned once per image into multi-class labels assigned for every region in an image, based on a machine annotator. The machine annotator is a strong classifier trained on a large extra source of visual data. We also proposed a novel scheme for training a classifier with the localized multi-class labels (LabelPooling). We experimentally verified significant performance gains induced by our labels and the corresponding training technique. ReLabel results in a consistent gain across tasks, including the ImageNet benchmarks, transfer-learning tasks, and multi-label classification tasks. We will open-source the localized multi-labels from ReLabel and the corresponding pre-trained models.
Acknowledgement We thank NAVER AI Lab members for valuable discussion and advice. NAVER Smart Machine Learning (NSML) has been used for experiments.
References
Appendix
Appendix A ReLabel Algorithm
We present the pseudo-codes of ReLabel in Algorithm A1. We assume the minibatch size is for simplicity. First, an input image and its saved label map are loaded from the dataset. Then the random crop augmentation is conducted on the input image. We then perform RoIAlign on the label map with the random crop coordinates [,,,]. Finally softmax function is conducted on the pooled label map to get a multi-label ground-truth in . The multi-label ground-truth is used for updating the model with the standard cross-entropy loss.
Appendix B Re-labeling ImageNet: Detailed Procedure and Examples
We utilize EfficientNet-L2 as our machine annotator classifier whose input size is . For all training images, we resize them into without cropping and generate label maps by feed-forwarding them. The spatial size of label map is , number of channel is , and the number of classes is .
Appendix C Results on ImageNetV2
We present full ImageNetV2 results in Table A1. Three metrics “Top-Images”, “Matched Frequency”, and “Threshold 0.7” are reported with two baselines Label smoothing and Label cleaning . ReLabel obtained , , and accuracies on ImageNetV2 “Top-Images”, “Matched Frequency”, and “Threshold 0.7”, where the gains are , , and pp against the vanilla ResNet-50, respectively.
Appendix D Implementation details
We present the implementation details in this section.
In most experiments, we have trained the models with SGD optimizer with learning rate and weight decay . For further improved performance, we utilized AdamP optimizer with learning rate , and weight decay when applying additional tricks such as CutMix regularizer or extra training data (ImageNet-21K).
D.2 EfficientNet on ImageNet
We utilize an open-source pytorch codebase to train EfficientNet variants on ImageNet. We utilize AdamP optimizer and set training epochs , minibatch size , learning rate , and weight decay with four NVIDIA V100 GPUs. Dropout and drop path regularizers are used with dropout rate and drop path rate , respectively. We also utilize Random erasing , RandAugment , and Mixup augmentations as suggested in . All training settings are samely used for both vanilla training and ReLabel training of EfficientNet variants.
D.3 Knowledge Distillation
Training with knowledge distillation is also conducted on the pytorch codebase . For teacher network, we use official EfficientNet (B1-B7) weights trained with noisy student techniques. We utilize outputs of networks after soft-max layer and the cross-entropy loss between teacher and students is only used for the distillation loss . The temperature and cross-entropy with ground truth were not used. Since the EfficientNet teachers are trained with large-size images (), we put the large-size image for teacher network and resize it to for inputs of student network. We adopt SGD with Nesterov momentum for the optimizer and the standard setting with long epochs: learning rate , weight decay , batch size , training epochs and cosine learning rate schedule with four NVIDIA V100 GPUs.
D.4 COCO Multi-label Classification
Appendix E ReLabel Examples on ImageNet
We present ReLabel exsamples on ImageNet training data in Figure A2. We show the full training images (left) and the random cropped patches (right). The random crop coordinates are denoted by blue bounding boxes. We also present the original ImageNet label and the new multi-labels by ReLabel. As shown in the examples, ReLabel can generate location-specific multi-labels with more precise supervision than the original ImageNet label.