ImageNet-21K Pretraining for the Masses

Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, Lihi Zelnik-Manor

Introduction

ImageNet-1K dataset, introduced for the ILSVRC2012 visual recognition challenge , has been at the center of modern advances in deep learning . ImageNet-1K serves as the main dataset for pretraining of models for computer-vision transfer learning , and improving performances on ImageNet-1K is often seen as a litmus test for general applicability on downstream tasks . ImageNet-1K is a subset of the full ImageNet dataset , which consists of 14,197,122 images, divided into 21,841 classes. We shall refer to the full dataset as ImageNet-21K, following (although other papers sometimes described it as ImageNet-22K ). ImageNet-1K was created by selecting a subset of 1.2M images from ImageNet-21K, that belong to 1000 mutually exclusive classes.

Even though some previous works showed that pretraining on ImageNet-21K could provide better downstream results for large models , pretraining on ImageNet-1K remained far more popular. A main reason for this discrepancy is that ImageNet-21K labels are not mutually exclusive - the labels are taken from WordNet , where each image is labeled with one label only, not necessarily at the highest possible hierarchy of WordNet semantic tree. For example, ImageNet-21K dataset contains the labels "chair" and "furniture". A picture, with an actual chair, can sometimes be labeled as "chair", but sometimes be labeled as the semantic parent of "chair", "furniture". This kind of tagging methodology complicates the training process, and makes evaluating models on ImageNet-21K less accurate. Other challenges of ImageNet-21K dataset are the lack of official train-validation split, the fact that training is longer than ImageNet-1K and requires highly efficient training schemes, and that the raw dataset is large - 1.3TB.

Several past works have used ImageNet-21K for pretraining, mostly in comparison to larger datasets, which are not publicly available, such as JFT-300M . and used ImageNet-21K and JFT-300M to train expert models according to the datasets hierarchies, and combined them to ensembles on downstream tasks; and compared pretraining JFT-300M to ImageNet-21K on large models such as ViT and ResNet-50x4. Many papers used these pretrained models for downstream tasks (e.g., ). There are also works on ImageNet-21K that did not focus on pretraining: used extra (unlabled) data from ImageNet-21K to improve knowledge-distillation training on ImageNet-1K; used ImageNet-21k for testing few-shot learning; tested efficient softmax schemes on ImageNet-21k; tested pooling operations schemes on animal-oriented subset of ImageNet-21k.

However, previous works have not methodologically studied and optimized a pretraining process specifically for ImageNet-21K. Since this is a large-scale, high-quality, publicly available dataset, this kind of study can be highly beneficial to the community. We wish to close this gap in this work, and make efficient top-quality pretraining on ImageNet-21K accessible to all deep learning practitioners.

Our pretraining pipeline starts by preprocessing ImageNet-21K to ensure all classes have enough images for a meaningful learning, splitting the dataset to a standardized train-validation split, and resizing all images to reduce memory footprint. Using WordNet semantic tree , we show that ImageNet-21K can be transformed into a (semantic) multi-label dataset. We thoroughly analyze the advantages and disadvantages of single-label and multi-label training. Extensive tests on downstream tasks show that multi-label pretraining does not improve results on downstream tasks, despite having more information per image. To effectively utilize the semantic data, we develop a novel training method, called semantic softmax, which exploits the hierarchical structure of ImageNet-21K tagging to train the network over several semantic softmax layers, instead of the single layer. Using semantic softmax pretraining, we consistently outperform both single-label and multi-label pretraining on downstream tasks. By integrating semantic softmax into a dedicated semantic knowledge distillation loss, we further improved results. The complete end-to-end pretraining pipeline appears in Figure 1.

Using semantic softmax pretraining on ImageNet-21K we achieve significant improvement on numerous downstream tasks, compared to standard ImageNet-1K pretraining. Unlike previous works, which focused on pretraining of large models only , we show that ImageNet-21K pretraining benefits a wide variety of models, from larger models like TResNet-L , through medium-sized models like ResNet50 , and even small mobile-dedicated models like OFA-595 and MobileNetV3 . Our proposed pretraining scheme also outperforms previous ImageNet-21K pretraining schemes that were used to trained MLP-based models like Vision-Transformer (ViT) and Mixer .

The paper’s contribution can be summarized as follows:

We develop a methodical preprocess procedure to transform raw ImageNet-21K into a viable dataset for efficient, high-quality pretraining.

Using WordNet semantic tree, we convert each (single) label to semantic multi labels, and compare the pretrain quality of two baseline methods: single-label and multi-label pretraining. We show that while a multi-label approach provides more information per image, it can have significant optimization drawbacks, resulting in inferior results on downstream tasks.

We develop a novel training scheme called semantic softmax, which exploits the hierarchical structure of ImageNet-21K. With semantic softmax pretraining, we outperform both single-label and multi-label pretraining on downstream tasks. We further improve results by integrating semantic softmax into a dedicated semantic knowledge distillation scheme.

Via extensive experimentations, we show that compared to ImageNet-1K pretraining, ImageNet-21K pretraining significantly improves downstream results for a wide variety of architectures, include mobile-oriented ones. In addition, our ImageNet-21K pretraining scheme consistently outperforms previous ImageNet-21K pretraining schemes for prominent new models like ViT and Mixer.

Dataset Preparation

Our preprocessing stage consists of three steps, as described in Figure 1 (leftmost image): (1) invalid classes cleaning, (2) creating a validation set, (3) image resizing. Details are as follows: Step 1 - cleaning invalid classes: the original ImageNet-21K dataset consists of 14,197,122 images, each tagged in a single-label fashion by one of 21,841 possible classes. The dataset has no official train-validation split, and the classes are not well-balanced - some classes contain only 1-10 samples, while others contain thousands of samples. Classes with few samples cannot be learned efficiently, and may hinder the entire training process and hurt the pretrain quality . Hence we start our preprocessing stage by removing infrequent classes, with less than 500500 labels. After this stage, the dataset contains 12,358,688 images from 11,221 classes. Notice that the cleaning process reduced the number of total classes by half, but removed only 13%13\% of the original pictures. Step 2 - validation split: we allocate 5050 images per class for a standardized validation split, that can be used for future benchmarks and comparisons. Step 3 - image resizing: ImageNet-1K training usually uses crop-resizing which favours loading the original images at full resolution and resizing them on-the-fly. To make ImageNet-21K dataset more accessible and accelerate training, we resized during the preprocessing stage all the images to 224224 resolution (equivalent to squish-resizing ). While somewhat limiting scale augmentations, this stage significantly reduces the dataset’s memory footprint, from 1.3TB to 250GB, and makes loading the data during training faster.

After finishing the preprocessing stage, we kept only valid classes, produced a standardized train-validation split, and significantly reduced the dataset size. We shall name this processed dataset ImageNet-21K-P (P for Processed).

2 Utilizing Semantic Data

We now wish to analyze the semantic structure of ImageNet-21K-P dataset. This structure will enable us to better understand ImageNet-21K-P tagging methodology, and employ and compare different pretraining schemes.

Each image in the original ImageNet-21K dataset was labeled with a single label, that belongs to WordNet synset . Using the WordNet synset hyponym (subtype) and hypernym (supertype) relations, we can obtain for each class its parent class, if exists, and a list of child classes, if exists. When applying the parenthood relation recursively, we can build a semantic tree, that enables us to transform ImageNet-21K-P dataset into a multi-label dataset, where each image is associated with several labels - the original label, and also its parent class, parent-of-parent class, and so on. Example is given in Figure 1 (middle image) - the original image was labeled as ’swan’, but by utilizing the semantic tree, we can produce a list of semantic labels for the image - ’animal, vertebrate, bird, aquatic bird, swan’. Notice that the labels are sorted by hierarchy: ’animal’ label belongs to hierarchy , while ’swan’ label belongs to hierarchy 44. A label from hierarchy kk has kk ancestors.

Understanding the inconsistent tagging methodology

The semantic structure of ImageNet-21K enables us to understand its tagging methodology better. According to the stated tagging methodology of ImageNet-21K , we are not guaranteed that each image was labeled at the highest possible hierarchy. An example is given in Figure 2.

Two pictures, that contain the animal cow, were labeled differently - one with the label ’animal’, the other with the label ’cow’. Notice that ’animal’ is a semantic ancestor of ’cow’ (cow →\rightarrow placental →\rightarrow mammal →\rightarrow vertebrate →\rightarrow animal). This kind of incomplete tagging methodology, which is common in large datasets , hinders and complicates the training process. A dedicated scheme that tackles this tagging methodology will be presented in section 3.3.

Semantic statistics

By using WordNet synsets, we can calculate for each class the number of ancestors it has - its hierarchy. In total, our processed dataset, ImageNet-21K-P, has 1111 possible hierarchies. Example of classes from different hierarchies appears in Table 1. In Figure 4 in appendix A we present the number of classes per hierarchy. We see that while there are 1111 possible hierarchies, the vast majority of classes belong to the lower hierarchies.

Pretraining Schemes

In this section, we will review and analyze two baseline schemes for pretraining on ImageNet-21K-P: single-label and multi-label training. We will also develop a novel new scheme for pretraining on ImageNet-21K-P, semantic softmax, and analyze its advantages over the baseline schemes.

The straightforward way to pretrain on ImageNet-21K-P is to use the original (single) labels, apply softmax on the output logits, and use cross-entropy loss. Our single-label training scheme is similar to common efficient training schemes on ImageNet-1K , with minor adaptations to better handle the inconsistent tagging (Full training details appear in appendix B.1). Since we aim for an efficient scheme with maximal throughput, we don’t incorporate any tricks that might significantly increase training times. To further shorten the training times, we propose to initialize the models from standard ImageNet-1K training, and train on ImageNet-21K-P for 8080 epochs. On 8xV100 NVIDIA GPU machine, mixed-precision training takes 4040 minutes per epoch for ResNet50 and TResNet-M architectures (∼\sim 5000imgsec5000\frac{\textrm{img}}{\textrm{sec}}), leading to a total training time of 5454 hours. Similar accuracies are obtained when doing random initialization, but training the models longer - 140140 epochs. Pros of using single-label training

Well-balanced dataset - with single-label training on ImageNet-21K-P, the dataset is well-balanced, meaning each class appears, roughly, the same number of times.

Single-loss training - training with a softmax (a single loss) makes convergence easy and efficient, and avoids many optimization problems associated with multi-loss learning, such as different gradient magnitudes and gradient interference .

Inconsistent tagging - due to the tagging methodology of ImageNet-21K-P, where we are not guaranteed that an image was labeled at the highest possible hierarchy, ground-truth labels are inherently inconsistent. Pictures, containing the same object, can appear with different single-label tagging (see Figure 2 for example).

No semantic data - during training, we are not presenting semantic data via the single-label ground-truth.

2 Multi-label Training Scheme

Using the semantic tree, we can convert any (single) label to semantic multi labels, and train our models on ImageNet-21K-P in a multi-label fashion, expecting that the additional semantic information per image will improve the pretrain quality. As commonly done in multi-label classification , we reduce the problem to a series of binary classification tasks. Given NN labels, the base network outputs one logit per label, znz_{n}, and each logit is independently activated by a sigmoid function σ(zn)\sigma(z_{n}). Let’s denote yny_{n} as the ground-truth for class nn. The total classification loss, LtotL_{\text{tot}}, is obtained by aggregating a binary loss from the NN labels:

Eq. 1 formalizes multi-label classification as a multi-task problem. Since we have a large number of classes (11,22111,221), this is an extreme multi-task case. For training, we adopted the high-quality training scheme described in , that provided state-of-the-art results on large-scale multi-label datasets such as Open Images . Full training details appear in appendix B.2. Pros of using multi-label training

More information per image - we present for each image all the available semantic labels.

Tagging and metrics are more accurate - if an image was originally given a single label at hierarchy k, with multi-label training we are guaranteed that all ground-truth labels at hierarchies 0 to k are accurate. Hence, multi-label training partly mitigates the inconsistent tagging problem, and makes training metrics more accurate and reflective than single-label training.

Extreme multi-tasking - with multi-label training, each class is learned separately (sigmoids instead of softmax). This extreme multi-task learning makes the optimization process harder and less efficient, and may cause convergences to a local minimum .

Extreme imbalancing - as a multi-label dataset with many classes, ImageNet-21K-P suffers from a large positive-negative imbalance . In addition, due to the semantic structure, multi-label training is hindered by a large class imbalance - on average, classes from a lower hierarchy will appear far more frequent than classes from a higher hierarchy.

In appendices C.2 and E we show that for multi-label training, ASL loss , that was designed to cope with large positive-negative imbalancing, significantly outperforms cross-entropy loss, both on upstream and downstream tasks. This supports our analysis of extreme imbalancing as a major optimization challenge of multi-label training. Notice that we also list extreme multi-tasking as another optimization pitfall of multi-label training, and a dedicated scheme for dealing with it might further improve results. However, most methods that tackle multi-task learning, such as GradNorm and PCGrad , require computation of gradients for each class separately. This is computationally infeasible for a dataset with a large number of classes, such as ImageNet-21K-P.

3 Semantic Softmax Training Scheme

Our goal is to develop a dedicated training scheme that utilizes the advantages of both the single-label and the multi-label training. Specifically, our scheme should present for each input image all the available semantic labels, but use softmax activations instead of independent sigmoids to avoid extreme multi-tasking. We also want to have fully accurate ground-truth and training metrics, and provide the network direct data on the semantic hierarchies (this is not achieved even in multi-label training, the hierarchical structure there is implicit). In addition, the scheme should remain efficient in terms of training times.

To meet these goals, we develop a new training scheme called semantic softmax training. As we saw in section 2.2, each label in ImageNet-21K-P can belong to one of 1111 possible hierarchies. By definition, for each hierarchy there can be only one ground-truth label per input image. Hence, instead of single-label training with a single softmax, we shall have 1111 softmax layers, for the 1111 different hierarchies. Each softmax will sample the relevant logits from the corresponding hierarchy, as shown in Figure 1 (rightmost image). To deal with the partial tagging of ImageNet-21K-P, not all softmax layers will propagate gradients from each sample. Instead, we will activate only softmax layers from the relevant hierarchies. An example is given in Figure 3 - the original image had a label from hierarchy 55. We transform it to 66 semantic ground-truth labels, for hierarchies 0-5, and activate only the 66 first semantic softmax layers (only activated layers will propagate gradients). Compared to single-label and multi-label schemes, semantic softmax training scheme has the following advantages:

We avoid extreme multi-tasking (11,22111,221 uncoupled losses in multi-label training). Instead, we have only 1111 losses, as the number of softmax layers.

We present for each input image all the possible semantic labels. The loss scheme even provides direct data on the hierarchical structure.

Unlike single-label and multi-label training, semantic softmax ground-truth and training metrics are fully accurate. If a sample has no labels at hierarchy k, we don’t propagate gradients from the kth softmax during training, and ignore that hierarchy for metrics calculation (A dedicated metrics for semantic softmax training is defined in appendix C.3).

Calculating several softmax activations instead of a single one has negligible overhead, and in practice training times are similar to single-label training.

Weighting the different softmax layers

For each input image we have K losses (11). As commonly done in multi-task training , we need to aggregate them to a single loss. A naive solution will be to sum them: Ltot=∑k=0K−1LkL_{\text{tot}}=\sum_{k=0}^{K-1}L_{k} where LkL_{k}, the loss per softmax layer, is zero when the layer is not activated. However, this formulation ignores the fact that softmax layers at lower hierarchies will be activated much more frequently than softmax layers at higher hierarchies, resulting in over-emphasizing of classes from lower hierarchies. To account for this imbalancing, we propose a balancing logic: let NjN_{j} be the total number of classes in hierarchy j (as presented in Figure 4). Due to the semantic structure, the relative number of occurrences of hierarchy k in the loss function will be:

Hence, to balance the contribution of different hierarchies we can use a normalization factor Wk=1OkW_{k}=\frac{1}{O_{k}}, and obtain a balanced aggregation loss, that will be used for semantic softmax training:

4 Semantic Knowledge Distillation

Knowledge distillation (KD) is a known method to improve not only upstream, but also downstream results . We want to combine our semantic softmax scheme with KD training - semantic KD. In addition to the general benefit from KD training , for ImageNet-21K-P semantic KD has an additional benefit - it can predict the missing tags that arise from the inconsistent tagging. For example, for the left picture in Figure 2, the teacher model can predict the missing labels - ’cow, placental, mammal, vertebrate’. To implement semantic KD loss, for each hierarchy we will calculate both the teacher and the student the corresponding probability distributions {Ti}i=0K−1\left\{T_{i}\right\}_{i=0}^{K-1}, {Si}i=0K−1\left\{S_{i}\right\}_{i=0}^{K-1}. The KD loss of hierarchy ii will be:

where KDLoss is a standard measurement for the distance between distributions, that can be chosen as Kullback-Leibler divergence , or as MSE loss . We have found that the latter converges faster, and used it. A vanilla implementation for the total loss will be a simple sum of the losses from different hierarchies: LKD=∑i=0K−1LKDiL_{\text{KD}}=\sum_{i=0}^{K-1}L_{\text{KD}{i}}. However, this formulation assumes that all the hierarchies are relevant for each image. This is inaccurate - usually higher hierarchies represent subspecies of animals or plants, and are not applicable for a picture of a chair, for example. So we need to determine from the teacher predictions which hierarchies are relevant, and weigh the different losses accordingly. Let’s assume that for each hierarchy we can calculate the teacher confidence level, PiP_{i}. A confidence-weighted KD loss will be:

Eq. 5 is our proposed semantic KD loss. In appendix F we present a method to calculate the teacher confidence level, PiP_{i}, from the teacher predictions, similar to .

Experimental Study

In this section, we will present upstream and downstream results for the different training schemes, and show that semantic softmax pretraining outperforms single-label and multi-label pretraining. We will also demonstrate how semantic KD further improves results on downstream tasks.

In appendix C we provide upstream results for the three training schemes. Since each scheme has different training metrics, we cannot use these results to directly compare (pre)training quality.

2 Downstream Results

To compare the pretrain quality of different training schemes, we will test our models via transfer learning. To ensure that we are not overfitting a specific dataset or task, we chose a wide variety of downstream datasets, from different computer-vision tasks. We also ensured that our downstream datasets represent a variety of domains, and have diverse sizes - from small datasets of thousands of images, to larger datasets with more than a million images. For single-label classification, we transferred our models to ImageNet-1K , iNaturalist 2019 , CIFAR-100 and Food 251 . For multi-label classification, we transferred our models to MS-COCO and Pascal-VOC datasets. For video action recognition, we transferred our models to Kinetics 200 dataset . In appendix D we provide full training details on all downstream datasets.

In Table 2 we compare downstream results for three pretraining schemes: single-label, multi-label and semantic softmax.

We see that on 66 out of 77 datasets tested, semantic softmax pretraining outperforms both single-label and multi-label pretraining. In addition, we see from Table 2 that single-label pretraining performs better than multi-label pretraining (scores are higher on 55 out of 77 datasets tested).

These results support our analysis of the pros and cons of the different pretraining schemes from Section 3: with multi-label training, we have more information per input image, but the optimization process is less efficient due to extreme multi-tasking and extreme imbalancing. All-in-all, multi-label training does not improve downstream results. Single-label training, despite its shortcomings from the partial tagging methodology and the minimal information per image, provides a better pretraining baseline. Semantic softmax scheme, which utilizes semantic data without the optimization pitfalls of extreme multi-label training, outperforms both single-label and multi-label training.

Semantic KD

In Table 2 we also compare the downstream results of semantic softmax pretraining, with and without semantic KD. We see that on all tasks and datasets tested, adding semantic KD to our pretraining process improves downstream results. Indeed the ability of semantic KD to fill in the missing tags and provide a smoother and more informative ground-truth is translated to better downstream results. In appendix G we compare single-label pretraining with KD, to semantic softmax pretraining with semantic KD, and show that the latter achieves better results on downstream tasks.

Results

In the previous chapters we developed a dedicated pretraining scheme for ImageNet-21K-P dataset, semantic softmax, and showed that it outperforms two baseline pretraining schemes, single-label and multi-label, in terms of downstream results. Now we wish to compare our semantic softmax pretraining on ImageNet-21K-P to other known pretraining schemes and pretraining datasets.

We want to compare our proposed training scheme to other ImageNet-21K training schemes from the literature. However, to the best of our knowledge, no previous works have published their upstream results on ImageNet-21K, or shared thorough details about their training scheme or preprocessing stage. Recently, prominent new models called ViT and Mixer were published, and official pretrained weights were released . In Table 3 we compare downstream results when using the official ImageNet-21K weights, and when using weights from semantic softmax pretraining.

We see from Table 3 that our pretraining scheme significantly outperforms the official pretrain, on all downstream tasks tested. Previous works have observed that MLP-based models can be harder and less stable to use in transfer learning since they don’t have inherent translation inductive bias . When using the official weights, we also noticed this phenomenon on some datasets (Pascal-VOC, for example). Using semantic softmax pretraining, the transfer learning training was more stable and robust, and reached higher accuracy.

2 Comparison to ImageNet-1K Pretraining

In Table 4 we compare downstream results, for different models, when using ImageNet-1K pretraining (taken from ), and when using our ImageNet-21K-P pertraining. We can see that our pretraining scheme significantly outperforms standard ImageNet-1K pretraining on all datasets, for all models tested. For example, on iNaturalist dataset we improve the average top-1 accuracy by 2.9%2.9\%.

Notice that some previous works stated that pretraining on a large dataset benefits only large models . MobileNetV3 backbone, for example, has only 4.24.2M parameters, while ViT-B model has 85.685.6M parameters. Previous works assumed that a large number of parameters, like ViT has, is needed to properly utilize pretraining on large datasets. However, we show consistently and significantly that even small mobile-oriented models, like MobileNetV3 and OFA-595, can benefit from pretraining on a large (publicly available) dataset like ImageNet-21K-P. Due to their fast inference times and reduced heating, mobile-oriented models are used frequently for deployment. Hence, improving their downstream results by using better pretrain weights can enhance real-world products, without increasing training complexity or inference times.

3 ImageNet-1K SoTA Results

In Table 10 in appendix H we bring downstream results on ImageNet-1K for different models, when using ImageNet-21K-P semantic softmax pretraining. To achieve top results, similar to previous works , we added standard knowledge distillation loss into our ImageNet-1K training. To the best of our knowledge, for all the models in Table 10 we achieve a new SoTA record (for input resolution 224224). Unlike previous top works, which used private datasets , we are using a publicly available dataset for pretraining. Note that the gap from the original reported accuracies is significant. For example, MobileNetV3 reported accuracy was 75.2%75.2\% - we achieved 78.0%78.0\%; ResNet50 reported accuracy was 76.0%76.0\% - we achieved 82.0%82.0\%.

4 Additional Comparisons and Results

In appendix J we bring additional comparisons: (1) Comparison to Open Images pretraining; (2) Downstream results comparison on additional non-classification computer-vision tasks; (3) Impact of different number of training samples on upstream results.

Conclusion

In this paper, we presented an end-to-end scheme for high-quality efficient pretraining on ImageNet-21K dataset. We start by standardizing the dataset preprocessing stage. Then we show how we can transform ImageNet-21K dataset into a multi-label one, using WordNet semantics. Via extensive tests on downstream tasks, we demonstrate how single-label training outperforms multi-label training, despite having less information per image. We then develop a new training scheme, called semantic softmax, which utilizes ImageNet-21K hierarchical structure to outperform both single-label and multi-label training. We also integrate the semantic softmax scheme into a dedicated knowledge distillation loss to further improve results. On a variety of computer vision datasets and tasks, different architectures significantly and consistently benefit from our pretraining scheme, compared to ImageNet-1K pretraining and previous ImageNet-21K pretraining schemes.

In the past, pretraining on ImageNet-21K was out of scope for the common deep learning practitioner. With our proposed pipeline, high-quality efficient pretraining on ImageNet-21K will be more accessible to the deep learning community, enabling researchers to design new architectures and pretrain them to top results, without the need for massive computing resources or large-scale private datasets. In addition, our findings that even small mobile-oriented models significantly benefit from large-scale pretraining can be used to enhance real-world products. Finally, our improved pretraining scheme on ImageNet-21K can support prominent MLP-based models that require large-scale pretraining, like ViT and Mixer.

References

Appendix A Number of Classes in Different Hierarchies

Appendix B Training Details

To better handle the ground-truth inconsistencies of ImageNet-21K-P, we increase the label-smooth factor from the common value of 0.10.1 to 0.20.2. As explained in section 2.1, we also use squish-resizing instead of crop-resizing. We trained the models with input resolution 224224, using an Adam optimizer with learning rate of 3e-4 and one-cycle policy . When initializing our models from standard ImageNet-1K pretraining (pretraining weights taken from ), we found that 8080 epochs are enough for achieving strong pretrain results on ImageNet-21K-P. For regularization, we used RandAugment , Cutout , Label-smoothing and True-weight-decay . We observed that the common ImageNet statistics normalization does not improve the training accuracy, and instead normalized all the RGB channels to be between and 11. Unless stated otherwise, all runs and tests were done on TResNet-M architecture. On an 8xV100 NVIDIA GPU machine, training with mixed-precision takes 4040 minutes per epoch on ResNet50 and TResNet-M architectures (∼\sim 5000imgsec5000\frac{\textrm{img}}{\textrm{sec}}).

B.2 Multi-label ImageNet-21K-P Training Details

For multi-label training, we convert each image single label input to semantic multi labels, as described in section 2.2. Multi-label training details are similar to single-label training (number of epochs, optimizer, augmentations, learning rate, models initialization and so on), and training times are also similar. The main difference between single-label and multi-label training relies in the loss function: for multi-label training we tested 33 loss functions, following : cross-entropy (γ−=γ+=0\gamma_{-}=\gamma_{+}=0), focal loss (γ−=γ+=2\gamma_{-}=\gamma_{+}=2) and ASL (γ−=4,γ+=0\gamma_{-}=4,\gamma_{+}=0). For ASL, we tried different values of γ−\gamma_{-} to obtain the best mAP scores.

Appendix C Upstream Results

As we have a standardized dataset with a fixed train-validation split, the training metrics for each pretraining method can be used for future benchmark and comparisons.

For single-label training, regular top-1 accuracy metric becomes somewhat irrelevant - if pictures with similar content have different ground-truth labels, the network has no clear "correct" answer. Top-5 accuracy metric is more representative, but still limited. Upstream results of single-label training are given in Table 5. We can see that the top-1 accuracies obtained on ImageNet-21K-P, 37%−46%37\%-46\%, are significantly lower than the ones obtained on ImageNet-1K, 75%−85%75\%-85\%. This accuracy drop is mainly due to the semantic structure and inconsistent tagging methodology of ImageNet-21K-P. However, as we take bigger and better architectures, we see from Table 5 that the accuracies continue to improve, so we are not completely hindered by the inconsistent tagging.

C.2 Multi-label Upstream Results

For multi-label training, we will use the common micro and macro mAP accuracy as training metrics. However, due to the missing labels in the validation (and train) set, this metric also is not fully accurate. In Table 7 we compare the results for three possible loss functions for multi-label classification - cross-entropy, focal loss and ASL.

We see that ASL loss , that was designed to cope with large positive-negative imbalancing, outperform cross-entropy and focal loss. This is in agreement with our analysis in section 3.2, where we identify extreme imbalancing as one of the optimization challenges that stems from multi-label training.

C.3 Semantic Softmax Upstream Results

With semantic softmax training, we can calculate for each hierarchy its top-1 accuracy metric. We can also calculate the total accuracy by weighting the different accuracies by the number of classes in each hierarchy (see Figure 4). Notice that we are not using classes above the maximal hierarchy for our metrics calculation. Hence, and unlike single-label and multi-label training, with semantic softmax our training metrics are fully accurate.

In Figure 5 we present the top-1 accuracies achieved by different models on different hierarchy levels, when trained with semantic softmax (with KD).

Appendix D Downstream Datasets Training Details

For single-label classification, our downstream datasets were ImageNet-1K , iNaturalist 2019 , CIFAR-100 and Food-251 . For multi-label classification, our downstream datasets were MS-COCO and Pascal-VOC . For video action recognition, our downstream dataset was Kinetics-200 .

To minimize statistical uncertainty, for datasets with less than 150,000150,000 images (CIFAR-100, Food-251, MS-COCO, Pascal-VOC), we report result of averaging 33 runs with different seeds.

All results are reported for input resolution 224224.

For all downstream datasets we used cutout of 0.50.5, rand-Augment and true-weight-decay of 1e-4.

All single-label datasets are trained with label-smooth of 0.10.1

Unless stated otherwise, dataset was trained for 4040 epochs with Adam optimizer, learning rate of 3e-4, one-cycle policy and and squish-resizing.

Specific dataset details:

ImageNet-1K - Since the dataset is bigger than the others, we finetuned our networks for 100100 epochs using SGD optimizer, and learning rate of 4e-4. We used crop-resizing with the common minimal crop factor of 0.080.08.

MS-COCO - We used ASL loss with γ−=4\gamma_{-}=4.

Pascal-VOC - We used ASL loss with γ−=4\gamma_{-}=4, and learning rate of 5e-5.

Kinetics-200 - we trained for 3030 epochs with learning rate of 8e-5. We used the training method described in , with simple averaging of the embedding from each sample along the video.

Appendix E Downstream Results for Different Multi-label Losses

In Table 7 we compare downstream results when using multi-label pretraining with vanilla cross-entropy (CE) loss and ASL loss. We see that on all downstream datasets, pretraining with ASL leads to significantly better results

Appendix F Calculating Teacher Confidence

Using the teacher prediction for hierarchy i and the semantic ground-truth, we want to evaluate the teacher confidence level, PiP_{i}, so we can weight properly the contribution of different hierarchies in the KD loss. Our proposed logic for calculating the teacher’s (semantic) confidence is simple: - If the ground-truth highest hierarchy is higher than i, set PiP_{i} to 1. - Else, calculate the sum probabilities of the top 5%5\% classes in the teacher prediction (we deliberately don’t take only the probability of the highest class, to account for class similarities).

In Figure 6 we present the teacher confidence level for different hierarchies, averaged over an epoch.

We can see that lower hierarchies have, in average, higher confidence levels. This stems from the fact that not all hierarchies are relevant for each image. For the picture in Figure 3, for example, only hierarchies 0-5 are relevant, so we expect the teacher will have low confidence for hierarchies higher than 5.

Appendix G Semantic KD Vs Regular KD

Appendix H ImageNet-1K Transfer Learning Results

Appendix I ImageNet-21K-P - Winter21 Split

For a fair comparison to previous works, the results in the article are based on the original ImageNet-21K images, i.e. we are using Fall11 release of ImageNet-21K (fall11-whole.tar file), which contains all the original images and classes of ImageNet-21K. After we processed this release to create ImageNet-21K-P, we are left with a dataset that contains 11221 classes, where the train set has 11797632 samples and the test set has 561052 samples. We shall name this variant Fall11 ImageNet-21K-P.

Recently, the official ImageNet sitewww.image-net.org used our pre-processing methodology to offer direct downloading of ImageNet-21K-P, based on a new release of ImageNet-21K - Winter21 (winter21-whole.tar file). Compared to the original dataset, the Winter21 release removed some classes and samples. The Winter21 variant of ImageNet-21K-P is a dataset that contains 10450 classes, where the train set has 11060223 samples and the test set has 522500 samples. We shall name this variant Winter21 ImageNet-21K-P.

For enabling future comparison and benchmarking, we report the upstream accuracies also on this new variant of ImageNet-21K-P:

Note that the Winter21 variant of ImageNet-21K-P contains 10% fewer classes and 6% fewer images. In Table 12 we compare downstream results when using Winter21 and Fall11 variants of ImageNet-21K-P

We can see that compared to Fall11 variant, using Winter21 variant leads to a minor reduction in performances on downstream tasks.

Appendix J Additional Ablation Tests

In this section we will bring additional ablation tests and comparisons.

Open Images (v6) is a large scale multi-label dataset, which consists of 99 million training images and 96009600 labels. In Table 13 we compare downstream results when using two different datasets for pretraining: ImageNet-21K (semantic softmax training) and Open Images (multi-label training).

As we can see, ImageNet-21K pretraining consistently provides better downstream results than Open Images. A possible reason is that Open Images, as a multi-label dataset with large number of classes, suffers from the same multi-label optimization pitfalls we described in section 3.2.

J.2 Comparison on Additional Non-Classification Computer-Vision Tasks

In Table 14 and Table 15 we compare 1K and 21K pretraining on two additional computer-vision tasks: object detection (MS-COCO dataset) and image retrieval (INRIA holidays dataset).

We can see that also on non-classification tasks such as object detection and image retrieval, pretraining on ImageNet-21K translates to better downstream results than ImageNet-1K pretraining.

J.3 Impact of Different Number of Training Samples

In Figure 7 we test the impact of the number of training samples on on the upstream accuracies. As we can see, there is no saturation - more training images lead to better semantic accuracies.

Appendix K Pseudo-code

In the following sections we will bring pseudo-code (PyTorch-style) to some components in our semantic softmax training scheme: logits sampling, KD calculation and estimating teacher confidence.

K.2 KD Logic

K.3 Teacher Confidence

Appendix L Limitations

In this section we will discuss some of the limitations of our proposed pipeline for pretraining on ImageNet-21K:

1) While our work did put a large emphasis on the efficiency of the proposed pretraining pipeline, for reasonable training times we still need an 8-GPUs machine (1 GPU training will be quite long, 2-3 weeks).

2) For creating an efficient pretraining scheme, and also to stay within our inner computing budget, we did not incorporate training tricks that significantly increase training times, although some of these tricks might give additional benefits and improve pretraining quality.

An example - techniques for dealing with extreme multi-tasking, such as GradNorm and PCGrad , that would probably improve the pretrain quality of multi-label training, but would significantly increase training times.

Another example of methods from the literature we have not tested - general "semantic" techniques that can be used for training neural networks ( for example). We found that most of these techniques are not feasible for large-scale efficient training. In addition, we believe that since our novel method, semantic softmax, is designed and tailored to the specific needs and characterizations of ImageNet-21K, it will significantly outperform general semantic methods.