Concept Generalization in Visual Representation Learning

Mert Bulent Sariyildiz, Yannis Kalantidis, Diane Larlus, Karteek Alahari

Introduction

There has been an increasing effort to tackle the need for manually-annotated large-scale data in deep models via transfer learning, i.e., by transferring representations learned on resourceful datasets and tasks to problems where annotations are scarce. Prior work has achieved this in various ways, such as, imitating knowledge transfer in low-data regimes , exploiting unlabeled data in a self- or weakly- supervised manner.

The quality of the learned visual representations for transfer learning is usually determined by checking whether they are useful for, i.e., generalize to, a wide range of downstream vision tasks. Thus, it is imperative to quantify this generalization, which has several facets, such as generalization to different input distributions (e.g., from synthetic images to natural ones), to new tasks (e.g., from image classification to object detection), or to different semantic concepts (e.g., across different object categories or scene labels). Although the first two facets have received much attention recently , we observe that a more principled analysis is needed for the last one.

As also noted by , the effectiveness of knowledge transfer between two tasks is closely related to the semantic similarity between the concepts considered in each task. However, assessing this relatedness is not straightforward, as the semantic extent of a concept may depend on the task itself. In practice, models consider an exhaustive list of downstream tasks that cover a wide range of concepts in order to test their transfer learning capabilities. Previous attempts discussing this issue have been limited to intuition . We still know little about the impact of the semantic relationship between the concepts seen during training visual representations and those seen during their evaluation (seen and unseen concepts, respectively).

In this paper, we study the generalization capabilities of visual representations across concepts that exist in a large, popular, and broad ontology, the subset of WordNet used to build ImageNet-21K (IN-21K), while keeping all the other generalization facets fixed. Starting from a set of seen concepts, the concepts from the popular ImageNet-1K (IN-1K) dataset, we leverage semantic similarity metrics based on this ontology crafted by experts to measure the semantic distance between IN-1K and every unseen concept (i.e., any concept from IN-21K that is not in IN-1K). We rank unseen concepts with respect to their distance to IN-1K and define a sequence of five, IN-1K-sized concept generalization levels, each consisting of a distinct set of unseen concepts with increasing semantic distance to the seen ones. This results in a large-scale benchmark that consists of five thousand concepts, that we refer to as the ImageNet Concept Generalization benchmark, or ImageNet-CoG in short. The benchmark construction process is illustrated in Fig. 1.

Given a model trained on IN-1K, the evaluation protocol for ImageNet-CoG consists of two phases: it first extracts features for images of IN-1K and of the five concept generalization levels, and then learns individual classifiers, for each level, using a varying amount of samples per concept. By defining the set of seen concepts for our benchmark to be IN-1K classes, we are able to evaluate models trained on IN-1K out-of-the box. We therefore use publicly available pretrained models and analyse a large number of popular models under the prism of concept generalization. Our contributions are as follows.

We propose a systematic way to study concept generalization, by defining a set of seen concepts along with sets of unseen concepts that are semantically more and more distant from the seen ones.

We design ImageNet-CoG, a large-scale benchmark, which embodies this systematic way. It is designed to evaluate models pretrained on IN-1K out-of-the-box and draws unseen concepts from the rest of the IN-21K dataset. We measure concept generalization performance on five, IN-1K-sized levels, by learning classifiers with a few or all the training images from the unseen concepts.

We conduct a large-scale study benchmarking 31 state-of-the-art visual representation learning approaches on ImageNet-CoG and analyse how different architectures, levels of supervision, regularization techniques and additional web data impact the concept generalization performance, uncovering several interesting insights.

Related Work

Generalization has been studied under different perspectives such as regularization and augmentation techniques, links to human cognition , or developing quantitative metrics to better understand it, e.g., through loss functions or complexity measures . Several dimensions of generalization have also been explored in the context of computer vision, for instance, generalization to different visual distributions of the same concepts (domain adaptation) , or generalization across tasks . Generalization across concepts is a crucial part of zero-shot and few-shot learning. We study this particular dimension, concept generalization, whose goal is to transfer knowledge acquired on a set of seen concepts, to newly encountered unseen concepts as effectively as possible. Different from existing work, we take a systematic approach by considering the semantic similarity between seen and unseen concepts when measuring concept generalization.

Towards a structure of the concept space. One of the first requirements for rigorously evaluating concept generalization is structuring the concept space, in order to analyze the impact of concepts present during pretraining and transfer stages. However, previous work rarely discusses the particular choices of splits (seen vs. unseen) of their data, and random sampling of concepts remains the most common approach . A handful of methods leverage relations designed by experts. The WordNet graph for instance helps build dataset splits in and a domain-specific ontology is used to test cross-domain generalization . These splits are however based on heuristics, instead of principled mechanisms built on semantic relationship between concepts as we do in this paper.

Transfer learning evaluations. When it comes to evaluating the quality of visual representations, the gold standard is to benchmark models by solving diverse tasks such as classification, detection, segmentation and retrieval on many datasets . The most commonly used datasets are IN-1K , Places , SUN , Pascal-VOC , MS-COCO . Such choices, however, are often made independently from the dataset used to train the visual representations, ignoring their semantic relationship.

In summary, semantic relations between pretraining and transfer tasks have been overlooked in evaluating the quality of visual representations. To address this issue, we present a controlled evaluation protocol that factors in such relations.

Our ImageNet CoG Benchmark

Transfer learning performance is highly sensitive to the semantic similarity between concepts in the pretraining and the target datasets . Studying this relationship requires carefully constructed evaluation protocols: i) controlling which concepts a model has been exposed to during training (seen concepts), and ii) the semantic distance between these seen concepts and those considered for the transfer task (unseen concepts). As discussed earlier, current evaluation protocols severely fall short on handling these aspects. To fill this gap, we propose ImageNet Concept Generalization (CoG)—a benchmark composed of multiple image sets, one for pretraining and several others for transfer, curated in a controlled manner in order to measure the transfer learning performance of visual representations to sets of unseen concepts with increasingly distant semantics from the ones seen during training.

While designing this benchmark, we considered several important points. First, in order to exclusively focus on concept generalization, we need a controlled setup tailored for this specific aspect of generalization. In other words, we need to make sure that the only change between the pretraining and the transfer datasets is the set of concepts. In particular, we need the input image distribution (natural images) and the annotation process (which may determine the statistics of images ) to remain constant.

Second, to determine the semantic similarity between two concepts, we need an auxiliary knowledge base that can provide a notion of semantic relatedness between visual concepts. It can be manually defined with expert knowledge, e.g., WordNet , or automatically constructed, for instance by a language model, e.g., word2vec .

Third, the choice of the pretraining and target datasets is crucial. We need these datasets to have diverse object-level images and to be as less biased as possible, e.g., towards canonical views .

Conveniently, the IN-21K dataset fulfills all these requirements. We therefore choose it as the source of images and concepts for our benchmark. IN-21K contains 14,197,122 curated images covering 21,841 concepts, all of which are further mapped into synsets from the WordNet ontology, which we use to measure semantic similarity.

In the rest of this section, we first define the disjoint sets of seen and unseen concepts, then present our methodology to build different levels for evaluating concept generalization, and describe the evaluation protocol.

We make a natural choice and use the 1000 classes from the ubiquitous IN-1K dataset as the set of our seen concepts. IN-1K is a subset of the IN-21K . It consists of 1.28M images and has been used as the standard benchmark for evaluating novel computer vision architectures , regularization techniques as well as self- and semi-supervised models .

Choosing IN-1K as the seen classes further offers several advantages. Future contributions, following standard practice, could train their models on IN-1K, and then simply evaluate generalization on our benchmark with their pretrained models. It also enables us to benchmark visual representations learned on IN-1K out-of-the box, using publicly available models (as shown in Sec. 4).

2 Selecting eligible unseen concepts

We start from the Fall 2011 version of the IN-21K dataset Note that the recently released Winter 2021 ImageNet version shares the same set of images for all the unseen concepts selected in our benchmark with the Fall 2011 one. We refer the reader to the supplementary for further discussion on both the recent Winter 2021 release as well as a newer, blurred version of IN-1K.. Since we are interested in concepts that are not seen during training, we explicitly remove the 1000 concepts of IN-1K. We also remove all the concepts that are ancestors of these 1000 in the WordNet hierarchy. For instance, the concept “cat” is discarded since its child concept “tiger cat” is in IN-1K. It was recently shown that a subset of IN-21K categories might exhibit undesirable behavior in downstream computer vision applications . We therefore discard all the concepts under the ‘person’ sub-tree. In addition, we chose to discard a small set of potentially offensive concepts (see supplementary material for details). We follow IN-1K and keep only concepts that have at least 782 images, ensuring a relatively balanced benchmark. Finally, we discard concepts that are not leaf nodes in the WordNet subgraph defined by all so-far-eligible concepts. Formally, for any c1c_{1} and c2c_{2} in the set of unseen concepts, we discard c1c_{1} if c1c_{1} is a parent of c2c_{2}. These requirements reduce the set of eligible unseen IN-21K concepts to 5146 categories.

3 Concept generalization levels

Our next step is defining a sequence of unseen concept sets, each with decreasing semantic similarity to the seen concepts in IN-1K. We refer to each one of these as a concept generalization level. They allow us to measure concept generalization in a controlled setting, i.e., to consider increasingly difficult transfer learning scenarios.

Recall that IN-21K is built on top of the word ontology WordNet, where distinct concepts or synsets are linked according to their semantic relationships drafted by linguists. This enables the use of existing semantic similarity measures that exploit the graph structure of WordNet to capture the semantic relatedness of pairs of concepts. Following prior work , we use Lin similarity to define a concept-to-concept similarity. The Lin similarity between two concepts c1c_{1} and c2c_{2} is given by:

where LCS denotes the lowest common subsumer of two concepts in the WordNet graph, and IC(c)=−log⁡p(c)\text{IC}(c)=-\log p(c) is the information content of a concept with probability p(c)p(c) of encountering an instance of concept cc in a specific corpus (in our case the subgraph of WordNet including all IN-21K concepts and their parents till the root node of WordNet: ‘entity’). Following , we define p(c)p(c) as the number of concepts that exist under cc divided by the total number of concepts in the corpus. An example of five concepts from IN-21K ranked by decreasing Lin similarity to the IN-1K concept “Tiger cat” is shown in Fig. 1(a).

We extend the above formulation to define the asymmetric similarity between the set of seen concepts from IN-1K, CIN-1K\mathcal{C}_{\text{IN-1K{}}}{}, and any unseen concept cc as the maximum similarity between any concept from IN-1K and cc:

While designing our benchmark, we considered different semantic similarity measures before choosing Lin similarity. We explored other measures defined on the WordNet graph , such as the path-based Wu-Palmer and the information content-based Jiang-Conrath . We also considered semantic similarities based on Word2Vec representations of the titles and textual descriptions of the concepts. Our experiments with these alternative measures led to observations similar to the ones presented in Sec. 4 for Lin similarity. We refer the curious reader to the supplementary material for additional results with some of these measures.

With the similarity measure defined, our goal now is to group all eligible unseen concepts into multiple evaluation sets, which are increasingly challenging in terms of generalization. To ensure this, we would like the concepts contained in each consecutive set to be of decreasing semantic similarity to any concept from IN-1K. We achieve this by first ranking all unseen concepts with respect to their similarity to IN-1K using Eq. (2). Then, we split the ranked list into groups of consecutive concepts as shown in Fig. 2; each group corresponds to a concept generalization level.

We design our levels to be comparable to IN-1K , and therefore choose 1000 concepts per level. With 5146 eligible unseen concepts, we populate five sets. For increased diversity, we utilize the full span of the ranked list and end up with small gaps between levels (see supplementary material for more details). We denote the five concept generalization levels as L1/2/3/4/5L_{1/2/3/4/5}. Similar to , we further limit the maximum number of training images per concept to 1300. This brings the total number of training images per level to 1.10 million, which is close to the 1.28 million training images of IN-1K.

4 Evaluation protocol

We now present the protocol for ImageNet-CoG, and summarize the metrics for the different experiments presented in Sec. 4. The benchmark consists of two phases. First, a feature extraction phase, where the model trained on IN-1K is used to extract features, followed by the evaluation phase that is conducted on each level independently. An overview of the benchmark is presented in the gray box.

We base our protocol on the assumption that good visual representations should generalize to new tasks with minimal effort, i.e., without fine-tuning the backbones. Therefore, our benchmark only uses the pretrained backbones as feature extractors and decouples representation from evaluation. Concretely, we assume a model learned on the training set of IN-1K. We use this model as an encoder to extract features for images of IN-1K and of all the five levels L1/2/3/4/5L_{1/2/3/4/5}.

4.2 Phase 2: Evaluation

We learn linear logistic regression classifiers for each level using all available training images. Since each level is by design a dataset approximately as big as IN-1K, we also learn linear classifiers on IN-1K with the same protocol; this allows us to compare performance across seen and unseen concepts. We also evaluate how efficiently models adapt when learning unseen concepts, i.e. how many samples they need to do so, by performing few-shot concept classification.

4.3 Metrics and implementation details

We report top-1 accuracy for all the experiments. Absolute accuracy numbers are comparable across IN-1K and each level by construction, since all the levels share the same number of concepts and have training sets of approximately the same size. However, we mostly plot accuracy relative to a baseline model, for two reasons: (i) it makes the plots clearer and the differences easier to grasp, (ii) the performance range at each level is slightly different so it helps visualizing the trends better.

To create the train/test split, we randomly select 50 samples as the test set for each concept and use the remaining ones (at least 732, at most 1300) as a training set. We use part of the training data to optimize the hyper-parameters of the logistic regression for each level; see details in Sec. 4.

We use Optuna to optimize the learning rate and weight decay hyper-parameters for every model and every level; we use 20%20\% of the training sets as a validation set to find the best configuration and then re-train using the complete training set. We report results only on the test sets. We repeat the hyper-parameter selection 5 times with different seeds, and report the mean of the final scores; standard deviation is also presented in all figures.

Evaluating models on ImageNet-CoG

We now present our large-scale experimental study which analyzes how different CNN-based and transformer-based visual representation models behave on our benchmark, following the evaluation protocol defined in the previous section. For clarity, we only highlight a subset of our experiments and provide additional results in the supplementary material.

We choose 31 models to benchmark and present the complete list in Tab. 1. To ease comparisons and discussions, we split the models into the following four categories.

Architecture. We consider several architectures including CNN-based (a-VGG19 , a-Inception-v3 , ResNet50, a-ResNet152 ), transformer-based (a-DeiT-S , a-DeiT-S-distilled, a-DeiT-B-distilled, a-T2T-ViT-t-14 ) and neural architecture search (a-NAT-M4 , a-EfficientNet-B1 , a-EfficientNet-B4 ) backbones with varying complexities. We color-code the models in this category into two groups, depending on whether their number of parameters are comparable to ResNet50 (red) or not (orange); If they do, they are also directly comparable to all models from the following categories.

Self-supervision. ResNet50-sized self-supervised models (in blue) include contrastive (s-SimCLR-v2 , s-MoCo-v2 , s-InfoMin , s-MoCHi , s-BYOL ), clustering-based (s-SwAV , s-OBoW , s-DINO ), feature de-correlation (s-BarlowTwins ), and distilled (s-CompReSS ) models.

Regularization. ResNet50-sized models with label regularization techniques (in purple) applied during the training phase include distillation (r-MEAL-v2 ), label augmentation (r-MixUp , r-Manifold-MixUp , r-CutMix and r-ReLabel ) and adversarial robustness (r-Adv-Robust ) models.

Use of web data. Models pretrained using additional web data with noisy labels are color-coded in green. This includes student-teacher models d-Semi-Sup and d-Semi-Weakly-Sup , which are first pretrained on YFCC-100M (100x the size of IN-1K) and IG-1B (1000x) and then fine-tuned on IN-1K. We also consider cross-modal d-CLIP pretrained on WebImageText (400x) with textual annotations, and noise tolerant tag prediction model d-MoPro pretrained on WebVision-V1 (2x). As it is not clear if YFCC-100M, IG-1B, WebImageText or WebVision-V1 contain images of the unseen concepts we selected in the levels, models in this category are not directly comparable.

We use publicly available models provided by the corresponding authors for all these approaches. All the models, with the exception of those in the use-of-web-data category, are only pretrained on IN-1K. We also use the best ResNet-50 backbones released by the authors for all the ResNet-based models. We use the vanilla ResNet50 (the version available in the torchvision package) as a reference point, which makes cross-category comparisons easier. We prefix models’ names with the category identifiers for clarity.

2 Results

We measure image classification performance on IN-1K and each of the concept generalization levels L1/2/3/4/5L_{1/2/3/4/5} of ImageNet-CoG for the 31 models presented above, using a varying number of images per concept. These experiments allow us to study (i) how classification performance changes as we semantically move away from the seen concepts (Sec. 4.2.1), and (ii) how fast models can adapt to unseen concepts (Sec. 4.2.2). We refer the reader to Sec. 3.4 for the justification of our protocol and the choice of metrics.

We report the performance of linear classifiers learnt with all the training data in Fig. 3. In Fig. 3(a) we report top-1 accuracy for all models and levels, while Fig. 3(b)-(e) present performance relative to the baseline ResNet50 across the 4 model categories. Our main observations are as follows.

* It is harder to generalize to semantically distant concepts. The absolute performance of all models monotonically decreases as we move away semantically from IN-1K. This implies that transfer learning becomes more and more challenging on levels from L1L_{1} to L5L_{5}, i.e., as we try to distinguish concepts that are further from the training ones.

* Self-supervised models excel at concept generalization. Many recent self-supervised models (s-DINO, s-SwAV, s-BYOL, s-OBoW and s-SimCLR-v2) outperform ResNet50 on all levels. In general, we see that the performance gaps between ResNet50 and self-supervised models progressively shift in favor of the latter (Fig. 3(b)). Surprisingly, from Fig. 3(a) we also see that a ResNet50 trained with s-DINO competes with the top-performing models on L5L_{5} across all categories and model sizes. This shows that augmentation invariances learned by the model transfer well to images of unseen concepts.

* Visual transformers overfit more to seen concepts (for models with as many parameters as ResNet50). The top-performing model of the study overall is a-DeiT-B-distilled, a large visual transformer. However, for the same number of parameters as ResNet50, we see that the large gains that visual transformers like a-DeiT-S and a-T2T-ViT-t-14 exhibit on IN-1K are practically lost for unseen concepts (red lines in Fig. 3(e)). In fact, both end up performing slightly worse than ResNet50 on L5L_{5}.

* Using noisy web data highly improves concept generalization. Weakly-supervised models d-Semi-Sup, d-Semi-Weakly-Sup and d-CLIP pretrained with roughly 100x, 1000x, and 400x more data than IN-1K exhibit improved performance over ResNet50 on all levels (Fig. 3(d)). It is worth re-stating, however, that since their datasets are web-based and much larger than IN-1K, we cannot confidently claim that concepts in our levels are indeed unseen during training. Results on this model category should therefore be taken with a pinch of salt.

* Model distillation generally improves concept generalization performance. We see that distilled supervised models r-MEAL-v2 and a-DeiT-S-distilled consistently improve over their undistilled counterparts on all levels (Figs. 3(c) and (e)). However, these gains decrease progressively, and for L5L_{5} performance gains over the baseline are small. It is also worth noting that adversarial training (r-Adv-Robust) does not seem to hurt concept generalization.

* Neural architecture search (NAS) models seem promising for concept generalization. All NAS models we evaluate (a-EfficientNet-B1, a-EfficientNet-B4 and a-NAT-M4) exhibit stable gains over the baseline ResNet50 on all levels (Fig. 3(e)), showing good concept generalization capabilities. Among them, a-NAT-M4, a NAS model tailored for transfer learning with only 7.6M parameters achieves particularly impressive performance over all levels including IN-1K.

* Label-associated augmentation techniques deteriorate concept generalization performance. Although methods like r-MixUp, r-Manifold-MixUp, r-ReLabel and r-CutMix exhibit strong performance gains over ResNet50 on IN-1K, i.e., for concepts seen during training, Fig. 3(c) shows that such gains do not transfer when generalizing to unseen ones. They appear to overfit more to the seen concepts.

* What are the top-performing models overall for concept generalization? From Fig. 3(a) we see that better and larger architectures and models using additional data are on top for L3L_{3}-L5L_{5}. However, it is impressive how s-DINO, a contrastive self-supervised model, is among the top methods, outperforming the vast majority of models at the most challenging levels.

2.2 How fast can models adapt to unseen concepts?

We now study few-shot classification, i.e., training linear classifiers with N={2,4,8,16,32,64,128}N=\{2,4,8,16,32,64,128\} samples per concept. For clarity, we selected a subset of the models and in Fig. 4 we present their performance on L1L_{1}, L3L_{3} and L5L_{5}. The complete set of results for all models and levels is given in the supplementary material. We discuss the most interesting observations from Fig. 4 below.

* Transformer-based models are strong few-shot learners. Transformer-based models exhibit consistent gains over ResNet50 on all levels when N≤128N\leq 128. Despite the fact that performance gains from transformers diminish when using all available images on L5L_{5}, they exhibit a consistent 3-4% accuracy gain over ResNet50 for N≤128N\leq 128 (Fig. 4(f)).

* Model Distillation and Neural Architecture Search (NAS) exhibit consistent gains also in low-data regimes. The NAS-based a-EfficientNet-B4 model exhibits consistently higher performance than ResNet50 on all levels for all NN. The same stands for the distilled r-MEAL-v2 and a-DeiT-S-distilled that are also consistently better than their undistilled counterparts for all NN and all levels.

* Bigger models and additional web data help at few-shot learning. This is an observation from the extended set of figures (see supplementary material). Bigger models have consistent gains in low-data regimes. The same stands for models with additional web data. Moreover, as we go towards semantically dissimilar concepts, a-NAT-M4 outperforms all other methods and it even challenges the much bigger a-DeiT-B-distilled model.

Conclusion

In this paper, we studied concept generalization through the lens of our new ImageNet-CoG benchmark. It is designed to be used out-of-the-box with IN-1K pretrained models. We evaluated a diverse set of 31 methods representative of the recent advances in visual representation learning.

Our extensive analyses show that self-supervised learning produces representations that generalize surprisingly better than any supervised model with the same number of parameters. We see that the current transformer-based models appear to overfit to seen concepts, unlike neural architecture-search-based models. The latter outperform several other supervised learning models with far less parameters.

We also studied how fast models can adapt to unseen concepts by learning classifiers with only a few images per class. In this setting, we verify that visual transformers are strong few-shot learners, and show how distillation and neural architecture search methods achieve consistent gains even in low-data regimes.

We envision ImageNet-CoG to be an easy-to-use evaluation suite to study one of the most important aspects of generalization in a controlled and principled way.

This work was supported in part by MIAI@Grenoble Alpes (ANR-19-P3IA-0003), and the ANR grant AVENUE (ANR-18-CE23-0011).

References

Supplementary Material

.tocmtappendix \etocsettagdepthmtappendixsubsection \etocsettagdepthmtchapternone

This supplementary material is structured as follows. Sec. A details the design choices we made in creating the concept generalization levels of ImageNet-CoG (and is an extension of Secs 3.2 and 3.3 of the main paper). Sec. B describes the preprocessing pipeline for the models we benchmark in Sec. 4 of the main paper and also provides implementation details of our evaluation protocol (extending Sec. 3.4 of the main paper). Sec. C presents the complete set of results for the 31 models we evaluate in Sec. 4 of the main paper. Sec. D discusses how creating concept generalization levels with WordNet ontology and Lin similarity compares to creating them using the textual descriptions of concepts and a pretrained language model (and is an extension of Sec. 3.3 of the main paper). Finally, Sec. E discusses the impact of a recent update of ImageNet (impacting both ImageNet-21K and ImageNet-1K) on the results and future evaluation of our benchmark.

Appendix A Details of ImageNet-CoG levels

We begin by detailing the steps to create the concept generalization levels of ImageNet-CoG. They include the selection of eligible unseen concepts in ImageNet-21K (IN-21K) and the implementation details for Lin similarity . We then briefly discuss the potential noise from missing labels in ImageNet-CoG. At the end of the section, we provide basic statistics about the ImageNet-CoG levels, i.e., the exact number of images per concept in each ImageNet-CoG level.

As described in Sec. 3.2 of the main paper, prior to creating the concept generalization levels, we determine a set of eligible unseen concepts in IN-21K. To determine these concepts, we implemented the following steps.

We started with the whole set of IN-21K concepts (21,841) of the Fall 2011 release and excluded the ones from IN-1K, as they are the seen concepts.

In order to create levels whose size is comparable to IN-1K, following the design choices made for IN-1K, we removed concepts with fewer than 782 images (note that any concept in IN-1K contains at least 782 images and 50 of those are used within the test set).

It was shown that some of the concepts under the “person” sub-tree in IN-21K can be offensive or visually inappropriate, which may lead to undesirable behavior in downstream applications . We thus excluded the entire “person” sub-tree.

We also excluded all concepts that are parents in the ontology of eligible concepts, as this could lead to issues with the labeling strategy. Concretely, for any c1c_{1} and c2c_{2} in IN-21K, we exclude c1c_{1} if c1c_{1} is a parent of c2c_{2}.

Finally, we manually inspected the remaining unseen concepts and found 70 potentially problematic concepts, which may be considered to be offensive, or too ambiguous to distinguish. Examples of such concepts include the very generic “People” (any group of human beings, men or women or children, collectively) or “Orphan” (a young animal without a mother) concepts. The list of such manually discarded concepts is given in Tab. 2.

After sequentially applying these steps, we are left with 5146 eligible unseen concepts. The complete list of eligible unseen concepts, along with the concepts in each level L1/2/3/4/5L_{1/2/3/4/5} can be found on our project website.

A.2 Implementation details for Lin similarity

As described in Sec. 3.3 of the main paper, to produce the concept generalization levels, we sort the eligible unseen concepts by decreasing semantic similarity to the seen concepts in IN-1K. We used Lin similarity as the semantic measure, which computes the relatedness of two concepts defined in a taxonomy.

Computing this similarity between two concepts requires their information content in the taxonomy. Following and , we define the information content of a concept as −log⁡p(c)-\log p(c), where p(c)p(c) is the probability of encountering concept cc in the taxonomy. In our study, the taxonomy is a fragment of WordNet including all the concepts in IN-21K and their parents till the root node of WordNet: “Entity”. Probability of a concept ranges between $suchthatifsuch that ifc_{2}isaparentofis a parent ofc_{1}thenthenp(c_{1}),andtheprobabilityof“Entity”becomes, and the probability of “Entity” becomes1$.

In order to get superior-subordinate relationships between the concepts, we use WordNet-3.0 (the version ImageNet is built on) implementation in the NLTK library .

A.3 Potential label noise in ImageNet-CoG

It has been shown recently that ImageNet-1K (IN-1K) has missing-label noise. We can assume this extends to ImageNet-21K. Unfortunately, this type of noise is really difficult to correct and beyond the scope of our benchmark. However, we devise an experiment to get a sense of how much this noise could be. We take ResNet-50 classifiers for LKL_{K} and apply them to all the images of the IN-1K val set and vice versa (IN-1K classifiers on L5L_{5} val). After inspecting samples that are predicted with very high confidence (>0.99>0.99, about 2.7% of the images), we observe several cases where an unseen concept has (arguably) been seen during training without its label. Some examples are shown in Fig. 5. Given the low percentage of very confident matches and the fact that does not show a big change in performance after re-training with the noise corrected, we believe that this type of labeling noise does not significantly affect our findings.

A.4 Statistics for ImageNet-CoG

Number of images in each level. After selecting 1000 concepts for each level, we ensured that the image statistics are similar to those of IN-1K , i.e., we cap the number of images for each concept to a maximum of 1350 (1300 training + 50 testing). Note that we kept the same set of selected images per concept for all the experiments. We provide the complete list of image filenames in our code repository for reproducibility. In Fig. 9, we plot the number of images per concept for each of the five levels and for IN-1K. We note a minor class imbalance in all the generalization levels from these plots. To investigate if this imbalance had any effect on the observations of our benchmark, we further evaluated a subset of the models analyzed in Sec. 4 of the main paper on a variant of the benchmark, where we randomly sub-sampled images from all the selected concepts to result in the same number of 732 training images, i.e., on class-balanced levels. Apart from the overall reduced accuracy as a result of smaller datasets, this experiment produced similar results to the ones shown in the main text, and all our observations continue to hold. We attribute this to the fact that imbalance is minimal.

Appendix B Evaluation protocol of ImageNet-CoG

In this section, we provide additional implementation details of our evaluation protocol, thus extending Sec. 3.4 of the main paper.

We establish evaluation protocols for ImageNet-CoG with image features extracted from pretrained visual backbones. To extract these features, we first resize an image such that its shortest side becomes SS pixels, then take a center crop of size S×SS\times S pixels. To comply with the testing schemes of the models, for all the backbones we set S=224S=224, except a-Inception-v3 (S=299S=299), a-DeiT-B-distilled (S=384S=384), a-EfficientNet-B1 (S=240S=240) and a-EfficientNet-B4 (S=380S=380).

We also adapt their normalization schemes to be compatible with the data augmentation pipeline of the pretrained models. Concretely, we normalize each image by first dividing them by 255255 (so that each pixel value is in $),thenapplyingmeanandstandarddeviationnormalizationtothepixels,i.e.,subtracting), then applying mean and standard deviation normalization to the pixels, i.e., subtracting[0.485,0.456,0.406]fromtheRGBchannelsanddivingthembyfrom the RGB channels and diving them by[0.229,0.224,0.225],respectively.Notethatford−CLIPweusemean, respectively. Note that for d-CLIP we use mean[0.481,0.457,0.408]andstdand std[0.268,0.261,0.275]$, and do not apply normalization for s-SimCLR-v2 .

Tab. 3 lists the set of unique backbone architectures considered in our study, and the dimensionality of the produced feature representations. For all the architectures trained in a supervised way, we extract features from the penultimate layers, i.e., before the last fully-connected layers making class predictions. For self-supervised learning methods, we follow the respective papers and extract features from the layer learned for transfer learning.

B.2 Training classifiers

In ImageNet-CoG, we perform two types of transfer learning experiments on each set of concepts, i.e., IN-1K or our concept generalization levels L1/2/3/4/5L_{1/2/3/4/5} (see Sec. 4.2 of the main paper): (i) linear classification with all the available data, (ii) linear classification with a few randomly selected training samples. Both sets of experiments use the same test set, i.e., all the test samples.

In each of these experiments, we train a classifier with the features extracted using a given model. In order to evaluate each model in a fair manner in each setting, it is important to train each classifier in the best possible way.

We perform SGD to train classifiers, with momentum=0.9 updates, using batches of size 1024, and apply weight decay regularization to parameters. We choose learning rate and weight decay hyper-parameters on a validation set randomly sampled from the training set of each concept domain (20%20\% of the training set is randomly sampled as a validation set for each concept domain). We sample 30 (learning rate, weight decay) pairs using Optuna with a parzen estimator . We then train the final classifier (with the hyper-parameters chosen from the previous validation step) on the full training set and report performance on the test set. We repeat this process 5 times with different seeds. This means that, in each repetition, we take a different random subset of the training set as a validation set and start hyper-parameter tuning with different random pairs of hyper-parameters. Despite this stochasticity, the overall pipeline is quite robust, with standard deviation in most cases less than 0.2, therefore, not clearly visible in figures. We will release training configurations along with our benchmark on our project website.

Appendix C Extended results

In Sec. 4 of the main paper, we evaluate concept generalization performance for 31 models (listed in Tab. 1 of the main paper) on ImageNet-CoG. Figs. 3 and 4 of the main paper report the results of training logistic regression classifiers with all the available training data for each concept (discussed in Sec. 4.2.1), and training it with a few samples per concept (discussed in Sec. 4.2.2), respectively. Due to space constraints, although Fig. 3 includes the results for all the models on all concept generalization levels, Fig. 4 provides only a selection of the few-shot results. In this section, we present the full set of results for all the methods when training with few and all data samples in table form. We also present the full set of figures for all the methods and levels when training with a few training samples per concept.

How fast can models adapt to unseen concepts. For completeness, we present the scores of all the models for N={1,2,4,8,16,32,64,128,All}N=\{1,2,4,8,16,32,64,128,\text{All}\} on IN-1K and L1/2/3/4/5L_{1/2/3/4/5} in Fig. 10 (raw scores) and Fig. 11 (relative scores). These results, grouped by levels (i.e., for IN-1K and for L1/2/3/4/5L_{1/2/3/4/5} separately) are also presented in Tabs. 5, 6, 7, 8, 9, 10 respectively. These additional results complement Sec. 4.2.2 of the main paper.

Generalization to unseen concepts. To access the raw numbers of the results discussed in Sec. 4.2.1 of the main paper, we refer the reader to Tabs. 5, 6, 7, 8, 9, 10 and the N=AllN=\text{All} columns, which correspond to the scores shown in Fig. 3(a) of the main paper.

C.2 What if we fine-tuned the backbones?

Our benchmark and evaluation protocol are based on the assumption that good visual representations should generalize to different tasks with minimal effort. In fact, we explicitly chose to decouple representation learning and only consider frozen/pretrained backbones as feature extractors. We then evaluate how well pretrained representations transfer to concepts unseen during representation learning. Fine-tuning the models would therefore go against the main premise of our benchmark: after fine-tuning all concepts are “seen” during representation learning, i.e., the feature spaces can now be adapted. It would then be unclear: are we measuring the generalization capabilities of the pretraining strategy or of the fine-tuning process? How much does the latter affect generalization? We consider such questions out of the scope of our study. In fact, learning linear classifiers on top of pre-extracted features additionally allows us to exhaustively optimize hyper-parameters for all the methods and levels (see Sec. B), making sure that comparisons are fair across all models.

Measuring performance relative to fine-tuning, would however verify that the observed performance drops are due to increasing semantic distance and not variabilities across the levels. To this end, we fine-tune ResNet50 (pretrained on IN-1K) on IN-1K and on levels L1/2/3/4/5L_{1/2/3/4/5} separately. Then we compare their performance with the protocol we chose for our benchmark, i.e. the case where we learn linear classifiers on top of pre-extracted features. In Fig. 6, we show the relative scores of the linear classifiers on top of pre-extracted (labeled as “Pre-extracted”) against fine-tuned ResNet50s (labeled as “Fine-tuned”).

We observe that pre-extracted features become less and less informative for unseen concepts as we move from IN-1K to L5L_{5}, supporting our main assumption that semantically less similar concepts are harder to classify.

Appendix D An alternative semantic similarity

One of the requirements for studying concept generalization in a controlled manner is a knowledge base that provides the semantic relatedness of any two concepts. As IN-21K is built on the concept ontology of WordNet , in Sec. 3.3 of the main paper we leverage its graph structure, and propose a benchmark where semantic relationships are computed with the Lin measure .

As mentioned in Sec. 3 of the main paper, the WordNet ontology is hand-crafted, requiring expert knowledge. Therefore similarity measures that exploit this ontology (such as Lin) are arguably reliable in capturing the semantic similarity of concepts. However, it could also be desirable to learn semantic similarities automatically, for instance, using other knowledge bases available online such as Wikipedia. In this section, we investigate if such knowledge bases could be used in building our ImageNet-CoG.

With this motivation, we turn our attention to semantic similarity measures that can be learned over textual data describing the IN-21K concepts. Note that each IN-21K concept is provided with a namehttp://www.image-net.org/archive/words.txt and a short descriptionhttp://www.image-net.org/archive/gloss.txt. The idea is to use this information to determine the semantic relatedness of any two concepts.

To this end, we leverage language models to map the textual description of any concept into an embedding vector, such that the semantic similarity between two concepts can be measured as the similarity between their representations in that embedding space. We achieve this through the skip-gram language model , which has been extensively used in many natural language processing tasks, to extract “word2vec” representations of all concepts. However, we note that the name of many IN-21K concepts are named entities composed of multiple words, yet the vanilla skip-gram model tokenizes a textual sequence into words. We address this issue following that learns a skip-gram model by taking into account such named entities. Specifically, we use the skip-gram model trained on WikipediaApril 2018 version of the English Wikipedia dump. by the Wikipedia2Vec software .

We compute the word2vec embeddings of IN-21K concepts as follows. Firstly, we combine the names and descriptions of all concepts and learn tf-idf weights for each unique word. Secondly, for each concept, we compute two word2vec representations: one for the concept name, and one for the concept description, by averaging the word2vec representations of the words that compose them. These two average vectors are added and used as the final word2vec representation of the concept. Finally, as the semantic similarity measure, we simply use the cosine similarity between the word2vec representations of two concepts:

where ωc\mathbf{\omega}_{c} denotes the word2vec representation of concept cc.

Recall that in Sec. 3.3 of the main paper, first we rank the 5146 eligible unseen concepts in IN-21K (which remain after our filtering, as explained in Sec. 3.3 of the main paper and Sec. A.1), w.r.t. their Lin similarity to the concepts in IN-1K. Then, we sub-sample 5000 concepts to construct concept generalization levels. To create another benchmark based on the textual information of the concepts as described above, we could repeat this procedure by replacing Lin similarity with the cosine similarity we defined in Eq. (3). However, this could select a different sub-set of 5000 concepts, which, in turn, would produce two benchmarks with different sets of unseen concepts. To prevent this, we re-rank the 5000 concepts selected by the Lin similarity, based on their text-based cosine similarity to IN-1K concepts. Then we simply divide the re-ordered concepts into 5 disjoint sequential sets.Note that, given that the percentage of discarded concepts is very small (less than 3%, as 146 concepts are discarded from the 5146 eligible ones), this choice has minimal impact anyway.

We compare the two benchmarks constructed with different knowledge bases (i.e., using the WordNet graph vs. textual descriptions) for our baseline model ResNet50 that is pretrained on the seen concepts (IN-1K) for image classification, following our standard protocol. Concretely, first, we extract image features from the penultimate layer of the ResNet50, then we train linear classifiers on each concept domain separately.

We report results in Fig. 7 for the two benchmarks as well as randomly selected subsets of 1000 concept each. We see that the benchmark constructed using the WordNet ontology and the Lin similarity yield much more challenging concept generalization levels than the one obtained using textual data and a skip-gram language model pretrained on Wikipedia. This is especially visible when comparing classification performance on the levels L3/4/5L_{3/4/5} produced by each technique. We argue that this could be due to the fact that WordNet is an ontology hand-crafted by experts and is able to better approximate the semantic similarity of two concepts compared to the learned skip-gram model. We see that, for a given level LiL_{i}, WordNet combined with Lin similarity manages to gather concepts that are harder to discriminate and that the resulting classification performance is lower. This experiment, however, shows that it is possible to create a similar benchmark using automatically produced semantic similarity scores, the main alternative in the absence of any reliable hand-crafted ontology.

Appendix E The ImageNet 2021 release

The ImageNet team recently released a new version of the full ImageNet dataset (IN-21K) as well as the ILSVRC-2012 dataset (IN-1K)https://image-net.org/update-mar-11-2021.php. with this release, both the datasets are now available for download directly from the official websitehttps://image-net.org/download-images.php.

The 2021 Winter version of IN-21K. We built ImageNet-CoG on the 2011 Fall release of IN-21K, which was the only version available in 2020, when we started constructing our benchmark. The 2011 Fall version contained 21841 concepts, while the new release has only 19167 concepts–a subset of the concepts from the Fall 2011 release. This follows recent studies from the ImageNet team, which identify potentially problematic concepts . They were removed from the latest ImageNet version, including all the concepts under the “Person” sub-tree in WordNet.

With this modified version we successfully verified that: i) all the concepts of ImageNet-CoG are available in the new release, and ii) the images for all the 5000 concepts of ImageNet-CoG are identical in both releases. Consequently, all the results in our work can also be reproduced using the Winter 2021 version of IN-21K.

Blurred version of IN-1K. To protect the privacy of people present in some of the IN-1K images, the ImageNet team released a new version of this dataset, which we refer to as IN-1K-blurred . In this version, the faces of people are blurred in the images. The statistics of these two versions are compared in Tab. 4.

Although the models we evaluated in the paper were pretrained on IN-1K, with non-blurred images, for future reference, we performed our evaluation also on the blurred version of IN-1K (IN-1K-blurred) for all the models. Concretely, for each model, we follow our evaluation protocol on IN-1K-blurred by extracting features of the blurred images and training logistic regression classifiers on them. We report these results in Fig. 8. Note that Fig. 8 is the new version of Fig. 3 in the main paper, with results obtained on IN-1K-blurred instead of IN-1K. We observe that the scores drop on average 0.91%0.91\%, which is comparable to the 0.68%0.68\% drop observed on popular models .