ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

Chunyuan Li, Haotian Liu, Liunian Harold Li, Pengchuan Zhang, Jyoti Aneja, Jianwei Yang, Ping Jin, Houdong Hu, Zicheng Liu, Yong Jae Lee, Jianfeng Gao

Introduction

Visual recognition has become ubiquitous in our society , with applications in geo-localization , action recognition , street number transcription , satellite remote sensing , medical imaging , self-driving cars , etc. Core to these applications are visual recognition tasks such as image classification (IC) and object detection (OD). It is of high value to develop transferable visual models that perform well on a wide range of downstream applications. By leveraging large web crawled image-text corpora, recent advances in language-augmented visual models such as CLIP and ALIGN have demonstrated strong transfer performance, making this direction one of the most practical visual learning approaches. The reason is twofold: (i)(i) open-set recognition is made possible by reformulating classification tasks as retrieval; (ii)(ii) model generalization is improved as language supervision significantly increases the coverage of visual concepts for model training.

The success has immediately inspired many studies of large-scale model pre-training . However, these studies use their own evaluation settings based on customized sets of downstream tasks where the detailed process of adapting the models to these tasks is typically not accessible to the public. Thus, it is extremely difficult for researchers to fairly compare models and develop new models based on other people’s works. To fill this gap, we develop an open-source benchmark and toolkit, Elevater, to make the research results (e.g., model’s task-level transferability) more rigorous, and reproducible. Elevater is composed of three components.

Benchmark (Datasets and Knowledge). We build the first publicly available benchmark to evaluate the large-scale task-level transferability of language-augmented visual models. The benchmark consists of two challenges: Image Classification in the Wild (ICinW) with 20 IC datasets and Object Detection in the Wild (ODinW) with 35 OD datasets. A collected external knowledge base for each dataset which could be used for language data augmentation.

Comprehensive Metrics. To measure the cost of deploying models for real-world applications, we measure a model’s sample-efficiency in the zero-shot, few-shot and full-shot settings and parameter-efficiency in the linear probing and full model fine-tuning settings.

Reproducible Toolkit & Language-augmented Adaptation Methods. We develop an open-source software toolkit to support model adaptation and evaluation. Automatic hyper-parameter tuning is employed to avoid human-in-the-loop tuning, thus reducing human labor and ensuring a fair comparison among different model checkpoints. We also present a set of new model adaptation methods for pre-trained language-augmented visual models. Our methods significantly outperform the traditional vision-only adaptation methods. These methods serve as baselines for the development of more advanced adaptation methods.

In addition, our empirical study leads to interesting findings. (i)(i) Leveraging both text and vision in these models consistently yields better performance than vision-only in few-shot settings; In contrast, random initialization of the linear head in language-augmented visual models is sub-optimal. We also find that few-shot results are always better than zero-shot results, which is different from the results reported in . (ii)(ii) For language-augmented visual models, linear probing performs better than full model fine-tuning in the few-shot settings. As the task-specific training data increases, fine-tuning outperforms linear probing. (iii)(iii) Our study shows that the use of external knowledge, including explicit knowledge of human-compiled thesaurus/dictionaries/documents and implicit knowledge stored in GPT3 , can improve zero-shot and few-shot learning performance. We summarize the pipeline to use Elevater in Figure 1, and organize the paper to focus on the benchmark and toolkit.

Related Work: From Class-level to Task-level Transfer

Zero-shot learning in computer vision has been studied for decades. The research topic has witnessed two distinctive notions of zero-shot: the traditional class-level zero-shot that usually refers to the study of generalizing to unseen object categories , and the recently popular task-level zero-shot that refers to the study of generalizing to unseen datasets/tasks . In Table 1, we compare our benchmark against existing benchmarks. Existing zero-shot learning benchmarks are developed for the class-level zero-shot setting. They are usually in a single domain, with a manual split of categories to produce disjoint training and test categories, e.g., Animal with Attributes (AwA) , Caltech-UCSD Birds-200 (CUB) , SUN , aPY , and ZS-ImageNet . For the OD problem, COCO and LVIS represent two well established datasets to compare various OD methods in a single domain, while UODB is a multi-domain OD benchmark.

In contrast, our benchmark focuses on task-level transfer across domains, i.e., it aims to evaluate the transferability of models, by pre-training from their own large corpus, then evaluating zero-shot performance on a diverse set of downstream datasets. This setting has been recently studied , and is arguably more practical for real-world applications, as it brings the convenience towards the spirit of one-model-for-all. The well-known ImageNet-1K dataset was originally proposed as a large dataset for model training and testing. It has also recently been considered as one downstream task to study zero-shot transfer . Our work presents the first public benchmark to standardize the zero-shot task-level transfer setting. Note that visual task transfer has been previously explored in VTAB , which measures good visual representations as those that adapt to diverse, unseen tasks with an emphasis on few training examples. The pre-trained models and task adaptation in VTAB are considered for vision backbones only, and no language model/modality is involved. Our benchmark shares a similar spirit of task-level transfer to VTAB, but strives to analyze the vital role of language and knowledge in visual transfer. All of them are usually evaluated in full-shot settings, without considering task-level transfer. We have further made several novel contributions to consolidate the benchmark: (i)(i) We add external knowledge for each dataset to cultivate new research directions in knowledge-augmented visual models, inspired by the success of knowledge in traditional class-level transfer. (ii)(ii) We consider the full spectrum in measuring the sample-efficiency of task adaptation, including zero-shot, few-shot, and full-shot.

In this paper, we develop Elevater as a platform for “computer vision in the wild”, whose ultimate goal is to develop a transferable foundation model/system that can effortlessly adapt to a large range of visual tasks in the wild. It consists of two key factors: (i)(i) The task transfer cost is low., which is formally defined in Section 3.3, where our evaluation metrics is designed with efficiency considerations. (ii)(ii) The task transfer scenarios are broad. We illustrate and compare CVinW with other settings using a 2D chart in Figure 2, where the space is constructed with two orthogonal dimensions: input image and output concept. The 2D chart is divided into four quadrants, based on how the model evaluation stage is different from model development stage. Both training and evaluation distributions are consistent in both dimensions for the traditional close-set recognition. Open-set recognition allows new concepts in evaluation, while typically remains the same visual domain ; Domain shift allows new visual domain in evaluation, while typically remains the same concept pool . CVinW allows the flexibility in both dimensions, where any new tasks/datasets in the wild essentially fall into.

Benchmarks

As a proxy for performing unseen tasks in the wild, we collect a diverse set of public datasets from various domains in computer vision, as the basis of our benchmark. Specifically, we consider 20 datasets for IC and 35 datasets for OD. We exhibit the dataset names in Figure 3 (a), and the detailed statistics of each dataset in Table 5 and Table 6 in Appendix. It is recommended in that studying task-level zero-shot transfer is a way of measuring the task learning capabilities of machine learning systems. The task definition of each downstream recognition dataset is typically specified using category names. Adding user specification/note is a natural way to clarify the task definition e.g., the attribute or explanation of a visual concept. Importantly, a similar spirit has been implemented in traditional class-level zero-shot by adding individual domain-specific knowledge (see Table 1), and demonstrated promising zero-shot performance gains. In this paper, we generalize the notion of “zero-shot” to task-level, collecting external knowledge from general sources for our benchmark.

WordNet Hierarchy (def_path). The words along the traversal path from the query node in WordNet to the highest parent node is recorded as the hierarchy knowledge.

WordNet Definition (def_wn). The definition in WordNet synsets is used to explain the query.

Wiktionary Definition (def_wik). The definition of a query in Wiktionary is used.

GPT3 Knowledge (gpt3). For the above three knowledge sources, it is not always feasible to retrieve valid knowledge for any query. To enable full knowledge coverage, we propose to use GPT3 to generate “pseudo” knowledge using in-context-learning, where prompts are constructed with multiple pairs of class names and their Wiktionary definitions. We generate five GPT3 knowledge sequences for each class name, by constructing different context prompts with randomly sampled pairs. See details in Section C.4.

In Fig. 3 (b), we show examples to illustrate the knowledge sources. In practice, there is a trade-off between the knowledge quality and its coverage. For example, WordNet has relatively rich and precise knowledge, but the coverage is low; GPT3 knowledge has the full coverage (as it is generated via prompting a pre-trained neural language model), but it is hard to assess its quality. In the experiment section, we provide baseline results to demonstrate the benefits of external knowledge, and encourage the community to design advanced prompting techniques to leverage these knowledge sources.

2 Pre-trained Models for Transfer Learning

Our benchmark is an evaluation platform for pre-trained models, whose performance largely depends on the scale of the pre-training corpus. Larger corpus typically yields higher performance, but unfortunately results in a barrier to many participants, especially a majority of researchers from university labs. To increase inclusivity, we create two tracks with restrictions on the pre-training data scale: (i)(i) Academic track is a setting that limits the data in established public large datasets (i.e., ImageNet-21K , GCC3M & 12M , YFCC15M ). This track is more academic-friendly, aiming to encourage the exploration in data-efficient pre-training methodologies. (ii)(ii) Industry track has no limit on pre-training data scale, except that images in our benchmark are not allowed in pre-training when reporting zero-/few-shot performance. This track aims to explore the scaling limit. We encourage participants to report the pre-training datasets to enable reproducible research.

Pre-trained Models.

To establish baseline results on Elevater, we evaluate the pre-trained model checkpoints in Table 2. More details of the checkpoints are described in Appendix. Most existing visual models are language-free, where no text is used in model training. Till recently, visual models are trained in a language-augmented and/or knowledge-augmented manner using a language model , among which CLIP represents a strong baseline in the industry track. Please see the detailed taxonomy in Appendix Section G.1.

3 Evaluation Settings: Efficiency Considerations

One major advantage of pre-trained models is the promise that they can transfer to downstream tasks effortlessly. The cost is considered in two orthogonal dimensions: sample-efficiency and parameter-efficiency, as illustrated in Figure 4. The bottom-left corner and top-right corner is the most inexpensive and expensive adaptation strategy, respectively. One may interpolate and make combinations in the 2D space, to get different model adaptation methods with different cost.

Due to the high cost of annotating data, it is often desired to provide a small number of labeled image-label pairs in downstream datasets. Transferable models should be able to reach high performance in this data-limited scenario. To assess this ability, we vary the number of training set size NN per category in the downstream dataset. For IC, N=0,5,20,50N=0,5,20,50. For ODFor OD, NN-shot means providing at least NN images per category, N=0,1,3,5,10N=0,1,3,5,10. Three random seeds are chosen, each of which identifies a subset of samples from the full dataset in a deterministic manner. Once the random seed is given, the indices of training samples in few-shot settings are fixed to encourage reproducible research. We also consider the full-shot setting, where all samples of a given dataset are used.

Parameter-efficiency: Linear Probing vs Full Model Fine-tuning.

Maintaining a small number of dataset-specific model parameters is often favored for model maintenance, as it can be expensive to maintain a unique copy of large model checkpoints for each of the thousands of downstream applications. In IC, linear probing provides a simple strategy for training a dataset-specific linear embedding matrix, while keeping the pre-trained visual backbone frozen. It arguably represents the minimum cost solution for parameter-efficiency. In contrast, fine-tuning often updates the entire weights in backbone and linear head, representing the most expensive solution to model adaptation. In OD , linear probing means updating the linear heads for classification and localization tasks only, while fine-tuning means updating all model weights including the backbone and the detectors.

Toolkits

To ease the process to onboard new checkpoints for evaluation, we provide a software toolkit, including (i)(i) automatic hyper-parameter tuning and (ii)(ii) various strategies for model adaptation to downstream tasks. First, automatic hyper-parameter tuning pipeline is developed to avoid human-in-the-loop tuning, thus reducing human labor and ensuring fair comparisons of different model checkpoints. We follow CLIP to implement a simple grid-search style tuning pipeline, and leave more sophisticated methods like BOHB and DEHB as future work. Details are provided in Appendix. Second, we provide several model adaptation methods as strong baselines, which allow effective transfer learning of pre-trained visual models. The ideas are illustrated in Figure 5. For a downstream dataset, we first represent it in a triplet-wise data format D={(xn,tn,yn)}n=1N\mathcal{D}=\{(\boldsymbol{x}_{n},{\boldsymbol{t}}_{n},y_{n})\}_{n=1}^{N}, where x∈X\boldsymbol{x}\in\mathcal{X} is the image, t∈T{\boldsymbol{t}}\in\mathcal{T} is its corresponding language description, and y∈Yy\in\mathcal{Y} is a label indicating the index of the unique language description in the dataset. ∣B∣|\mathcal{B}| is batch size. In IC, the number of labels ∣Y∣=K|\mathcal{Y}|=K, i.e., the number of category names.

Language-augmented Visual Models.

Recent works that learn visual models with language supervision often employ a two-encoder architecture. Besides the image encoder model fθf_{\boldsymbol{\theta}}, a text encoder fϕ(t)f_{\boldsymbol{\phi}}({\boldsymbol{t}}) parameterized by ϕ\boldsymbol{\phi} is also used to encode text t{\boldsymbol{t}}. Additional projection layers Wv{{\bf W}}_{v} and Wt{{\bf W}}_{t} are introduced for image and language features, embedding them into a joint space with dimension PP, with projected features as u\boldsymbol{u} and v\boldsymbol{v} respectively. Note that lowercase u\boldsymbol{u} and v\boldsymbol{v} are single feature vectors while U{{\bf U}} and V{{\bf V}} are a batch with multiple feature vectors. As in Figure 5 (b), zero-shot learning can be directly performed in this space: the mean text feature v\boldsymbol{v} is first obtained for each category, by averaging text features of the category name in different language prompts. The image is predicted as the category yielding the highest similarity u⊤v\boldsymbol{u}^{\top}\boldsymbol{v}.

Discussion.

We highly recommend the proposed language-initialized methods as the standard to adapt language-augmented visual models for two reasons: (i)(i) This simple method yields surprisingly superior empirical performance, as demonstrated in our experiments. (ii)(ii) It provides an effective mechanism to leverage the external knowledge that is collected for a downstream task in our benchmark. Specifically, the knowledge can be concatenated with the original language prompt (with a simple “;” in our experiments), then encoded into contextualized text features. When multiple knowledge items exist (e.g., the case of GPT-3) for each concept, we concatenate one of its prompts and one of its knowledge items, and get the encoded text embedding of the concatenated sequence via the language encoder. This is performed for all the combinations between all prompts and knowledge items for this concept, then the averaged embedding is computed to represent the concept. In contrast, random initialization would ignore this knowledge source. The language-initialized method can serve as a strong baseline to encourage more effective knowledge-augmented adaptation methods.

In OD, GLIP is a language-augmented detector, whose overall architecture can be simply considered as adding a cross-modal module over the CLIP-like dual-encoder. In GLIP , its linear probing has been implemented via updating Wv{{\bf W}}_{v} and Wt{{\bf W}}_{t}. A prompt-tuning strategy was proposed, by initializing the language input of the cross-modal module as V{{\bf V}}, and only updating V{{\bf V}} during adaptation. This is similar to our language-initialized strategy.

Empirical Results and Findings

We present the experimental results with our benchmark to illustrate two points. Q1: The importance of language in visual model transfer in the adaptation stage. Q2: We present three playgrounds that our benchmark can help to cultivate research in, including sample-efficiency, parameter-efficiency and external knowledge for visual transfer. We also present novel empirical findings.

In Table 5.1, we compare the effectiveness of the proposed language-initialization methods with the checkpoint CLIP ViT-B32. The one-projection scheme is consistently better than two-projection scheme in all settings (though the gain is minor). This is because the former often has less parameters than the latter, as D=768>P=512D=768>P=512. To ensure fair comparisons with the random initialization of linear head in CLIP (i.e., # trainable parameters is the same), in the ensuing experiments, we consider the two-projection language-initialization scheme as the default, unless the one-projection scheme is specified.

As shown in Fig. 6, under both linear probe (LP) and fine-tuning (FT) settings, language-based initialization significantly outperforms random initialization. Notably, we show that even with very few shots (e.g., 2-shots), both our LP and FT is able to outperform the zero-shot CLIP. This is contradictory to the finding in the original CLIP paper , where zero-shot outperforms linear probing in the fewer shot (less than 4) settings. With the proposed language-init method, one can ensure that few-shot performance is always better than zero-shot, as we essentially reduce to zero-shot when zero iteration is updated in our language-init method. Moreover, we also find that with random initialization, FT performs significantly worse than LP under few-shot settings. However, with language-init, FT starts to outperform LP with more than 20 shots. Both findings demonstrate the proposed language-based initialization is consistently effective, suggesting that it is an important technique, and should be the standard adaptation method for language-augmented visual models like CLIP. Further, the correct adaptation methods for language-augmented visual models should leverage both the pre-trained visual and text encoder. It is not sufficient to solely transfer from the visual encoder, pre-trained language encoder plays an important role in task transfer.

We summarize the transfer performance of pre-trained models for IC in Table 5.1 and OD in Table 5.1. For IC, we also compare CLIP against other language-free visual models including MoCov3, MAE, ViT, DeiT in Appendix. We see that the language-augmented model (CLIP) outperforms language-free model (Supervised ViT) in the limited data settings. The gap is closed when more training examples are observed (e.g., , 50-shot and full-shot). This is probably because the pre-training power is gradually dominated by larger-scale downstream training. Further, language-augmented models are able to perform zero-shot task transfer, while traditional language-free models cannot. Similar conclusions can be drawn for OD in Table 5.1. Hence, we recommend the use of language-augmented visual models for task-level transfer.

We explore sample efficiency in Fig. 7 (a) for IC. First, we find that CLIP consistently outperforms supervised ViT (Sup-ViT), yielding a significant 5~10% gain in the 5-shot settings. This suggests that CLIP is more sample-efficient than supervised ViT . Furthermore, we find that fine-tuning CLIP yields better performance than linear probing in >20-shot settings, while being worse in the 5-shot setting. This is a bit surprising, as it is contradictory to the common convention that fine-tuning is always better than linear probing. We hypothesize this is because fine-tuning tends to over-fit in the scenarios with a large number of trainable parameters and a small number of training samples. Overall, it suggests that fine-tuning CLIP potentially has a better sample efficiency than linear probing, and a better adaptation strategy on fewer-shot settings can be explored in the future. For supervised ViT, FT is always better than LP, the performance gap becomes larger when more samples are used. In 5-shot settings, the gap is minor, which is similar to observations made for supervised CNNs . To compare pre-trained models, we suggest to report the evaluation results on the entire spectrum of sample-efficiency to fully study the behaviors of a pre-trained model. If compute resource is limited, zero-shot or few-shot evaluation can be used as a quick assessment.

We explore sample efficiency for OD in Fig. 7 (b). The conclusion is similar to IC in that the language-augmented visual model (GLIP) is more sample-efficient than the language-free visual model (DyHead), when the models are adapted using either LP or FT settings. The performance gap is large in the fewer-shot settings and is small in the full-shot settings. The difference is that FT consistently outperforms LP in all settings for OD. This is probably because there are a lot of boxes (training instances) per image in OD, which makes OD less likely to over-fit compared to IC.

We study the parameter efficiency in Fig. 7 (c) for IC. For CLIP, we experiment with two different settings of linear probing on whether to merge the last two linear projection layers (Wv\mathbf{W}_{v} and V\mathbf{V} in Fig. 5). Merging these two layers in CLIP allows 1.5×\times trainable parameters in the linear probe classifier as keeping them separated. First, it shows a trend that a larger number of trainable parameters leads to better performance, as demonstrated by three curves/scenarios: full-shot CLIP, full-shot Sup-ViT and 5-shot Sup-ViT. This also verifies that LP and FT provide the lower bound and upper bound, respectively, in terms of both #parameter and performance. Most existing parameter-efficient adaptation methods play a trade-off game. However, in the scenario of 5-shot CLIP, we do notice a slight drop in performance when we further increase the number of trainable parameters to full-model fine-tuning. It suggests that the scenario of adapting language-augmented visual models for data-limited settings is a more meaningful playground to explore the line of research in parameter-efficient adaptation methods, as the best performance may require an optimal number of trainable parameters, which has been less explored.

We study the parameter efficiency in Fig. 7 (d) for OD. The overall trend is similar in that better performance comes with more parameters. It turns out that prompt tuning an language-augmented OD model is an effective parameter-efficient approach. For example, prompting is better than linear probing in GLIP. Further, prompt tuning GLIP outperforms fine-tuning DyHead in the 1-shot setting, where the former has less than 0.1% parameters of the latter.

We investigate the effectiveness of external knowledge in Fig. 5.1, measured by zero-shot task transfer performance. The model K-Lite is evaluated, as its pre-training is knowledge-augmented. We find that leveraging external knowledge improves upon the knowledge-free pre-training counterparts (UniCL and GLIP). For example, UniCL is improved from 27.15 to 29.92~33.93, and GLIP-A is improved from 11.53 to 11.70~13.30. Further, for GPT3 knowledge, a larger number of generated knowledge items often leads to higher performance. When combining GPT3 knowledge with Wiktionary knowledge, we see a further performance boost. With an increasing number of GPT3 knowledge items, the gain is consistently improved for IC, but not for OD. In Table 5.1, we study the role of knowledge for task transfer in model adaptation. Initializing the linear head using features encoded with knowledge is an effective way to leverage the collected knowledge sources, especially for the fewer-shot settings.

One may wonder if the collected knowledge in Elevater benchmark can also benefit knowledge-free pre-trained models such as CLIP during model adaptation? We confirm its effectiveness in Appendix. In zero-shot transfer, external knowledge improves the baseline on four datasets. In few- and full-shot transfer, one may selectively choose whether to update the model using external knowledge, by observing the best performance in the auto-tuning stage. This selective strategy with the availability of external knowledge demonstrates a consistent improvement (or tie) on 15+ out of 20 datasets.

Conclusions

We have presented Elevater, a platform to evaluate the recently emerging language-augmented visual models for task-level transfer. It consists of 20 image classification datasets and 35 object detection datasets. All of them are collected from public domains, and are enriched with various external knowledge sources to enhance the language modality. We have developed open-source toolkits with an auto hyper-parameter pipeline and novel language-initialized adaptation methods to ensure easy utilization and fair comparisons. Strong baseline results are produced from the toolkit to cultivate research in a variety of topics, e.g., more transferable language-augmented visual models, advanced model adaption methods (sample-efficiency and parameter-efficiency), and external knowledge for task-level transfer. The question of how to design general-purpose task-level transferable visual models remains largely unanswered. Given benchmarks and tookits we have developed from the perspective of language-augmented visual models, we believe that Elevater can provide fertile soil for addressing this challenge.

The authors gratefully acknowledge Haotian Zhang for building the ODinW leaderboard on Eval AI, Pengcheng He for helpful discussions to have a separate track dedicated for users from academia, Baolin Peng and and Zhengyuan Yang for the inspirations of GPT3 to generate knowledge for dialogue and OK-VQA tasks, Bo Li for insights on the topic of domain generalization, Zhuowen Tu for the inspirations to make benchmark scope wider to measure all pre-trained vision models, Ce Liu for suggestions to compare the benchmark with well-established vision datasets such as ImageNet and COCO. The benchmark depends on publicly available datasets; we acknowledge all the original authors who made their datasets public. Please follow the original license of each dataset and keep this benchmark for academic purposes. This work was supported in part by NSF CAREER IIS-2150012, the Wisconsin Alumni Research Foundation, and Institute of Information & communications Technology Planning & Evaluation(IITP) grants funded by the Korea government(MSIT) (No. 2022-0-00871, Development of AI Autonomy and Knowledge Enhancement for AI Agent Collaboration) and (No. RS-2022-00187238, Development of Large Korean Language Model Technology for Efficient Pre-training).

Appendix A Societal Impact

We do not anticipate a specific negative impact, but, as with any Machine Learning method, we recommend to exercise caution. The existing knowledge bases such as Word-Net and Wiktionary are the results of crowd-sourcing various human knowledge or commonsense into a centered place. Elevater provides evidence to leverage such knowledge bases for AI research. It encourages the community to contribute more to improve the coverage and quality of knowledge items, which will further benefit AI research. We also leverage GPT3 to generate knowledge, which is stored as a part of benchmark for public academic use. The related societal impact on the usage of AI-generated content may apply to our work.

Appendix B Our Position

In this paper, we advocate our perspective on “Computer Vision in the Wild (CVinW)”, whose ultimate goal is to develop a transferable foundation model/system that can effortlessly adapt to a large range of visual tasks in the wild. We further illustrate two key factors as follows.

One major advantage of pre-trained/foundation models is the promise that they can transfer to downstream tasks effortlessly (or in an inexpensive manner). It means that model adaptation efficiency is an important factor to measure the performance of the pre-trained models. To concretely illustrate the notion of inexpensive adaptation, we provide a 2D chart on the model adaptation cost in Figure 4. The cost is considered in two orthogonal dimensions: sample-efficiency and parameter-efficiency. One may interpolate and make combinations in the 2D space, to get different model adaptation methods with different cost. This is design philosophy behind our comprehensive evaluation metrics. Two playgrounds with different efficiency considerations presented in the main paper are simplified settings to study model performance. As a north star, one foundation could with fixed weights should zero-shot transfer well on many downstream tasks, the most inexpensive regime in the bottom-left corner of Figure 4.

Factor II: The Task Transfer Scenarios are Broad.

We illustrate and compare the settings of CVinW using a 2D chart in Figure 2. It consists of two dimensions: the input visual content and output concept prediction. For the example provided in the standard setting, the natural image with concept “person, sheep, dog” is presented. We divide the the 2D chart into four quadrants

The Standard Close-Set Setting. The bottom-left quadrant is the standard setting, where most existing visual recognition lie in, training and evaluation are consistent in both their visual input distributions and output category sets. For example, only natural images with concept “person, sheep, dog” are presented in training and evaluation.

Open-Set/Vocabulary/World Setting. In the top-left quadrant, the recognition of new concepts is enabled, while the visual input distributions of training and evaluation are in the same domain. This research problem is usually tackled by traditional class-level zero-shot transfer, or some experimental settings in the open-set recognition. For example, natural images with concepts “person, sheep, dog” are presented in training, but natural images with concepts “border collie, running, while shirt” are presented in evaluation. Though the testing concepts are closely to training concepts, but they have not been observed by the models in training.

The Domain Shift Setting. In the bottom-right quadrant, the input image distributions are shifted between training and evaluation sets, while the output category sets are the same. This research problem is often tackled in the area of domain adaptation and out-of-distribution. For example, natural images with concepts “person, sheep, dog” are presented in training, but thermal images are presented in evaluation, though the concepts have been observed in training.

Computer Vision in the Wild Setting. In the top-right quadrant, the strong generalization ability to both new concepts and new visual distributions is required. Therefore, the model can perform well on new tasks of any customized set of concepts in any visual domains. This is a setting we advocate for computer vision in the wild, where any new downstream tasks can appear in this quadrant, and it requires models with a strong task-level visual transfer ability.

For the readers who are interested in the literature on Computer Vision in the Wild, we create an up-to-date CVinW reading list at https://github.com/Computer-Vision-in-the-Wild/CVinW_Readings.

B.2 Related Works in NLP: Benchmarks, Adaptation, and Knowledge

With a focused scope, our benchmark evaluates language-image models on two core CV problems: IC and OD. Though language-image models can also be deployed and evaluated in other scenarios, including joint visual-text evaluation (e.g., visual question answering , video-and-language understanding ) and the scenario of improving language encoders with vision . Our benchmark is complementary to them in its focus on evaluating vision encoders.

Our work takes major inspiration from the development of pre-trained language models in natural language processing (NLP) in several aspects: (i)(i) Benchmarks. Platforms with a suite of small datasets such as GLUE /SuperGLUE have been extensively used to evaluate the general language understanding ability of pre-trained models . Recently, there is a trend in NLP to develop task-agnostic models such as the GPT family that demonstrate task-level transfer learning ability, enabling zero-shot and few-shot transfer to downstream datasets. The success in NLP encourages us to build a generic benchmark to measure the similar transferability for visual models. (ii)(ii) Efficient adaptation. The democratization of large pre-trained models for efficient adaptation in downstream applications is an important topic in practice. Many algorithms have been developed for various efficiency considerations, including adapters and prompt tuning . In particular, natural language prompting is the method of reformatting NLP tasks in the format of a natural language response to natural language input, has attracted attentions in zero-shot and few-shot learning in NLP . It has inspired a few recent works for language-augmented visual models . Our benchmark can serve as a comprehensive playground to quantify the progress in the emerging field of visual model adaptation. We also propose to use external knowledge for prompt engineering, and a novel language/knowledge-initialized model adaptation method as a strong baseline. (iii)(iii) Knowledge. Knowledge-intensive tasks — those where a human can only be expected to perform the task with access to a knowledge source such as Wikipedia — are challenging for even cutting edge NLP and vision-and-language models, as it is infeasible to train large models to memorize everything. KILT is a benchmark that contains a suite of tasks/datasets for evaluating and analyzing knowledge-intensive NLP models. Similarly, we also add various external knowledge sources in each downstream dataset for our vision benchmark.

Appendix C Benchmark Suite

In Table 5 and Table 6, we list the basic statistics of 20 image classification datasets and 35 object detection datasets in the benchmark.

The benchmark may inherit data biases from the public datasets we have considered, both in the images and the annotations. Such biases might be reflected in the predictions of the systems trained on these data. Users should not completely rely on such systems for making real-world decisions.

C.2 Visualization Comparison with Established Vision Datasets

We also compare our benchmark with well established datasets in computer vision: ImageNet-1K for IC and COCO/LVIS for OD. Note that LVIS is much diverse than COCO in terms of concept coverage. The visualization of concept semantic space is Figure 9. The semantics is computed by extracting the CLIP text features from the category names. To quantitatively measure the diversity of different benchmarks, we compute the standard derivation (STD) over text features. The STD of ImageNet1-K and ICinW is 0.610 and 0.680, respectively. The STD of LVIS and ODinW is 0.533 and 0.619, respectively.

C.3 License

As per the original authors, the licenses of each dataset include CC BY-NC-SA 3.0https://creativecommons.org/licenses/by-nc-sa/3.0/, CC BY-NC-SA 4.0https://creativecommons.org/licenses/by-nc-sa/4.0/, CC BY 4.0https://creativecommons.org/licenses/by/4.0/, ODbL v1.0https://opendatacommons.org/licenses/odbl/1-0/, MIThttps://choosealicense.com/licenses/mit/, CC0 1.0https://creativecommons.org/publicdomain/zero/1.0/. Some datasets have published dedicated usage aggrements: Hateful Memeshttps://www.drivendata.org/competitions/64/hateful-memes/page/214/. All datsets allow the usage for research purposes. The images used in the datasets are from Internet, on non-offensive topics. The annotations in the datasets do not contain personally identifiable information.

For external knowledge collected on Elevater, we suggest the users to follow the corresponding licenses: WordNethttps://wordnet.princeton.edu/license-and-commercial-use, Wiktionaryhttps://en.wiktionary.org/wiki/Wiktionary:Main_Page, GPT-3https://openai.com/api/policies/sharing-publication/. For the GPT-3 generated knowledge, we have the approval from OpenAI to release it as a part of Elevater to encourage future research.

C.4 Generating GPT-3 Knowledge with In-Context-Learning

Wiktionary and WordNet do not provide a 100% coverage for all downstream concepts. As shown in , an incomplete knowledge coverage can lead to deteriorated model performance. In this paper, we show that GPT3 can be used for generating additional external knowledge and providing a full coverage for downstream concepts.

We use in-context-learning to prompt GPT-3. As an input to GPT3, we start by asking “Please explain the concept according to the context”. In addition, we provide multiple concept-explaining Q (concept)-A (explanation) pairs. Each pair of the concept and explanation are sampled from the concepts that have the Wiktionary knowledge available. Finally, we send a different concept to GPT3, and ask for the explanation. In this way, GPT3 is able to generate explanatory descriptions for the concepts even when its Wiktionary knowledge is missing. For example, as shown in Fig. 10, there is no Wiktionary knowledge available for “snowberg”, while “ship” and “storage tank” have their corresponding Wiktionary explanations. By providing the concept-explanation pairs of “ship” and “storage tank”, GPT3 recognizes this as a concept explaining task, and when a new concept “snowberg” is given, it explains the concept without the need for its external knowledge. By randomly sampling different Q-A groups from the concepts with Wiktionary knowledge, we are able to generate a diverse set of GPT3 responses.

C.5 Prompting and Knowledge

For each visual recognition dataset, there comes naturally with a set of category names. A specific set of natural language templates are created for each dataset, following . In our toolkit (vision_benchmark/datasets/prompts.py\mathtt{vision\_benchmark/datasets/prompts.py}), we maintain the mappings from a dataset to its specific category names and template sets, respectively. External knowledge for each dataset is maintained at the folder vision_benchmark/resources/knowledge\mathtt{vision\_benchmark/resources/knowledge}. To construct the language prompt, we suggest the following steps:

For a given dataset, choose one category from a set of its category names

Choose one template from a set of pre-defined dataset-specific language templates.

Fill in the category name into the template, which yields the constructed language prompt for this category.

(Optional) If external knowledge is preferred to add into the prompt construction, please select a knowledge source with non-empty value, and concatenate the knowledge sequence after the text sequence in Step 3, separated by “;”.

In Table 7, we provide examples to construct prompts with and without external knowledge, by following the above procedure.

Appendix D Evaluation

As demonstrated in Section 3.3, we advocate an evaluation setting with efficiency considerations, which decomposes the adaptation cost into two orthogonal dimensions: sample-efficiency and parameter-efficiency. To encourage future users compare their models with efficiency considerations, We build the public leaderboards on EvalAI:

https://eval.ai/web/challenges/challenge-page/1832/overview • Object Detection in the Wild (ODinW) https://eval.ai/web/challenges/challenge-page/1839/overview

D.2 A new metric with performance-efficiency trade-off

For parameter-efficiency track, to compare different methods with a single number that considers both prediction accuracy and parameter-efficiency, we define the performance-efficiency (PE) metric:

where score measures the prediction accuracy, while # trainable-parameters is the number of updated parameters in the model adaptation stage, and M0M_{0} is the normalization constant. We set M0=108M_{0}=10^{8} because most existing vision backbone model size are designed in this magnitude, for example, ViT-Base (80M parameters) and ViT-Large (300M parameters). With larger models designed in the future, one may increase M0M_{0} for sensible measurement.

Appendix E Toolkit

For a given dataset, we split its training set into training and validation with a ratio 80% vs 20%. At least one training sample per class is ensured for training and validation. Grid search is applied over learning rate η\eta and weight decay α\alpha. In the hyper-parameter search stage, the model is trained with a given configuration (η,α)(\eta,\alpha) for 10 epochs, the best hyper-parameter configuration is chosen as the one with the best validation performance along the entire process. After that, a final run is performed for 50 epochs to report the performance on the testing set.

Object Detection.

A validation set is chosen in the hyper-parameter search stage. We consider validation set size (1,1,1,3,full)(1,1,1,3,full) for N=1,3,5,10N=1,3,5,10, respectively. For each type of checkpoints (DyHead, GLIP) and each adaption method, we have a set of pre-selected hyper-parameters, i.e., batch size |B\mathcal{B}|, initial learning rate η0\eta_{0} and weight decay α\alpha, as shown in Table E.1 in Appendix. They are determined by either empirical rules or simple hyper-parameter tuning. For each setting and each train/val split, we evaluate on the val split after every training epoch to decrease the learning rate in a step-wise manner. More specifically, we use the PyTorch ReduceLROnPlateau with patience 3 and factor 0.1 to decrease the learning rate when there is no improvement on val. We terminate the fine-tuning process if we do not see improvements for continuously 9 epochs, return the checkpoint with the best score on val, and report its score on the test split. For each few-shot setting, we random sample the train/val split 3 times, and report the average score and standard deviation on the test split. For each type of checkpoints (DyHead, GLIP) and each adaption method, we have a set of pre-selected hyper-parameters, i.e., batch size |B\mathcal{B}|, initial learning rate η0\eta_{0} and weight decay α\alpha, as shown in Table E.1. They are determined by either empirical rules or simple hyper-parameter tuning.

To make a fair comparison between different methods in image classification, we conduct experiments with FP32 precision. Our preliminary experiments show that on average FP16 and FP32 yields similar zero-shot performance, while FP32 models outperform FP16 ones on 16 out of 20 datasets.

Object Detection.

For OD, one image could contain multiple classes. We run an algorithm to go over the images in the full training set one by one, and add the image to the NN-shot training set if the image contains some classes that do not have NN images yet. We stop if all classes have at least NN images or we have exhausted the full training set. Thus, the total number of images in the dataset could be between N∼N∗KN\sim N*K, where KK is the number of categories. We will release all the NN-shot samples we used for experiments . For OD full fine-tuning, the common practice is to freeze the bottom two layers of the backboneshorturl.at/AOZ13.

In Section 4, we have proposed language-initialized adaptation strategy, which consistently improves the linear probing and fine-tuning performance of language-image pre-trained models like CLIP. By initializing the linear head of CLIP model with the embeddings from the language encoder, it allows the model update and prediction of CLIP in few- / full-shot adaptation settings behaving in a similar way as in the zero-shot setting. This, in other words, narrows the gap between the pre-training CLIP objective and the downstream image classification objective (cross-entropy). In this section, we explore other factors that differs in the pre-training CLIP and downstream CLIP adaptations.

F.2 Training objective

Although the training objective is aligned between the pre-training and downstream adaptation already with the proposed language-initialization adaptation strategy, there are several factors that may cause a difference in the gradient flow between pre-training and downstream adaptation, which can potentially hurdle the model training.

In CLIP, a scaled pairwise cosine similarity is first computed between all image-text pairs, and the bidirectional cross entropy loss is then applied to the computed similarity score. Although the loss function of pre-training CLIP and downstream adaptation can be reduced to the same objective, one key difference is the size of the similarity matrix. For each image, the similarity is computed with all text embeddings. In CLIP, it is the number of all text samples in a large batch (e.g., ∣B∣|\mathcal{B}|=32,768); while in downstream, it is the number of text embeddings of all classes KK (which is typically less than 200). Such disparity can cause a significant change in the pattern of the gradient flow.

Temperature.

In CLIP, a trainable log-parameterized temperature τ\tau controls the range of the logits in the Softmax, which is typically not used in downstream adaptation. Although the temperature parameter does not alter the ranking of its predictions, it modifies the scale of the gradients when backward propagation is performed in downstream adaptation.

Experiment/Analysis.

Based upon the above analysis, we design experiments to explore the effect of these factors on the gradient flow and the downstream adaptations.

We compare the initialization of the temperature τ\tau and whether to keep it frozen during the adaptation in Table 9 (Row 1,5-7). First, setting it to trainable has minimal effect to the training process; as there are now only KK classes, it might not be as important in CLIP to have a learnable τ\tau. Second, initializing it with the pretrained checkpoint (after training with CLIP, exp⁡(τ)=100\exp(\tau)=100) yields a significant performance drop. We attribute this performance drop to the change in the size of Softmax from ∣BCLIP∣|\mathcal{B}_{\text{CLIP}}| to KK, where ∣BCLIP∣≫K|\mathcal{B}_{\text{CLIP}}|\gg K. Having a large temperature coefficient like exp⁡(τ)=100\exp(\tau)=100 dramatically increases the sharpness in the pattern of Softmax and its gradient flow, which is inappropriate for training.

F.3 Conclusion

The language-augmented initialization is the most critical componenet in aligning the training behavior of CLIP models (30%+ mean score improvement for 5-shot finetuning), without which the pre-trained capacity in the language encoder would be completely lost. Other factors like visual feature normalization, batch size, temperature, etc. have a much smaller effect to the training procedure. We choose to use the parameter-free batch normalization, keep the traditional batch size, and not bring in additional parameters like temperatures, for trading off between the performance and the simplicity of the model.

Appendix G Empirical Comparisons of Existing Pre-trained Vision Models

We provide the taxonomy for pre-trained vision models from the perspective whether language and/or is employed in pre-training, as shown in Table 10. The taxonomy is a two-level hierarchy.

In the 1st level hierarchy, given a visual recognition problem (IC or OD), the models are first categorized into language-augmented or language-free, depending on whether language is used or not in pre-training.

In the 2nd level hierarchy, the language-augmented models are further categorized into knowledge-augmented or knowledge-free, depending on whether the textual external knowledge is used or not in pre-training.

Note that our taxonomy is only related to pre-training, which is independent from how the model is adapted to a downstream task.

For knowledge-augmented pre-trained models such as K-LITE , the model is pre-trained with both natural language supervision and external knowledge supervision. The external knowledge is employed in the following manner: (1) For image-text pairs, query is identified using entity extraction on the text, (2) The relevant “knowledge text” of the query is retrieved from knowledge bases; (3) The retrieved “knowledge text” is appended to the original text. In the downstream adaptation stage, it follows the same prompting process with other pre-trained models, as described in Section C.5.

G.2 Baseline with Vision Pre-trained Models

We consider seven checkpoints to produce baseline results for IC. In the main paper, we report the following four checkpoints.

Supervised ViT represents a checkpoint for the traditional language-free visual models, where model training is performed on ImageNet-22K with cross-entropy loss.

CLIP ViT represents a checkpoint for the family of the language-augmented visual models, trained with 400M image-text pairs.

UniCL Swin represents knowledge-free language-augmented visual models with Swin as the visual backbone, trained in the academic setting with ImageNet-21K, which excludes ImageNet-1K categories from ImageNet-22K.

KLITE, UniCL Swin represents knowledge-enriched language-augmented visual models. Its pre-training setting is the same as UniCL Swin, but external knowledge such as Wiktionary is leveraged in model pre-training.

We also consider three popular language-free visual models in Appendix:

DeiT represents a checkpoint for the supervised visual backbone, where model training is performed on ImageNet-1K with cross-entropy loss and advanced data augmentation and training schedule.

MoCo represents a checkpoint for the family of augmented-view-based methods for image self-supervised learning, trained with images only in ImageNet-1K.

MAE represents a checkpoint for the family of recent masked region (visual token) modeling based methods for image self-supervised learning, trained with images only in ImageNet-1K.

CAE represents a checkpoint that benefits the separation of the representation learning (encoding) role and the pretext task completion role, trained with images only in ImageNet-1K.

Object Detection

We consider four checkpoints to produce baseline results for OD. They are for the academic track, as they are pre-trained on public datasets. All of them employ Swin-Tiny backbone .

DyHead represents a checkpoint for the traditional language-free object detector, where model is pre-trained on Object365 without leveraging the category name information.

GLIP represents a checkpoint for the family of the language-augmented object detector, trained with Object365 and Flicker phrase grounding data .

GLIP-A represents knowledge-free language-augmented object detector, where model is trained on Object365 and the semantics of category names is leveraged.

KLITE, GLIP-A represents knowledge-enriched language-augmented object detector. Its training setting is the same as GLIP-A, except that Wiktionary knowledge is leveraged in model pre-training.

In summary, among four checkpoints for each problem, the first two are used to compare the state-of-the-art in language-free and language-augmented models, and latter two are used to compare the knowledge-free and knowledge-augmented models (both belongs to language-augmented models, as knowledge is presented as a structured form of language).

G.3 Experimental Results of Different Model Checkpoints

In Table 11, we report IC performance with ViT-B16 pre-trained with representative methods, using different objectives and datasets. We present its breakdown experimental results in Table 12. Note that all of the models are adapted to downstream datasets, using the same automatic hyper-parameter tuning process in our toolkit, and no model- / dataset-specific tuning is employed. This ensures fairness in model adaptation process, but may not represent the best transfer performance of each pre-trained model, if more careful tuning efforts are paid. Nevertheless, we believe the results represent the model transferability with affordable efforts, and use them as baseline results for Elevater benchmark.

We found that the overall ranking of the models in the descending order: CLIP, ViT, DeiT, MoCo-v3, MAE. Surprisingly, we found that MAE performs worse than MoCo, and both of them are worse than supervised method DeiT, though all three of them are pre-trained on the same ImageNet-1K dataset. We note that an similar observation is made in , when evaluated these checkpoints on a large range of downstream datasets. This is perhaps because the region-based pre-training tasks in MAE is can better capture region-level dependency (thus benefits dense prediction tasks such as object detection), while view-based pre-training tasks in MoCo can better capture image-level dependency (thus benefits image classification). ViT outperforms DeiT probably due to the larger pre-training dataset. CLIP performs the best. To the best of our knowledge, language-augmented visual models such as CLIP enjoy the best scaling performance; In contrast, the scaling performance of language-free visual models are either less studied or less successful so far.

In Table 14, we presented the comparisons of random and language-augmented initialization for language-image model adaptation with more checkpoints under 5-shot settings. This includes ViT-Base and ViT-Large models of DeCLIP , OpenCLIP and CLIP .

In Table 15, we presented zero-shot results of more model checkpoints for both Industry and Academic Tracks. For Academic Tracks, we consider CLIP , DeCLIP , FILIP , SLIP , with network ViT-Base32 pre-trained on YFCC (15M). For Industry Tracks, we consdier DeCLIP, OpenCLIP and CLIP, with models ranging from ViT-Base to ViT-Large, and training data ranging from 88M to 400M image-text pairs.

G.4 Breakdown Experimental Results on CLIP

We show the individual linear probing and finetuning scores for comparing the random and language-augmented initialization in Table 13. Language initialization consistently outperforms random initialization across different domains: sample efficiency, parameter efficiency, and different datasets. See Sec. 4 for more discussions on the design and the effectiveness of the language-augmented initializations.

Appendix H Benefits of External Knowledge in Model Adaptation

We also explore the benefits of the external knowledge to models that are pre-trained without the external knowledge (e.g., CLIP). On CLIP, we compare the effect of adding different combinations of external knowledge (Wiktionary, the nubmer of GPT3 knowledge items). The results are summarized in 16, and detailed in Table 17.

In zero-shot settings, we find that when the external knowledge is available, CLIP demonstrates consistent improvement on four datasets and considerable gains on the other three datasets. This suggests that the knowledge can benefit language-image models (though varying between datasets) as a new language prompting technique for some datasets, even if the pre-trained model is trained without the external knowledge.

In few- / full-shot settings, we argue that the pre-trained model can selectively incorporate different knowledge sources to achieve the best adaptation performance. One simple strategy is to train the model with different knowledge sources, compare the split validation accuracy of checkpoints with different knowledge sources, and use the best one for testing. We called it as knowledge-augmented adaptation, in contrast to the baseline method knowledge-free adaptation, where no collected external knowledge is employed at all. We find such simple strategy is already effective for linear probing and fine-tuning CLIP. As shown in Table. 16, knowledge-based adaptation of CLIP consistently improves over knowledge-free adaptation both in terms of accuracy and the number of wins. Notably, by selectively incorporating the external knowledge, it shows a significant 1.8 improvement for 5-shot CLIP fine-tuning. Note that such gain comes for free, even when the base CLIP model is not pre-trained with the external knowledge. We believe more sophisticated knowledge adaptation strategy can yield even better performance and we leave that to future work.

These experiments show that the collected external knowledge on Elevater is a useful resource for improving the adaptation of language-augmented visual models.