Learning to Prompt for Vision-Language Models

Kaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei Liu

Introduction

A common approach for building state-of-the-art visual recognition systems is to train vision models to predict for a fixed set of object categories using discrete labels (He et al. 2016; Dosovitskiy et al. 2021). From a technical point of view, this is achieved by matching image features—produced by a vision model like ResNet (He et al. 2016) or ViT (Dosovitskiy et al. 2021)—with a fixed set of weights that are seen as visual concepts and initialized randomly. Although training categories often have a textual form, such as “goldfish” or “toilet paper,” they will be converted into discrete labels just for easing the computation of the cross-entropy loss, leaving the semantics encapsulated in texts largely unexploited. Such a learning paradigm limits visual recognition systems to closed-set visual concepts, making them unable to deal with new categories since additional data are required for learning a new classifier.

Recently, vision-language pre-training such as CLIP (Radford et al. 2021) and ALIGN (Jia et al. 2021) has emerged as a promising alternative for visual representation learning. The main idea is to align images and raw texts using two separate encoders—one for each modality. For instance, both CLIP and ALIGN formulate the learning objective as a contrastive loss, which pulls together images and their textual descriptions while pushes away unmatched pairs in the feature space. By pre-training at a large scale, models can learn diverse visual concepts and can readily be transferred to any downstream task through prompting (Radford et al. 2021; Jia et al. 2021; Fürst et al. 2021; Li et al. 2021; Singh et al. 2021; Yuan et al. 2021). In particular, for any new classification task one can first synthesize the classification weights by giving sentences describing task-relevant categories to the text encoder, and then compare with image features produced by the image encoder.

We observe that for pre-trained vision-language models, the text input, known as prompt, plays a key role in downstream datasets. However, identifying the right prompt is a non-trivial task, which often takes a significant amount of time for words tuning—a slight change in wording could make a huge difference in performance. For instance, for Caltech101 (Figure 1(a), 2nd vs 3rd prompt), adding “a” before the class token brings more than 5% increase in accuracy. Moreover, prompt engineering also requires prior knowledge about the task and ideally the language model’s underlying mechanism. This is exemplified in Figure 1(b-d) where adding task-relevant context can lead to significant improvements, i.e., “flower” for Flowers102, “texture” for DTD and “satellite” for EuroSAT. Tuning the sentence structure could bring further improvements, e.g., putting “a type of flower” after the class token for Flowers102, keeping only “texture” in the context for DTD, and adding “centered” before “satellite photo” for EuroSAT. However, even with extensive tuning, the resulting prompts are by no means guaranteed to be optimal for these downstream tasks.

Inspired by recent prompt learning research in natural language processing (NLP) (Shin et al. 2020; Jiang et al. 2020; Zhong et al. 2021), we propose a simple approach called Context Optimization (CoOp) CoOp is pronounced as /ku:p/. to automate prompt engineering, specifically for pre-trained vision-language models. Concretely, CoOp models a prompt’s context words with learnable vectors, which could be initialized with either random values or pre-trained word embeddings (see Figure 2). Two implementations are provided to handle tasks of different natures: one is based on unified context, which shares the same context with all classes and works well on most categories; while the other is based on class-specific context, which learns a specific set of context tokens for each class and is found to be more suitable for some fine-grained categories. During training, we simply minimize prediction errors using the cross-entropy loss with respect to the learnable context vectors while keeping the entire pre-trained parameters fixed. The gradients can be back-propagated all the way through the text encoder, distilling the rich knowledge encoded in the parameters for learning task-relevant context.

To demonstrate the effectiveness of CoOp, we benchmark on 11 datasets, which cover a diverse set of visual recognition tasks including classification on generic objects, scenes, actions and fine-grained categories, as well as specialized tasks like recognizing textures and satellite imagery. The results show that CoOp effectively turns pre-trained vision-language models into data-efficient visual learners, requiring as few as one or two shots to beat hand-crafted prompts with a decent margin. The performance can be further boosted by using more shots, e.g., with 16 shots the margin over hand-crafted prompts averages at around 15% and reaches over 45% for the highest. CoOp also outperforms the linear probe model, which is known as a strong few-shot learning baseline (Tian et al. 2020). Furthermore, CoOp demonstrates much stronger robustness than the zero-shot model (which uses manual prompts) to domain shifts, despite being a learning-based approach.

In summary, we make the following contributions:

We present a timely study on the adaptation of recently proposed vision-language models in downstream applications and identify a critical problem associated with the deployment efficiency, i.e., prompt engineering.

To automate prompt engineering specifically for pre-trained vision-language models, we propose a simple approach based on continuous prompt learning and provide two implementations that can handle different recognition tasks.

We for the first time show that the proposed prompt learning-based approach outperforms both hand-crafted prompts and the linear probe model in terms of downstream transfer learning performance and robustness under domain shifts for large vision-language models.

We open-source our project at https://github.com/KaiyangZhou/CoOp.

We hope the findings together with the open-source code can inspire and facilitate future research on efficient adaptation methods for large vision-language models—an emerging topic related to democratization of foundation models (Bommasani et al. 2021) i.e., making them easier and cheaper to adapt for the wider community.

Related Work

Vision-language models have recently demonstrated great potential in learning generic visual representations and allowing zero-shot transfer to a variety of downstream classification tasks via prompting (Radford et al. 2021; Jia et al. 2021; Zhang et al. 2020; Singh et al. 2021; Yuan et al. 2021).

To our knowledge, the recent developments in vision-language learning, particularly CLIP (Radford et al. 2021) and ALIGN (Jia et al. 2021), are largely driven by advances in the following three areas: i) text representation learning with Transformers (Vaswani et al. 2017), ii) large-minibatch contrastive representation learning (Chen et al. 2020; He et al. 2020; Hénaff et al. 2020), and iii) web-scale training datasets—CLIP benefits from 400 million curated image-text pairs while ALIGN exploits 1.8 billion noisy image-text pairs.

The idea of mapping images and text onto a common embedding space has been studied since nearly a decade ago (Socher et al. 2013; Frome et al. 2013; Elhoseiny et al. 2013), but with drastically different technologies. For text features extraction, early work has mainly utilized pre-trained word vectors (Socher et al. 2013; Frome et al. 2013) or the hand-crafted TF-IDF features (Elhoseiny et al. 2013; Lei Ba et al. 2015). Matching images and text features has been formulated as metric learning (Frome et al. 2013), multi-label classification (Joulin et al. 2016; Gomez et al. 2017), n-gram language learning (Li et al. 2017), and the recently proposed captioning (Desai and Johnson 2021).

Our work is orthogonal to recent research in vision-language models, aiming to facilitate the adaptation and deployment of such models in downstream datasets.

2 Prompt Learning in NLP

Knowledge probing for large pre-trained language models, formally defined by Petroni et al. 2019 as “fill-in-the-blank” cloze tests, has recently sparked interest in prompt learning research in NLP (Shin et al. 2020; Jiang et al. 2020; Li and Liang 2021; Zhong et al. 2021; Lester et al. 2021; Gao et al. 2020; Liu et al. 2021b).

The basic idea of knowledge probing is to induce pre-trained language models to generate answers given cloze-style prompts, which can benefit a number of downstream tasks, such as sentiment analysis. Jiang et al. 2020 propose to generate candidate prompts through text mining and paraphrasing, and identify the optimal ones that give the highest training accuracy. Shin et al. 2020 introduce a gradient-based approach, which searches for tokens with the largest gradient changes in the label likelihood.

Most related to our work are continuous prompt learning methods (Zhong et al. 2021; Li and Liang 2021; Lester et al. 2021) which optimize continuous vectors in the word embedding space. A drawback of such methods compared to searching discrete tokens is the lack of a clear way to visualize what “words” are learned for the vectors. We refer readers to Liu et al. 2021a for a comprehensive survey in the topic of prompt learning in NLP.

It is worth noting that we are the first to apply prompt learning to the adaptation of large vision-language models in computer vision—which we view as an important topic for democratizing foundation models (Bommasani et al. 2021)—and justify that prompt learning not only brings significant improvements to computer vision tasks in terms of transfer learning performance but also produces robust models that can handle domain shifts.

Methodology

We briefly introduce vision-language pre-training with a particular focus on CLIP (Radford et al. 2021). Our approach is applicable to broader CLIP-like vision-language models.

CLIP consists of two encoders, one for images and the other for text. The image encoder aims to map high-dimensional images into a low-dimensional embedding space. The architecture of the image encoder can take the form of a CNN like ResNet-50 (He et al. 2016) or a ViT (Dosovitskiy et al. 2021). On the other hand, the text encoder is built on top of a Transformer (Vaswani et al. 2017) and aims to generate text representations from natural language.

Specifically, given a sequence of words (tokens), such as “a photo of a dog,” CLIP first converts each one of the token (including punctuation) into a lower-cased byte pair encoding (BPE) representation (Sennrich et al. 2016), which is essentially a unique numeric ID. The vocabulary size in CLIP is 49,152. To facilitate minibatch processing, each text sequence is encompassed with the [SOS] and [EOS] tokens and capped at a fixed length of 77. After that, the IDs are mapped to 512-D word embedding vectors, which are then passed on to the Transformer. Finally, the features at the [EOS] token position are layer normalized and further processed by a linear projection layer.

Training

CLIP is trained to align the two embedding spaces learned for images and text respectively. Specifically, the learning objective is formulated as a contrastive loss. Given a batch of image-text pairs, CLIP maximizes the cosine similarity for matched pairs while minimizes the cosine similarity for all other unmatched pairs. To learn diverse visual concepts that are more transferable to downstream tasks, CLIP’s team collects a large training dataset consisting of 400 million image-text pairs.

Zero-Shot Inference

Since CLIP is pre-trained to predict whether an image matches a textual description, it naturally fits zero-shot recognition. This is achieved by comparing image features with the classification weights synthesized by the text encoder, which takes as input textual descriptions specifying classes of interest. Formally, let f\bm{f} be image features extracted by the image encoder for an image x\bm{x} and {wi}i=1K\{\bm{w}_{i}\}_{i=1}^{K} a set of weight vectors generated by the text encoder. KK denotes the number of classes and each wi\bm{w}_{i} is derived from a prompt that could have the form of “a photo of a [CLASS].” where the class token is replaced by the specific class name, such as “cat,” “dog” or “car.” The prediction probability is then computed as

where τ\tau is a temperature parameter learned by CLIP and cos⁡(⋅,⋅)\cos(\cdot,\cdot) denotes cosine similarity.

Compared with the traditional classifier learning approach where closed-set visual concepts are learned from random vectors, vision-language pre-training allows open-set visual concepts to be explored through a high-capacity text encoder, leading to a broader semantic space and in turn making the learned representations more transferable to downstream tasks.

2 Context Optimization

We propose Context Optimization (CoOp), which avoids manual prompt tuning by modeling context words with continuous vectors that are end-to-end learned from data while the massive pre-trained parameters are frozen. An overview is shown in Figure 2. Below we provide several different implementations.

We first introduce the unified context version, which shares the same context with all classes. Specifically, the prompt given to the text encoder g(⋅)g(\cdot) is designed with the following form,

where each [V]m[\text{V}]_{m} (m ⁣∈ ⁣{1,…,M}m\!\in\!\{1,\ldots,M\}) is a vector with the same dimension as word embeddings (i.e., 512 for CLIP), and MM is a hyperparameter specifying the number of context tokens.

By forwarding a prompt t\bm{t} to the text encoder g(⋅)g(\cdot), we can obtain a classification weight vector representing a visual concept (still from the [EOS] token position). The prediction probability is computed as

where the class token within each prompt ti\bm{t}_{i} is replaced by the corresponding word embedding vector(s) of the ii-th class name.

Other than placing the class token at the end of a sequence as in Equation (2), we can also put it in the middle like

which increases flexibility for learning—the prompt is allowed to either fill the latter cells with supplementary descriptions or cut off the sentence earlier by using a termination signal such as full stop.

Class-Specific Context

Another option is to design class-specific context (CSC) where context vectors are independent to each class, i.e., [V]1i[V]2i…[V]Mi≠[V]1j[V]2j…[V]Mj[\text{V}]^{i}_{1}[\text{V}]^{i}_{2}\ldots[\text{V}]^{i}_{M}\neq[\text{V}]^{j}_{1}[\text{V}]^{j}_{2}\ldots[\text{V}]^{j}_{M} for i ⁣≠ ⁣ji\!\neq\!j and i,j∈{1,…,K}i,j\in\{1,\ldots,K\}. As an alternative to unified context, we find that CSC is particularly useful for some fine-grained classification tasks.

Training

is performed to minimize the standard classification loss based on the cross-entropy, and the gradients can be back-propagated all the way through the text encoder g(⋅)g(\cdot), making use of the rich knowledge encoded in the parameters to optimize the context. The design of continuous representations also allows full exploration in the word embedding space, which facilitates the learning of task-relevant context.

3 Discussion

Our approach specifically addresses the emerging problem of the adaptation of recently proposed large vision-language models such as CLIP (Radford et al. 2021). There are some differences that distinguish our approach from the prompt learning methods developed in NLP for language models (e.g., GPT-3 (Brown et al. 2020)). First, the backbone architectures are clearly different for CLIP-like models and language models—the former take both visual and textual data as input and produce alignment scores used for image recognition, while the latter are tailored to handle textual data only. Second, the pre-training objectives are different: contrastive learning vs autoregressive learning. This would lead to different model behaviors and thus require different module designs.

Experiments

We select 11 publicly available image classification datasets used in CLIP: ImageNet (Deng et al. 2009), Caltech101 (Fei-Fei et al. 2004), OxfordPets (Parkhi et al. 2012), StanfordCars (Krause et al. 2013), Flowers102 (Nilsback and Zisserman 2008), Food101 (Bossard et al. 2014), FGVCAircraft (Maji et al. 2013), SUN397 (Xiao et al. 2010), DTD (Cimpoi et al. 2014), EuroSAT (Helber et al. 2019) and UCF101 (Soomro et al. 2012) (see Appendix A for their statistics). These datasets constitute a comprehensive benchmark, which covers a diverse set of vision tasks including classification on generic objects, scenes, actions and fine-grained categories, as well as specialized tasks like recognizing textures and satellite imagery. We follow the few-shot evaluation protocol adopted in CLIP (Radford et al. 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison.

Training Details

CoOp has four versions: positioning the class token in the end or middle; unified context vs CSC. Unless otherwise stated, ResNet-50 (He et al. 2016) is used as the image encoder’s backbone and the number of context tokens MM is set to 16. Investigations on other design choices are discussed in Section 4.3. All models are built on top of CLIP’s open-source code. https://github.com/openai/CLIP. CoOp’s context vectors are randomly initialized by drawing from a zero-mean Gaussian distribution with standard deviation equal to 0.02. Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot (except for ImageNet where the maximum epoch is fixed to 50). To mitigate explosive gradients observed in the early training iterations, we use the warmup trick by fixing the learning rate to 1e ⁣− ⁣51e\!-\!5, only during the first epoch.

Baseline Methods

We compare CoOp with two baseline methods. The first is zero-shot CLIP, which is based on hand-crafted prompts. We follow the guideline of prompt engineering introduced by Radford et al. 2021. For generic objects and scenes, “a photo of a [CLASS].” is adopted. For fine-grained categories, task-relevant context is added like “a type of pet” for OxfordPets and “a type of food” for Food101. When it comes to specialized tasks such as recognizing textures in DTD, the prompt is customized as “[CLASS] texture.” where the class names are adjectives like “bubbly” and “dotted.” See Appendix A for the details. The second baseline is the linear probe model. As suggested by Radford et al. 2021 and a recent study on few-shot learning (Tian et al. 2020), training a linear classifier on top of high-quality pre-trained models’ features (like CLIP) can easily achieve performance that is on a par with that of state-of-the-art few-shot learning methods, which are often much more sophisticated. We follow the same training method used by Radford et al. 2021 to train the linear probe model.

Comparison with Hand-Crafted Prompts

Figure 3 summarizes the results. Our default model is CLIP+CoOp with the class token positioned in the end. The two different ways of positioning the class token achieve similar performance as their curves highly overlap. From the average performance displayed in the top-left corner, we observe that CLIP+CoOp is a strong few-shot learner, requiring only two shots on average to obtain a decent margin over zero-shot CLIP. Given 16 shots for training, the average gap brought by CoOp can be further increased to around 15%.

Figure 4 ranks the absolute improvements obtained by CoOp at 16 shots over hand-crafted prompts. Huge improvements are observed on specialized tasks namely EuroSAT and DTD where the increase in performance reaches over 45% and 20% respectively. The jumps in performance are also significant (those more than 10%) on most fine-grained datasets including Flowers102, StanfordCars and FGVCAircraft, as well as on scene and action recognition datasets (i.e., SUN397 & UCF101). Since ImageNet is a challenging dataset that contains 1,000 classes, the 4.77% improvement is also noteworthy. In contrast, the increases on the two fine-grained datasets, OxfordPets and Food101, are less appealing. We find that the negative results on Food101, for learning-based models including CoOp and linear probe, are caused by the noisy training data with “intense colors and sometimes wrong labels” (Bossard et al. 2014). By digging into CLIP+CoOp’s curves on these two datasets in Figure 3, we find there is a loss of momentum in performance improvements even with more shots used, seemingly an overfitting problem. A potential solution is to impose higher regularization like increasing the weight decay. Nonetheless, the overall results are strong enough to serve as evidence of CoOp’s capability of learning task-relevant prompts in a data-efficient manner.

Comparison with Linear Probe CLIP

In terms of the overall performance (Figure 3, top-left), CLIP+CoOp demonstrates clear advantages over the linear probe model. The latter requires more than 4 shots on average to match the zero-shot’s performance while CoOp’s average gain at 4 shots is already impressive. It is also clear that the gaps in the extreme low-data regime such as one or two shots are much larger, suggesting that CoOp is much more effective than learning a linear classifier from scratch for few-shot learning. We also observe that the linear probe model is comparable to CLIP+CoOp on the two specialized tasks (DTD & EuroSAT) as well as on a couple of fine-grained datasets (Flowers102 & FGVCAircraft)—this is not too surprising as the pre-trained CLIP space has been proved powerful, making the linear probe model a strong competitor. Nevertheless, CoOp’s CSC version can beat the linear probe CLIP on the aforementioned datasets, and moreover, shows much better potential when more shots become available. We later show that CoOp obtains much stronger performance than the linear probe model in domain generalization.

Unified vs Class-Specific Context

On average, using unified context leads to better performance. In terms of when to apply CSC and when not to, we have the following suggestions. For generic objects (ImageNet & Caltech101), scenes (SUN397) and actions (UCF101), using unified context is clearly better. Unified context also works better on some fine-grained datasets including OxfordPets and Food101, but on others like StanfordCars, Flowers102 and FGVCAircraft the CSC version is preferred. CSC also yields better performance on the two specialized tasks, DTD and EuroSAT, at 16 shots in particular. However, CSC mostly underperforms unified context in challenging low-data scenarios (fewer than 8 shots), which makes sense because CSC has more parameters than unified context and needs more data for training.

2 Domain Generalization

Since CoOp requires training on a specific data distribution, it risks learning spurious correlations that are detrimental to generalization in unseen distributions (domains), as suggested in recent studies (Taori et al. 2020; Zhou et al. 2021). On the contrary, zero-shot CLIP is not tied to a specific data distribution and has exhibited strong robustness to distribution shifts (Radford et al. 2021). In this section, we aim to unveil how robust CoOp is to distribution shifts, in comparison to zero-shot CLIP and the linear probe model.

The source dataset is ImageNet. The target datasets are ImageNetV2 (Recht et al. 2019), ImageNet-Sketch (Wang et al. 2019), ImageNet-A (Hendrycks et al. 2021b) and ImageNet-R (Hendrycks et al. 2021a), all of which have compatible class names with ImageNet allowing seamless transfer for the prompts learned by CoOp. ImageNetV2 is a reproduced test set using different sources while following ImageNet’s data collection process. ImageNet-Sketch contains sketch images belonging to the same 1,000 ImageNet classes. Both ImageNet-A and -R contain 200 classes derived from a subset of ImageNet’s 1,000 classes. The former consists of real-world adversarially filtered images that cause current ImageNet classifiers to produce low results, whereas the latter features a rendition of the ImageNet classes in diverse image styles such as paintings, cartoons and sculptures.

Results

Table 1 summarizes the results (with a variety of vision backbones). It is surprising that CoOp enhances CLIP’s robustness to distribution shifts, despite the exposure to the source dataset. This suggests that the learned prompts are also generalizable. Moreover, it is interesting to see that using fewer context tokens leads to better robustness. In contrast, the linear probe model obtains much worse results on these target datasets, exposing its weakness in domain generalization. In Appendix B, we provide the domain generalization results on DOSCO-2k (Zhou et al. 2022b), a recently proposed benchmark focusing on contextual domain shift.

3 Further Analysis

How many context tokens should be used? And is it better to have more context tokens? The results in Section 4.2 suggest having a shorter context length benefits domain generalization (probably due to less overfitting as fewer parameters are learned). Here we study this hyperparameter for source datasets. Specifically, we repeat experiments on the 11 datasets by varying the context length from 4 to 8 to 16. The average results are shown in Figure 5(a), which indicate that having more context tokens leads to better performance and that positioning the class token in the middle gains more momentum with longer context length. To sum up, there is no golden rule for selecting perfect context length since one needs to balance between performance and robustness to distribution shift.

Vision Backbones

Figure 5(b) summarizes the results on the 11 datasets using a variety of vision backbones covering both CNNs and ViTs. The results are expected: the more advanced the backbone, the better the performance. The gap between CoOp and hand-crafted prompts is significant across all architectures.

Comparison with Prompt Ensembling

The authors of CLIP (Radford et al. 2021) have suggested that additional improvements can be obtained by ensembling over multiple zero-shot classifiers generated using different hand-crafted prompts, such as “a photo of the large [CLASS].”, “a bad photo of the [CLASS].” and “a origami [CLASS].”, which reflect a different scale, view and abstraction respectively for an image. We are interested to know whether the prompts learned by CoOp can still maintain advantages when compared with prompt ensembling. For fair comparison, we use the select prompts from Radford et al. 2021, which have been extensively tuned on ImageNet, to construct the ensemble classifier. Table 2 shows the comparison and justifies the superiority of CoOp. Given the potential of prompt ensembling, future work could investigate how to improve CoOp from the ensembling perspective.

Comparison with Other Fine-tuning Methods

We further compare CoOp with other fine-tuning methods: i) fine-tuning CLIP’s image encoder; ii) optimizing a transformation layer added to the text encoder’s output; iii) optimizing a bias term added to the text encoder’s output. The results are shown in Table 5. Obviously, fine-tuning the image encoder does not work well. Adding a transformation layer slightly improves upon the zero-shot model. Adding a bias term shows promising results, but still largely underperforms CoOp, which suggests that the gradients that went through the text encoder provide more useful information.

Initialization

We compare random initialization with manual initialization. The latter uses the embeddings of “a photo of a” to initialize the context vectors for the 11 datasets. For fair comparison, we also set the context length to 4 when using random initialization. Table 3 suggests a “good” initialization does not make much difference. Though further tuning of the initialization words might help, in practice we suggest using the simple random initialization method.

Interpreting the Learned Prompts

is difficult because the context vectors are optimized in a continuous space. We resort to an indirect way by searching within the vocabulary for words that are closest to the learned vectors based on the Euclidean distance. Note that CLIP (Radford et al. 2021) uses the BPE representation (Sennrich et al. 2016) for tokenization, so the vocabulary includes subwords that frequently appear in text, such as “hu” (subsumed by many words like “hug” and “human”). Table 4 shows the searched results on some datasets. We observe that a few words are somewhat relevant to the tasks, such as “enjoyed” for Food101, “fluffy” and “paw” for OxfordPets, and “pretty” for DTD. But when connecting all the nearest words together, the prompts do not make much sense. We also observe that when using manual initialization (like “a photo of a”), the nearest words for the converged vectors are mostly the ones used for initialization. We conjecture that the learned vectors might encode meanings that are beyond the existing vocabulary. Overall, we are unable to draw any firm conclusion based on the observations because using nearest words to interpret the learned prompts could be inaccurate—the semantics of the vectors is not necessarily correlated with the nearest words.

Conclusion, Limitations and Future Work

Large pre-trained vision-language models have shown surprisingly powerful capabilities in diverse downstream applications. However, these models, also called vision foundation models given their “critically central yet incomplete” nature (Bommasani et al. 2021), need to be adapted using automated techniques for better downstream performance and efficiency.

Our research provides timely insights on how CLIP-like models can be turned into a data-efficient learner by using prompt learning, and reveals that despite being a learning-based approach, CoOp performs much better in domain generalization than manual prompts. The results serve as strong evidence that prompt learning has potential for large vision models. It is worth noting that our paper presents the first comprehensive study about adapting large vision models with prompt learning.

Though the performance is excellent, the results of CoOp are relatively difficult to interpret, like other continuous prompt learning methods in NLP. The experiments also reveal that CoOp is sensitive to noisy labels given the weak performance on Food101.

Nevertheless, the simplicity of CoOp allows easy extension for future work and there remain many interesting questions to explore, such as cross-dataset transfer (Zhou et al. 2022a) and test-time adaptation (Wang et al. 2020). It would also be interesting to investigate more generic adaptation methods for mega-size vision models (Jia et al. 2022; Bahng et al. 2022; Gao et al. 2021). In summary, we hope the empirical findings and insights presented in this work could pave the way for future research on efficient adaptation methods for emerging foundation models, which is still a nascent research topic.

Appendix

Appendix A Datasets Details

The detailed statistics of the 11 datasets, as well as the four variants of ImageNet, are shown in Table 6. The hand-crafted prompts used for zero-shot CLIP are also detailed in the table. For Caltech101, the “BACKGROUND_Google” and “Faces_easy” classes are discarded. For the video dataset, UCF101, the middle frame of each video is used as input to the image encoder.

Appendix B Results on DOSCO-2k

The DOSCO (DOmain Shift in COntext) benchmark (Zhou et al. 2022b) contains 7 image recognition datasets, which cover a wide range of classification problems, such as generic object recognition, fine-grained recognition on aircraft models, and action recognition. Unlike existing domain generalization datasets where the domain labels are manually defined and often limited to image style variations, DOSCO-2k focuses on broader contextual domain shift, which is automatically detected by a neural network pre-trained on the Places dataset (Zhou et al. 2017). Following Zhou et al. 2022b, we use the 2k version where the training and validation splits in each dataset have 2,000 images in total (1,600 for training and 400 for validation).

Results

We study three methods’ domain generalization performance on DOSCO-2k: CLIP, CoOp and CoCoOp (Zhou et al. 2022a). All models are trained on the training set and the checkpoints with the best validation performance are used for final test in unseen domains. Table 7 shows the results of four different architectures. It is clear that the two learnable methods outperform the zero-shot method with a large margin, despite having only a small number of parameters to tune. CoCoOp beats CoOp on 4 out of 7 datasets but CoOp’s average performance is higher. In summary, the results suggest that efficient adaptation methods like CoOp and CoCoOp have great potential in tackling transfer learning problems.

References