Conditional Prompt Learning for Vision-Language Models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei Liu
Introduction
Recent research in large-scale vision-language pre-training has achieved striking performance in zero-shot image recognition radford2021learning; jia2021scaling; furst2021cloob; li2021supervision, demonstrating a potential in learning open-world visual concepts for such a paradigm. The key design lies in how visual concepts are modeled. In traditional supervised learning where labels are discretized, each category is associated with a randomly initialized weight vector that is learned to minimize the distance with images containing the same category. Such a learning method focuses on closed-set visual concepts, limiting the model to a pre-defined list of categories, and is unscalable when it comes to new categories unseen during training.
In contrast, for vision-language models We follow existing studies radford2021learning; jia2021scaling; furst2021cloob; li2021supervision to refer to CLIP-like models as vision-language models. like CLIP radford2021learning and ALIGN jia2021scaling, the classification weights are diametrically generated by a parameterized text encoder (e.g., a Transformer vaswani2017attention) through prompting liu2021pre. For instance, to differentiate pet images containing different breeds of dogs and cats, one can adopt a prompt template like “a photo of a {class}, a type of pet” as input to the text encoder, and as a result, class-specific weights for classification can be synthesized by filling in the “{class}” token with real class names. Compared to discrete labels, vision-language models’ source of supervision comes from natural language, which allows open-set visual concepts to be broadly explored and has been proven effective in learning transferable representations radford2021learning; jia2021scaling.
With the rise of such powerful vision-language models, the community has recently started to investigate potential solutions to efficiently adapt these models to downstream datasets zhou2021coop; yao2021cpt; gao2021clip; wortsman2021robust. To fit web-scale data, such as the 400 million pairs of images and texts used by CLIP, vision-language models are purposefully designed to have high capacity, entailing that the model size would be enormous, typically with hundreds of millions of parameters or even billions. Therefore, fine-tuning the entire model, as often adopted in deep learning research he2016deep, is impractical and might even damage the well-learned representation space.
A safer approach is to tune a prompt by adding some context that is meaningful to a task, like “a type of pet” for the pet dataset mentioned above, which has been found effective in improving performance radford2021learning. However, prompt engineering is extremely time-consuming and inefficient as it has to be based on trial and error, and does not guarantee an optimal prompt either. To automate prompt engineering, Zhou et al. zhou2021coop have recently explored the concept of prompt learning—a recent trend in NLP shin2020autoprompt; jiang2020can; li2021prefix; zhong2021factual; lester2021power; gao2020making—for adapting pre-trained vision-language models. Their approach, Context Optimization (CoOp), turns context words in a prompt into a set of learnable vectors, taking advantage of the differentiable nature of neural networks. With only a few labeled images for learning, CoOp achieves huge improvements over intensively-tuned manual prompts across a wide range of image recognition datasets.
In our study, we identify a critical problem of CoOp: the learned context is not generalizable to wider unseen classes within the same task. Figure 1 illustrates the problem: the context learned by CoOp works well in distinguishing the base classes like “arrival gate” and “cathedral” but suffers a significant drop in accuracy when it is transferred to the new (unseen) classes, such as “wind farm” and “train railway”—even though the task’s nature remains the same, i.e., recognizing scenes. The results suggest that the learned context overfits the base classes, thus failing to capture more generalizable elements that are vital for broader scene recognition. We argue that such a problem is caused by CoOp’s static design: the context, which is fixed once learned, is optimized only for a specific set of (training) classes. On the contrary, the manually-designed prompts adopted by the zero-shot method are relatively generalizable.
To address the weak generalizability problem, we introduce a novel concept: conditional prompt learning. The key idea is to make a prompt conditioned on each input instance (image) rather than fixed once learned. To make the model parameter-efficient, we introduce a simple yet effective implementation of conditional prompt learning. Specifically, we extend CoOp by further learning a lightweight neural network to generate for each image an input-conditional token (vector), which is combined with the learnable context vectors. We call our approach Conditional Context Optimization (CoCoOp). Pronounced as /kku:p/. An overview is shown in Figure 2. Interestingly, the paradigm of CoCoOp is analogous to image captioning vinyals2015show, which explains why instance-conditional prompts are more generalizable: they are optimized to characterize each instance (more robust to class shift) rather than to serve only for some specific classes.
We present comprehensive experiments on 11 datasets covering a diverse set of visual recognition tasks. Specifically, we design a base-to-new generalization setting where a model is first learned using base classes and then tested on completely new classes. Compared with the zero-shot method radford2021learning and CoOp zhou2021coop, our approach achieves the best overall performance (Table 1). Importantly, CoCoOp gains significant improvements over CoOp in unseen classes (Figure 3(a)), allowing the gap between manual and learning-based prompts to be substantially reduced.
In a more challenging scenario where the context learned for one task is directly transferred to other tasks with drastically different classes, CoCoOp still beats CoOp with a clear margin (Table 2), suggesting that instance-conditional prompts are more transferable and have the potential to succeed at larger scale. CoCoOp also obtains stronger domain generalization performance than CoOp (Table 3), further justifying the strengths of dynamic prompts.
In summary, our research provides timely insights into the generalizability problem in prompt learning, and crucially, demonstrates the effectiveness of a simple idea in various problem scenarios. We hope our approach and the findings presented in this work can pave the way for future research in generalizable—and transferable—prompt learning.
Related Work
We mainly review studies focused on aligning images and texts to learn a joint embedding space radford2021learning; jia2021scaling; zhang2020contrastive. The idea of cross-modality alignment is certainly not new and has been investigated since nearly a decade ago—though with dramatically different technologies than today.
A typical vision-language model consists of three key elements: two for image and text encoding while the third is related to the design of loss functions. In early days, models for processing images and texts are often designed and also learned independently, with their outputs connected by extra modules (losses) for alignment. Images are often encoded using hand-crafted descriptors socher2013zero; elhoseiny2013write or neural networks frome2013devise; lei2015predicting, while texts are encoded using, for instance, pre-trained word vectors socher2013zero; frome2013devise or the frequency-based TF-IDF features elhoseiny2013write; lei2015predicting. In terms of cross-modality alignment, common approaches include metric learning frome2013devise, multi-label classification joulin2016learning; gomez2017self, and n-gram language learning li2017learning. Recently, a study suggests that training the vision part with an image captioning loss can make the visual representation more transferable desai2021virtex.
Recent vision-language models radford2021learning; jia2021scaling; furst2021cloob; li2021supervision bridge the two modalities by learning two encoders jointly. Also, the models are now built with much larger neural networks. As discussed in Zhou et al. zhou2021coop, recent successes in vision-language models are mainly attributed to the developments in i) Transformers vaswani2017attention, ii) contrastive representation learning chen2020simple; he2020momentum; henaff2020data, and iii) web-scale training datasets radford2021learning; jia2021scaling. A representative approach is CLIP radford2021learning, which trains two neural network-based encoders using a contrastive loss to match pairs of images and texts. After consuming 400 million data pairs, the CLIP model demonstrates a remarkable zero-shot image recognition capability. Similar to CoOp zhou2021coop, our approach is orthogonal to the research of CLIP-like models radford2021learning; jia2021scaling; furst2021cloob; li2021supervision, aiming to offer an efficient solution for adapting pre-trained vision-language models to downstream applications.
Prompt Learning
This topic originates from the NLP domain. The motivation was to view pre-trained language models, such as BERT devlin2019bert or GPT radford2019language, as knowledge bases from which information useful to downstream tasks is elicited petroni2019language. Concretely, given a pre-trained language model, the task is often formulated as a “fill-in-the-blank” cloze test, such as asking the model to predict the masked token in “No reason to watch. It was [MASK]” as either “positive” or “negative” for sentiment classification. The key lies in how to design the underlined part, known as prompt (template), in such a format familiar to the model.
Instead of manually designing a prompt, research in prompt learning aims to automate the process with the help of affordable-sized labeled data. Jiang et al. jiang2020can use text mining and paraphrasing to generate a group of candidate prompts, within which the optimal ones are chosen to have the highest training accuracy. Shin et al. shin2020autoprompt propose AutoPrompt, a gradient-based approach that selects from a vocabulary the best tokens that cause the greatest changes in gradients based on the label likelihood. Our research is most related to continuous prompt learning methods zhong2021factual; li2021prefix; lester2021power, where the main idea is to turn a prompt into a set of continuous vectors that can be end-to-end optimized with respect to an objective function. See Liu et al. liu2021pre for a more comprehensive survey.
In computer vision, prompt learning is a nascent research direction that has only been explored very recently zhou2021coop; yao2021cpt; rao2022denseclip; ju2021prompting; zhang2021pointclip. Our research is built on top of CoOp zhou2021coop, which is the earliest work to bring continuous prompt learning to the vision domain for adaptation of pre-trained vision-language models. Crucially, our approach solves the weak generalizability problem of CoOp zhou2021coop, based on a simple idea of conditional prompt learning—which to our knowledge is also novel in the context of NLP and thus could be of interest to the NLP community as well.
Zero-Shot Learning (ZSL)
is another relevant research area where the goal is similar to ours, i.e., to recognize novel classes by training only on base classes wang2019survey; xian2017zero; chao2016empirical; yi2022exploring. Moreover, the generalization problem where a model trained on base classes often fails on novel classes is also linked to the “seen-class bias” issue raised in the ZSL literature xian2017zero. The most common approach to ZSL is to learn a semantic space based on auxiliary information such as attributes huynh2020fine or word embeddings frome2013devise; wang2018zero. Different from existing ZSL methods, our work addresses the emerging problem of adapting large vision-language models and uses drastically different techniques based on prompting.
Methodology
An overview of our approach is shown in Figure 2. Below we first provide brief reviews on CLIP radford2021learning, which is the base model used in this paper, and CoOp zhou2021coop. Then, we present the technical details of our approach as well as the rationale behind the design. Same as CoOp, our approach is applicable to broader CLIP-like vision-language models.
known as CLIP radford2019language, has well demonstrated the potential of learning open-set visual concepts. CLIP is built using two encoders, one for image and the other for text, as shown in Figure 2. The image encoder can be either a ResNet he2016deep or a ViT dosovitskiy2021image, which is used to transform an image into a feature vector. The text encoder is a Transformer vaswani2017attention, which takes as input a sequence of word tokens and again produces a vectorized representation.
During training, CLIP adopts a contrastive loss to learn a joint embedding space for the two modalities. Specifically, for a mini-batch of image-text pairs, CLIP maximizes for each image the cosine similarity with the matched text while minimizes the cosine similarities with all other unmatched texts, and the loss is computed in a similar fashion for each text too. After training, CLIP can be used for zero-shot image recognition. Let be image features generated by the image encoder and a set of weight vectors produced by the text encoder, each representing a category (suppose there are categories in total). In particular, each is derived from a prompt, such as “a photo of a {class}” where the “{class}” token is filled with the -th class name. The prediction probability is then
where denotes cosine similarity and is a learned temperature parameter.
Context Optimization (CoOp)
aims to overcome the inefficiency problem in prompt engineering for better adapting pre-trained vision-language models to downstream applications zhou2021coop. The key idea in CoOp is to model each context token using a continuous vector that can be end-to-end learned from data. Concretely, instead of using “a photo of a” as the context, CoOp introduces learnable context vectors, , each having the same dimension with the word embeddings. The prompt for the -th class, denoted by , now becomes where is the word embedding(s) for the class name. The context vectors are shared among all classes. CoOp has an alternative version that learns class-specific context, which is not considered here because it is not straightforward to transfer class-specific context to unseen classes. Let denote the text encoder, the prediction probability is then
To adapt CLIP to a downstream image recognition dataset, a cross-entropy loss can be used as the learning objective. Since the text encoder is differentiable, gradients can be propagated all the way back to update the context vectors. Note that the base model of CLIP is frozen in the entire training process (ours too).
2 CoCoOp: Conditional Context Optimization
CoOp is a data-efficient approach allowing the context vectors to be trained with only a few labeled images in a downstream dataset. However, as discussed CoOp is not generalizable to wider unseen classes within the same task. We argue that instance-conditional context can generalize better because it shifts the focus away from a specific set of classes—for reducing overfitting—to each input instance, and hence to the entire task.
A straightforward way to implement CoCoOp is to build neural networks to get context tokens. However, such a design would require the size of a neural network, which is much larger than having context vectors as in CoOp. Here we propose a parameter-efficient design that works very well in practice. Specifically, on top of the context vectors, we further learn a lightweight neural network, called Meta-Net, to generate for each input a conditional token (vector), which is then combined with the context vectors. See Figure 2 for a sketch of the architecture.
Let denote the Meta-Net parameterized by , each context token is now obtained by where and . The prompt for the -th class is thus conditioned on the input, i.e., . The prediction probability is computed as
During training, we update the context vectors together with the Meta-Net’s parameters . In this work, the Meta-Net is built with a two-layer bottleneck structure (Linear-ReLU-Linear), with the hidden layer reducing the input dimension by . The input to the Meta-Net is simply the output features produced by the image encoder. We leave exploration of more advanced designs for future work.
Experiments
Our approach is mainly evaluated in the following three problem settings: 1) generalization from base to new classes within a dataset (Section 4.1); 2) cross-dataset transfer (Section 4.2); 3) domain generalization (Section 4.3). All models used in our experiments are based on the open-source CLIP radford2021learning. https://github.com/openai/CLIP. Before discussing the results, we provide the details of the experimental setup below.
For the first two settings, i.e., base-to-new generalization and cross-dataset transfer, we use the 11 image recognition datasets as in Zhou et al. zhou2021coop, which cover a diverse set of recognition tasks. Specifically, the benchmark includes ImageNet deng2009imagenet and Caltech101 fei2004learning for classification on generic objects; OxfordPets parkhi2012cats, StanfordCars krause20133d, Flowers102 nilsback2008automated, Food101 bossard2014food and FGVCAircraft maji2013fine for fine-grained classification; SUN397 xiao2010sun for scene recognition; UCF101 soomro2012ucf101 for action recognition; DTD cimpoi2014describing for texture classification; and finally EuroSAT helber2019eurosat for satellite imagery recognition. For domain generalization experiments, we use ImageNet as the source dataset and four other variants of ImageNet that contain different types of domain shift as the target datasets, namely ImageNetV2 recht2019imagenet, ImageNet-Sketch wang2019learning, ImageNet-A hendrycks2021natural and ImageNet-R hendrycks2021many.
Following Zhou et al. zhou2021coop, we randomly sample for each dataset a few-shot training set while using the original test set for testing. We only evaluate the highest shot number studied in Zhou et al. zhou2021coop, i.e., 16 shots, which is sufficient to justify our approach. For learning-based models, the results are averaged over three runs.
Baselines
The direct rival to our approach is CoOp zhou2021coop, which essentially learns static prompts (in comparison to our dynamic prompts). The zero-shot method, i.e., CLIP radford2021learning is also compared, which is based on manual prompts. It is worth mentioning that the manual prompt for each dataset was intensively tuned using all classes in the test data radford2021learning.
Training Details
Our implementation is based on CoOp’s code. https://github.com/KaiyangZhou/CoOp. Throughout the experiments, we use the best available vision backbone in CLIP, i.e., ViT-B/16. Zhou et al. zhou2021coop have suggested that a shorter context length and a good initialization can lead to better performance and stronger robustness to domain shift. Therefore, we fix the context length to 4 and initialize the context vectors using the pre-trained word embeddings of “a photo of a” for both CoOp and CoCoOp. Due to the instance-conditional design, our approach is slow to train and consumes much more GPU memory than CoOp. Therefore, to ensure the model can fit into a GPU and meanwhile reduce the training time, we train CoCoOp with batch size of 1 for 10 epochs. Such a limitation is discussed in more detail in Section 5.
1 Generalization From Base to New Classes
Solving the weak generalizability problem of CoOp is the main focus in this research. On each of the 11 datasets, we split the classes equally into two groups, one as base classes and the other as new classes. Learning-based models, i.e., CoOp and CoCoOp, are trained using only the base classes while evaluation is conducted on the base and new classes separately to test generalizability. The detailed results are shown in Table 1.
The split does not guarantee that the two class groups are equally difficult, as evidenced in CLIP’s bumpy results: the base and new accuracy numbers are dramatically different. For convenience, we refer to base accuracy as the performance in base classes; and similarly for new accuracy. Nonetheless, CoOp’s new accuracy is consistently much weaker than the base accuracy on nearly all datasets, leaving a huge gap of almost 20% on average (82.69% vs 63.22%). Despite maintaining an advantage over CLIP in terms of average performance, CoOp’s gains in the base classes are nearly zeroed out by the catastrophic failures in the new classes, highlighting the need to improve generalizability for learning-based prompts.
CoCoOp Significantly Narrows Generalization Gap
As shown in Table 1(a), CoCoOp improves the accuracy in unseen classes from 63.22% to 71.69%, which largely reduces the gap with manual prompts. The results confirm that instance-conditional prompts are more generalizable. A more detailed breakdown of per-dataset improvement is visualized in Figure 3(a) where we observe more than 10% increases in accuracy on 5 out of 11 datasets. Notably, on the challenging ImageNet dataset, CoCoOp’s surge from 67.88% to 70.43% represents a non-trivial progress (the 70.43% accuracy even surpasses CLIP’s 68.14%).
CoCoOp’s Gains in Generalization Far Outweigh Losses in Base Accuracy
In comparison to CoOp, performance drops in the base classes occur for CoCoOp on most datasets (see Figure 3(b)). This is reasonable because CoOp optimizes specifically for base classes whereas CoCoOp optimizes for each instance in order to gain more generalization over an entire task. But it is worth noting that on the 9 datasets where CoCoOp’s base accuracy drops below CoOp’s, most losses are under 3% (precisely on 6 out of 9 datasets), which are far outweighed by the gains in unseen classes shown in Figure 3(a); even for those where CoCoOp suffers the biggest losses, the boosts in generalization are mostly significant enough to turn the averages into positives, e.g., StanfordCars sees the worst base accuracy drop of -7.63% but has the third-highest accuracy gain of +13.19% in the new classes, which together bring a 5.56% positive improvement for CoCoOp.
CoCoOp Is More Compelling Than CLIP
When taking into account both the base and new classes, CoCoOp shows a gain of more than 4% over CLIP (75.83% vs 71.70), suggesting that instance-conditional prompts have a better potential in capturing more generalizable elements that are relevant for a recognition task. Theoretically, learning-based prompts have a much higher risk of overfitting base classes than manual prompts. Therefore, CLIP is a strong competitor to beat in unseen classes. Different from CoOp, we obtain promising results for CoCoOp: the new accuracy is even better than CLIP’s on 4 out of 11 datasets (i.e., ImageNet, OxfordPets, Food101 and SUN397) and not too far away from CLIP’s on the rest except FGVCAircraft where the gap between manual and learning-based prompts is generally large. In the ablation study on context length, we find that FGVCAircraft benefits from longer context, which is aligned with the findings in Zhou et al. zhou2021coop. To close or even overturn the gaps between manual and learning-based prompts in unseen classes, more efforts are required and we hope the insights presented in this research can help the community tackle the generalizability issue in prompt learning.
2 Cross-Dataset Transfer
Having demonstrated CoCoOp’s generalizability within a dataset, we further show that CoCoOp has the potential to transfer beyond a single dataset, which is a much more challenging problem because the fundamentals can be totally changed across different datasets (e.g., from object recognition to texture classification). We only consider prompt learning methods in this setting.
We compare CoCoOp with CoOp by transferring context learned from ImageNet, with all 1,000 classes used, to each of the other 10 datasets. The results are detailed in Table 2. On the source dataset, the two models perform similarly. Whereas on the target datasets, CoCoOp mostly outperforms CoOp by a clear margin. Since the ImageNet classes mainly contain objects, as well as a fair amount of dog breeds, it is reasonable to see high accuracy for both models on the relevant target datasets including Caltech101 and OxfordPets.
By comparison, the performance on other datasets with distant—and more fine-grained or specialized—categories is much lower, such as FGVCAircraft and DTD (containing various textures) where the accuracy numbers are well below 50%. Nonetheless, CoCoOp exhibits much stronger transferability than CoOp on the two mentioned datasets as well as on most other fine-grained or specialized datasets.
3 Domain Generalization
Generalization to out-of-distribution data is a capability essential for machine learning models to succeed in practical applications taori2020measuring; zhou2021domain. Zhou et al. zhou2021coop have revealed that their learnable prompts are more robust than manual prompts to domain shift. We are also interested to know if instance-conditional prompts still maintain the advantages as in previous experiments.
Following Zhou et al. zhou2021coop, we evaluate CoCoOp’s domain generalization performance by transferring the context learned from ImageNet to the four specially designed benchmarks. We also include the comparison with CLIP. Table 3 shows the results. Both prompt learning methods clearly beat CLIP on all target datasets. Compared to CoOp, CoCoOp performs slightly worse on ImageNetV2 but better on the other three. The results confirm that instance-conditional prompts are more domain-generalizable.
4 Further Analysis
We consider a practical problem scenario where the recognition targets originally composed of base classes are expanded to include completely new classes. This problem is relevant to the existing continual learning literature parisi2019continual but different in that the model here does not have access to any training data from new classes and needs to perform zero-shot recognition on them. We compare CLIP, CoOp and CoCoOp using the 11 datasets. The average results are reported in Table 4. Clearly, CoOp loses competitiveness against CLIP as their performance is similar but the former needs training data. Again, CoCoOp beats the two competitors with a significant margin.
Initialization
To understand the impact of initialization, we conduct an ablation study by comparing word embeddings-based initialization and random initialization while keeping all other parameters identical. For random initialization, we follow Zhou et al. zhou2021coop to sample from a zero-mean Gaussian distribution with 0.02 standard deviation. Figure 4(a) shows the base-to-new generalization results averaged over the 11 datasets, which suggest that a proper initialization is more beneficial to both the base and new classes. Note that the findings from Figure 4 only represent the overall trend while each individual dataset might have a different result.
Context Length
The ablation study on context length is also carried out in the base-to-new generalization setting. Following Zhou et al. zhou2021coop, we study 4, 8 and 16 context tokens. For fair comparison, we use random initialization for all context tokens. Figure 4(b) summarizes the results on the 11 datasets. The differences in the base classes are fairly small whereas in the new classes the models with a longer context length clearly perform better. From Figure 4(a) and (b) we observe that using 8 randomly initialized context tokens is marginally better than using 4 properly initialized context tokens, suggesting that a further boost might be possible if we initialize 8 context tokens with word embeddings.
CoCoOp vs a Bigger CoOp
Since CoCoOp introduces more parameters than CoOp, namely the Meta-Net, one might question if the improvements simply come from an increased learning capacity. To clear the doubt, we remove the Meta-Net part and increase the number of context tokens in CoOp to the maximum such that CoOp’s and CoCoOp’s sizes are similar. The results in Table 5 show that increasing the parameter size is not the key.
Limitations
The first limitation is about training efficiency: CoCoOp is slow to train and would consume a significant amount of GPU memory if the batch size is set larger than one. The reason is because CoCoOp is based on an instance-conditional design that requires for each image an independent forward pass of instance-specific prompts through the text encoder. This is much less efficient than CoOp that only needs a single forward pass of prompts through the text encoder for an entire mini-batch of any size.
The second limitation is that on 7 out of the 11 datasets (see Table 1), CoCoOp’s performance in unseen classes still lags behind CLIP’s, indicating that more efforts are needed from the community to fully close or overturn the gaps between manual and learning-based prompts.
Discussion and Conclusion
Our research addresses an important issue that arises with the availability of large pre-trained AI models, i.e., how to adapt them to downstream applications. These models, also called foundation models bommasani2021opportunities, have received increasing attention from academia and industry in both the vision and NLP communities because they are so powerful in terms of their capabilities for diverse downstream tasks. However, foundation models are costly to pre-train in terms of data scale and compute resources; and typically contain an enormous number of parameters in order to develop sufficient capacity. For instance, the CLIP model radford2021learning based on ViT-B/16 used in our experiments has a whopping 150M parameter size. These factors together highlight the need for research of efficient adaptation methods for democratizing foundation models.
Our studies, which follow the line of parameter-efficient prompt learning zhou2021coop, provide timely insights into the generalizability issue of static prompts, and more importantly, demonstrate that a simple design based on conditional prompt learning performs superbly in a variety of problem scenarios, including generalization from base to new classes, cross-dataset prompt transfer, and domain generalization.
In terms of future work, one direction is to further develop conditional prompt learning with potentially a more efficient implementation that can accelerate the training, as well as enhance generalizability. The cross-dataset transfer experiments indicate that instance-conditional prompts are more transferable—compared to static prompts—across tasks of varying natures. Therefore, it would be interesting to see if such an idea can scale to, e.g., bigger model size for the Meta-Net, larger-scale training images, and even heterogeneous training data mixed with different datasets.
This work is supported by NTU NAP, and under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).
Appendix
Appendix A Results on DOSCO-2k
The DOSCO (DOmain Shift in COntext) benchmark zhou2022device contains 7 image recognition datasets, which cover a wide range of classification problems, such as generic object recognition, fine-grained recognition on aircraft models, and action recognition. Unlike existing domain generalization datasets where the domain labels are manually defined and often limited to image style variations, DOSCO-2k focuses on broader contextual domain shift, which is automatically detected by a neural network pre-trained on the Places dataset zhou2017places. Following Zhou et al. zhou2022device, we use the 2k version where the training and validation splits in each dataset have 2,000 images in total (1,600 for training and 400 for validation).
Results
We study three methods’ domain generalization performance on DOSCO-2k: CLIP, CoOp and CoCoOp. All models are trained on the training set and the checkpoints with the best validation performance are used for final test in unseen domains. Table 6 shows the results of four different architectures. It is clear that the two learnable methods outperform the zero-shot method with a large margin, despite having only a small number of parameters to tune. CoCoOp beats CoOp on 4 out of 7 datasets but CoOp’s average performance is higher. In summary, the results suggest that efficient adaptation methods like CoOp and CoCoOp have great potential in tackling transfer learning problems.