CLIP-Adapter: Better Vision-Language Models with Feature Adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, Yu Qiao
Introduction
Visual understanding tasks, such as classification Krizhevsky et al 2012; He et al 2016; Howard et al 2017; Dosovitskiy et al 2021; Touvron et al 2021; Gao et al 2021a; Mao et al 2021, object detection Ren et al 2015; Carion et al 2020; Gao et al 2021b, and semantic segmentation Long et al 2015, have been improved significantly based on the better architecture designs and large-scale high-quality datasets. Unfortunately, collecting large-scale high-quality datasets for every visual task is labor-intensive and too expensive to scale. To solve the problem, the “pretraining-finetuning” paradigm, namely pretraining on large-scale datasets like ImageNet Krizhevsky et al 2012 and then fine-tuning on a variety of downstream tasks, has been widely adopted in vision domain. However, such approaches still need a huge amount of annotations for fine-tuning on many downstream tasks. Recently, Contrastive Language-Image Pretraining (CLIP) Radford et al 2021 was proposed for solving vision tasks by exploiting contrastive learning with large-scale noisy image-text pairs. It achieves inspirational performances on various visual classification tasks without any annotations (i.e., zero-shot transfer) by putting visual categories into suitable hand-crafted template as prompts.
Although prompt-based zero-shot transfer learning showed promising performances, designing good prompts remains an engineering problem that demands substantial time and domain knowledge. To address the issue, Context Optimization (CoOp) Zhou et al 2022 further proposed to learn continuous soft prompts with few-shot examples for replacing the carefully-chosen hard prompts. CoOp brings about significant improvement on few-shot classification over both zero-shot CLIP and linear probe CLIP settings, exhibiting the potential of prompt tuning on large-scale pretrained vision-language models. In this paper, we propose a different approach for better adapting vision-language models with feature adapters instead of prompt tuning. Different from CoOp that performs soft prompt optimization, we simply conduct fine-tuning on the light-weight additional feature adapters. Because of the over-parameterization of CLIP and lack of enough training examples, naive finetuning would lead to overfitting on specific datasets and the training process would be very slow owing to the forward and backward propagations across all CLIP layers.
Motivated by the adapter modules in parameter-efficient transfer learning Houlsby et al 2019, we propose CLIP-Adapter, which only finetunes a small number of additional weights instead of optimizing all parameters of CLIP, as shown in Figure 1. CLIP-Adapter adopts a lightweight bottleneck architecture to prevent the potential overfitting problem of few-shot learning by reducing the number of parameters. Meanwhile, CLIP-Adapter is different from Houlsby et al 2019 in two important aspects: CLIP-Adapter only adds two additional linear layers following the last layer of vision or language backbone. This is because the original pretrained encoder of CLIP has already been equipped with strong representation capabilities, so it only requires a lightweight adaption in the form of residuals. In contrast, the original adapter modules are inserted into all layers of the language backbone; In addition, CLIP-Adapter mixes the original zero-shot visual or language embedding with the corresponding finetuning feature via residual connection. Through such a “residual-style blending”, CLIP-Adapter can simultaneously exploit the knowledge stored in the original CLIP and the freshly learned knowledge originated from the few-shot training examples. Figure 2 gives an intuitive illustration of the differences between our CLIP-Adapter and other visual classification architectures. Overall, our contributions can be summarized as follows:
We propose CLIP-Adapter that conducts residual-style feature blending to achieve efficient few-shot transfer learning via fine-tuning.
Compared with CoOp, CLIP-Adapter achieves better few-shot classification performance while having a much simpler design, demonstrating that CLIP-Adapter is a promising alternative to prompt tuning.
We perform extensive ablation studies of CLIP-Adapter on eleven classification datasets to analyze its characteristics.
Related Work
Model Fine-Tuning. Deep neural network is data-hungry. However, collecting and annotating large amount of high-quality data is costly and even impossible for some special domains. The “pretraining-finetuning paradigm” offers a good solution to different computer vision Krizhevsky et al 2012; Simonyan and Zisserman 2015; He et al 2016 and natural language processing Devlin et al 2019; Dong et al 2019; Conneau et al 2020 tasks and has been widely adopted for many years. For data-efficient finetuning over downstream tasks, adapter modules Houlsby et al 2019 is proposed to freeze the weight of backbones and insert learnable linear layers to each Transformer layer. Subsequent work such as Parallel Adapter He et al 2022 and VL-Adapter Sung et al 2022b further extend Houlsby et al 2019’s ability to language and multimodal tasks. Some other methods, e.g., black-box tuning Sun et al 2022, ladder side-tuning Sung et al 2022a, sparse structure searchHu et al 2022, and Scaling&Shifting Lian et al 2022 also introduce different techniques for adapting large language and vision models in a parameter-efficient manner. Different from existing approaches, the proposed CLIP-Adapter applies a simple residual transformation layer over the feature embedding or classifier weight generated by CLIP. Thanks to the residual connection and bottleneck linear layer, CLIP-Adapter can improve the performance of CLIP on few-shot learning setting and achieve superior performance than the recently proposed CoOp. To alleviate the performance gap under distribution shifting, WiSE-FT Wortsman et al 2022 proposes a post-ensemble method for improving CLIP’s out-of-distribution robustness. While our CLIP-Adapter adopts a learnable gating ratio to dynamically balance and mix the knowledge from the original features and CLIP-Adapter’s outputs throughout the training stage.
Prompt Design. Prompt design Liu et al 2023 are popularized by the success of GPT series Radford et al 2019; Brown et al 2020. GPT-3 showed that a huge autoregressive language model trained on a large-scale dataset can perform any NLP tasks in a zero-shot or few-shot style without finetuning the base architecture. Following the brand new “pretrain, prompt, and predict” paradigm Liu et al 2023, various prompt design approaches are proposed recently. As the earliest attempt, one type of them focus on tuning pretrained language or vision-language models with natural language discrete prompts Tsimpoukelli et al 2021; Alayrac et al 2022; Wang et al 2022b; Yao et al 2022; Yao et al 2021; Gao et al 2021c or prompt engineering by mining or generating proper natural language discrete prompts Jiang et al 2020; Shin et al 2020; Gao et al 2021c. In contrast, continuous prompts circumvent the restriction from pretrained language models and are adopted by various approaches such as Gu et al 2022; Li and Liang 2021; Liu et al 2021; Lester et al 2021 on NLP tasks. Recently, they have also been introduced to vision tasks Jia et al 2022. Motivated by GPT-3, CLIP trains a large contrastive learning model over 400 million image-text pairs and demonstrates the potential for prompt-based zero-shot visual classification. With CLIP as the backbone, CoOp Zhou et al 2022 further shows that optimizing continuous prompts can largely surpass manually-designed discrete prompts on vision tasks. In this paper, we demonstrate that prompt tuning is not the only path to better vision-language models. Fine-tuning with a small portion of parameters can also achieve comparable or even better performance on vision tasks yet with much simpler design.
Vision-Language Models. Exploring the interaction between vision and language is a core research topic in artificial intelligence. Previously, attention-based approaches such as bottom-up top-down attention Anderson et al 2018, BAN Kim et al 2018, Intra-Inter Gao et al 2019, and MCAN Yu et al 2019 had dominated visual-language tasks. Inspired by the success of BERT Devlin et al 2019, ViLBERT Lu et al 2019, LXMERT Tan and Bansal 2019, UNITER Chen et al 2020, Oscar Li et al 2020, ALBEF Li et al 2021, and BEiT Wang et al 2022a further push the boundary of multimodal reasoning. Recently, CLIP Radford et al 2021 and ALIGN Jia et al 2021 demonstrates the power of visual-language contrastive representation learning. They achieve astonishing results on a wide spectrum of vision tasks, including 3D Zhu et al 2022; Zhang et al 2023b, video Lin et al 2022, and depth Zhang et al 2022a understanding. To further close the gap between CLIP and supervised training, CoOp proposes a continuous prompt optimization method for improving the performance on visual classification tasks. While CoOp improves vision-language models from the perspective of prompt design, our CLIP-Adapter explores simple finetuning with lightweight feature adapters.
Our Approach
In this section, we introduce the proposed CLIP-Adapter. In Section 3.1, we first revisit CLIP and CoOp from the perspective of classifier weight generation. In Section 3.2, we elaborate the details of the proposed CLIP-Adapter. In Section 3.3, we provide several variants of CLIP-Adapter.
where stands for the temperature of Softmax, represents the prototype weight vector for class , and denotes the probability of category .
Different from supervised training, in the paper, we are interested in image classification with few-shot examples. Training the backbone and classifier together from scratch with a small number of samples is prone to overfit certain datasets and might suffer from severe performance drop on the test split. Typically, the representative paradigm on few-shot learning is to first pretrain the backbone on a large-scale dataset, and then transfer the learned knowledge to downstream tasks by either conducting zero-shot prediction directly or further fine-tuning on few-shot examples.
CLIP adheres to the zero-shot transfer style – it first pretrains the visual backbone and textual encoder through contrastive learning on large-scale noisy image-text pairs, and then after pretraining, CLIP directly performs image classification without any finetuning. Given an image classification downstream dataset that contains categories with their natural language name , CLIP constructs to place each category name into the pre-defined hard prompt template . Then the language feature extractor encodes the resulting prompt as a classifier weight . We denote the classifier weight generation process as below:
For both CLIP and CoOp, with the generated classifier weight , where , we can thus calculate the prediction probability for class by the previously mentioned Eq. (1).
2 CLIP-Adapter
Unlike CoOp’s prompt tuning, we present an alternative framework for achieving better vision-language models on few-shot image classification by fine-tuning additional feature adapters. We claim that the previous widely-adopted “pretrain-finetuning” paradigm would fail in finetuning the whole CLIP backbone under the few-shot setting due to the enormous amount of parameters and the shortage of training examples. Hence, we propose CLIP-Adapter, which only appends a small number of additional learnable bottleneck linear layers to CLIP’s language and image branches while keeping the original CLIP backbone frozen during few-shot fine-tuning. However, naive fine-tuning with additional layer may still fall into overfitting on the few-shot examples. To deal with overfitting and improve the robustness of CLIP-Adapter, we further adopt residual connections to dynamically blend the fine-tuned knowledge with the original knowledge from CLIP’s backbone.
Specifically, given the input image and a set of categories’ natural language names , the image feature and classifier weight from the original CLIP backbone are computed with Equations (1) and (2). Afterwards, two learnable feature adapters, and , each of which contains two layers of linear transformations, are integrated to transform and , respectively. We adopt a residual connection for the feature adapter to avoid forgetting the original knowledge encoded by the pretrained CLIP. Two constant values and are employed as “residual ratio” to help adjust the degree of maintaining the original knowledge for better performance. In summary, the feature adapters can be written as
where , and , are the weights of bottleneck linear layers for visual branch and text branch, respectively. The new knowledge captured via finetuning is added with the original features via residual connections:
After obtaining new image feature and classifier weight , we also adopt Equation (1) to calculate the category probability vector and predict the image category by selecting the class that has the highest probability: .
During the few-shot training, the weights of and are optimized with the contrastive loss following original CLIP Radford et al 2021 as below:
where is the total number of training examples; represents all learnable parameters of CLIP-Adapter.
3 Variants of CLIP-Adapter
Our CLIP-Adapter has three structural variants: 1) only fine-tuning the feature adapter for the image branch while keeping the text branch frozen; 2) only fine-tuning the feature adapter for the text branch while keeping the image branch frozen; 3) fine-tuning both the image and text branches of CLIP backbone. In terms of the hyperparameters and , we observe that different datasets have different optimal and values. Choosing the hyperparameters manually is time-consuming and laborious. Thus we also explore learning and in a differentiable manner by setting them as learnable parameters. In this way, and can be dynamically predicted from either visual feature or classifier weight via a hypernetwork : .
Experiments
Datasets. Following CLIP and CoOp, we select 11 image classification datasets to validate CLIP-Adapter’s effectiveness, namely ImageNet Deng et al 2009, StanfordCars Krause et al 2013, UCF101 Soomro et al 2012, Caltech101 Fei-Fei et al 2004, Flowers102 Nilsback and Zisserman 2008, SUN397 Xiao et al 2010, DTD Cimpoi et al 2014, EuroSAT Helber et al 2019, FGVCAircraft Maji et al 2013, OxfordPets Parkhi et al 2012, and Food101 Bossard et al 2014. Specifically, we train our CLIP-Adapter under the few-shot setups of , , , , shots and then test the tuned models on full test splits. Considering the randomness of few-shot training, we run every setting three times and report the average accuracy. We conduct all experiments on a single NVIDIA A100 GPU.
Implementation Details. The first variant of CLIP-Adapter is adopted by default if not specified, which finetunes the image feature while freezing the classifier weight. In other words, it only implements CLIP-Adapter for the visual adapter. The results of other variants that activate text adapter are presented in Section 4.2.5. We use the same training hyperparameters as CoOp, including a batch size of and a learning rate of for all datasets except for the residual ratio . We perform hyperparameter searching over different value selections of for each dataset and report the best performance among all searching spaces. We use ResNet-50 He et al 2016 as the visual backbone (visual encoder) and 12-layer Transformer as classifier weight generator (textual encoder). The hidden embedding dimensionality of both visual and text bottleneck layers is set to , which is a quarter of the original embedding dimensionality. In contrast to the learnable continuous prompts in CoOp, simple hand-crafted hard prompts are utilized as the text inputs of CLIP-Adapter, which is the same as CLIP. For generic-category image datasets, such as ImageNet, we adopt “a photo of a {class}” as the hard prompt template. For fine-grained classification datasets, we specify its corresponding domain keyword in the template for a better performance, for instance, “a centered satellite photo of {class}” for EuroSAT, and similarly for other fine-grained datasets.
Notes on Image Pre-processing. There are two image pre-processing methods adopted by existing methods. The first one is adopted by CLIP and the second one is reported in CoOp. We denote them as CLIP-style and CoOp-style preprocessings, respectively. They are both composed of random cropping, resizing, and random horizontal flip transformations. Their differences lie in the resizing. The CLIP-style pre-processing resizes the cropped image’s short side to while keeping its original aspect ratio. In contrast, the CoOp-style resizes an image’s both sides to . By default, we follow CoOp-style preprocessing. In Section A of the Appendix, we present the result comparison under the CLIP-style preprocessing which preserves the original aspect ratios of the cropped images.
2 Comparison on Few-Shot Learning
We compare our CLIP-Adapter with three baseline models – the Zero-shot CLIP Radford et al 2021, Linear probe CLIP Radford et al 2021, and CoOp Zhou et al 2022. In our implementation, CLIP-Adapter shares the same hand-crafted hard prompts with Zero-shot CLIP Radford et al 2021 for fair comparisons. CoOp Zhou et al 2022 substitutes discrete tokens with learnable continuous vectors. Thus there are multiple candidate positions to place the class token in the prompt template, namely at the front, in the middle, or at the end. Here, we choose CoOp’s best-performance variant – placing the class token at the end of the -token soft prompt and shares such a context among different classes. Linear probe CLIP Radford et al 2021 trains an additional linear classifier on top of its visual encoder and follows a few-shot training manner. It is different from our bottleneck adapter that finetunes both the image feature and classifier weight in a dynamic and residual fashion.
2.2 Performance Comparison & Analysis
The main results are presented in Figure 3. From the average accuracy over the 11 datasets shown at the top-left corner, CLIP-Adapter clearly outperforms the other three baseline models on all different shot setups, demonstrating its superior few-shot learning capacity. It is especially worth noticing that, under extreme conditions such as -shot or -shot training setup, CLIP-Adapter achieves larger performance improvements against the baselines, which indicates a better generalization ability in data-deficient training circumstances.
Compared with Zero-shot CLIP Radford et al 2021, our CLIP-Adapter achieves significant performance gains over all 11 datasets. The ranked absolute performance improvements for all 11 datasets under the 16-shot training setup are shown in Figure 4. For the first five fine-grained datasets, from EuroSAT to FGVCAircraft, CLIP-Adapter achieves huge performance boosts ranging from to . The improvements become smaller on more challenging and generic datasets, such as Caltech101 and ImageNet. As for OxfordPets and Food101, CLIP-Adapter shows relatively limited improvements, since the original results of Zero-shot CLIP are already quite decent.
Compared with Linear probe CLIP Radford et al 2021, which follows a similar style to finetune the pretrained vision-language models, CLIP-Adapter also shows comprehensive performance advantages. Under 1-shot and 2-shot training setups, Linear probe CLIP barely reaches the performance of Zero-shot CLIP, but CLIP-Adapter can always surpass Zero-shot CLIP and exceed Linear probe CLIP by a large margin. For instance, the absolute margin of 1-shot and 2-shot training setups are and for OxfordPets, and and for ImageNet, respectively.
Compared with CoOp Zhou et al 2022, although it has already gained huge improvements over Zero-shot CLIP, CLIP-Adapter still outperforms CoOp on all datasets and different shot settings. Note that CLIP-Adapter handles few-shot learning from a totally different perspective (i.e., fine-tuning) instead of CoOp’s prompt tuning. This suggests finetuning lightweight adapters with residual connections for prompt-fixed pretrained vision-language models can achieve better performance than prompt engineering Liu et al 2023.
2.3 Efficiency Comparison & Analysis
In Table 1, we provide the comparison of parameters, training budget, and inference speed for different methods. Compared to the baselines, CLIP-Adapter achieves the best accuracy while maintaining the best trade-off between parameters and efficiency. Specifically, CLIP-Adapter uses only half amount of parameters as Linear probe CLIP does. Although CLIP-Adapter costs 37 more minutes than Linear probe CLIP, CLIP-Adapter achieves a large 7.89% accuracy boost. Moreover, CLIP-Adpater takes less training time and faster inference speed than CoOp. Although CoOp only uses a small number of parameters, its learnable prompts are equipped in front of the text encoder and require both the forward and backward propagation, as shown by green dotted lines in Figure 1. Calculating the weights and gradients for the large-scale text encoder cost much more GPU memory and training time. In contrast, our CLIP-Adapter utilizes the hand-crafted prompts and only back-propagates the gradients through the lightweight adapters, achieving superior computational efficiency.
2.4 Observation on Optimal Residual Ratio
Interestingly, we observe the best residual ratio , to some extent, reflects the characteristics of different datasets under the “pretrain-finetuning” paradigm. A larger semantic gap between pretrained and finetuning datasets requires CLIP-Adapter to learn a higher portion of knowledge from the newly adapted feature compared to the original CLIP’s output, thus resulting in a larger optimal residual ratio, and vice versa. For fine-grained datasets on specialized domains, like EuroSAT of satellite images and DTD of detailed textures, the optimal residual ratio is usually located within the range from to . By contrast, the best value of comprehensive and generic image datasets (e.g., Caltech-101 and ImageNet) is often around .
2.5 Variants with Text Adapter
Here, we investigate the other two variants of CLIP-Adapter mentioned in Section 3.3 – finetuning the text adapter while keeping the visual adapter frozen and finetuning both the text and visual adapters. Rather than manually selecting the residual ratios for each dataset, we utilize learnable parameters and since it is time-efficient and can also achieve satisfactory performance. We compare their performances on four datasets that can be divided into two categories – fine-grained datasets (EuroSAT & DTD) and generic datasets (Caltech101 & ImageNet). As shown in Figure 5, we can conclude that the text adapter and visual adapter perform comparably and both improve the classification accuracy greatly over Zero-shot CLIP. In addition, adopting visual adapter only is better than text adapter only. This indicates that it is more important to conduct image feature adaption than text feature adaption for few-shot image classification, since the semantic gap between visual features in pretrained and finetuning datasets is larger than that of text features. Surprisingly, combining both adapters together does not observe a better performance than visual adapter only. This demonstrates that the text and visual adapters might capture redundant information or even conflict with each other.
2.6 Where to Insert CLIP-Adapter?
By default, we insert our residual-style adapters at the end of CLIP’s encoder. We also investigate other positions in the visual backbone to equip our adapter. In Table 2, we adopt ViT-B/16 as the visual backbone and respectively add the visual adapter after its 2nd, 4th, 6th, 8th, 10th, and 12th layers, where the 12th-layer variant denotes our final solution. As shown, our approach achieves superior performance with minimal computational cost when inserted at the end. Inserting at earlier layers requires more computation resources for back-propagating the gradients and, to some degree, harms the pretrained knowledge in CLIP. Inserting adapters in all layers yields a total of 5.20M parameters, which is much heavier-weight than inserting only at the 12th layer (0.52M). The latter can well alleviate the over-fitting issue on few-shot training data, and largely preserve the pre-trained knowledge of CLIP by inserting the adapter at the end.
2.7 Comparison with other Adapter Methods
In Table 3, we equip CLIP with other existing adapter-based methods Houlsby et al 2019; He et al 2022 and compare with our CLIP-Adapter. As shown, our approach outperforms them in terms of both performance and efficiency. This is because we largely preserve the knowledge obtained from CLIP pretraining by inserting adapters at the end with residual connections, while others adopt non-residual forms and insert them densely in the middle of the backbone, which adversely influences CLIP’s pretrained knowledge and results in overfitting.
2.8 Comparison with ELEVATER Benchmark Baselines
In Table 4, we compare CLIP-Adapter with the newly proposed ELEVATER benchmark Li et al 2022 for pretrained vision-language models. ELEVATER includes three baselines: Random-Init with Two-Projection, Language-Init with Two-Projection, and Language-Init with One-Projection. As shown in the table, compared to training additional projection layers initialized by randomness on the text encoder, our CLIP-Adapter achieves higher classification accuracy, since the residual design can largely preserve the pretrained knowledge in CLIP.
3 Visualization of Learned Manifold
We use t-SNE Van der Maaten and Hinton 2008 to visualize the manifold of CLIP, CoOp, CLIP-Adapter without residual connections, and CLIP-Adapter with residual connections after training them on the EuroSAT dataset. The t-SNE visualization results are presented in Figure 6, where the numbers to stand for the categories of AnnualCrop, Forest, Herbaceous Vegetation Land, Highway or Road, Industrial Buildings, Pasture Land, Permanent Crop Land, Residential Buildings, River, Sea or Lake, respectively. It is clearly illustrated that in high-dimensional classification space, the CLIP-Adapter with residual connections in sub-figure (d) shows much more obvious separation of image features belong to different categories. As for the confusing categories such as Highway or Road (red points), Permanent Crop Land (pink points), and Pasture Land (brown points), compared with other methods, our CLIP-Adapter is more effective in detecting the similarities among the image manifolds from the same class. In summary, the visualization results prove that CLIP-Adapter is good at learning better feature manifolds under few-shot setups.
4 Ablation Studies
In this section, we perform several ablation studies for CLIP-Adapter. We choose the best-performance variant which only activates the visual adapter, and select two datasets – DTD & ImageNet, serving as the representatives of fine-grained and generic datasets, to perform the ablation studies.
Dimension of Bottleneck Layer. We first conduct ablations by varying the hidden dimension of bottleneck layers. The results are shown in Table 5, where represents the dimension of the original image feature. By reducing the hidden dimension from to , we observe that either too small or too large intermediate dimensionality will deteriorate the performance significantly and the best bottleneck dimension is , which is able to preserve enough semantics without redundancy.
Residual Ratio . Moreover, we perform ablation study of the residual ratio . From Table 6, we can see that the best residual ratio of fine-grained dataset DTD is , and that of generic dataset ImageNet is . This verifies our observation in Section 4.2.4 that adapting fine-grained dataset requires more new knowledge than old knowledge, and the case is opposite to the generic dataset. Note that when equals to , it is equivalent to Zero-shot CLIP since no new knowledge is learned. When is set to , the classification is fully rely on the adapted feature (CLIP-Adapter w/o Res). However, this is not optimal because CLIP-Adapter tends to overfit in such condition. Combining Table 6 and Figure 6, we can also conclude the advantages of residual connections in CLIP-Adapter: 1) avoids overfitting on few-shot examples and improves the generalization ability of CLIP-Adapter with the help of zero-shot knowledge; 2) preserves the freedom for learning better image feature or classifier weight through few-shot fine-tuning.
Influence of Prompt Styles. In this section, we investigate the influence of different prompt styles on few-shot performance. For ImageNet dataset, the default hard prompt used as text inputs of CLIP-Adapter is simply “a photo of a {class}”. Besides, we also try prompt ensembling Zhou et al 2022 of 7 hard prompts. The 7 hard prompt templates are: “itap of a {class}”, “a bad photo of the {class}”, “a origami {class}”, “a photo of the large {class}”, “a {class} in a video game”, “art of the {class}” and “a photo of the small {class}”. Another candidate prompt style is the mixture of hard prompt and learnable soft prompt Zhou et al 2022. As shown in Table 7, the prompt ensembling strategy slightly outperforms hard prompt and achieves the best performance among all three prompt styles. The experimental results prove that raw text descriptions contain helpful knowledge which is effective and robust under different situations. In contrast, soft prompts don’t have clear meaning and are not a ideal source for zero-shot knowledge.
Ablation of Visual Backbones. We also study the influence of visual backbones on few-shot learning performance (16 shots). There are four candidate visual backbones including ResNet-50, ResNet-101, ViT-B/32, and ViT-B/16. As reported in Table 8, CLIP-Adapter consistently outperforms CoOp when we vary the visual backbones on both DTD and ImageNet datasets.
Robustness under Distribution Shift. To further validate the robustness of CLIP-Adpater, we also perform experiments to observe performance variation by shifting the distribution. We train our CLIP-Adapter on ImageNet Deng et al 2009 and respectively evaluate on four out-of-distribution datasets: ImageNetV2 Recht et al 2019, ImageNet-Sketch Hendrycks et al 2021b, ImageNet-A Wang et al 2019, and ImageNet-R Hendrycks et al 2021a. As shown in Table 9, CLIP-Adapter consistently outperforms other baselines and demonstrates enough robustness against distribution shift.
Finetuning Whole CLIP vs. CLIP-Adapter. To verify the claim that finetuning the whole CLIP would lead to overfitting. We perform ablation experiments on finetuning different components of CLIP-Adapter ( denotes unfrozen for training). For finetuning CLIP’s encoders, we adopt early stopping as suggested to obtain the highest accuracy. From the results presented in Table 10, we observe that finetuning either CLIP’s visual or textual encoder would hurt the performance and take more training time. This indicates the overfitting of the huge-parameter CLIP on the few-shot dataset and the effectiveness of the proposed adapter.
Conclusions and Future Work
We present CLIP-Adapter as an alternative of prompt-based approaches for few-shot image classification. The CLIP-Adapter revives the “pretrain-finetuning” paradigm by only fine-tuning a small number of additional bottleneck layers. To further improve the generalization ability, we adopt residual connections parameterized by a residual ratio to dynamically blend zero-shot knowledge with new adapted features. According to the experimental results, CLIP-Adapter outperforms competitive baselines on eleven image classification datasets under different few-shot setups. Extensive ablation studies confirm our design and prove CLIP-Adapter’s ability in learning better feature manifolds. In the future, we plan to extend CLIP-Adapter to more vision-language applications and tasks. We will also combine CLIP-Adapter with cache models Zhang et al 2022b; Zhu et al 2023 and enhanced prompts Zhang et al 2023a to further unleash the power of CLIP backbone.
Data Availability Statement
No new data were created during the study. All experiments of this manuscript were conducted (training and evaluation) on 11 publicly available image classification datasets Deng et al 2009; Krause et al 2013; Soomro et al 2012; Fei-Fei et al 2004; Nilsback and Zisserman 2008; Xiao et al 2010; Cimpoi et al 2014; Helber et al 2019; Maji et al 2013; Parkhi et al 2012; Bossard et al 2014.
Acknowledgement
This project is funded in part by the National Natural Science Foundation of China (No.62206272), by the National Key R&D Program of China Project (No.2022ZD0161100), by the Centre for Perceptual and Interactive Intelligence (CPII) Ltd under the Innovation and Technology Commission (ITC)’s InnoHK, and by the General Research Fund of Hong Kong RGC Project 14204021. Hongsheng Li is a PI of CPII under the InnoHK.
Appendix A Appendix
In Figure 7, we present the result comparison under CLIP-style preprocessing of few-shot learning on 11 datasets. Compared to CoOp-style preprocessing, the performances of all methods are improved under CLIP-style preprocessing. Similar to Figure 3 of the main body, CLIP-Adapter still outperforms other baselines across different shot settings.