Unified Vision and Language Prompt Learning
Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, Chen Change Loy
Introduction
Vision-language (VL) models (e.g., CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021)) pre-trained on millions of image-text pairs have shown promising transferability on various downstream tasks such as few-shot learning (Zhou et al., 2022a; b; Ju et al., 2021) and open-vocabulary perception (Gu et al., 2022; Zhou et al., 2022c; Zang et al., 2022; Ghiasi et al., 2021). When adapting large VL models to downstream tasks, it is often prohibitive to fine-tune the entire model directly due to their huge parameter size. Therefore, many studies (Gao et al., 2021; Li & Liang, 2021; Lester et al., 2021; Zhou et al., 2022a; Lu et al., 2022; Ju et al., 2021; Yao et al., 2021; Jia et al., 2022; Bahng et al., 2022) explore prompt tuning, i.e., freezing the parameters of VL models and only fine-tuning extra learnable parameters, known as prompts, to adapt VL models for downstream tasks efficiently and effectively.
A typical VL model consists of two sub-networks—an image encoder and a text encoder—to extract representations for visual and text modalities. Correspondingly, existing prompt tuning approaches can be grouped into two types: text prompt tuning and visual prompt tuning. For text prompt tuning methods, such as CoOp (Zhou et al., 2022a)), extra text prompts, treated as learnable parameters, are applied on the text encoder in CLIP (Fig. 1(a)), to mitigate the potentially sub-optimal hand-crafted text prompt templates (e.g., ``a photo of a [CLASS].''). On the contrary, visual prompt tuning approaches focus on modulating the image encoder (Fig. 1(b)). For example, VPT (Jia et al., 2022) injects learnable parameters, known as visual prompts, into multiple layers of vision Transformers. In general, these prompt approaches consider tuning the representations of visual and text modalities independently.
Despite significant improvements that have been achieved, we observe that existing prompt tuning approaches (Zhou et al., 2022a; Jia et al., 2022) cannot obtain consistent performance improvements due to inherent variances in visual features and text embedding in downstream tasks. That is to say, using the single-modal prompt may acquire good results on one task while performing poorly on other tasks. To analyze this phenomenon, we measure the discrepancy in data distribution via the intra-class variance of visual features and inter-class variance of text embedding and study the correlation between data statistics and performance improvements. As shown in Fig. 1(d), when the intra-class variance of image features are large (bottom right), we observe that CoOp struggles to learn suitable text prompts for improving the text classifier. As for visual prompt tuning, VPT faces difficulties when inter-class variance of text embeddings are low, as shown in bottom left of Fig. 1(e). That is, if the text classifiers are established based on text embeddings of low separability, tuning visual prompts would lend little help to improve the final performance. Moreover, intra-class visual variance and inter-class text variance are typically orthogonal. Thus, we can observe that single-modal prompt tuning methods vary widely across different datasets: CoOp beats VPT with 8.1% on Flowers102 (Nilsback & Zisserman, 2008), while VPT outperforms CoOp by 8.4% on EuroSAT (Helber et al., 2019). Choosing which prompt modality to tune becomes a dilemma given different downstream tasks. Clearly, the key is to simultaneously adapt both text and visual prompts to overcome the vast differences across different data distributions. A straightforward solution is to introduce both text and visual prompts to the model and jointly optimize two modality-specific prompts together. However, we find that such a naïve joint training leads to sub-optimal performance due to the intrinsic discrepancy between text and image modalities. In particular, the performance is occasionally worse than tuning modality-specific prompts as shown in our experiments.
Solving the issues above requires modality-agnostic optimization to bridge the isolated prompts. To this end, we present a unified prompt tuning for both the text and visual modalities, dubbed as Unified Prompt Tuning (UPT), see Fig. 1(c). Specifically, we start with a shared initial prompt and propose a lightweight self-attention network to generate the prompt for CLIP text and visual encoders. We empirically show that such a design can preserve the benefit of individual modality.
Our contributions are summarized as follows: 1) We provide a comprehensive analysis for existing text or visual prompt tuning strategies; 2) We present a unified prompt method for VL models to tune both the visual and text modality representations; 3) We conduct extensive experiments to show that unified prompt tuning is a viable strategy that outperforms previous single-modal prompt tuning methods, especially under the few-shot learning and domain generalization settings. We hope our work can motivate future research using multi-modal prompts for VL models.
Methodology
We first introduce vision-language models focusing on CLIP (Radford et al., 2021), in company with text/visual prompt tuning approaches for visual recognition in Sec. 2.1. We then analyze the limitations of previous single-modal prompt tuning approaches in Sec. 2.2. Finally, we present technical details of our proposed unified prompt learning in Sec. 2.3.
where denotes the cosine similarity and is a fixed temperature value (e.g., ). Conceptually, such a decision process for the input image in Eq. (1) is formulated in a way that the text encoder takes a role of generating dynamic classifiers from open-set categories , with the image encoder producing encoded visual features . In practice, it is generally infeasible to fine-tune the millions of parameters (i.e., and ) in a VL model for transfer learning in every downstream task.
Visual Prompt Tuning. Conversely, visual prompt tuning methods focus on extracting more transferable visual features while keeping the visual encoder unchanged. Following the success of text prompt tuning approaches, recent Visual Prompt Tuning (VPT) (Jia et al., 2022) introduces a similar prompt tuning recipe for the visual encoder . Suppose the image encoder contains Vision Transformer layers, the output of -th layer, , where , is given by:
where stands for the length of visual prompts. Two VPT variants are proposed: VPT-shallow and VPT-deep. For VPT-shallow, the visual prompts are only inserted into the first Transformer layer (). Whereas for VPT-deep, visual prompts are introduced at every layer. The learnable visual prompts are data-independent, which once learned, can modulate the visual features of input images for better downstream transfer learning.
2 Analysis
We conduct a series of probing studies to analyze the characteristics of text/visual prompt tuning. First, when adapting the CLIP model with two representative text and visual prompt tuning approaches (CoOp (Zhou et al., 2022a) and VPT (Jia et al., 2022)), we measure the variance of both visual features and text embeddings (i.e., classifiers) for all 11 downstream vision datasets (see Appendix A for detailed implementations). For text prompt tuning, as shown in Fig. 1(d), we observe that CoOp performs well on datasets with low intra-class variance between visual features, such as Flowers102, but fails on Food101 dataset with high intra-class feature variance. As for visual prompt tuning, VPT succeeds in improving performance on SUN397 dataset with large inter-class text embeddings, while being less effective on Food101 and Flowers102 with relatively smaller inter-class text embedding variance. The performance improvements of text/visual prompt tuning are highly correlated with the variance of visual features or text embeddings in downstream datasets.
In conclusion, the single-modal prompt tuning approaches (CoOp and VPT), face the dilemma that consistent improvements over Zero-shot CLIP are hard to achieve due to inherent variances of visual features and text embedding in downstream tasks. Our observation motivates us to present a unified prompt tuning method that tunes the and at the same time.
3 Unified Prompt Tuning
Experiments
In this section, we conduct experiments under two problem settings, i.e., (i) few-shot image classification (Sec. 3.1) and (ii) domain generalization (Sec. 3.2). We also present ablation studies in Sec. 3.3 on several design choices.
Baselines. We compare our approach against the following methods: (1) Zero-shot CLIP. This baseline uses hand-crafted text prompt templates and does not involve any prompt-learning strategies. (2) Single-modal Prompt Tuning methods, including CoOp (Zhou et al., 2022a), for the text modality, and VPT (Jia et al., 2022), for the visual modality. In the domain generalization setting, we further compare with CoCoOp (Zhou et al., 2022b), which improves CoOp's generalization performance with an input-conditional design. For VPT, we report the results of both the shallow and deep variants, as described in Sec. 2.1.
In this section, we measure a model's generalization ability by conducting prompt tuning using different strategies, with just a limited amount of labeled examples per-class in the specific downstream task. Detailed implementation is presented in Appendix B.
Datasets. We follow Zhou et al. (2022b) to use 11 datasets (ImageNet (Deng et al., 2009), Caltech101 (Fei-Fei et al., 2004), OxfordPets (Parkhi et al., 2012), StanfordCars (Krause et al., 2013), Flowers102 (Nilsback & Zisserman, 2008), Food101 (Bossard et al., 2014), FGVC-Aircraft (Maji et al., 2013), SUN397 (Xiao et al., 2010), UCF101 (Soomro et al., 2012), DTD (Cimpoi et al., 2014), EuroSAT (Helber et al., 2019)) as our benchmarks. Following Zhou et al. (2022a), we use the few-shot evaluation protocol selecting 1/2/4/8/16 shots for training and the whole test set for evaluation. We report averaged results over three runs with different random seeds to reduce the variance. The detailed results are shown in Fig. 4.
Limitation of Single-modal Baselines. Figure 4 shows that the performance improvements of existing text prompt tuning method CoOp and visual prompt tuning method VPT are not consistent across different datasets. In particular, CoOp obtains better performance than VPT on some datasets, such as StanfordCars and SUN397. However, for other datasets with high intra-class visual variances, VPT is much more effective than CoOp. For instance, on the EuroSAT dataset, VPT-deep beats CoOp by over 12%. The discrepancy of previous single-modal baselines is also consistent with our motivation in Fig. 1(d) and Fig. 1(e). According to VPT (Jia et al., 2022), VPT-deep is more effective than VPT-shallow, and our experimental results also verify this point. We later show that VPT-shallow obtains much stronger performance than the VPT-deep in the domain generalization setting (Sec. 3.2).
UPT vs. Single-modal Baselines. Our UPT achieves clear advantages over the single-modal prompt-tuning counterparts CoOp and VPT, as suggested by the averaged performance (top-left of Fig. 4). In general, the average performance gap between UPT and baselines increases with the shot number available for prompt tuning. Specifically, UPT obtains 0.48/1.36/1.29/2.46/3.19(%) accuracy improvements compared with the text prompt tuning method CoOp on 1/2/4/8/16 shots settings. Similarly, UPT achieves 0.89/2.70/2.03/2.40/2.01(%) accuracy gains over the visual prompt tuning approach VPT-deep. Notably, UPT significantly boosts the performance over CoOp and VPT-deep on challenging large datasets, such as ImageNet with 1,000 classes and SUN397 with 397 categories. UPT also surpasses CoOp and VPT-deep on fine-grained datasets such as StanfordCars and FGVC Aircraft. We also observe that UPT shows less improvement on the two datasets (OxfordPets and Food101), possibly caused by the noisy training data (Zhou et al., 2022a; Bossard et al., 2014). Overall, the experimental results in Fig. 4 demonstrate the effectiveness of our proposed UPT.
2 Domain Generalization
Pre-trained VL models like CLIP have shown strong generalization ability. However, the prompt tuned on a specific downstream dataset may hinder the generalization ability on categories outside the training set. In this section, we evaluate the generalization ability of different prompt tuning methods to out-of-distribution (OOD) data.
Datasets. We follow (Zhou et al., 2022a) to use five datasets (ImageNet (Deng et al., 2009), ImageNet V2 (Recht et al., 2019), ImageNet-Sketch (Wang et al., 2019), ImageNet-A (Hendrycks et al., 2021b) and ImageNet-R (Hendrycks et al., 2021a)) for evaluation. Following the protocol, we train a model on ImageNet and evaluate it on four other variants of ImageNet with their domains shifted.
Results. Table 1 summarizes the results. We report the average accuracy on both the source and target datasets (penultimate column), and the OOD average accuracy on target datasets (last column). The results show that VPT-shallow (row #2) achieves higher OOD accuracy than VPT-deep (row #3), and text prompt tuning methods outperform visual prompt tuning approaches. Furthermore, the proposed UPT (row #5) is generally a better option than single-modal baselines (rows #1-#4) and obtains comparable performance with CoCoOp. Our UPT achieves the best results four times on five target datasets, showing that UPT is a reliable prompt tuning method among its competitors in the domain generalization setting.
3 Ablation Studies
Comparison with the Joint Training Baseline. As shown in Fig. 5(a), a straightforward approach for multi-modal prompts is tune the text prompt (using CoOp) and visual prompt (using VPT) jointly. We investigate the effectiveness of such joint training scheme, and report its results in Table 2 row #4. From the results, we see that such a joint training solution performs slightly worse than the visual prompt tuning method VPT-deep (78.70% vs 79.39%), and shows inferior performance to our UPT.
Shared Prompts for Text and Visual Modalities. We also investigate the results of directly sharing prompts for different modalities. As shown in Fig. 5(b), the shared prompts will be optimized for both text and visual modalities. This scheme differs from the proposed UPT, where the shared prompts are transformed with self attention. Experimental results are presented in Table 2 row #5, and we observe that such a prompt sharing strategy achieves worst performance among all the methods.
MLP Baseline. For our proposed UPT, we use a Transformer layer with the self-attention operator to partially share the hyper-parameters for different modalities. Here, we study a simpler design that generates the unified prompts with two MLP layers. Results are presented on Table 2 row #6. The MLP baseline is still competitive, yielding best performance on two datasets. Nonetheless, the average results on 11 datasets is still poorer than the proposed self-attention based approach.
4 Qualitative Results
While it is hard to visualize what have been learned during text prompt tuning, it is possible to visualize the visual prompts learned by VPT and UPT following the self-supervised learning method, DINO (Caron et al., 2021). In particular, for each layer of the Vision Transformer (ViT), we can compute the self-attention response map of visual prompts and image patch tokens. Figure 6 compares such response maps by VPT and the proposed UPT. We find that UPT shows stronger self-attention responses compared with VPT. This could be the possible reason why UPT achieves better performance on the few-shot learning and the OOD generalization settings.
Related Work
Vision-Language Models. Recent vision-language pre-trained models (Radford et al., 2021; Jia et al., 2021) use the contrastive loss to align an image encoder (e.g., ViT (Dosovitskiy et al., 2021)) and a text encoder (e.g., BERT (Kenton & Toutanova, 2019)) in a common feature space. These vision-language models are trained on web-scale image-text pairs and are transferable across various downstream tasks such as point cloud classification (Zhang et al., 2022a), video classification (Qian et al., 2022), object detection (Gu et al., 2022; Du et al., 2022; Zhou et al., 2022c; Zang et al., 2022) and semantic segmentation (Ghiasi et al., 2021). In this work, we aim to explore how to adapt the CLIP model to the downstream few-shot recognition task.
Text Prompt Tuning. The concept of prompt tuning was first proposed in the NLP area (Liu et al., 2021; Gao et al., 2021; Li & Liang, 2021; Lester et al., 2021). In particular, a text prompt refers to a task-specific template for language models. For example, in sentiment analysis, the template might be ``I [MASK] the movie.'' where the mask placeholder will be filled with either ``love'' or ``hate.'' Common practices in text prompt tuning include (i) searching for a specific word in the dictionary, known as hard prompt learning (Gao et al., 2021), or (ii) turning masked tokens into learnable vectors, known as soft prompt learning (Li & Liang, 2021; Lester et al., 2021). Text prompt tuning has also been applied in computer vision after the emergence of large vision-language models (e.g., CLIP (Radford et al., 2021)), which are too big to fine-tune. A representative work is CoOp (Zhou et al., 2022a), which turned the input context tokens in CLIP's text branch into learnable vectors for adapting CLIP to downstream image recognition. Other follow-ups of CoOp include CoCoOp (Zhou et al., 2022b), DualCoOp (Sun et al., 2022), ProGrad (Xing et al., 2022), and ProDA (Lu et al., 2022).
Visual Prompt Tuning. The idea of visual prompt tuning is to adapt large pre-trained Vision Transformers (Dosovitskiy et al., 2021) by adding learnable parameters in the visual input space, which is analogous to text prompt tuning in NLP. VPT (Jia et al., 2022) and Visual Prompting (Bahng et al., 2022) both add trainable tokens to the input of Transformer models. A recent work, NOAH (Zhang et al., 2022b), uses neural architecture search algorithms to identify the optimal configuration of prompt modules. In comparison to the unimodal prompt learning methods discussed above, our paper provides a timely study on how to achieve a better trade-off using multimodal prompt learning.
Conclusion
With the rapid scaling of vision models along the size dimension, efficient downstream adaptation methods have become essential for facilitating large-scale deployment of vision models in the wild. Our paper provides a timely and comprehensive study on how to adapt large vision-language models like CLIP from the prompt learning perspective. In particular, our study unveils that the previous unimodal prompt tuning methods do not work consistently well across a wide range of vision datasets. In contrast, the proposed UPT method, despite having a simple design, achieves a better trade-off compared with the unimodal counterparts. The results suggest that one should exploit correspondences between different modalities for prompt learning.
On the other hand, the results achieved by UPT are by no means perfect: in the ablation studies we observe that some alternative designs, such as using MLP instead of Transformer, might sometimes give better performance. Overall, we believe multimodal prompt learning is a promising framework to build upon, and we expect more improvements can be achieved with more advanced (and efficient) designs.
References
Appendix A Intra-/Inter- Class Variance
In this section, we provide the implementation details about how we compute the intra-class visual variance and inter-class text variance for different datasets (Fig.1 (d)(e) in the main paper.
Intra-class Visual Variance. Given one dataset with classes in total, for each image that belongs to class , we first use the CLIP image encoder to extract the corresponding image feature . Then we get the intra-class variance of class as:
where denotes to the set of images that have the ground-truth class label , and refers to the mean values of class . Then we can compute the intra-class variance for all the classes as
Inter-class Text Variance. For each dataset, we first compute the CLIP text features of classs , and the mean value of all the classes. Then we get the inter-class text variance as:
Appendix B Implementation Details
Our implementation is based on the source code of CoOp (Zhou et al., 2022a). We use ViT-B/16 as the CLIP backbone (Radford et al., 2021). Following Zhou et al. (2022b), we set the context length of CoOp as (same for VPT). For Zero-shot CLIP and VPT, we use the default prompt template, ``a photo of a [CLS].'' We use SGD as the optimizer, with an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The batch size is set to 32 for all datasets.