Texts as Images in Prompt Tuning for Multi-Label Image Recognition

Zixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai, Yiwen Guo, Wangmeng Zuo

Introduction

Recent few years have witnessed rapid progress in large vision-language (VL) pre-trained models as well as their remarkable performance on downstream vision tasks. A VL pre-trained model generally involves data encoders and it is becoming increasingly popular to exploit image-test contrastive loss to align the embedding of images and texts into a shared space. When adapting to downstream tasks in relatively data-limited or label-limited settings, it is often ineffective to fine-tune the entire model, due to its high complexity. Then, prompt tuning as a representative parameter-efficient learning paradigm has emerged as an efficient way to adapt VL model to downstream tasks.

Albeit considerable achievements have been made, existing prompt tuning methods generally require visual data to learn prompts (as shown in Fig. 1(a)). For example, CoOp learns from annotated images. CoCoOp further introduces generalizable input-conditional prompts. DualCoOp adapts CLIP to multi-label recognition tasks by training pairs of positive and negative prompts with partial-labeled images. Nonetheless, the performance of these prompting methods may be limited when it is infeasible to obtain sufficient image data or annotate the required images.

In this paper, we advocate treating Texts as Images for prompt tuning, i.e., TaI prompting. It is considered feasible as the image encoder and text encoder in many pre-trained VL models encode images and texts into a shared space. Given an image and its caption, the visual features produced by the image encoder will be close to the text feature of the caption produced by the text encoder. Therefore, in addition to extracting visual features from images, it is also feasible to extract text features as alternatives form, for example, descriptive sentences and captions, for prompt tuning (see Fig. 1(b)). TaI prompting has several interesting properties and merits. Taking a downstream image recognition task as an example, given a set of object categories, one can easily crawl a large set of text descriptions that contain object names from these categories. Text descriptions are easily accessible in this way, and class labels can be directly derived from text descriptions, which means, in contrast to prompting from images, TaI prompting may suffer less from the data-limited and label-limited issues.

We use multi-label image recognition to verify the effectiveness of our TaI prompting in this paper. To begin with, we crawl the captions from public image caption datasets (e.g., MS-COCO ) and localized narratives from object detection datasets (e.g., Open Images ) to form the training set of text descriptions. For any specific multi-label recognition task, we adopt a noun filter to map the nouns in the text descriptions to the corresponding object categories, and then only keep the text descriptions that contain one or more classes of target objects. To better cope with multi-label classification, we introduce double-grained prompt tuning (i.e., TaI-DPT) which involves: (i) a set of global prompts to generate embeddings for classifying whole sentences or images, and (ii) a set of local prompts to extract embeddings for discriminating text tokens or image patches. Given a set of text descriptions, global and local prompts can be tuned by minimizing the ranking loss . Note that, though these prompts are learned from text descriptions solely, they can be readily deployed to classify whole images as well as image patches during testing (see Fig. 1(c)). Experimental results show that, without using any labeled images, our TaI prompting surpasses zero-shot CLIP by a large margin on multiple benchmarks, e.g., MS-COCO, VOC2007, and NUS-WIDE.

Moreover, when images are also available during training, our TaI prompting can be combined with existing methods of prompting from images to improve its performance. In particular, given a few annotated images, our TaI-DPT can be integrated with CoOp as a prompt ensemble for improving classification accuracy. With partially labeled training data being provided, we may also combine TaI-DPT and DualCoOp to improve multi-label recognition accuracy consistently. Extensive results verify the effectiveness of our TaI-DPT, using in isolation or in combination, in comparison to state-of-the-arts.

To sum up, the contribution of this work include:

We propose Texts as Images in prompt tuning (i.e., TaI prompting) to adapt VL pre-trained models to multi-label image recognition. Text descriptions are easily accessible and, in contrast to images, their class labels can be directly derived, making our TaI prompting very compelling in practice.

We present double-grained prompt tuning (i.e. TaI-DPT) to extract both coarse-grained and fine-grained embeddings for enhancing multi-label image recognition. Experiments on multiple benchmarks show that TaI-DPT achieves comparable multi-label recognition accuracy against state-of-the-arts.

The prompts learned by TaI-DPT can be easily combined with existing methods of prompting from images in an off-the-shelf manner, further improving multi-label recognition performance.

Related Work

Multi-label image recognition aims to recognize all the object categories or concepts in an input image. To cope with multi-label images that are content-rich, various modules have been introduced to better represent the inter-class relationships and modern classification losses have been used to make model learning easier.

To model the label dependencies, CNN-RNN introduces recurrent neural networks, e.g., RNN and LSTM, to predict appeared classes in a sequential manner. use graph convolution modules to learn the correlation between class labels. CHAMP measures the severity of misclassification by building a domain-specific hierarchy tree according to the relation of categories, where each class are related to a tree node, to improve the robustness of the model. Albeit effective, these methods requires a considerable number of labeled images to let the models learn the category relationships sufficiently. While in data-limited or label-limited regimes, e.g., few-shot or partial-label data, it will be difficult for these models to learn well as expected. Specifically designed loss functions also struggle to obtain significant improvements when learning with limited data.

Multi-Label Recognition from Few-shot Samples. To better exploit the small number of samples, LaSO synthesizes samples by manipulates the features of paired training images. Different ways of manipulating label sets are used to train the model, resulting in generalizable discriminative features. introduces a meta-learning framework for better learning of past tasks and generalization to new tasks, and leverages the number of labels as useful information for learning.

Multi-Label Recognition from Partial-label Data. Partial-label refers to the scenarios where some labels are unknown. propose a normalized BCE loss to balance the proportion of known labels. learns to complement unknown labels by utilizing within-image and cross-image semantic correlations. blends the representation of training images and class proxies to compensate the loss of information due to unknown labels.

Albeit significant progress has been made, it remains a challenging issue for learning multi-label image recognition in image-limited or label-limited regimes. Built upon VL pre-trained models, this paper suggests to generate prompts from text descriptions instead of images, thereby offering a novel yet complementary perspective for handling low resource multi-label image recognition.

2 Prompt Tuning for Vision-Language Models

To transfer pre-trained knowledge to downstream tasks in data-limited settings, prompt tuning has become a popular parameter-efficient way to achieve the goal, due to its flexibility and ease of use. CoOp learns the prompts by using (a few) annotated images of each class from target dataset. CoCoOp further proposes to improve CoOp by formulating the prompts in an image-conditional way to maintain better generalization to unseen classes. To avoid overfitting, ProGrad leverages predictions from zero-shot CLIP to regularize gradients in prompt learning process. TPT suggests to optimize test-time prompts by promoting the consistency of augmented test images. ProDA uses multiple pieces of prompts to estimate the distribution of classifier weights for better handle of varying visual features. DualCoOp firstly adapts CLIP to multi-label image recognition with partially labeled data by learning pairs of positive and negative prompts for each class to ensure independent binary classification for each class.

Albeit existing prompt tuning approaches have achieved significant improvements in downstream tasks, images as well as a portion of class labels are prerequisite to supervise the optimization of the learnable prompts. In this paper, we propose to treat texts as images in prompt tuning, which, compared to labeled images, are much easier to collect with existing caption datasets and modern search engines. Our proposed TaI-DPT surpasses zero-shot CLIP by a large margin, and can be combined with the prompts learned by existing methods of prompting from images to further boost recognition performance.

Proposed Method

In this section, we present our proposed Text-as-Image prompting, i.e., TaI prompting, for adapting pre-trained VL models to multi-label image recognition. Our TaI prompting uses only easily-accessed free-form texts as training data to learn effective prompts for downstream multi-label recognition tasks. To begin with, We present an overview of TaI prompting in Sec. 3.1. Then, we introduce our preparation of training texts in Sec. 3.2. We further explain the design of the double-grained prompt tuning (i.e., TaI-DPT) and the training and testing procedure in Sec. 3.3, and provide the loss function used to train the model in Sec. 3.4. Finally, we combine TaI-DPT with the existing methods of prompting from images to improve multi-label recognition performance further. CLIP is used to introduce our methdod.

Fig. 2 illustrates the design of our proposed TaI-DPT framework, including the training and testing phases. During training, we learn prompts with only supervision from texts. Two identical copies of the text encoder EncT{\rm Enc_{T}} from the pre-trained CLIP are used to encode the prompts and text data, respectively. We introduce two sorts of trainable prompts (i.e., the global prompts and local prompts) to obtain global and regional class embeddings. A noun filtering strategy is used to generate classification pseudo-labels for each text description, which is applied to supervise the classification scores obtained by calculating the cosine similarity of class embeddings and text features. Only the parameters in prompts are optimized in the training phase, while the text encoders are both kept frozen. During testing, the class embeddings are obtained by encoding the two sets of learned prompts with the text encoder EncT{\rm Enc_{T}} as in training, while the other input source changes from text descriptions to test images. Pre-trained image encoder EncI{\rm Enc_{I}} from CLIP is used to extract global and dense features of each test image, then computing global and local classification scores with class embeddings generated by the global prompts and the local prompts via cosine similarity. The final classification result is obtained by fusing the global and local classification scores. In the following, we explain the details of the main components of our proposed method.

2 Preparation of Text Descriptions

To obtain sufficient category information from the language that helps in image recognition, we have to ensure that: 1) the collected text descriptions should contain rich contents that describe a relatively complete scene of an image, and 2) the contents of all text descriptions need to cover the category set of the target dataset so that the prompts can learn the discriminative features of each class well and thus obtain better recognition performance. With an aim of ensuring reproducibility, we use captions from public image caption datasets (e.g., MS-COCO ) and localized narratives from object detection datasets (e.g., OpenImages ) as our language data source, while avoiding the workloads associated with randomly crawling texts from the Internet in this paper. Note that although each caption is paired with a corresponding image and human-annotated labels, we only use the captions, and no information from the pictures and labels are disclosed during training.

For a target multi-label recognition dataset X\mathcal{X} that has a category set S\mathcal{S} = {s1\rm s_{1}, s2\rm s_{2}, s3\rm s_{3}, …, sC\rm s_{\textit{C}}}, where C denotes the number of categories and si\rm s_{i} denotes particular class name like “dog”, “plane”, etc., we search for sentences that contain at least one class name si\rm s_{i} in S\mathcal{S}. Since multiple words or phrases usually exist to represent the same meaning for each class, searching solely for exact match of category names in texts may lead to many false negatives in the obtained pseudo ground-truth labels, which is harmful to prompt tuning. Towards tackling this issue, we introduce a noun filter to map nouns with similar meanings into the corresponding class label. Specifically, we construct a synonym dictionary D\mathcal{D} by including common synonyms of each class name in the target dataset. If a word in a text description matches any synonym of a specific class name, it is considered to contain a description of that category. Several examples of synonyms are shown as follows:

More details of the synonym dictionary D\mathcal{D} are provided in the Suppl.

Then we conduct noun filtration by the following steps. First, for each text description, we use the tokenizer and lemmatizer from NLTK to recover the stem of each word in the sentences. Next, for all keywords in D\mathcal{D}, which contains all synonyms of the category set S\mathcal{S}, we search in our language data source for sentences that contains at least one class name. For the text descriptions that do not match any synonym of any class name, we simply drop it away to ensure each piece of data has at least one concerned label. Finally, for each retained text description, we convert the class names it contains into binary pseudo-ground-truth vectors by setting classes that appear as positive and other classes as negative, following the order of class labels in the target dataset X\mathcal{X}.

The word-level filtered labels may not be precisely correct since our searching strategy mentioned above is rather simple considering the diversity of free-form texts, where complex paraphrases and misspellings that widely exist in the corpus are not fully addressed. However, such a simple noun filtration can guarantee reproducibility of this work and already leads to satisfactory results of our TaI, as will be shown. And our experiments also demonstrate that this simple and efficient data preparation lead to practical prompt tuning and compelling multi-label recognition accuracy.

3 Text-as-Image for Dual-grained Prompt Tuning

where i∈{1,2,...,C}i\in\{1,2,...,C\} is the class index, si\boldsymbol{s}_{i} denotes word embedding of the ii-th class name si\rm s_{i}. For j∈{1,…,M}j\in\{1,\ldots,M\}, vj\boldsymbol{v}_{j} is a learnable word embedding whose dimension is the same as the dimension of normal word embeddings in the vocabulary. Just like in previous methods, e.g. CoOp , the prompts are learned by maximizing the probability of classifying each image into its ground-truth class:

where x\boldsymbol{x} denotes the image and ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle calculates the cosine similarity.

After large-scale pre-training with image-text contrastive loss, text features have been well-aligned to the image features of the same semantic meanings. Therefore, based on the aligned VL representation, we advocate considering the feature of a piece of text description that describes a specific category, as an alternative to an image feature. Given a piece of text description, optimizing the similarity between its feature representation produced by a VL model and some class embeddings is considered, for guiding the learning of prompts towards achieving categorical discriminative information.

Apart from using the global sentence representation (i.e., the coarsest-grained text feature), we find that the sequential feature of word tokens from CLIP also possesses rich fine-grained information which is very similar to the region feature of dense image feature. In CLIP , cosine similarity between global image features, obtained by visual attention pooling, and global text features, obtained by projecting the feature of the last token, are directly supervised with contrastive loss. In general, the global feature is sufficient for single-label classification because the target object usually is prominent in the picture. However, in multi-label recognition, the global feature is usually dominated by major objects, suppressing the recognition of non-significant objects concurrently existing in the image. Thus, it motivates us to explore fine-grained features and avoid the domination of the overly prominent object.

To achieve this goal, we propose double-grained prompt tuning (i.e., TaI-DPT) that uses two sets of prompts to handle global (i.e., the coarsest-grained level) and local (i.e., the fine-grained level) features, respectively, in two parallel branches. The global prompts achieve discrimination by learning from the global feature directly learned in CLIP, while the local prompt learns from localized features. Formally, the double-grained prompt is defined as follows:

where vj\boldsymbol{v}_{j} and vj′,j∈{1,…,M}\boldsymbol{v}_{j}^{\prime},j\in\{1,\ldots,M\} are learnable embeddings that are concatenated with word embedding si\boldsymbol{s}_{i} of the ii-th class to obtain the global prompt tiG\boldsymbol{t}^{G}_{i} and local prompt tiL\boldsymbol{t}^{L}_{i}, respectively. The sequences in Eq. (3) are fed to a copy of the text encoder EncT\rm Enc_{T} of CLIP to generate global and local class embeddings for each class, i.e. Gi=EncT(tiG)\boldsymbol{G}_{i}={\rm Enc_{T}}(\boldsymbol{t}^{G}_{i}) and Li=EncT(tiL)\boldsymbol{L}_{i}={\rm Enc_{T}}(\boldsymbol{t}^{L}_{i}), G={Gi}i=1C\boldsymbol{G}=\{\boldsymbol{G}_{i}\}^{C}_{i=1} and L={Li}i=1C\boldsymbol{L}=\{\boldsymbol{L}_{i}\}^{C}_{i=1} are encouraged to be correlated with global and local features, respectively. Note that the proposed double-grained prompts are different from dual prompts , which include a pair of contrastive positive and negative prompts for each class (More discussion about the differences between our method and DualCoOp is provided in the Suppl).

To preserve the fine-grained region features for the input image, we maintain the feature map before attention pooling layer of CLIP. As for the input text description, we preserve the sequential token features of the entire sentence instead of only the token features. So we have:

Then, the global and local similarities are computed by:

where u\boldsymbol{u} denotes either language feature h\boldsymbol{h} in training or visual feature f\boldsymbol{f} in testing, and U\boldsymbol{U} denotes H\boldsymbol{H} or F\boldsymbol{F} coordinately. Information in local branch P\boldsymbol{P} (visualized in Fig. 3 and Fig. 4) can be aggregated in a spatially weighted manner:

where τs\tau_{s} accommodates the extent of focusing on a specific location. pi\boldsymbol{p}_{i} and pi′\boldsymbol{p}_{i}^{\prime} are optimized by the loss terms Lglobal\mathcal{L}_{global} and Llocal\mathcal{L}_{local}, respectively, which we will discuss in Sec. 3.4. And in the testing phase, p\boldsymbol{p} and p′\boldsymbol{p}^{\prime} are combined to obtain the final classification score.

The visualization results in Fig. 3 and Fig. 4 show that the learned local class embedding L\boldsymbol{L} can focus on each specific location where corresponding class appears, both in text descriptions and images, even if the fine-grained visual and language features are not explicitly supervised in the training of CLIP.

4 Learning Objective

We briefly discuss the loss terms used during the training of TaI-DPT. The overall learning objective is defined as L=Lglobal+Llocal\mathcal{L}=\mathcal{L}_{global}+\mathcal{L}_{local}, where Lglobal\mathcal{L}_{global} and Llocal\mathcal{L}_{local} are loss terms for global text embedding and local text tokens, respectively. We adopt the ranking loss to measure the discrepancy between classification scores and ground-truth labels, instead of a commonly used binary cross-entropy loss. The binary cross-entropy loss is generally accompanied with a sigmoid function σ(x)=1/(1+exp⁡(−x))\boldsymbol{\sigma}(x)=1/(1+\exp(-x)) to convert model outputs to probabilities. Nevertheless, we observe that the value of cosine similarities between image and text CLIP features p\boldsymbol{p} are not evenly distributed on either side of 0. Directly constraining the probability σ(p)\boldsymbol{\sigma}(\boldsymbol{p}) makes the optimization more difficult in this case, and this is why we employ a different ranking loss function . There may exist other options, e.g., the asymmetric loss as in .

Specifically, Lglobal\mathcal{L}_{global} and Llocal\mathcal{L}_{local} are formulated as follows:

where p\boldsymbol{p} and p′\boldsymbol{p}^{\prime} are global and aggregated local similarities described in Sec. 3.3, mm is the margin controlling how much higher the similarity score with the positive classes is than with the negative classes. During training, we minimize the overall objective L\mathcal{L} with frozen text encoders, by optimizing the global and local prompts.

5 Incorporating with Prompting from Images

Though our TaI-DPT is very different from existing methods of prompting from images, it is also complementary to them. To show this, we utilize an off-the-shelf prompt ensemble strategy to combine our TaI-DPT with existing methods in this section. As illustrated in Fig. 5, using CoOp as an example, we can simply combine the scores of CoOp and that of our TaI-DPT in a weighted sum manner. In particular, our TaI-DPT can be integrated with CoOp when a few annotated images are provided and integrated with DualCoOp when partially labeled training data are available.

We ensemble prompts by fusing the predicted scores, rather than averaging the class embeddings generated by different prompts, since the image encoder used in different methods may be different (e.g. we conduct our experiments with ResNet50, while DualCoOp uses ResNet101 for partial-label prompting). So ensembling with the classification score is more convenient. In Sec. 4.3, we also empirically show that our prompt ensemble strategy is effective in advancing multi-label recognition performance in the few-shot and partially labeled settings.

Experiments

Architecture. In our experiments, we adopt CLIP ResNet-50 as the visual encoder, and use the corresponding CLIP Transformer as the text encoder. During training, the parameters of both the two encoders are kept frozen, and only learnable prompts are optimized.

Learnable Prompts. Our learnable prompts are shared among classes of all datasets. Class-specific prompting (i.e., an individual set of parameters for each category) has also been explored, but brings limited benefits. Hence, we adopt the shared prompts and initialize the value of each parameter with the Gaussian noise sampled from N(0,0.02)\mathcal{N}(0,0.02). In our experiments, the length of both the global prompts and local prompts are set to MM = 16, while a longer sequence brings trivial improvements.

Datasets. To evaluate our TaI-DPT, we conduct the experiments on VOC2007 , MS-COCO , and NUS-WIDE . VOC2007 contains 20 common categories, and following , we form the training/test set based on the official trainval/test split (5,011 images/4,952 images). MS-COCO includes 80 categories, and following the official split, we take 82,081 images to form the training set and 40,504 images to form the validation set. NUS-WIDE includes 81 concepts, which have certain inclusion relationships. We adopt its test set (107,859 images) to evaluate our method. For zero-shot experiments in Sec. 4.2, the training sets of the datasets are not used, and we use only text data to learn the prompts as mentioned in Sec. 3.2. Besides, for VOC2007 and MS-COCO, the language data sources are captions from MS-COCO. For NUS-WIDE, we introduce localized narratives from OpenImages , which have a broader range of content, to cover all the concepts in NUS-WIDE. In Sec. 4.3 and Sec. 4.4, for each dataset, the corresponding training data is used to conduct the experiments of partial-label and few-shot multi-label classification.

Training Details. We adopt SGD optimizer to learn our prompts, and the training epochs is set to 20 for all datasets. The learning rates for MS-COCO, VOC2007, and NUS-WIDE are empirically initialized with 1e-4, 1e-4, and 1e-3, and decay by the cosine annealing rule during training. For ranking loss, we choose m=1m=1, and scale the p\boldsymbol{p} and p′\boldsymbol{p}^{\prime} by a factor of 4. τs\tau_{s} is set as 0.02 via validation.

2 Comparison with Zero-Shot Methods

To demonstrate the effectiveness of our proposed TaI and DPT, we first compare it with the zero-shot CLIP (ZSCLIP). For fair comparison, we also introduce the DPT to ZSCLIP. Specifically, we adopt two identical default prompts “a photo of a [CLASS]” to separately deal with global and local features as DPT does.

Table 1 lists the comparison results on VOC2007 , MS-COCO , and NUS-WIDE datasets. From the table, our TaI prompting surpasses ZSCLIP by a large margin of 9.8%, 13.8%, and 8.5% mAP on VOC2007, MS-COCO, and NUS-WIDE, respectively, showing the effectiveness of our TaI. Furthermore, after training with fine-grained token features extracted from texts, our proposed DPT demonstrates a more powerful capability of discriminating local features than the default hand-crafted prompts and single global prompts.

3 Comparison with Few-Shot Methods

We further compare with multi-label few-shot learning methods to verify the effectiveness of our TaI-DPT. In contrast to the well-studied single-label few-shot classification problem, few works tackle the multi-label few-shot scenario. Existing methods often deploy models trained on seen classes to few-shot novel classes. In Table 3, we compare our zero-shot TaI-DPT to few-shot methods on 16 novel classes (we refer readers to for details about data split). Our TaI-DPT is comparable to the methods trained on 5-shot samples.

Besides, we consider a new multi-label few-shot setting where all the classes are regarded as novel classes. We select 1, 2, 4, 8, and 16-shot samples for each category following the strategy in . For fair comparison, we train CoOp and our TaI in the same settings, and we also extend them with DPT for a more comprehensive comparison. For CoOp-DPT, we set two sets of learnable prompts, to deal with global and local features, respectively. The results are illustrated in Fig. 7. One can see that, even without any image information regarding novel classes, our TaI can achieve comparable results to CoOp trained on 16-shot. Similar trends with the MS-COCO dataset and the DPT setting support our observation that the discriminative feature of text data can be used as images for prompting. Moreover, benefiting from the flexibility of prompts, we can easily integrate our TaI-DPT with CoOp-DPT by utilizing prompt ensembles. As illustrated in Fig. 7, though CoOp-DPT has achieved a high accuracy, combining our prompts learned with text data still brings further improvement on recognition performance. This also proves that texts and images are complementary to each other to some extent.

4 Integration with Partially Labeled Methods

Following , we conduct the experiments of multi-label recognition with partial-labeled images. We reproduce DualCoOp on partial-labeled VOC2007 and MS-COCO with the same experimental setting as reported (reproduced results are marked with *) and explore the enhancement brought by integration with TaI-DPT. The results are reported in Table 2. With no prior knowledge from pre-trained models, previous forefront method like SARB struggles to learn from incomplete labels. While DualCoOp achieves promising performance by prompting with images, TaI-DPT can still bring further improvements.

5 Ablation Study

To thoroughly investigate the effect of each component, we conduct a series of ablation studies on the quantity of texts, training loss, ensemble weight, and texts v.s. images for prompting. More details are shown in the Suppl.

Quantity of texts. Here, we mainly discuss the the effect of the number of text descriptions used in training on the performance of TaI-DPT on VOC2007. Following the data preparation procedure in Sec. 3.2, we end up with a total number of 66087 pieces of text that contain descriptions for 20 categories involved in VOC2007. We test the performance of TaI-DPT with different numbers of randomly selected texts, and the results are shown in Fig. 7. When no collected texts are available, 80 templates of hand-crafted prompts from , like “a cropped photo of a [CLASS]”, are used for training (all templates are shown in the Suppl), and each template sentence correlates with one positive label corresponding to the class name inserted in [CLASS]. The increasing number of texts gradually forms a complete description of target categories, and the relationship between classes is also better characterized, which results in ascending performance.

Conclusion

In this paper, we propose a new view of treating texts as images in prompt tuning (i.e. TaI), which learns the prompt from discriminative features of text descriptions. Compared to prior prompt tuning methods trained with images, our TaI benefits from the easy accessibility of scalable content-rich texts, which enables prompt tuning for vision tasks (e.g., multi-label image recognition) even without downstream image data. Double-grained prompting is further introduced to utilize both the global and fine-grained features for better multi-label recognition ability. Nonetheless, when few-shot image samples or partial-labeled images are available, our TaI-DPT can conveniently integrate with existing prompting methods. Experiments on MS-COCO, VOC2007, and NUS-WIDE show the validity of our proposed method.

References

Appendix A Appendix Overview

Here we provide more information of our TaI-DPT and experimental results. The appendix is organized as follows. In Appendix B, we present more details about our prepared text data used for training. In Appendix C, we display more ablation study on the training loss, texts v.s. images for prompting and the coefficients used in the prompt ensemble. In Sec. 2, we discuss the connection and distinction between our prompt design and existing methods.

Appendix B More Details about Text Descriptions

To extract the category labels from texts exhaustively, we construct synonym dictionaries for classes involved in VOC2007 , MS-COCO , and NUS-WIDE by gathering the expressions of the classes from different sources. We use the WordNet interface provided by to get a relatively comprehensive list of synonyms and then manually select words with specific meanings for inclusion in the synonym dictionary. In addition, we also collect expressions for categories from standard online dictionaries. Besides, some words exist in the corpus in simple and compound forms, like “cellphone” and “cell phone”, and we prioritize compound word matches. Since the 80 categories of MS-COCO cover the categories of VOC2012 , for these two datasets, we filtered the captions from MS-COCO using the same synonym dictionary (shown in “synonyms_COCO.txt”) to obtain the texts and labels as the training data. For NUS-WIDE , we introduce localized narratives from OpenImages , which have a broader range of content, to cover all the concepts in NUS-WIDE. The synonym dictionary for NUS-WIDE is shown in “synonyms_NUSWIDE.txt”.

B.2 Hand-craft Prompt Templates

Using the noun filtration strategy above, we end up with 66,087, 100,543, and 456,759 pieces of texts for VOC2007, MS-COCO and NUS-WIDE, respectively. Even for some common categories, the amount of texts is relatively sufficient, but we still find that there are few occurrences of certain categories in the texts. Especially for objects that are not prominent on which the text descriptions tended not to focus. So to process these categories better, we also added the hand-crafted prompt templates for each class as training data. The used templates are listed in “prompt_templates.txt”.

Appendix C More Ablation Studies

As explained in Sec. 3.4 of our main paper, we discussed the loss function used to train our TaI-DPT. Here, we provide the results on the three datasets when training with common binary cross-entropy loss (BCE), asymmetric loss (ASL) and ranking loss (RL) . Formally, the binary cross-entropy loss is defined as:

where p\boldsymbol{p} and p′\boldsymbol{p}^{\prime} are global and local classification score. And the asymmetric loss is defined as:

where qm=max(q−m,0)\boldsymbol{q}^{m}={\rm max}(\boldsymbol{q}-m,0) and hyperparameters γ+\gamma_{+}, γ−\gamma_{-} and mm are set as 1, 2 and 0.05, respectively, according to . The training results with different losses are shown in Table 4.

C.2 Texts v.s. Images for Prompting

To directly compare the difference between prompting with texts and prompting with images, we train our double-grained prompt with images (I-DPT) from trainval set and compare it with TaI-DPT on the test set of VOC2007 . The results are shown in Table 5. It’s obvious that we can learn the prompts well with sufficient labeled images, improving the mAP of zero-shot CLIP from 77.3 to 93.9. However, when no image data is available, our TaI-DPT can reach 88.3 mAP, demonstrating the effectiveness of our zero-shot prompt tuning scheme.

C.3 Summation Coefficient in Prompt Ensemble

As illustrated in Sec. 3.5 of our main paper, our TaI-DPT can easily combine with existing prompting methods learned with images and yield complementary improvements. Here, we explore the coefficient used to fuse the classification score produced by different models. For example, let p1\boldsymbol{p}_{1} denotes the score provided by CoOp and p2\boldsymbol{p}_{2} denotes the score yielded by our TaI-DPT. The merged score is obtained by weighted summation p=λ⋅p1+(1−λ)⋅p2\boldsymbol{p}=\lambda\cdot\boldsymbol{p}_{1}+(1-\lambda)\cdot\boldsymbol{p}_{2}.

From Fig. 8 we can see the change of mAP of p\boldsymbol{p} relative to coefficient λ\lambda. So we set λ=0.6\lambda=0.6 for the ensemble of TaI-DPT and CoOp-DPT learned from few-shot samples, which gives better results in various few-shot settings. Similarly, we set λ=0.9\lambda=0.9 when combining our TaI-DPT with DualCoOp when partially annotated images are available.

Appendix D Comparison with DualCoOp

As the first approach to adapt pre-trained CLIP to multi-label recognition tasks, DualCoOp proposes to use a pair of contrastive positive and negative prompts to generate binary classification probability for each class. However, the negative prompt may not be a property way to adapt CLIP. In Table 6 we show zero-shot recognition results of CLIP with hand-crafted positive and negative templates. We use a positive template, ”a photo of a [CLASS]” and a negative template, ”a photo without [CLASS]”. It seems that the negative prompt is dominated by the [CLASS] token and still gives rise to considerable recognition accuracy as the positive prompt does, which can make it reluctant to analyze the effect of a negative prompt.

But for our proposed double-grained prompt tuning (DPT), the two prompts are all positive and focus on global and local features separately. Intuitively, the global prompt can be seen as a hand-crafted prompt like “a photo of a [CLASS]”, and the local prompt can be seen as “a cropped photo of a [CLASS]”. The two positive prompts can be learned flexibly with ranking loss , without relying on each other to produce a classification score for each class.

Besides, DualCoOp uses all images from the training set with partial labels to learn the prompts. Our TaI-DPT advocates using descriptive texts as an alternative when there is no image data, and the pseudo-label for each text derived with noun filtration can be regarded as incomplete categorical labels. As such, our prepared text data is somewhat homogeneous with the partial-labeled image data, which leads to gentle improvements when combining our method with DualCoOp. However, in the case of few-shot image samples available, our TaI-DPT brings considerable enhancements by ensemble with the few-shot approach like CoOp as shown in Fig. 8.