DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

Ximeng Sun, Ping Hu, Kate Saenko

Introduction

Image recognition has become a very popular and successful research area in recent years, due to the development of large-scale datasets and advanced model architectures . However, the majority of image recognition approaches have focused on single-label prediction, which ignores the intrinsic multi-label nature of images. Unlike single-label recognition , multi-label image recognition aims to recognize all semantic labels present in an image , providing a more comprehensive understanding and benefiting applications like image retrieval, video analysis, and recommendation systems.

Multi-label recognition typically deals with images of complex scenes and diverse objects. Collecting multi-label annotations becomes difficult to scale up, for two reasons: (i) annotating images with the full semantic label set is laborious and (ii) samples of particular categories can be hard to find. The first challenge can be addressed by multi-label recognition with partial labels, where merely some of the categories are annotated for each training image. Recent works proposed solutions to partial-label MLR based on semi-supervised learning , normalized training objectives , or label correlations . The second setting involves zero-shot MLR, where novel unseen categories are recognized by transferring knowledge from seen categories, with solutions like principal image features , knowledge graphs , and attention mechanisms . Despite significant progress on the two settings, existing approaches are not designed to handle both. We propose to unify these settings as limited-annotation MLR and design a solution that can handle practical scenarios with either partial or missing labels.

Successful solutions to the above problems transfer knowledge from fully-annotated categories to partially-labeled and novel categories by learning an alignment between images and category names . Recently, vision-language pretraining models are bridging the visual-textual gap via large-scale pretraining, e.g., CLIP is trained with 400 million image-text pairs. In this work, we draw inspiration from the recent success of prompt learning for such models . Prompt learning provides a convenient way to transfer pretrained vision-language models to other tasks. It designs additional templated or learnable prompt tokens for textual input to “inform” the model about downstream tasks and avoids finetuning the entire model which can be inefficient and data-hungry. By doing so, recent works like CoOp have demonstrated CLIP’s remarkable generalisation to various zero-shot image tasks . However, these methods mainly focus on matching each image with a single label, hence they are not able to handle the multi-label setting.

To adapt the knowledge learned in CLIP to multi-label image recognition, we propose the DualCoOp framework. As shown in Fig. 1 (c), DualCoOp learns a pair of differentiable prompts to provide positive and negative contexts for the target class. Instead of using hand-crafted thresholding to determine positive labels , the dual prompts naturally result in a positive and a negative classifier, so the existence of the target class in the image can be easily decided by comparing their scores. Unlike prior models, shown in Fig. 1 (a)(b), we avoid fine-tuning the full vision-language model and only learn the prompts, which are much smaller compared to the entire model. Therefore, our simple framework achieves much higher efficiency when adapting to different datasets. Additionally, we modify the attention mechanism of CLIP to better model spatial information in images, improving its ability to recognize multiple objects in MLR. With these design choices, we achieve a unified framework for addressing the general challenges of multi-label recognition with limited annotations.

We summarize our contributions as follows:

We propose DualCoOp to quickly adapt powerful vision-language models to solve multi-label recognition tasks using limited annotations.

We propose dual (positive and negative) prompts to drastically reduce the number of learnable parameters, and improve the spatial modeling of the visual encoder to better distinguish multiple objects.

We conduct extensive experiments on partial-label MLR (on MS-COCO and VOC2007 ) and zero-shot MLR (on MS-COCO and NUS-WIDE ). Notably, DualCoOp improves mAP by 6.8%6.8\% with 10%10\% of labels on VOC2007, and F1-score at Top-3 Prediction by 10.8%10.8\% for zero-shot MLR on NUS-WIDE.

Related Works

Multi-Label Recognition with Limited Annotations. Multi-label image recognition has drawn increasing attention in past years. One straightforward solution to this problem is to individually learn a binary classifier for each category , which however does not consider correlations among labels. Hence, recent works have focused on incorporating semantic dependencies among labels via graph neural networks or RNN/LSTM . Some work also considers the spatial distribution of labels in the image, and exploits object proposals or attention mechanism as a regularization to rectify the prediction. However, despite achieving significant progress, these methods require a large-scale and complete annotated dataset to train models . This limits their application to more practical scenarios where data is partially annotated for training and unseen (zero-shot) categories may appear during testing .

With partially labeled data, where merely some labels of each sample are known, Mahajan et al. and Joulin et al. attempt to use web supervision to automatically generate the pseudo labels, which unfortunately leads to poor performance as the web supervision is noisy and incomplete . To avoid external noise, Durand et al. exploit the proportion of annotated samples for different labels and propose a normalized BCE loss to train models based on the given partial labels. More recent works explicitly transfer information from known labels to complement unknown labels by utilizing category-specific feature blending or label co-occurrences at both instance-level and prototype-level.

Unlike partial annotation of the same label set for training and testing, zero-shot multi-label image recognition needs to handle novel categories during testing, hence inspiring a different route based on a joint visual-label embedding space . Zhang et al. propose to find a principal direction that ranks related labels first in the joint embedding space optimized via a tailored zero-shot ranking loss. Cohen et al. further improve the idea by learning multiple principal vectors to support the semantic diversity. Huynh et al. consider the spatial regularization and propose a shared multi-attention model and obviate the need for explicit region proposals . Narayan et al. propose to enhance the region-based features so as to minimize inter-class feature entanglement.

Though significant progress has been made in each of the directions, existing methods still require a lot of MLR data and complex architectures/losses. Our approach reduces the need for hard-to-get MLR data by pretraining on unsupervised text-image pairs. While it may seem unfair to compare existing MLR methods with ones based on such pretraining, we point out that the pretraining data is unsupervised and thus easier to obtain. We also provide experiments comparing DualCoOp to baselines using the same pretraining. Importantly, previous methods are designed for only one task, hence have limitations in practical applications. In contrast, our proposed framework can be easily adapted with small data and can address both partial and zero-shot tasks at the same time.

Prompt Learning for Vision-Language Models. Vision-Language Models based on contrastive learning have demonstrated impressive ability to learn generic visual representations. As a milestone, CLIP is trained with 400 million curated image-text pairs, and shows remarkable transfer capability for over 30 classification datasets. With such powerful vision-language models, several follow-ups have been proposed to explore the training strategies for training downstream classification tasks. Instead of fine-tuning the entire model , which may damage the learned representation space, recent approaches adopt the prompt-based paradigm that formalizes NLP tasks as masked language modeling (prompt templates) . Zhou et al. propose to tune prompts for downstream classification tasks, and further introduce input-conditional prompts for better generalization ability . Lu et al. learn the distribution of diverse prompts to handle the varying visual representations. Huang et al. generate pseudo labels for images to learn prompts in an unsupervised way. Though achieving promising improvements for downstream tasks, these methods address the multi-class zero-shot image recognition, assuming each image has one label, hence lacking the ability to handle the multi-label setting. In this paper, we present a novel framework to efficiently transfer VLMs to address multi-label image recognition with limited annotations.

Method

Problem Definition. We formally define multi-label recognition with limited annotations as follows: Consider MM as the set of categories which describe objects or attributes in images. Given a training image II, the existence of a category m∈Mm\in M can be positive, negative or unknown, corresponding to the label ym=1,−1y_{m}=1,-1 or respectively. During inference, we predict each label of interest for an input image.

Many existing MLR problems fit into this broad definition. In this paper, we consider the settings with partial or missing labels: (1) Partial-label MLR , in which only a subset of labels are known (+1+1 or −1-1) for one training image and we are interested in predicting all existing labels during inference. (2) Zero-shot MLR , in which each label is either known (seen) or unknown (unseen) for all images during training and we are interested in predicting either all labels or only unknown (unseen) labels during inference. In this paper, we propose a unified framework to address the limited-annotation MLR in both scenarios.

Approach Overview. To compensate for insufficient or missing image labels, it is important to learn how the meanings of category names are related to each other, so we can transfer knowledge between related categories. This is usually done by learning an alignment between the visual and textual spaces. However, our dataset is too limited to learn a broad and generalizable mapping. We propose DualCoOp to instead leverage the strong alignment of visual and textual feature spaces learned by large-scale vision-language pretraining (CLIP ) with a light-weight learnable overhead which quickly adapts to the MLR task with limited semantic annotations. Figure 2 provides an overview of our proposed approach. DualCoOp learns a pair of “prompt” contexts in the form of two learnable sequences of word vectors, to provide positive and negative contextual surroundings of a given category name mm. This generates positive and negative textual features (Ftm)+(F_{t}^{m})_{+} and (Ftm)−(F_{t}^{m})_{-} that are fed into the pretrained text encoder. Furthermore, to better recognize multiple objects, which can be located at different locations in the image, the spatial aggregation step is modified. We first compute the similarity score of each projected visual feature FviF_{v}^{i} at location ii with (Ftm)+(F_{t}^{m})_{+}/(Ftm)−(F_{t}^{m})_{-} to obtain prediction logits over regions. For each class, we perform aggregation of all spatial logits, in which the weight for each logit is determined by its relative magnitude. We call this Class-Specific Region Feature Aggregation. During training, we optimize the learnable prompts via the ASL loss while keeping all other network components frozen. During inference, we directly compare the final positive and negative logits to make a prediction for each label ymy_{m}.

Dual Learnable Prompts. Instead of learning a single prompt for a class , we propose Dual Context Optimization (DualCoOp) which learns two contrastive prompts’ contexts for each class. The learnable part in dual prompts carries positive and negative contextual surroundings individually and can be optimized end-to-end from data via binary classification loss. Specifically, we define the pair of prompts given to the text encoder as follows:

where each VV is a learnable word embedding vector (e.g. with dimension 512 in CLIP ) and CLS is the given category name. N+N^{+} and N−N^{-} are the numbers of word tokens learned in the positive and negative prompts respectively. For simplicity, we set N+=N−N^{+}=N^{-} in our experiments. We learn a pair of positive and negative prompts for each class (i.e. class-specific prompt pair) when solving MLR with partial labels, and learn a pair of prompts shared for all classes in zero-shot MLR. With a pair of prompts, we compute the binary classification output pp with the following form:

where <⋅,⋅><\cdot,\cdot> represents cosine similarity and pp is the predicted probability for a given (image, label) pair as positive example. Ev(⋅)E_{v}(\cdot) and Et(⋅)E_{t}(\cdot) are the visual and textual encoders from the vision-language pretraining. A(⋅)A(\cdot) is our new aggregation function to adaptively reduce the spatial dimension of visual features for each class, which will be discussed next.

Class-Specific Region Feature Aggregation. In multi-label image recognition, it is common that multiple objects appear in different regions of the image. Pooling to produce a single image-level feature vector for all classes gives sub-optimal performance since spatial information is reduced and different objects are mixed. In this work, we reformulate the last multi-headed attention pooling layer of the visual encoder in CLIP and apply class-specific pooling to adaptively aggregate region features in the multi-label setting. The original attention pooling layer in CLIP pools the visual feature map first, and then projects the global feature vector into text space as follows:

where qq, vv and kk are independent linear embedding layers and x=Ev(I)x=E_{v}(I) is the output feature map of the visual encoder. By removing the pooling operation, we can project the visual feature xix_{i} of each region ii to the textual space :

For each region ii and each class mm, we compute cosine similarity between FviF_{v}^{i} and (Ftm)+=Et(Prompt⁡+)(F_{t}^{m})^{+}=E_{t}(\operatorname*{\text{Prompt}}^{+}) as Si,m+=<Fvi,(Ftm)+>S_{i,m}^{+}=<F_{v}^{i},(F_{t}^{m})^{+}>, and compute Si,m−S_{i,m}^{-} in the same way. In order to make a single prediction for the whole image, we aggregate Si,m+S_{i,m}^{+} and Si,m−S_{i,m}^{-} into Sm+S_{m}^{+} and Sm−S_{m}^{-} according to the magnitude of Si,m+S_{i,m}^{+}, i.e.:

Notably, we do not introduce any new parameters in our re-formulation of the spatial aggregation function. All parameters used to project visual features to the textual space are inherited from the original multi-headed attention pooling layer in CLIP.

Optimization. We apply the Asymmetric Loss (ASL) to handle the inherent positive-negative imbalance in the optimization of multi-label recognition. Specially, we compute losses for a positive (image, label) pair L+\mathcal{L}_{+} and a negative (image, label) pair L−\mathcal{L}_{-} as follows:

where pc=max⁡(p−c,0)p_{c}=\max(p-c,0) is the probability for negative examples shifted by hard thresholding via the margin cc. We set the hyper-parameters γ−≥γ+\gamma_{-}\geq\gamma_{+}, so that ASL down-weighs and hard-thresholds easy negative samples. The pair of learnable prompts are updated by back-propagating ASL through the frozen text encoder.

Experiments

Datasets. We conduct experiments on MS-COCO and VOC2007 to evaluate multi-label recognition with partial labels. MS-COCO contains 80 common object categories and we use the official train2014 (82K images) and val2014 (40K images) splits for training and test. VOC2007 contains 20 object categories and we use the official trainval (5K images) and test (5K images) splits for training and test. To create the training set with partial labels, we randomly mask out labels from the fully annotated training setThe difference in performance is within 1.0%1.0\% of independent runs. and use the remaining labels for training by following standard practice . In this work, we vary the proportion of kept labels from 10%10\% to 90%90\% .

Evaluation. On MS-COCO and VOC2007 datasets, we follow to report the mean average precision (mAP) for each proportion of labels available for optimization (from 10%10\% to 90%90\%) and its average value for all proportions. We count the learnable parameters (#P) of each baseline and DualCoOp to measure the complexity of optimizationFor baselines without public released implementation, we only measure the major part of the learnable parameters based on description in their papers. (indicated as #P ≥\geq [a value] in Table 1-3). We also report the per-class and the average overall precision (CP and OP), recall (CR and OR), and F1 (CF1 and OF1) of DualCoOp under different proportions of labels for training in the supplementary material due to the page limit.

Implementation. We adopt ResNet-101 as the visual encoder in all baselines and DualCoOp for input resolution 448×\times448, and use the same Transformer in CLIP as the text encoder. The visual and text encoders are initialized from the CLIP pretrained model and kept frozen during optimization. For each class/label, we learn two independent context vectors with 16 context tokens (N = 16) following , which is the only learnable part in DualCoOp. We use the SGD optimizer with an initial rate of 0.002 which is decayed by the cosine annealing rule. We train context vectors for 50 epochs with a batch-size 32/8 for MS-COCO/VOC2007, respectively. For ASL loss, we choose γ+=1\gamma_{+}=1, γ−=2\gamma_{-}=2 and c=0.05c=0.05 via validation. The training is done with one RTX A6000.

Baselines. To evaluate the effectiveness of DualCoOp, we compare with the following baselines: (1).SSGRL , GCN-ML and KGGR adopt graph neural networks to model label dependencies. We follow to report their performance in the partial-label setting. (2). Curriculum labeling and SST generate pseudo labels for unknown labels. (3). Partial BCE uses a normalized BCE loss to better exploit partial labels. (4).SARB blends category-specific representation across different images to transfer information of known labels to complement unknown labels.

Results. Table 1 shows the comparison of mAP between DualCoOp and all baselines optimized with 10%10\% to 90%90\% of labels. For the two most recent works (SST and SARB ), we further substitute the ImageNet pretrained weights with the CLIP pretrained weights when initializing of their visual encoders, which results in SST∗\text{SST}^{*} and SARB∗\text{SARB}^{*} in Table 1. Since we learn class-specific prompts, DualCoOp on MS-COCO adopts more learnable parameters than VOC2007. Our proposed DualCoOp achieves the best performance across all proportions of labels available during the training with the smallest learnable overhead (1.3M vs. 29.6M in SARB∗\text{SARB}^{*} on MS-COCO and 0.3M vs. 29.6M in SARB∗\text{SARB}^{*} on VOC2007). Notably, DualCoOp yields a great improvement over the second-best method, 3.2%3.2\% on MS-COCO and 6.8%6.8\% on VOC2007, especially when only providing 10%10\% of labels during the training. This indicates that DualCoOp can quickly adapt to the multi-label recognition task with a few labels by taking advantage of the powerful vision-language pretraining.

2 Zero-shot Multi-Label Recognition

Datasets. Following , we conduct experiments on MS-COCO and NUS-WIDE to perform zero-shot multi-label recognition. On MS-COCO, we follow to split the dataset into 48 seen classes and 17 unseen classes. NUS-WIDE dataset includes 270K images. Following we use 81 human-annotated categories as unseen classes and an additional set of 925 labels obtained from Flickr tags as seen classes.

Evaluation. We follow and report precision, recall, and F1 score at Top-3 predictions in each image on MS-COCO. We also follow to report mAP over all categories as well as precision, recall, and F1 score at Top-3 and Top-5 predictions in each image on NUS-WIDE. We evaluate all methods with both zero-shot setting (test only on unseen classes) and generalized zero-shot setting (test on both seen and unseen classes).

Implementation. We adopt ResNet-50 similar to as the visual encoder in DualCoOp for input resolution 224. Instead of learning class-specific prompts, we learn the class-agnostic context vectors with 64 context tokens (N = 64) for all classes, which is the only learnable part in DualCoOp. We optimize context vectors for 50 epochs with a batch-size 32/192 for MS-COCO/NUS-WIDE, respectively. Other implementation details are the same with Sec. 4.1

Baselines. To evaluate the effectiveness of DualCoOp in the zero-shot setting, we compare with the following baselines: (1). CONSE adopts an ensemble of classifiers for unseen classes. (2). LabelEM learns a joint image-label embedding. (3). Fast0Tag and SDL estimate one or multiple diverse principal directions of the input images. (4). Deep0Tag and LESA estimate the relevant regions via region proposals and attention techniques respectively. (5). BiAM enhances the region-based features to minimize inter-class feature entanglement.

Results. Table 2-3 shows the comparison between DualCoOp and all SOTA methods of zero-shot learning and generalized zero-shot learning on MS-COCO and NUS-WIDE datasets. DualCoOp achieves the best F1 score in all cases with a very light learnable overhead (0.02M) and improves the performance of zero-shot learning (unseen labels) with a significant margin: F1 score improves by 12.5 @Top-3 on MS-COCO, and by 10.8 @Top-3 and 10.9 @Top-5 on NUS-WIDE. This shows the power of exploiting the pretrained alignment of textual and visual spaces in CLIP via DualCoOp to solve multi-label recognition.

3 Ablation Studies

Effectiveness of Text Supervision. To show the effectiveness of text supervision from label space, we compare the model learned with discrete label space (“Discrete Label”) with three methods (SST , SARB ] and DualCoOp) which introduce the textual space to utilize the contextual correlation of labels in Table 4. We find that methods with text supervision usually perform better than the method only using discrete labels. However, when the semantic annotations are limited, text supervision sometimes yields worse performance (e.g. mAP of SST is 1.5%1.5\% lower than Discrete Labels with only 10%10\% of labels). By adopting the well-pretrained visual-textual alignment, DualCoOp achieves a great performance (e.g. 7.8%7.8\% higher than Discrete Labels with 10%10\% of labels) and quickly adapts to the dataset even with limited labels.

Ablation of Prompt Design. We compare our proposed dual learnable prompts with two hand-crafted prompts and one prompt learning method on the MS-COCO dataset with the zero-shot setting (see Table 5). Hand-crafted prompts can use either contextless class names or manually designed prompt templates. In our experiments, we carefully choose the positive and negative prompt templates as “a photo of a [classname]” and “a photo without a [classname]”. In contrast with performing the binary classification for each class with dual learnable prompts as the input, we also experiment with learning a single prompt of positive or negative contexts and use a chosen threshold (0.5 in our experiment) to make the prediction for each class. As we can see, the single positive prompt learning method (M2) performs better than non-learnable methods (M0 and M1), and a single negative learnable prompt (M3) achieves much worse accuracy than its positive counterpart (M2). However, when we include both positive and negative prompts, dual prompts (M4) performs even better than a single prompt, which indicates that DualCoOp learns complementary and beneficial information in the dual prompt pair. To keep the same amount of learnable parameters as in single prompt settings, we also halve the token size (M5), and find that DualCoOp still outperforms two single prompts in M2 and M3 by large gaps, demonstrating the effectiveness of our dual-prompt design.

Multi-Headed Attention vs. Class-Specific Region Aggregation. In Table 6, we compare the adaptive ability of these two visual aggregation methods when training/testing with a larger resolution (see Table 6), which is crucial in multi-label recognition as spatial details matter. For a fair comparison, we only replace the class-specific region aggregation in DualCoOp with the original multi-headed attention layer in CLIP at the end of the visual encoder. We adaptively resize the input feature map to match the input dimension of the multi-headed attention layer.

As shown in Table 6, multi-headed attention is bonded to the pre-training image resolution (224 in CLIP), while our class-specific region aggregation benefits from the increased input resolution either during training or in inference. Our class-specific feature aggregation uses original weights, but actually performs better than finetuning the original multi-headed attention layer.

Ablation of Aggregation Function. We experiment with different functions to aggregate the regional logits for each class in Fig. 3. We compute final logits in three ways: (1) taking the average of logits at all spatial locations (“Ave”), (2) taking the region with the largest positive logit (“Max”), and (3) generating aggregating weights for all spatial locations via a softmax function over the positive logits (“Ours”). “Max” performs better than “Ave”, which indicates the regional feature is more informative than the global feature in multi-label recognition. Furthermore, by taking account of both the regional and the global features, “Ours” gives the best performance.

Conclusion

In this paper, we propose a unified framework, DualCoOp, for two types of multi-label recognition with limited annotations. It utilizes the powerful vision-language pretraining from a large-scale dataset. By introducing a lightweight learnable overhead, it can quickly adapt to solve multi-label recognition after receiving a small amount of labels. In DualCoOp, we learn a pair of positive and negative prompts followed by the target class name as the linguistic input. Furthermore, to better aggregate visual region features for each class, we reformulate the original visual attention in the pretraining model as a class-specific region feature aggregation. We conduct extensive experiments for both partial-label MLR and Zero-Shot MLR across MS-COCO, VOC2007, and NUS-WIDE datasets showing the efficacy of our proposed approach over state-of-the-art methods.

Limitations. Since the vision-language pretraining adopts a large Transformer-based language model and all labels need to be feed-forward through the text encoder, the large language model limits the size of the label set. Also, compared to training the model with both seen and unseen labels, we still get worse performance for the zero-shot unseen classes even though we have used 400M auxiliary samples in the pretraining. This highlights the difficulty of zero-shot MLR.

Negative Societal Impacts. Negative impacts of our research are difficult to predict, however, it shares many of the pitfalls associated with deep learning models. These include susceptibility to adversarial attacks and data poisoning, dataset bias, and lack of interpretability. Other risks associated with the deployment of computer vision systems include privacy violations when images are captured without consent, or used to track individuals for profit, or increased automation resulting in job losses. While we believe that these issues should be mitigated, they are beyond the scope of this paper. Furthermore, we should be cautious of the result of failures of the system which could impact the performance/user experience of the high-level AI systems based on our research.

References

Appendix A Different Prompt Length

We have provided the comparison of the performance of DualCoOp with different lengths of prompt context (i.e. N=2,4,6,8,16,32,64N=2,4,6,8,16,32,64) in all three different experiment scenarios (see Fig. 4 and 5). In MLR with partial labels, we learn class-specific prompts and thus DualCoOp performs good when NN is small, such as 8, 16. For zero-shot learning in MLR, we learn uniform prompts shared by all classes and it requires larger NN (e.g. 32 or 64) for good performance. In the main paper, we use N=16N=16 for all experiments of MLR with partial labels and use N=32N=32 for experiments in zero-shot learning.

Appendix B Performance on the Full Dataset

We also finetune the visual backbone on the full multi-labeled recognition dataset MS-COCO and achieve mAP 85.8%85.8\% with ResNet 101 and input resolution 448, comparing to 85.0%85.0\% achieved by the same setting using ASL .

Appendix C Full performance of MLR with Partial Labels

In this section, we provide the average per-class and average overall precisions (CP and OP), recalls (CR and oR) and F1 scores (CF1 and OF1) of DualCoOp in the experiment of MLR with Partial Labels on MS-COCO and VOC2007 (see Table 7 and 8 in supplementary material) as a supplementary for Table 1 in the main paper.

Appendix D Visualization of Class-Specific Region Feature Aggregation

We have visualized the class-specific region feature aggregation on MS-COCO dataset (in Fig. 6). We can see DualCoOp generates the high attention score at the correct objects.