DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, Jiwen Lu
Introduction
The “pre-training + fine-tuning” paradigm is recognized as one of the key discoveries that has largely pushed the state-of-the-art for various downstream computer vision tasks, including image classification , object detection , semantic segmentation , and action recognition . Due to the high annotation and computation cost of the per-pixel prediction, pre-training is even more critical for dense prediction tasks. As illustrated in Figure 1 (a), the pre-training step is usually accomplished via supervised classification or self-supervised learning of the backbone model on large-scale datasets like ImageNet . Then, a task-specific module like a detector or a segmentation decoder is added to the backbone and the whole model is fine-tuned on the target dataset with less training data .
Different from conventional supervised and self-supervised pre-training methods only based on images, Contrastive Language-Image Pre-training (CLIP) is a new framework to learn high-quality visual representation by exploring contrastive learning with large-scale noisy image-text pairs. By exploiting the semantic relationships between the images and the associated texts, this new framework benefits from rich and semantic level supervision from texts while enjoying a broader and cheaper source of data. Thanks to the language supervision, models pre-trained via CLIP achieve impressive results on various visual classification tasks with no or very limited annotations .
Very recently, several efforts have been made to adopt the prompt engineering from NLP community to better transfer the CLIP models to the downstream visual classification tasks. Several learning-based prompting methods are proposed to modify the output of the language model to better adapt to the new tasks. However, they mainly focus on transferring the CLIP model to classification tasks by performing image-text matching, which is much close to the original pre-training task. The problem of transferring the knowledge learning from image-text pairs to more complex dense prediction tasks and a more generic setting has barely been visited.
In this paper, we study how to fine-tune the pre-trained CLIP models to dense prediction tasks. Compared to conventional ImageNet pre-trained models, one distinct challenge is the gap between the upstream contrastive pre-training task and the downstream per-pixel prediction task, where the former involves instance-level representation of both images and texts, and the latter is only based on the visual information at the pixel level. To tackle this problem, we present a new language-guided dense prediction framework named DenseCLIP. As shown in Figure 1 (b), it is designed for various Dense prediction tasks by implicitly and explicitly leveraging the pre-trained knowledge from CLIP models. An implicit way to exploit the pre-trained knowledge is to directly fine-tune the models on the downstream datasets. Our results show that the CLIP models can outperform the conventional ImageNet pre-trained models with some modifications on hyper-parameters (see the CLIP result in Figure 2). But the straightforward way cannot fully exploit the potential of the CLIP models. Inspired by the original contrastive learning framework in CLIP, we propose to convert the original image-text matching problem in CLIP to a pixel-text matching problem and use the pixel-text score maps to guide the learning of dense prediction models explicitly. By further using the contextual information from the image to prompt the language model with a Transformer module, we are able to facilitate our model to better exploit the pre-trained knowledge by optimizing the text embeddings.
Our method can be a plug-and-play module to improve the fine-tuning of CLIP pre-trained models on off-the-shelf dense prediction methods and tasks. By applying our method to the popular semantic segmentation framework semantic FPN on the challenging ADE20K dataset, we exhibit +4.9%, +4.7% and +2.3% mIoU improvement compared over ImageNet pre-trained models and +3.9%, +2.4% and +1.2% mIoU improvement compared to vanilla fine-tuning of a CLIP models based on ResNet-50, ResNet-101 and ViT-B respectively. We also observe significant improvements in object detection and instance segmentation tasks. Notably, we show a ResNet-101 model equipped with our method and a lightweight semantic FPN decoder can achieve 46.5% mIoU on ADE20K, which outperforms state-of-the-art solutions like DeepLabV3+ and UperNet with only 1/3 computation.
Moreover, our framework can also be applied to any backbone models by using the pre-trained language model to guide the training of dense prediction tasks. We observe significant improvements by applying DenseCLIP to ImageNet pre-trained ResNets and recent Swin Transformers with slight computation overhead. We expect our method to be a new and generic paradigm to improve dense prediction models with guidance from pre-trained language models.
Related Work
Pre-training and fine-tuning. The revolution of computer vision in the past decade has been driven by the “pre-training + fine-tuning” paradigm. Specifically, it first pre-trains models on large-scale datasets (e.g., ImageNet , JFT , Kinetics , etc.) in a supervised learning or self-supervised learning manner , and then fine-tunes the models on various downstream tasks. In NLP community, this framework has also been similarly and widely used and recently evolves into a prompt paradigm , in which downstream tasks are reformulated to simulate the solved tasks in original pretraining process. Inspired by these works, we explore to transfer the knowledge in large-scale vision-language pre-trained models to the downstream dense prediction tasks.
Vision-language models. There have been a series of works on the interaction of computer vision and natural language processing fields, e.g., text-to-image retrieval , image caption , visual question answering , referring segmentation and so on. Among these works, vision-language pre-training has attracted growing attention during the past few years . As a milestone, Radford et al. devise a large-scale pretraining model, named CLIP , which employs a contrastive learning strategy on a huge amount of image-text pairs, and shows impressive transferable ability over 30 classification datasets. Motivated by this work, a number of follow-ups have been proposed to improve the training strategy (e.g., CoOp , CLIP-Adapter , Tip-adapter ) or apply it to other domains (e.g., ActionCLIP ). However, there are very few attempts on performing dense prediction tasks via the CLIP model. The work most related to ours is CPT , which reformulates dense predictions into a fill-in-the-blank problem by jointly marking co-referential parts of both image and text in color. Differently, we consider a standard dense prediction setting in this paper, where we use pixel-text relationships to guide the training of dense prediction models and further optimize language embedding with image context with a context-aware prompting method.
Dense prediction. Compared with conventional instance-level classification problem, dense prediction tasks (e.g., semantic segmentation , instance segmentation , object detection ) are more challenging as they requires to model the finer-grained representation at the pixel level or region level. Following the “pre-training + fine-tuning” paradigm, previous literatures have developed various dense prediction models like FCN , PSPNet , FPN , UperNet , and many others. To alleviate the heavy annotation cost in previous supervised pre-training settings, a number of self-supervised pre-training approaches have been proposed for dense prediction . Orthogonal to these prior arts, we introduce a new fine-tuning strategy that leverages the knowledge in the large-scale vision-language pre-trained model and uses the language information to guide the learning process.
Approach
We begin by reviewing the Contrastive Language-Image Pre-training (CLIP) framework to illustrate the motivation of our method. CLIP consists of two encoders, including an image encoder (ResNet or ViT ) and a text encoder (Transformer ). The goal of CLIP is to align the embedding spaces of visual and language during pre-training through a contrastive objective.
To learn more transferable pre-trained knowledge, CLIP collects 400 million image-text pairs for model training. To transfer knowledge of CLIP to downstream classification task, a simple yet effective way is to construct a set of text prompts based on a template such as “a photo of a [CLS].”, where [CLS] can be replaced by the actual class names. Then given an image, one can use CLIP to compute the similarities between the image and the text prompts in the embedding space and the class with the highest score is regarded as the final prediction. Recently, several works have shown CLIP can obtain strong classification performance with few examples. Therefore, it raises an interesting question: whether the impressive ability of CLIP can be transferred to more complex vision tasks like dense prediction?
However, the extension is nontrivial. Firstly, how to leverage the visual-language pre-trained model in dense prediction tasks is a barely visited question. Although a simple solution is to only use the image encoder like a pre-trained 2D backbone, we argue that the language priors contained in the text encoder are also of great importance. Secondly, unlike the classification considered in , transferring the knowledge from CLIP to dense prediction is more difficult due to the substantial gap between the upstream contrastive pre-training task and the downstream per-pixel prediction task, where the former considers instance-level representation of both images and texts, and the latter is only based on the visual information but expects pixel-level outputs.
2 Language-Guided Dense Prediction
In the standard training process of CLIP, the global feature is used as the output of the image encoder while the other outputs are usually neglected. However, we find has two interesting properties: (1) still retains sufficient spatial information thus can serve as a feature map. (2) since the MHSA is symmetric to each input element, might behave similarly to , which aligns well with the language features. Based on the above observations, we can use as a language-compatible feature map. It is also noted that for architectures like ViT , can be obtained similarly by excluding the class token of outputs.
3 Context-Aware Prompting
Previous efforts have already proved that mitigating the domain gaps in visual or language can significantly improve the performance of CLIP models on downstream tasks. Therefore, instead of using the vanilla human pre-defined templates, we seek for other methods to improve the text features .
Language-domain prompting. Different from the original CLIP that uses human-designed templates like “a photo of a [CLS].” as text prompts, CoOp introduces learnable textual contexts to achieve better transferability in downstream classification tasks by directly optimizing the contexts using back-propagation. Inspired by CoOp , we also use learnable textual contexts in our framework as a baseline, which only includes language-domain prompting. The input of the text encoder then becomes:
Vision-to-language prompting. Including descriptions of visual contexts can make the text more accurate. For example, “a photo of a cat in the grass.” is more accurate than “a photo of a cat.”. Therefore, we investigate how to use visual contexts to refine the text features. Generally, we can use the cross-attention mechanism in Transformer decoder to model the interactions between vision and language.
We propose two different strategies of context-aware prompting, which is shown in Figure 4. The first strategy we consider is the pre-language-model prompting, or pre-model prompting for short. We pass the features to a Transformer decoder to encode visual contexts:
Another choice is to refine the text features after the text encoder, namely post-model prompting. In this variant, we use CoOp to generate text features and directly use them as the queries of the Transformer decoder:
This implementation encourage the text features to find most related visual clues. We then update the text features through a residual connection:
Although the two variants target the same goal, we prefer the post-model prompting for mainly two reasons: (1) The post-model prompting is efficient. The pre-model prompting requires extra forward passes of the text encoder during inference since its input is dependent on the image. In the case of post-model prompting, we can store the extracted text features after training and thus can reduce the overhead brought by the text encoder during inference. (2) Our empirical results show the post-model prompting can achieve better performance than pre-model prompting.
4 Instantiations
where is a temperature coefficient following and is the ground truth label. The auxiliary segmentation loss can help the feature map to recover its locality faster, which is beneficial to dense prediction tasks for both segmentation and detection.
Applications to any backbone models. Another interesting usage of our framework is that we can replace the image encoder of CLIP with any backbones (e.g., ImageNet pre-trained models and self-supervised models). Although there might be no strong relation between the outputs of the visual backbone and the text encoder, the backbone can learn better and faster with language guidance. In other words, we can leverage the language priors from the pre-trained text encoder to improve the performance of any pre-trained image backbone, which makes DenseCLIP a more generic framework to improve dense prediction with the natural language priors learned from large-scale pre-training.
Experiments
To evaluate the effectiveness of our DenseCLIP, we conduct extensive experiments on dense prediction tasks including semantic segmentation, object detection and instance segmentation. The following subsections describe the details of the experiments, results and analyses.
Setups. We start by evaluating our DenseCLIP on ADE20K , a challenging large-scale semantic segmentation dataset that covers a broad range of 150 categories. ADE20K contains 20K images for training and 2K images for validation. Following common practice , we report the mIoU on the validation set. For fair comparisons, we also include the FLOPs and the number of parameters.
Implementation details. We experiment with the popular Semantic FPN framework to evaluate our DenseCLIP. Specifically, we apply the pre-trained image encoder of the CLIP as the segmentation backbone, and directly use the Semantic FPN as the decoder. We consider three kinds of image backbones including ResNet-50 , ResNet-101 , and ViT-B . For language-domain prompting, we use a context length of 8. The Transformer decoder to extract visual contexts consists of 6 layers and we set the number of heads as 4. We fix the text encoder during training to preserve the natural language knowledge learned from large-scale pre-training. To reduce the computational costs, we project both the image embeddings and the text embeddings to a lower dim (256) before the Transformer module. We empirically find that directly fine-tuning CLIP models to dense prediction with the default training strategies in will lead to unsatisfactory results (only 21.9% mIoU on ADE20K, which is 15.6% lower than its ImageNet pre-trained counterpart). Therefore, two key modifications are made compared to the default configurations: (1) we use AdamW instead of the default SGD inspired by recent progress in vision Transformers ; (2) to better preserve the pre-trained weights, we set the learning rate of the image encoder as of the other parameters. We also adopt the above training strategies to our baselines in ablation studies for fair comparisons (+1.1% mIoU over the ImageNet pre-trained ResNet-50 with the default settings in ).
Main results. We report the semantic segmentation results of our DenseCLIP with three different backbones on ADE20K in Table 1. We include the FLOPs, the number of parameters, and the mIoU in both single-scale (SS) and multi-scale (MS) testings. The experiments results show that for the same backbone, our DenseCLIP with a simple Semantic FPN can outperform the state-of-the-art methods that use more sophisticated decoders by large margins. Unlike previous works that use dilated backbones (ResNet-D8 ), the ResNet encoder in DenseCLIP is more close to standard ResNet thus our DenseCLIP has much fewer FLOPs. Besides, our DenseCLIP is +4.9%, +4.7%, and +2.3% mIoU (SS) higher than the original ImageNet pre-trained baselines on ResNet-50, ResNet-101 and ViT-B backbones with acceptable extra computation cost. DenseCLIP is also +3.9%, +2.4%, and +1.2% mIoU higher than the vanilla fine-tuning strategy (CLIP + Semantic FPN).
Ablation studies. To further demonstrate the effects of different components of our DenseCLIP, we perform detailed ablation studies with the ResNet-50 backbone and the results are shown in Table 2. Firstly, we show by adopting a better training strategy aforementioned the ResNet-50 baseline we implemented has a higher mIoU than (38.6% vs. 37.5%). Secondly, we find that CLIP pre-trained ResNet-50 outperforms the ImageNet pre-trained one by 1%, which indicates that large-scale vision language pre-trained model can be better transferred to downstream vision tasks. To better leverage the language priors, we adopt our language-guided with language-domain prompt and witness a significant performance boost (+2.5% mIoU). Finally, we compare the two methods to perform vision-language prompting to incorporate visual contexts. We find both the pre-model and post-model prompting can improve the performance, while the post-model prompting is better and more computationally efficient. Therefore, we choose the post-model prompting as the default configuration in all the rest experiments.
Effects of language-guided pre-training and fine-tuning. We compare the performance on ADE20K of different pre-training and fine-tuning strategies to better reveal the potential of language-guided paradigm, which is shown in Figure 2. We consider supervised pre-training on ImageNet1K and ImageNet21K , self-supervised pre-training via MoCoV2 and DenseCL , and the vision-language pre-training. We show that the vision-language pre-trained model (CLIP) can outperform ImageNet1K pre-trained model by vanilla fine-tuning. Furthermore, through the language-guided fine-tuning with context-aware prompting, our DenseCLIP surpasses even the ImageNet21K pre-trained model. These promising results demonstrate that language-priors can largely facilitate vision models in downstream dense prediction tasks.
2 Object Detection and Instance Segmentation
Setups. We also conduct experiments to apply our DenseCLIP to object detection and instance segmentation tasks on COCO , which contains 118K training images and 5K validation images. We adopt two widely used frameworks, RetinaNet and Mask R-CNN . Following , we report the standard AP, AP at IoU=0.5/0.75, and cross-scale AP. For Mask R-CNN, we report both the mAPs for object detection and instance segmentation since these two tasks are performed simultaneously.
Results analysis. The results using the RetinaNet and the Mask R-CNN are summarized in Table 3 and Table 4, respectively. For object detection with RetinaNet, we compare DenseCLIP with ImageNet1K pretrained model and vanilla CLIP fine-tuning on detection task. One can observe that DenseCLIP outperforms the ImageNet1K pretrained model by +1.5% and +2.6% AP. Meanwhile, it also improves the vanilla fine-tuning strategy by +0.9% and +0.6% AP on both ResNet-50 and ResNet-101 backbones.
For Mask R-CNN, we observe that DenseCLIP achieves consistent improvement on both object detection and instance segmentation tasks within an affordable computational budget. Especially for instance segmentation, our DenseCLIP outperforms the ImageNet1K pre-trained model with +2.9% and +2.5% mask AP on both ResNet50 and ResNet101 backbones and also outperforms the vanilla fine-tuning strategy with +0.8% and +0.7% mask AP. The significant improvements of DenseCLIP on the instance segmentation task suggest that our pixel-text matching is conceptually suitable for segmentation.
3 DenseCLIP for Any Visual Backbone
Previous experiments have demonstrated the effectiveness of our DenseCLIP framework. However, since DenseCLIP is specifically designed to leverage the visual-language relation contained in the pre-trained CLIP models, the generalization ability of DenseCLIP might be somehow doubted: Is DenseCLIP only suitable to CLIP image encoders? To answer this question, we perform experiments to verify whether our DenseCLIP can also perform well with other backbones. The extension is actually straightforward: we can simply replace the CLIP image encoder with any given 2D pre-trained image model. Although there are no strong correlations between the feature maps of the new backbone and the text features output by the CLIP text encoder, we hypothesize that if we preserve the language priors by freezing the text encoder as before, the text encoder will guide the backbone to better adapt to downstream tasks.
To verify the above assumption, we choose two representative 2D models including ResNet , the most widely used CNN model, and Swin , the recent state-of-the-art vision Transformer. Following the standard setting in and , we use the Semantic FPN framework for ResNet models and the UperNet framework for Swin models. The experimental results are summarized in Table 5, where we report the mIoU on ADE20K of both the single-scale and multi-scale testing. We demonstrate that our DenseCLIP can consistently improve all the baseline models notably. Specifically, DenseCLIP can bring single-scale mIoU improvement for ResNet-50/101 with semantic FPN , and improvement for Swin-T/S with UperNet . These results clearly show that our DenseCLIP can successfully guide any pre-trained 2D backbone by language priors to boost performance. Since the text encoder can be removed after training, our method provides a low-cost solution to improve arbitrary dense prediction models. Although these performances still lag behind our models with CLIP image encoders, the findings in this section provide a solution to generalize human knowledge learned from large-scale vision-language pre-training to a wider range of models. We expect this could be an interesting direction to connect vision and language researches in the future.
4 Visualization
To better demonstrate the superiority of DenseCLIP, we provide several qualitative results in Figure 5. We compare the segmentation maps of our method and the baselines and find DenseCLIP is better at identifying holistic objects.
Conclusion and Discussion
In this paper, we have presented a new framework, DenseCLIP, to transfer the knowledge from the vision-language pre-trained model (CLIP) to the downstream dense prediction tasks. DenseCLIP is a model-agnostic framework to use the pre-trained vision-language knowledge with the context-aware prompting strategy. The framework can be applied to various dense prediction tasks including semantic segmentation, object detection, and instance segmentation. We conducted extensive experiments to demonstrate the superior performance of our method.
Limitations & societal impact. Although our method has achieved substantial improvement in segmentation, we find the improvements on detection are not such significant. We conjecture that it is because the pre-trained CLIP image encoder lacks locality since there is no such constraint during the pre-training of CLIP while object-centered tasks can only provide less dense supervision. We believe DenseCLIP can be further improved by introducing the dense supervision during pre-training or better recovering the locality after pre-training. We develop a general method for dense prediction in this paper. Since our method is not for a specific application, it does not directly involve societal issues.
This work was supported in part by the National Natural Science Foundation of China under Grant 62125603, Grant U1813218, and in part by a grant from the Beijing Academy of Artificial Intelligence (BAAI).
References
Appendix: More Analysis
We provide more analyses of both the design of our model and the training strategies in detail in the section.
Effects of learning rate multipliers. As discussed in Section 4.1, we found that the optimal learning rate for CLIP models and conventional ImageNet pre-trained models are different. Here we further investigate the effects of learning rate multiplier for image encoder and text encoder in Table 6. We see both fixing the text encoder and using a lower learning rate for image encoder is beneficial to train the dense prediction model. Note that we observe a much lower performance (30% mIoU) when directly fine-tuning CLIP models with learning rate for the image encoder, which suggests our language guided method can largely stabilize the training process and make the final results less sensitive to the learning rate configuration.
Effects of optimization of the textual contexts. Previous works on transferring CLIP models to downstream classification tasks have clearly shown the importance of adapting the textual contexts for different datasets and tasks. We show the effects of optimizing the textual contexts compared to the original prompting strategy proposed in in Table 7. We see that although the learnable contexts will introduce additional computation during training (gradient computation for the text encoder), this strategy can bring notable improvement over the baseline. Therefore, we choose to add the learnable textual contexts for our models.
Effects of . Table 8 shows the effects of . We see a learnable initialized with small values can improve the final performance.