CLIP-DIY: CLIP Dense Inference Yields Open-Vocabulary Semantic Segmentation For-Free
Monika Wysoczańska, Michaël Ramamonjisoa, Tomasz Trzciński, Oriane Siméoni
Introduction
The task of semantic segmentation, which aims at predicting the class of every pixel in an image has been widely tackled using fully-supervised approaches , which require tedious and therefore expensive per-pixel annotations. Moreover, semantic segmentation has typically been performed with a finite set of classes describing the types of objects that should be discovered in images. However, using a fixed number of classes is limiting for real-world applications as interesting object classes may vary in time and per application – having to perform annotation and re-training on new classes is expensive and sub-optimal. In this context, recent advances in Visual Language Models (VLMs) have paved the way to open-vocabulary perception. Indeed, VLMs, trained with cheap and widely available image-text pairs, e.g. captions, offer new possibilities to describe images with a large and open vocabulary. Using such models can help alleviate both the problem of supervision and finite vocabulary.
In particular, the popular VLM Open-CLIP has been exploited to perform open-vocabulary perception. While it has high performance on image classification, applying OpenCLIP on dense tasks is more challenging . In order to improve CLIP segmentation abilities, different methods have been proposed to modify the architecture , to add new modules or to train new specifically designed models from scratch. Instead, we propose CLIP-DIY, a new zero-shot open-vocabulary semantic segmentation approach which makes direct use of the high-performance image classification properties of CLIP, does not need architecture changes or additional training. In particular, our method applies CLIP to a multi-scale grid of patches and aggregates the information into a single prediction map.
Moreover, in order to further improve the quality of the localization of our predicted maps, we propose with CLIP-DIY to leverage the recent efforts in unsupervised object localization. This task aims at discovering every object depicted in images in a class-agnostic fashion and thus without manual annotation. Recent methods achieve impressive localization results by leveraging self-supervised features . While some methods discover one object per image , more recent ones try to highlight all objects in an image .
We therefore propose to leverage those capabilities in CLIP-DIY, by guiding CLIP predictions with a very lightweight unsupervised foreground/background strategy which greatly improves the predictions’ quality.
To summarize, our novel approach CLIP-DIY best leverages the open-world classification capabilities of CLIP and the high-quality of unsupervised object localization approaches yielding the following contributions:
We introduce CLIP-DIY, a novel, simple technique for open-vocabulary semantic segmentation which does not require additional training or any pixel-level annotation but instead leverages strong self-supervised features with good localization properties combined with CLIP.
Our multi-scale approach which uses simply CLIP as it was designed —for image classification— enables CLIP-DIY to produce well-localized predictions.
We demonstrate that unsupervised foreground/background methods can be effectively used to provide spatial guidance to CLIP predictions.
We achieve a new state-of-the-art zero-shot open-vocabulary semantic segmentation on PASCAL VOC dataset and perform on par with the best methods on COCO.
We perform an extensive validation of the design of our method, and show that it is robust as it can be directly applied to in-the-wild open-world segmentation.
Related work
In this section, we discuss previous work related to ours, starting in Sec. 2.1 with zero-shot open-vocabulary semantic methods, then following with unsupervised object localization methods in Sec. 2.2. Finally, in Section 2.3 we focus more specifically on works that leverage a combination of self-supervised learned features and CLIP to perform open-world segmentation.
With the aim to build generalizable models, zero-shot methods for semantic segmentation propose to extend models trained in a fully supervised fashion on a set of seen classes to new unseen classes. Many leverage relationships encoded in pre-trained word embeddings to discover new unseen concepts. Such methods require annotation for the seen classes while we aim to use no pixel-level annotation.
Alternatively, open-vocabulary approaches exploit image-text alignment without needing to pre-define vocabulary. Several build on top of the popularCLIP model which showed impressive global text-image alignment properties but lacks localization quality . Using class-agnostic object masks, it is possible to learn to align the embeddings of selected pixels with text , but at the cost of pixel-level annotations. Without extra supervision, MaskCLIP alters the last pooling layer of CLIP to produce dense predictions and use them as pseudo-labels to train a segmentation model, forming MaskCLIP+. Using only image captions—cheap to acquire and widely available—as supervision, learn local alignment between image regions and paired text with contrastive objectives. Regions are formed using a learnt hierarchical mechanism , cross-attention based clustering , using a clustering head trained with diverse view or slot-attention . PACL adds an embedder module that learns affinity between patches and the global text token, TCL builds a new local contrastive objective which directly aligns captions with pre-selected patches and ViewCO proposes a multi-view consistent learning approach. CLIPpy proposes to fully re-train CLIP with a few well-designed modifications to directly obtained denser features. Alternatively, ReCO builds prototypes of the desired vocabulary (using CLIP-based retrieval) which are then used for co-segmentation.
Rather than modifying the architecture of CLIP or training a new module specifically designed to densify its outputs, we propose to directly use as is the good classification ability of the model. Indeed, we perform a dense multi-scale patch classification. By doing so, our method can easily be adapted to any new dataset, with any new vocabulary.
2 Unsupervised object localization
Interestingly, recent works have shown that ViT features trained in a self-supervised fashion on images–with no human-made annotations–have good localization properties . Such properties have been exploited to tackle the problem of unsupervised object localization which requires to localize objects–any object–depicted in images, and such without any cue. A set of methods exploit the good correlation properties of the feature and find an object as the set of patches which highly differs to the other patches. Alternatively , exploits the attention mechanisms with different queries and produces maps that are ranked and filtered and proposes to look for the background instead of the objects in order to avoid single object discovery and to need priors about objects. The coarse object localization results obtained using those methods can be used as pseudo-labels to train large instance or segmentation models in a class-agnostic fashion . Recent FOUND is a very light model–a single conv1x1–trained to produce foreground/background segmentation of good quality. When self-trained FOUND achieves even better results and discovers more objects per image . In this work, we propose to leverage the good object localization properties of unsupervised object localization models, which make no hypotheses about object classes and remain therefore open. In particular, we take advantage of the efficiency of FOUND to guide our zero-shot segmentation.
3 Combining self-supervised features & CLIP
Combining self-supervised learning with VLMs has been previously explored by different open-vocabulary segmentation methods . Correlation qualities are used to perform co-segmentation when pre-training properties are directly leveraged to initialize the visual encoder backbone . Related to our work, ZGS builds potential masks using clustering strategies on self-supervised features and assigns them a class. Contrary to all previous approaches, it explores the task without predefined text prompts. Although ZGS currently obtains lower results than other baselines on the task, it opens an interesting direction for future work.
In this work, we do not perform co-segmentation (which expects to know classes of interest), nor retrain a model from scratch. Instead, we propose to guide CLIP prediction with an unsupervised foreground/background segmentation method, which to the best of our knowledge has not yet been explored.
CLIP-DIY
We tackle the problem of open-vocabulary semantic segmentation with no supervision. Let us consider a set of queries formulated in natural language. Our goal is to localize each query if present in the image, yielding one mask per query, i.e. a segmentation map. Our approach, summarized in Fig. 2, consists of two stages. In Sec. 3.1, we describe our first step, where we run our proposed multi-scale dense inference to obtain coarse semantic maps by running CLIP on image patches at different scales. Our second step, which we cover in Sec. 3.2, consists of refining the initial segmentation using an off-the-shelf foreground-background extractor.
Our method leverages CLIP as a backbone. Contrary to most CLIP-based approaches for zero-shot semantic segmentation (as discussed in Sec. 2.1), we do not rely on patch tokens within the image encoder. Instead, we leverage CLIP’s zero-shot capabilities by running the model densely on image partitions. Thus, we calculate the alignment between each of the multi-scale patches and the considered textual queries.
Given our multi-scale partitions, and a text query , we build a dense similarity map for each scale , such that:
where is a merging operator that puts patches back onto a 2D grid, denotes the inner product computed between the visual and text embeddings and is a bilinear up-sampling operator which upsamples its input to the resolution .
Having obtained similarity maps for each text query and each scale we aggregate the multi-scale predictions into a map such that:
2 Guided segmentation
Finally, we propose to refine the multi-scale segmentation maps of Eq. 2 using an objectness map produced by an off-the-shelf unsupervised foreground-background segmentation method , noted . As discussed in related work, such methods exploit self-supervised features, e.g. , to discover the pixels likely depicting objects.
such that similarities with the background class are down-weighted for pixels deemed as salient by .
Finally, we compute the output of CLIP-DIY as:
where denotes the Hadamard product and is the softmax operator computed over text queries. In Fig. 3 we show how the aggregation of all scales paired with the guidance of results in accurate object segmentation.
Experiments
In this section, we present the experiments conducted to evaluate our method and justify particular design choices. First, in Sec 4.1 we give details about our experimental setup. In Sec. 4.2, we discuss how our method compares against other open-vocabulary semantic segmentation approaches, both quantitatively and qualitatively. We then give more insight into our method with a series of ablations (Sec. 4.3), failure mode analysis (Sec. 4.4) and finally real open-world evaluation (Sec. 4.5).
We evaluate our method on two common semantic segmentation benchmarks: PASCAL VOC 2012 and COCO , comprising of 20 and 80 foreground classes respectively. PASCAL VOC has an additional background class, and we adopt a unified protocol considering a background class in all datasets. We evaluate results with the mean Intersection-over-Union (mIoU) metric. For evaluation, we resize input images to have the shorter side of length 448 following .
If not otherwise specified, we use the CLIP ViT-B/32 model OpenCLIP version trained with LAION . The input images are resized to and the patch size is . We empirically find that running our model on 3 different scales i.e. with patch sizes of gives the best results for both evaluated datasets. We discuss this later in Sec. 4.3.
We compare our method with existing state-of-the-art zero-shot open-vocabulary methods. In particular, we evaluate against methods including self-trained MaskCLIP+ , learning grouping strategies: GroupViT , SegCLIP , ViL-Seg , OVSegmentor , ViewCo , using class prototypes: ReCo† , text-grounding strategy: TCL and with CLIPpy which uses a T-5 backbone and improves dense abilities of CLIP. We do not compare against due to lack of comparable results and open source code. We detail in Tab. 1 the VLM backbones used per method and the if additional training data was. Every method (except for MaskCLIP) requires training a specific module/model used to get denser predictions; instead, we use the vanila CLIP model.
Following TCL , we use a unified evaluation protocol corresponding to an open-world scenario where prior access to the target data before evaluation is not allowed. In particular, we do not consider query expansion, e.g. class name expansion or rephrasing. As discussed in exploring language biases can greatly improve the overall segmentation performance. However, we only use the original class names from the compared datasets. We also report best-reported scores for all methods. It is to be noted that TCL uses a post-processing technique, namely PAMR while other methods do not.
2 Results
In this section, we compare our method to previous work both quantitatively and qualitatively.
We summarize in Tab. 1, the comparison of our CLIP-DIY to baselines. We report results of CLIP-DIY with both CLIP ViT-B/32 and CLIP ViT-B/16, given that most methods use the latter.
First, comparing our results with the two different backbones, we remark that better scores are obtained with CLIP ViT-B/32 which patch inputs are larger, and are therefore less expensive at inference time. We believe this could be explained by an existing upper bound on CLIP accuracy w.r.t the level of granularity; patches too small might induce noisy classification—as we observed in Fig. 3.
Compared to baselines, we observe that our method achieves the best mIoU result on PASCAL VOC dataset and outperforms all previous works by more than 4 mIoU pts, and such without post-processing. This result is particularly interesting given that our method does not require dedicated training to improve CLIP segmentation abilities but instead leverages the unsupervised object localization method.
Moreover, we obtain mIoU on COCO when the best-performing method CLIPpy achieves , so just 1 mIoU pt better than our approach. We can observe in Fig.4 that CLIPpy discovers more queries per image than us, even though its segmentation outputs appear to be noisier. Such results might suit a better COCO benchmark and less PASCAL VOC.
We qualitatively compare here our method with best performing TCL and CLIPpy in Fig. 4. Our method consistently produces better masks with better object boundaries, which we attribute to the high-quality saliency maps. Moreover, we observe that CLIP-DIY produces correct semantic results on the foreground objects, with fewer artefacts than CLIPpy. Our method seems also less sensitive to biases, for instance, both TCL and CLIPpy hallucinate the sheep class on the grass (on the very left image) and CLIPpy also predicts aeroplane in the sky (in most left image) and zebra next to the elephants (middle image in COCO dataset).
3 Ablations
In this section, we perform ablation studies to validate the individual choices in the design of CLIP-DIY.
First, we study in Tab. 2 the impact of the different elements of our method. In particular, we investigate the impact of using a multi-scale mechanism and leveraging the objectness produced by the foreground/background segmenter. We notice that by removing our multi-scale scheme we drop results by 3.9 and 5.1 mIou pts on Pascal and COCO respectively, showing the benefit of considering patches of different sizes. Additionally, the largest drop is observed when removing the foreground/background saliency guidance showing the effectiveness of combining CLIP with the current lightest unsupervised object localization model, FOUND.
We conduct an ablation study on the scales used in our multi-scale scheme to generate the predictions. We progressively add inner scales in Eq. 2 and report the resulting accuracy on all datasets. The results, detailed in Tab. 3 show an optimum when using three fine scales. Adding more does not appear to improve results nor downgrade them, showing the stability of our method. Interestingly, we also see that removing information from the global scale greatly reduces performance on all datasets (-5.7/-3.5 mIoU pts on PASCAL VOC/COCO respectively). In Fig. 5 we visualize the individual contribution of each scale to our final prediction. As in Fig. 3, we observe that coarse scales capture the global context, while finer scales capture more local one, such that objects can be separated.
We compare different foreground-background segmenters and the overall performance of our method using each one of them. The results, summarized in Tab. 4, show that our method performs the best when using FOUND, more specifically the FOUND model that has been re-trained with self-training in the original work. We also experiment with CutLER , which performs unsupervised instance segmentation. We use the predicted instance masks or compute a saliency. In both cases, we obtain slightly worse results. We give more details in Sec. 2.2 of Supplementary material.
4 Failure cases
We qualitatively analyse failure cases of our method by showing a couple of examples in Fig. 6. We observe that some of the failures are due to inaccurate annotations: in (a) bear is only partially annotated and in (b) a mask for an elephant is annotated too coarsely. Our method, benefiting from FOUND’s accurate saliency predictions is able to produce better segmentation masks. We also observe that our method is limited by the quality of the saliency (c, d), which we comment more on below.
In the examples (c) and (d) of Fig. 6, we observe that our method can fail to segment objects with significant overlap, resulting in ambiguity regarding the foreground class, which is especially harmful when annotations are too coarse or incomplete. In (c), only the bench is annotated, while our method segments only the orange, which would result in a low IoU score. In (d), the opposite behaviour occurs, where CLIP-DIY discovers more classes than the annotated ones. Finally, similarly to most of the open-vocabulary methods based on CLIP, our method suffers from sensitivity to text ambiguities. We show more failure cases in Sec. A.3.3 of the Supplementary material.
5 Our method in the wild
We also test our method in the wild. We randomly download a set of images and provide textual queries we find most suitable. We show in Fig. 1 that our method exhibits off-the-shelf open-world segmentation capabilities, being able to produce masks for specific prompts. More results, including comparisons against other methods of in-the-wild open-world segmentation, are presented in Sec. A.3.2 of the Supplementary material.
Conclusions
We introduce a new method for open-vocabulary semantic segmentation, namely CLIP-DIY, which exploits CLIP’s open-vocabulary classification abilities. As opposed to recent approaches we run CLIP densely at multiple scales to obtain coarse semantic mask proposals. When further guided by the quick fully unsupervised object localization method FOUND, which estimates foreground saliency, our model obtains state-of-the-art results on PASCAL VOC and performs on par with baselines on COCO dataset. Since our method does not require any specific training, it can be used as an off-the-shelf method for open-world segmentation, and could therefore serve as a tool to help dataset annotators. While CLIP-DIY already yields competitive results, we believe future work could make it more diverse and efficient.
Acknowledgments
We would like to thank Georgy Ponimatkin for the interesting discussions. This work was supported by the National Centre of Science (Poland) Grant No.2022/45/B/ST6/02817 and by the grant from NVIDIA providing one RTX A5000 24GB used for this project.
References
Appendix A Supplementary material
In this Supplementary material we consider broader impact in Sec. A.1, give more details on different foreground-background segmenters used in this work and their adaptations in Sec. A.2. We then show more qualitative results in Sec. A.3, including comparisons with other methods on the datasets used in the evaluation as well as the examples in the wild. We conclude with an in-depth analysis of the failure case of our method.
Semantic segmentation plays a crucial role across a wide range of fields, including healthcare, medicine, self-driving cars, and many more. While this technology can foster many applications with a positive impact, there still exists a risk of negative misuse. Additionally, since CLIP-DIY builds on foundation models that were trained on large-scale data, our method is not free from biases present in the datasets. Overall CLIP-DIY has a broad range of applications. Being training-free, CLIP-DIY can especially serve as an off-the-shelf image annotator in computing or budget-limited environments.
A.2 Foreground-background segmenters
In this section, we provide more details on the foreground-background segmenters we use in our work. We consider the two variants of FOUND : we first use the coarse saliency maps produced without self-training, noted FOUND-bkg in the paper, which corresponds to the set of similar pixels to the least salient background seed pixel in the self-supervised feature space. We also use the quick conv1x1 model, named FOUND in the main paper, which is self-trained on only 10,553 images and produces more pixel-aligned results. We refer the reader to the original paper for more details.
Second, for a fair comparison, we also use CutLER off-the-shelf as a mask extractor. We use previously described masks and with each one of them, we create an image mask where the background is masked out. Each masked image is then fed separately to CLIP to obtain a CLIP prediction. We denote this approach as CutLER mask in Tab. 4 of the main paper.
Overall, CLIP-DIY achieves the best performances with the light self-trained FOUND as discussed in the main paper.
A.3 More qualitative results
We provide in this section more qualitative examples produced with CLIP-DIY. We first compare our method against other state-of-the-art approaches in Fig. 7 for PASCAL VOC and Fig. 8 for COCO.
In Fig. 9, we then present more in-the-wild examples and conclude this section by discussing failure cases and limitations of our method in detail.
We first present comparisons with other methods on the two segmentation datasets used in this work, namely PASCAL VOC and COCO Object .
Fig. 7 shows randomly sampled images from PASCAL VOC dataset and the results of CLIP-DIY and our baselines. Our method produces accurate masks for all of the images and the result of CLIP-DIY is the closest to the ground truth compared to other methods. We observe that the two other methods, TCL and CLIPpy , produce masks that are too coarse, with the latter frequently even assigning most of the image to one segment.
Fig. 8 shows the examples from COCO dataset. While generating masks with mostly the correct category, TCL produces very noisy boundaries compared to CLIP-DIY. CLIPpy not only generates noisy masks but also produces a lot of clutter, assigning often wrong labels to background pixels.
A.3.2 In-the-wild examples
We provide more in-the-wild examples to showcase the open-vocabulary abilities of our method. In Fig. 9 we present a couple of randomly mined images from the Web in comparison with TCL . Both of the methods correctly assign queries to proper segments, even very specific types of objects, such as traditional dishes e.g. polish dumplings and pasteis de nata; monuments Eiffel tower and Sacré coeur. Moreover, thanks to CLIP backbone both methods can distinguish between different colours, e.g. grey elephant against pink elephant. However, we observe that the quality of masks produced by TCL again is not as detailed as ours. Note that TCL uses PAMR post-processing technique thus we would expect the generated masks to be more precise.
A.3.3 Failure cases
We analyze failure cases of our method in Fig. 10. We can see that CLIP-DIY suffers from producing incomplete masks column (a) and missing objects (c). This happens due to the saliency produced by the foreground-background segmenter, which in the case of complex, multi-object scenes focuses on certain aspects of a scene. Moreover, our method has limited performance in the case of overlapping objects, such as dog and chair in column (d). Finally, we find more failures due to inaccurate annotations, such as the one in column (b), where bowl is misclassified by what is inside, i.e. carrot.
A.3.4 Detailed quantitative results
We present detailed quantitative results on PASCAL VOC in Tab. 5 for each class. We observe that the worst performance is obtained on classes which are typically only partially visible in images, such as furniture (chair, sofa and table). This is mostly due to the decreased performance of the saliency detector in those classes, which is biased towards object-centric images. The high performance (87.6 IoU) for the background class confirms the efficacy of our saliency detector.