CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
Monika Wysoczańska, Oriane Siméoni, Michaël Ramamonjisoa, Andrei Bursuc, Tomasz Trzciński, Patrick Pérez
Introduction
Semantic segmentation is a key visual perception task for many real-world systems, e.g., self-driving cars, and industrial robots. Typically tackled in a dataset-oriented manner, best methods require a training dataset which is manually annotated for a specific and finite set of classes. The advent of powerful Vision-Language Models (VLM) is stimulating a shift from a closed-vocabulary paradigm to an open-world one. Such models are trained with a simple but scalable objective: to align pairs of image and coarse text captions that can be obtained in large amounts with limited manual supervision. VLMs excel at associating global image content with arbitrary text inputs with remarkable generalization capabilities , but struggle to provide dense open-vocabulary features . Obtaining such an alignment between pixels and language can lead to open-vocabulary extensions for multiple other modalities, such as point clouds , 3D scenes , 3D shapes , radiance fields , inter-modality alignment , with multiple potential applications for which the construction of training datasets is even more challenging and where CLIP-derived models are showing promising results.
Different strategies have been recently proposed towards improving CLIP’s patch-level feature extraction abilities by modifying the original CLIP architecture for dense pooling and retraining or finetuning on an annotated segmentation dataset with pre-defined classes . The former requires long training and/or large collections of annotated data, while the latter leads to an alteration of the vision-language associations of the CLIP features. An alternative line of approaches freezes the CLIP encoder and directly densifies its features with different heuristics often with multiple forward passes , but are less practical due to the extensive computational overhead. MaskCLIP arises as a computationally efficient dense CLIP extractor. It converts CLIP’s global self-attention layer into a convolutional one to produce patch features with original vision-language qualities. If such features are local, they appear to be too noisy for high-quality segmentation mask extraction (see Fig. 4).
Meanwhile, recent self-supervised learning (SSL) approaches produce strong visual representations displaying object localization properties, and such without requiring any manual annotation. DINO stands out with its visual concept-aware features which have been exploited for unsupervised object discovery . DINO features prove useful also for zero-shot semantic segmentation , but require expensive sliding window sampling or building concept-specific prototypes and ensemble strategies .
In this work, we aim for unaltered patch-level CLIP features with minimal runtime overhead. To this end, we re-examine the localization properties of MaskCLIP features and observe that it is possible to easily refine them with guidance from SSL models. In detail, we train a simple convolutional layer to produce pooling weights to perform concept-aware dense feature pooling from CLIP without distorting the vision-language association. This layer is optimized to mimic the patch correlations of DINO that indicate likely layouts of visual concepts in the images. Furthermore, we show that the unsupervised objectness information given by FOUND from DINO features, can be also directly learnt from CLIP features with a single convolutional layer and help improve the segmentation of the ill-defined ‘background’ prompt. With CLIP-DINOiser, we are able to obtain high-quality masks with a single forward pass on CLIP. CLIP-DINOiser is amenable to producing dense semantic maps or object-focused ones.
To summarize, our contributions are: (1) We propose a light pooling mechanism to refine MaskCLIP features leveraging guidance from SSL features without degrading its original open-vocabulary properties. CLIP-DINOiser does not require any annotations, nor retraining CLIP from scratch, only a single CLIP forward pass. (2) We show that CLIP already contains good localization properties which can be exploited. We leverage simple convolutional layers to emphasize visual concept layouts from dense CLIP features. We believe that this finding could be further exploited in different contexts. (3) Our method achieves state-of-the-art results on complex semantic segmentation datasets such as COCO , Pascal Context , Cityscapes and ADE20K .
Related Work
In this section, we discuss different methods related to ours.
Zero-shot semantic segmentation. This task has been typically approached by methods which aim at generalizing from seen classes to unseen ones . Such strategies train models with full supervision on the set of seen classes and propose different solutions to extend them to unseen ones without new images (labelled or unlabeled), e.g., by exploiting class information and relationships encapsulated in popular word embeddings . While they produce fine segmentations without computational overhead, these methods require pixel-level annotations for the seen classes.
The surge of VLMs with aligned image-language representations brought back into the spotlight the zero-shot classification task. However, the extension to zero-shot segmentation is not obvious as the CLIP architecture is not equipped to yield dense vision-language features . To produce dense CLIP features, several approaches fine-tune or train from scratch pixel-aligned CLIP-like models with additional modules, mechanisms or supervision objectives on datasets with annotations of varying granularity and quality: dense annotations , class-agnostic object masks , coarse captions or pseudo-labels . Recent works leverage image-level captions to align text to regions (obtained without supervision): PACL trains an embedder module to learn patch-to-text affinity, TCL proposes a local constrative objective to align well-selected patches to the text and ViewCO leverages multi-view consistency. On the downside, such models require long training on millions of images or specific types of annotations that are highly costly. Also, fine-tuning CLIP with a defined vocabulary is more computationally appealing , but alters the open-vocabulary properties of the features .
Most related to us is a line of works that investigate how to directly densify CLIP features in order to obtain per patch features. Such densification can be performed by aggregating features from multiple views or from sliding windows , however at the extra-cost of multiple forward passes. MaskCLIP drops the global pooling layer of CLIP and matches the projected features directly to text via a convolution layer. By doing so they achieve dense predictions, however noisy.
With a concept-driven perspective, some methods build codebooks of visual prototypes per concept, including negative prototypes , and then perform co-segmentation. This gets good results, however at the cost of building expensive class-specific prototypes therefore diverging from open-vocabulary scenarios. Instead, we aim to remain open to avoid retraining a model or building new expensive prototypes whenever a new concept is considered. To this effect, we devise a dense feature extraction method of CLIP that preserves the open-vocabulary quality.
Leveraging self-supervised models & CLIP.
Recent self-supervised ViTs allow us to produce features with good localization properties . Such features have also been exploited in the context of open-vocabulary segmentation methods: pre-training for the visual backbone backbone , co-segmentation , clustering patches into masks , representing object prototypes . Related to us is the recent CLIP-DIY which computes patch-level representations from CLIP features from different image crops with guidance from an unsupervised saliency segmenter . We also leverage the FOUND segmenter but require a single pass of CLIP and mitigate the limits of FOUND in cluttered scenarios by integrating an uncertainty constraint. Finally, we leverage the informative patch correlation properties of DINO and show that it is possible to teach CLIP to produce DINO-like features through light convolutional layers.
Method
We present in this section CLIP-DINOiser, a simple and efficient strategy to improve MaskCLIP using localization information extracted from CLIP—with a lightweight model trained to mimic some of DINO’s properties. An overview of the method is presented in Fig. 5. We first set the goal in Sec. 3.1. We present in Sec. 3.2 how dense text-aligned 2D feature maps can be generated with CLIP following MaskCLIP . We then introduce our strategies which leverage self-supervised features localization information to consolidate MaskCLIP features in Sec. 3.3 and a way to improve the ‘background’ in Sec. 3.4. We finally show in Sec. 3.5 that such information can be successfully learnt directly from CLIP, allowing our method to run in a single pass of CLIP with no external backbone.
2 Preliminaries on MaskCLIP
We also extract CLIP textual features for each text query with . Segmentation maps are then generated by computing the cosine similarity between each of the visual patch features and of the textual prompts, after L2-normalization. The most similar prompt is assigned to each patch. Note that a query ‘background’ can be added in order to obtain negative patches. Using MaskCLIP allows us to produce dense segmentation maps with a single forward pass of the classic CLIP model, but outputs are noisy as visible in Fig. 4.
3 Leveraging self-supervised features to improve MaskCLIP features
In this work, we aim to improve MaskCLIP’s open-vocabulary features described above. To do so, we propose to leverage the known good localization properties of self-supervised features already exploited for the task of class-agnostic unsupervised object localization .
with denoting the outer product. We compare in Fig. 2 the patch-similarities obtained for different patch seeds with MaskCLIP and DINO features (left and middle columns) and observe that the self-supervised features are more densely and accurately correlated than those of CLIP.
We then produce the segmentation maps , by comparing the new features to each textual queries in . When using such consolidated features, we obtain more stable and accurate outputs as visible in Fig. 4 and the high-frequency predictions observed in MaskCLIP are smoothed out, therefore showing the benefit of the pooling.
4 Producing a strong background detection.
Moreover, as discussed earlier, a ‘background’ query may be added to the set of textual queries in order to help filter out patches falling in the background and not corresponding to any objects. We argue that relying solely on the textual prompt ‘background’ to catch all non-salient patches is underperforming and, similarly to , we propose to use a very light-weight unsupervised foreground/background segmentation method, namely FOUND which also relies on DINO self-supervised features. As opposed to , we apply FOUND a single time for the entire image and extract a prediction mask in which a patch is assigned the value if falling into the foreground and otherwise. We also observe that saliencies produced by FOUND can be too restrictive and discard objects which are partially visible or in a clutter. In order to mitigate this behaviour, we propose to relax the background selection by integrating an additional uncertainty constraint. To this end, we leverage the best of both worlds and assign the ‘background’ class to patches which are both uncertain, e.g. have low confidence score , with the softmax operation, and which falls in the background in . We observe improved results and indeed a background better segmented as visible in Fig. 5 (more examples in Fig. 11).
5 Teaching CLIP a few DINO tricks
We have shown in the previous section that self-supervised correlation information can successfully be used to improve the dense quality of open-vocabulary features. If the difficulty of densifying CLIP is well-known, we show here that CLIP features already contain good localization information which can be extracted with light models. We indeed predict both DINO correlation and FOUND predictions (described in Sec. 3.3) from CLIP with dedicated convolutional layers.
to be closed to the binarized correlations , using the binary cross-entropy loss :
We visualize in Fig. 2 examples of and observe their similarity to DINO-based correlations. We use the CLIP-produced correlations to replace in Eq. 2 to weight the pooling and observe a similar boost over MaskCLIP, therefore showing that good patch correlation can indeed be extracted directly from CLIP. We can now discard DINO and we name CLIP-DINOiser the guided-pooling strategy which uses CLIP-based correlation. Our method runs with a single forward pass of the CLIP model (and a small additional convolutional layer).
We show examples of predicted CLIP-based objectness in Fig. 6 and observe their very high similarity to those produced with DINO. Moreover, we can now replace in Sec. 3.4 with the binarized CLIP-based scores , with the sigmoid operation, and observe a minimal drop in performances.
Experiments
In this section, we present experiments performed to evaluate our method CLIP-DINOiser. We detail in Sec. 4.1 the experimental setup used in our evaluation. We produce the ablation studies in Sec. 4.2 and state-of-the-art results on the task of zero-shot semantic segmentation in Sec. 4.3.
Technical details. We use in all experiments a frozen CLIP ViT-B/16 pre-trained by OpenCLIP . Our method CLIP-DINOiser uses two convolutional layers to extract DINO-like information from CLIP layer (the 3rd before last): has a kernel and output dimension and a kernel with . The first is trained to match the correlation information extracted from the value embeddings of the last layer of a ViT-B/16 model trained following DINO we discuss using different features in supplementary. The second layer is trained to replicate the unsupervised object localization predictions of FOUND –which also uses DINO model. We train both layers with a binary cross-entropy loss and train the model on PASCAL VOC training set consisting of 1464 images. We train for 20 epochs with a batch size of 32 images, which takes approximately 40 mins on 1 Nvidia A5000 GPU card. We decrease the learning rate by a factor of 0.1 after 15 epochs for FOUND head. For the correlations head, we stop the training after 5 epochs as we observe the model stops improving onwards. We apply data augmentations during training (random scale and cropping, flipping and photometric distortions). Overall, we binarize the correlations with and use a confidence score of . We ablate the parameters in Sec. A.1 and show that our method is rather stable.
Datasets and metric. We evaluate our method on eight benchmarks typically used for zero-shot semantic segmentation . Following , we split them into two groups. The first consist in datasets with a ‘background’ query: PASCAL VOC (noted ‘VOC’), PASCAL Context (noted ‘Context’), and COCO Object (noted ‘Object’) and the second without: PASCAL VOC20 (noted ‘VOC20’), PASCAL Context59 (noted ‘C59’), COCO-Stuff (noted ‘Stuff’), Cityscapes (noted ‘City’), and ADE20K (noted ‘ADE’). We evaluate results with the standard mIoU metric. We also follow the evaluation protocol of , use the implementations provided by MMSegmentation , employ a sliding window strategy, resize the input image to have a shorter side of 448. We also do not perform text expansions of the class names and use only the standard ImageNet prompts following (more discussion on prompts in Sec. A.3 of supplementary).
Baselines. We compare our method against state-of-the-art methods on open-vocabulary zero-shot semantic segmentation. For a fair comparison between methods, we report results without any post-processing step. We split our comparison in three categories: MaskCLIP+ is vocabulary specific, then ReCO , OVDiff , NamedMask construct prototypes, and finally others which either learn patch-level representation, e.g. GroupViT , ZeroSeg , SegCLIP , TCL , CLIPpy , OVSegmentor or refine frozen CLIP features: CLIP-DIY and MaskCLIP .
2 Ablation study
In this section, we conduct an ablation study of the different components of CLIP-DINOiser by first investigating the impact of our proposed feature pooling mechanism as well as the background detection.
The impact of the pooling mechanism. We propose with CLIP-DINOiser to combine MaskCLIP features with a well-defined linear combination and compare different solutions in 1(a). In , the authors proposed to refine the predictions with a combination weighted by CLIP key embeddings (noted ‘CLIP keys (preds.)’ in the table) and boost MaskCLIP results by more than + mIoU on VOC and VOC20, + and + and + mIoU on the other datasets. However, we show that working directly on the features allows us to achieve better results; we obtain consistent improvements ranging from + to + mIoU on all datasets when using DINO-based weight and further improve when using trained CLIP-based weights .
The impact of the background detection. We now discuss the improvement provided by our background refinement strategy, which is applied when stuff-like background patches need to be detected. We report such results in 1(b) when employing our pooling strategy (either using DINO features, noted ‘w. DINO ’ or those extracted from CLIP, noted ‘w. trained ’). When using solely ‘FOUND’ for background detection, as in , we improve by + mIoU on VOC (achieving mIoU), but when relaxing FOUND (see Sec. 3.4) with an uncertainty condition, we boost scores up to 62.1 on VOC, showing the limitation of using FOUND alone. We also achieve similar results when using CLIP-based predictions both with DINO-based and trained CLIP-based correlations, although we observe that best results are obtained with trained . We visualize CLIP-based mask in Fig. 6 and see high similarity to DINO-based predictions, therefore showing the localization quality of CLIP.
3 Zero-shot semantic segmentation
We now discuss state-of-the-art results on the task of zero-shot semantic segmentation.
Evaluation with no ‘background’ class. We first compare in Tab. 2 results on datasets which aim at the segmentation of most of the pixels in an image and do not consider a ‘background’ class. We observe that our method CLIP-DINOiser achieves the best results on four datasets yielding +, +, + and + mIoU over the second best performing method. Interestingly, we outperform methods which build expensive prototypes per visual concept on fine-grained datasets, showing the benefit of our lightweight and generalizable method. The only drop (- mIoU) is observed on VOC20 which is object-centric and depicts few and large objects. We believe that considering an adaptive granularity of features correlation could help mitigate this drop and we leave this for future work.
Evaluation with ‘background’ class. We now compare our method on datasets which include a ‘background’ query in Tab. 3. In this setup, we also apply our background detection mechanism (detailed in Sec. 3.4) on VOC and Object in order to improve the stuff-like background detection. We observe that CLIP-DINOiser significantly outperforms all methods which do not construct prototypes. Moreover, we note that we outperform OVDiff on Context and Object datasets by + and + mIoU respectively. We also reach on-par results on VOC, in particular when considering versions without ensembling. We would also like to highlight that OVDiff requires the construction of a ‘background’ prototype per concept, otherwise losing - mIoU on VOC as reported in the paper. Finally, our method is computed in a single pass of CLIP at inference with the light addition of two convolutional layers, while remaining fully open-vocabulary as it does not require any class-specific construction.
Qualitative results. We qualitatively compare in Fig. 7 CLIP-DINOiser with high-performing TCL and CLIP-DIY (two recent methods which provide code) on images taken from the datasets considered in the evaluation. We observe that our method generates predictions accurate both in terms of localization and assignment. Indeed we obtain fined-grained results on the challenging Cityscapes and ADE20k datasets, the query ‘car’ and ‘fountain’ are accurately located when CLIP-DIY and TCL produce much coarser results.
Conclusions
In this work, we propose to make the most out of CLIP features and show that the features already contain useful localization information. Indeed with light convolutional layers, we are able to learn both good patch-correlation and objectness information by using DINO self-supervised model as a guide. With such information, our method CLIP-DINOiser can perform zero-shot open-vocabulary semantic segmentation in a single pass of CLIP model and with two light extra convolutional layers. CLIP-DINOiser reaches state-of-the-art results on complex semantic segmentation datasets.
Limitations. Despite yielding strong results on open-vocabulary semantic segmentation, CLIP-DINOiser is still bounded by the capability of the CLIP model to separate classes, as it inherits its granularity. We believe that better prompt engineering, paired with better image-text models, could further boost CLIP-DINOiser.
Acknowledgments
This work was supported by the National Centre of Science (Poland) Grant No. 2022/45/B/ST6/02817 and by the grant from NVIDIA providing one RTX A5000 24GB used for this project.
References
A More ablations
In this section, we present additional ablations conducted on our proposed method. In particular, we discuss in Sec. A.1 the sensitivity to different parameter values and show that results are rather stable. We also present in Sec. A.2 a study of the impact of the features used to ‘teach’ CLIP the tricks.
We investigate in this section the impact of the different parameters of CLIP-DINOiser on the performance of our method.
To first evaluate the stability of our training, we perform runs with different random seeds. The results reported in the main paper on all datasets of CLIP-DINOiser correspond to and are highlighted in Tab. 5. We observe that in all cases, the standard deviation equals or lower, therefore showing the stability of our training.
Correlation threshold γ𝛾\gamma.
We study here the impact of the correlation threshold applied to the DINO affinity maps , which can be adjusted to control how many patches are used in the weighted pooling (Eq. 2 ). The case of corresponds to including in the weighted pooling all patches that are positively correlated to the seed. As increases towards , the pooling is more selective as fewer, yet better-correlated patches are considered. We report results in Tab. 4 and observe that scores stay rather stable on Context59, COCO Stuff and ADE20k. We also notice that the best results are obtained on VOC20 with , while more restrictive thresholds benefit Cityscapes. This can be explained by the fact that VOC mostly depicts large objects while each Cityscapes image contains multiple smaller objects from different classes. In all our experiments we use , which corresponds to the value used in .
Input layer l𝑙l used for training.
We present in Tab. 5 the results when applying our convolutional layers and on a different CLIP layer . We note the last CLIP layer which is aligned with the text. We report scores averaged over 3 runs along with their standard deviation. We observe that worse results are obtained with the last two layers, and , which are closer to the final alignment task. Instead, results obtained with layer are close, with a variation range below points of mIoU. Moreover, we can observe that the absolute standard deviation is in all cases below mIoU, showing the stability of our training.
Confidence threshold δ𝛿\delta in the background.
We report in Fig. 11 results for different confidence score thresholds in our background refinement (see Sec. 3.4 for details). This threshold controls how ‘uncertain’ a patch must be in order to be considered for FOUND ‘background’ segmentation. Using corresponds to applying the standard FOUND . We observe that using a threshold around achieves good results on both VOC and Object and that integrating an uncertainty consideration always helps vs. using FOUND alone (). We use in all our experiments.
A.2 Self-supervised features discussion
We investigate here the impact of the type of features used to ‘teach’ CLIP the self-supervised tricks. We present in Fig. 9 visualizations of correlation obtained with different DINO embeddings extracted from the last attention layer, namely ‘query’, ‘key’ and ‘value’. Most unsupervised localization methods use the ‘key’ embeddings which allow the easy separation of foreground from background. However, we observed in this work that using instead the value features allows us to separate better elements in the background, as visible in the figure. Indeed, patches in the background correlate to fewer background patches and regions are therefore better separated.
We also depict the final segmentation when using each type of feature, and observe the best result with ‘value’. We use following . We observe that more objects in the background are well segmented and labelled with ‘value’ embeddings e.g. ‘tree’ and ‘sky’.
We also tried using DINOv2 and its artefact-free version , but we observed that correlation maps were harder to exploit, thus leading to worse performance on our task.
A.3 Text template discussion
We discuss here the impact of using the text query templates which are proposed in the original CLIP repositoryhttps://github.com/openai/CLIP/ following the implementation from related works , versus the single template “a photo of a { }”. We report in Tab. 6 results for both CLIP-DINOiser and MaskCLIP when using both the single template, noted ‘single’ and the 80 ImageNet templates, noted ‘IN’. We observe that in the case of MaskCLIP, using multiple templates either does not significantly improve (between + and + mIoU) or even hurts performances (on VOC20), whereas using ImageNet templates always helps for CLIP-DINOiser (between + and + mIoU).
B More qualitative results
In this section, we illustrate the benefits of our method through more comparative qualitative results.
We show more examples of the application of our method CLIP-DINOiser and compare it to MaskCLIP results in Fig. 10. We observe that in all cases, our pooling reduces the noise in the predictions and helps produce good-quality segmentation.
Our background.
By visualizing more results with and without the background refinement step in Fig. 11, we observe that the background refinement step helps remove uncertain segmentation such as the snow area (which was classified as ‘snowboard’) in the left image, or on the cabinet which is not annotated in VOC (right image).
B.2 Failure cases
We discuss here the known failure modes of our method CLIP-DINOiser, which are visualized in Fig. 12.
We first observe some CLIP biases, which for instance produce similar features for ‘train’ and ‘train tracks’ (left figure), likely due to their frequent co-occurrence across images. We have observed other instances of this bias e.g. for ‘boat’ and ‘sea’ queries. Second, although CLIP-DINOiser can produce rather fine-grained segmentations (in terms of object sizes and classes), it can miss small or far-away objects as in Cityscapes (middle image). Finally, as with other open-vocabulary semantic segmentation methods, CLIP-DINOiser is not robust to the ambiguities of the text queries. The example from ADE20K (right image) is such a case, where ’house’ is mistaken for ’building’. In our experiments, we observed multiple segmentation ambiguities and we believe that the redefinition of evaluation metrics could help address the issue. We stress that the current evaluation setup, which is taken directly from fully supervised settings, might be limiting in an open-vocabulary paradigm.
B.3 More state-of-the-art results
We present more visual comparisons against state-of-the-art results in Fig. 14 and Fig. 15. We observe that CLIP-DINOiser produces fine segmentation results and outperforms baselines.
B.4 In the wild examples
We present more in-the-wild examples in Fig. 13, where we compare CLIP-DINOiser against MaskCLIP. MaskCLIP produces very noisy masks, especially when multiple false positive queries are considered (we define such false positive queries as prompted ones which are not represented in the image). Instead, CLIP-DINOiser is robust to such false positives and produces good quality masks.