A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model

Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, Xiang Bai

Introduction

Semantic segmentation is a fundamental computer vision task that assigns every pixel of an image with category labels. Accompanied by the development of deep learning , the semantic segmentation has also evolved tremendously under the supervised learning paradigm . However, unlike common image-level datasets such as ImageNet-1K/ImageNet-22K image classification which are easily scaled up to tens of thousands of categories, existing semantic segmentation tasks involve usually up to tens or hundreds of categories due to the significantly higher annotation cost, and thus limit the segmentors’ capability in handling rich semantics.

Zero-shot semantic segmentation is an attempt to break the bottleneck of limited categories. However, the narrowly defined zero-shot semantic segmentation usually only takes a small amount of labeled segmentation data and refuses to make use of any other data/information, consequently resulting in poor performance. In this work, we focus on another more practical setting: open-vocabulary semantic segmentation, as a generalized zero-shot semantic segmentation, concentrates more on establishing a feasible method to segment arbitrary classes and allows the use of additional data/information except the segmentation data. Specifically, we propose to leverage a recent advance of image-level vision-language learning model, i.e., CLIP .

While the vision-language learning model has learnt a strong vision-category alignment model using rich image-caption data, how to effectively transfer its image-level recognition capability to pixel-level is unclear. An natural idea is to integrate the vision-language model with a fully convolutional networks (FCN) , an architecture widely used for fully supervised semantic segmentation. A main difficulty of the integration is that the CLIP model is learnt at image-level, which differs from the granularity of FCN that models semantic segmentation as a pixel classification problem, where a linear classifier is applied on each pixel feature to produce the classification results, with each column of the linear classifier weight matrix representing each category. Empirically, we found the granularity inconsistency lead unsatisfactory performance.

To better leverage the strong vision-category correspondence capability involved in the image-level CLIP model, we pursue mask proposal based semantic segmentation approaches such as MaskFormer , which first extracts a set of class-agnostic mask proposals and then classifies each mask proposal into a different category. This two-stage approach decouples the semantic segmentation task into two sub-tasks of class-agnostic mask generation and mask category classification. Both sub-tasks prove well adaptation to handle unseen classes: firstly, the class-agnostic mask proposal generation trained using seen classes is observed well generalizable to unseen classes; secondly, the second mask proposal classification stage is at a same recognition granularity than that used in a CLIP model. To further bridge the gap with a CLIP model, the masked image crop of each proposal is used as input to the CLIP model for unseen classes classification. In addition, we employ a prompt-learning approach to further improve the unseen classes classification accuracy given a pre-trained CLIP model.

We evaluate the proposed approach under two different settings: 1)Cross-dataset setting where the model is trained on one dataset and evaluated on other datasets without fine-tuning. Under this setting, our two-stage framework demonstrate well generalization capability. It outperforms FCN approach by +13.1 mIoU on Cityscapes, +19.6 mIoU on Pascal Context, +5.6 mIoU on ADE20k with 150 classes and +2.9 mIoU on ADE20k with 847 classes. 2) Zero-shot setting where the model is trained on a part of seen class of a dataset and evaluated on all classes (including seen and unseen classes). We use this setting for comparing with other zero-shot semantic segmentation methods. We show that the proposed approach, though simple and straightforward, can surpass previous state-of-the-arts zero-shot segmentation approaches by a large margin. On Pascal VOC 2012 , this approach outperforms previous best methods that w/o self-training by +37.8 hIoU , and by +29.5 hIoU when an additional self-training process is involved. On COCO Stuff , the approach outperforms previous best methods that w/o self-training by +19.6 hIoU and by +8.9 hIoU when an additional self-training process is involved. We hope our simple but effective approach can encourage more study in this direction.

Related Works

Vision-language pre-training focuses on how to connect visual concepts and language concepts. Early approaches were performed on some cleaned datasets with relatively small data scale. Therefore, those models usually need to be fine-tuned on some specific downstream tasks. Some recent works have explored the benefits of large-scale noisy data obtained from web pages for vision-language pre-training. CLIP , as a representative work, employs a contrastive learning approach to distinguish the correct image-text pair in each training batch. Because many vision/language concepts are covered in large-scale data, the CLIP illustrates surprisingly strong capability on zero-shot/open-vocabulary image classification and image-text retrieval. This work introduces the CLIP model as a strong vision-category correspondant for open-vocabulary semantic segmentation.

0.2 Semantic Segmentation.

Semantic segmentation is a fundamental task in computer vision that aims to assign a category to each pixel. Fully convolutional network and its variants , as a practical and straightforward approach to model the semantic segmentation as a pixel-wise classification problem, have dominated this field in the past few years. Recently, MaskFormer explored to model the semantic segmentation as two sub-tasks: segment generation and segment classification and has shown competitive performance compared to FCN based approaches.

0.3 Zero-Shot Learning and Open-vocabulary Learning.

Zero-shot learning has been widely studied in recent years. A narrowly defined zero-shot learning focuses on learning transferable representations from the annotated data of seen classes to represent unseen classes. For example, proposed to learn a joint embedding space between the images and the name/description of the category for image classification, and explored taking the advantages of mid-level semantic representation. Recently, the open-vocabulary learning has attracted more attentions. As a generalized zero-shot learning, the open-vocabulary learning is more concerned with establishing a feasible method for arbitrary class recognition and allows the use of any additional information. For example, Visual N-Grams and CLIP explored the use of web-crawled data for image classification and introduced the vision-language pre-training model for the open-vocabulary object detection and showed it could significantly improve the long-tile object detection .

0.4 Zero-shot Semantic Segmentation.

Some pioneer works to study the zero-shot learning for semantic segmentation. ZS3Net uses generative models to synthesize pixel-level features by word embeddings of unseen classes. CSRL further incorporating the structural relation in feature synthesize. CaGNet introduce a contextual module for better feature generation. Different from , SPNet attempt to mapping vision feature to the semantic space via word embedding. JoEm a joint embedding strategy between the vision encoder and semantic encoder. In , variational mapping is used to learn semantic features. In , the uncertainty-aware losses are proposed to eliminate noisy samples. Other works explored other directions or aspects of zero-shot semantic segmentation. In , the super-pixel pooling is utilized to improve the region grouping generalization. In , the self-training for zero-shot semantic segmentation are carefully studied. In , the transductive learning setting are explored. In , they explore the utilization of image caption. However, all those methods have not explored the utilization of the vision-language pre-training model in zero-shot semantic segmentation. There are two concurrent work try to utilize the vision-language pre-training model in semantic segmentation. However, LSeg is an FCN-based approach focus on few shot setting. Openseg , which is a similar work to ours, utilizes external grounding dataset while we don’t. In addition, Openseg is based on ALIGN while we adopt CLIP .

Preliminary

In this section, we first introduce the setting of open-vocabulary semantic segmentation and revisit CLIP as preliminary.

Open-vocabulary is an generalized zero-shot task, so the zero-shot semantic segmentation protocol can also evaluate open-vocabulary semantic segmentation. In this setting, model predicts masks for unseen classes Cunseen{\cal C}^{\text{unseen}} by learning from some labeled data of seen classes Cseen{\cal C}^{\text{seen}}, and the seen classes and unseen classes are disjoint, i.e., Cunseen∩Cseen=∅{\cal C}^{\text{unseen}}\cap{\cal C}^{\text{seen}}=\varnothing. Usually, Cseen{\cal C}^{\text{seen}} and Cunseen{\cal C}^{\text{unseen}} are often represented with semantic words like dog, cat, apple, and sometimes the description of the classes are also provided.

1.2 Cross-Dataset Setting.

In this setting, the model is trained on one dataset and evaluated on another dataset without fine-tuning. This is a more challenging setting than the zero-shot setting, where the model not only deals with the unseen classes, but also has to address the domain gap among different datasets.

2 Revisiting CLIP

CLIP is a powerful pre-trained vision-language model, which shows surprisingly strong performance in associating the visual and textual concepts. CLIP is a two-stream method: it contains an image encoder Eimage{\cal E}_{\text{image}} and a text encoder Etext{\cal E}_{\text{text}}. For any given image-text paired data {I,T}\{{\cal I},{\cal T}\} , their semantic similarity can be estimated by computing the cosine distance between Eimage(I){\cal E}_{\text{image}}({\cal I}) and Etext(T){\cal E}_{\text{text}}({\cal T}).

The pre-trained CLIP model can be used to classify images by a given set of classes without fine-tuning, which is also known as zero-shot/open-vocabulary image classification. Specifically, the class names are injected into the pre-defined prompt template and fed into CLIP’s text encoder to generate the class embeddings, e.g., a typical prompt template is ‘a photo of [CLASS]’, where [CLASS] is replaced by the specific class name such as ‘person’ and ‘cat’. The generated class embeddings are used as the classifier and the similarity with image embedding is computed for classification.

In this work, we extend the compatibility of CLIP from image-level zero-shot/open-vocabulary classification to pixel-level open-vocabulary semantic segmentation, by exploring the use of a pre-trained CLIP model as a strong vision-category correspondent.

Two-Stage Open-Vocabulary Semantic Segmentation

Figure. 1 shows an overview of our two-stage framework. Given an image, a set of mask proposals are first generated, and then each proposals is fed into an image encoder and compared with the class weights obtained by applying text encoder on the prompt class description to perform the classification. Finally, the mask prediction are assembled together to produce the final segmentation results. We will describe each component of our framework in the following.

We first introduce the mask proposal generation. In our work, we try three different methods to generate the mask proposals {Mkp}\{{\cal M}^{p}_{k}\}:

GPB-UCM . This is a classical method to generate hierarchical segments by considering multiple low-level cues, e.g., brightness, color, texture, and local gradients. The generated segments of this approach are usually well aligned with the contour of objects.

Selective Search . This method can also generate hierarchical segments. Since this method can effectively localize objects, it is widely used in object detection systems .

MaskFormer . This is a recently proposed method for supervised semantic segmentation. Unlike a fully convolution network that models the semantic segmentation as the pixel-wise classification problem, MaskFormer disentangles the semantic segmentation into two sub-tasks: predicting the segments at first and then classifying the category of each segment. We observe that the predicted segments by MaskFormer can be used as the mask proposals, and we empirically demonstrate (see Sec. 6.4) that the MaskFormer trained on seen classes can produce high-quality mask proposals on the unseen classes. Therefore, we take this advantage of MaskFormer as our default mask proposal generator.

2 Region Classification via CLIP

There are two strategies to perform the region classification by utilizing the pre-trained CLIP:

The first strategy is to directly apply the CLIP image encoder on each mask proposals for classification. Specifically, given an image I{\cal I} and a mask proposal Mp{\cal M}^{p}, the mask proposals are first binarized with a threshold of 0.5, and then apply the binarized Mp{\cal M}^{p} to image I{\cal I}, erase the unused background and only crop foreground area. The masked image crop is resized to 2242224^{2} and then fed into CLIP for classification. However, since there is no extra training process, the training data of seen classes cannot be utilized, resulting in inferior performance on seen classes in the inference (see Sec. 6.5.1).

To utilize the training data of seen classes, another approach is to retrain an image encoder. However, if we simply learn a set of new classifiers on the training data of seen classes, the retrained image encoder has no generalization ability on unseen classes since these classes have no corresponding classifiers. Therefore, we propose to use the features generated from the text encoder of the pre-trained CLIP model as the fixed classifier weights for the retrained image encoder. In this approach, the image encoder has a certain generalization ability to the unseen classes since the image encoder is encouraged to embed the vision features into the same embedding space of the text encoder through the seen classes. Notably, this approach can be easily integrated into the training process of the MaskFormer, by simply using the CLIP generated text features as the classifier weights of the MaskFormer, thus avoiding the need of training an additional image encoder.

The two strategies complement each other (see Sec. 6.5.1), therefore we ensemble the results of these two strategies by default. Given a mask proposal Mp\mathcal{M}^{p}, we crop the foreground area Afg=crop(Mp,I)A_{fg}=\text{crop}(\mathcal{M}^{p},\mathcal{I}) (See Appendix for details), and compute its classification probability via CLIP vision encoder EvisionE_{\text{vision}} and text encoder EtextE_{\text{text}}:

,where Ci\mathcal{C}_{i} is name of i-th class and temperature τ\tau=100. The classification probability of CLIP can be ensembled with supervised model trained on seen classes and then generate final mask results according to Sec.4.3.

2.2 Prompt Design.

The original CLIP is not designed for open-vocabulary semantic segmentation. How to design feasible text prompts need to be explored.

A simple approach is to re-use the hand-crafted prompts provided by CLIP which is originally designed for image classification on ImageNet-1K . There are 80 different prompts, each consisting of a natural sentence with a blank position for injecting the category names. Since these prompts are not originally designed for semantic segmentation, some of them may have a adverse effect. So we evaluate each of these prompts on training data to select one most helpful prompt for open-vocabulary semantic segmentation.

Prompt learning recently showed great potential for adapting the pre-trained language/vision-language models on specific downstream tasks. We also explore this technique. Specifically, a prompt is a sequence of tokens. Each token belongs to one of the two types: [P][P] indicates the prompt token and [CLS][CLS] indicates the class token. A generalized prompt can be formulated as [P]0...[P]m[CLS][P]_{0}...[P]_{m}[CLS], where mm is the number of prompt token. In prompt learning, the prompt tokens [P]0...[P]m[P]_{0}...[P]_{m} are set as learnable parameters that can be trained on the seen classes and generalized to the unseen classes.

3 Mask Prediction Assembly

Since the mask proposals may overlap each other, resulting in the possibility of some pixels being covered by several different mask proposals. Therefore, we employ a simple aggregation mechanism to generate semantic segmentation results from the mask predictions. Specifically, for a given pixel qq, its predicted probability of being ii-th category is defined as:

where Mkp(q){\cal M}_{k}^{p}(q) denotes the predicted probability of pixel qq in kk-th mask proposal Mkp{\cal M}_{k}^{p}, and Ckp(i){C}_{k}^{p}(i) is the predicted probability of mask proposals Mkp{\cal M}_{k}^{p} belonging to ii-th category. Note that the sum of Ci(q)C_{i}(q) over all categories is not guaranteed to be 1, and pixel qq is classified to the category with highest predicted value.

Fully Convolution Network Approach

In addition to our proposed two-stage framework, a more conventional approach is to use the widely-used fully convolution network (FCN). As a dominant method in supervised semantic segmentation, FCN formulates the semantic segmentation as a pixel-wise classification problem. Specifically, given an image, FCN generates a high-resolution feature map, and a set of learned classifiers is applied on each pixel to produce segmentation predictions. Similar to our proposed two-stage framework, there are also two strategies to apply the CLIP on FCN framework:

Directly using the feature map generated by the CLIP vision encoder to perform pixel-wise classification. Note that in the original CLIP model, the feature of an image are represented by the feature of [CLS] token, not the feature map, and this difference may lead to performance degradation. In addition, the original CLIP model uses the image size of 224×224224\times 224 during pre-training, while semantic segmentation usually requires a higher image resolution (e.g., shorter size is 640640). Therefore, the direct use of high-resolution image during inference may lead to inferior performance due to inconsistency in image size. To alleviate this problem, we try to use the sliding window technique, which is widely used in previous works for performing multi-scale inference. We empirically found that it can improve performance and thus use it by default.

The training data of the seen classes cannot be utilized in the first strategy. Instead, we retrain an FCN-based vision encoder on seen classes via the similar method introduced in Sec. 4.2.1. Specifically, we use the CLIP text encoder to generate a fixed classifier weight. Therefore, the retrained model can obtain a certain generalization ability to the unseen classes.

As the same as the two-stage framework, we also ensemble the prediction of these two strategies by default if not specified.

Experiments

We conduct extensive experiments on five challenging datasets to evaluate our method: COCO Stuff , Pascal VOC 2012 , Pascal Context , Cityscapes , and ADE20K .

COCO Stuff is a large-scale dataset that contains 117k training images and 5k validation images. It contains annotations of 171 classes, 80 thing classes and 91 sutff classes respectively.

Pascal VOC 2012 contains 11,185 training images and 1,449 validation images from 20 classes. The provided augmented annotations are used.

Cityscapes is a scene parsing dataset collected on urban streets, containing 5,000 finely annotated images and 20,000 coarsely annotated images. According to the common practices , we use 1,525 images of 19 classes in the finely annotated set for validation.

Pascal Context is an extensive dataset of Pascal VOC 2010, containing 4,998 training images and 5,005 validation images. We use the frequent 59 classes for validation.

ADE20K contains 20k training images, 2k validation images, and 3k testing images. There are two settings of 150 classes and 857 classes.

1.2 Data Split.

For Cross-dataset setting, we train our model on the COCO Stuff dataset and test on the validation set of the others. For Zero-shot setting, we evaluate our method on COCO Stuff and Pascal VOC 2012. Following , we divide the COCO Stuff dataset into 156 seen classes and 15 unseen classes and the Pascal VOC 2012 dataset into 15 seen classes and 5 unseen classes.

1.3 Evaluation Protocol.

For cross-dataset setting, we use the mean of class-wise intersection over union (mIoU) as major metric. For zero-shot setting, we use harmonic mean IoU (hIoU) among the seen classes and unseen classes as major metric by following previous works (see Appendix for detail definition). We also report the pixel-wise classification accuracy (pAcc) as a reference.

2 Implementation Details

We conduct all experiments on 8×\timesNvidia V100 GPUs. We train a MaskFormer model on the COCO Stuff dataset with ResNet-101 as the default backbone. An AdamW optimizer with the initial learning rate of 1e-4, weight decay of 1e-4 and a backbone multiplier of 0.1, and a poly learning rate policy with a power of 0.9 are used. The batch size is set to 32 for each GPU, and the total training iteration is 60K/120K for zero-shot setting and cross-dataset setting,respectively. If not specified, the MaskFormer model is only trained on seen classes, and we use 100 mask proposals for both training and testing. For all other settings and hyper-parameters, we keep the original setting of MaskFormer without changes. CLIP with ViT-B/16 backbone is used by default if not specified. In text prompt tuning, the prompts are randomly initialized, and a SGD optimizer is used to train the learnable prompts. The learning rate is set to 0.02 and decayed according to the cosine learning rate policy, and the batch size is set to 32. We train 50 and 100 epochs for Pascal VOC and COCO Stuff, respectively. For Pascal VOC 2012 dataset, we use a batch size of 16 and a total training iteration of 20K, and keep all other setting as the same as the COCO Stuff dataset.

3 Comparison in Cross-Dataset Setting

We first evaluate our method on the cross-dataset setting. The model is trained on the COCO Stuff dataset and then evaluated on other datasets without fine-tuning. Table. 1 clearly shows that our two-stage approach outperforms the FCN approach by a noticeable margin, demonstrating that our two-stage approach can better leverage the pre-trained CLIP model than the FCN approach. We do not list the result on Pascal VOC as its categories overlap much with the COCO Stuff dataset, and our method can achieve 88.4 mIoU.

4 Comparison in Zero-Shot Setting

We then compare our method with previous state-of-the-arts on Pascal VOC 2012 dataset and COCO Stuff dataset. Since some works reported the performance by applying the self-training techniques (denoted as “ST”), we follow this practice and report the performance with or without self-training.

COCO Stuff. Sec. 6.2 shows the results. Compared with Pascal VOC 2012 dataset, COCO Stuff is more challenging. However, our approach still outperforms state-of-the-arts by a large margin. Specifically, without using the self-training, our method achieves 37.8 hIoU and 36.3 mIoU-unseen, outperforming the previous best method CaGNet by +19.5 hIoU and +24.1 mIoU-unseen. By further employing the self-training, our method achieves 41.5 hIoU and 43.6 mIoU-unseen, outperforming the previous best method STRICT by +8.9 hIoU and +13.3 mIoU-unseen. The qualitative results are shown in Figure. 2.

Pascal VOC 2012. The results are shown in Sec. 6.2. Without using the self-training, our method achieves 77.5 hIoU and 72.5 mIoU-unseen, outperforming the previous best method CaGNet by a huge margin of +37.7 hIoU and +46.8 mIoU-unseen. By further employing the self-training, our method achieves 79.3 hIoU and 78.1 mIoU-unseen, outperforming the previous best method STRICT by +29.5 hIoU and +42.5 mIoU-unseen.

While our method outperforms other state-of-the-art zero-shot semantic segmentation methods, how the larger pre-trained data and image encoder affects the performance is still unclear. To study these impacts, we design a new implementation that enables our approach to only leverage ImageNet-1K classification data. Specifically, we train a vision-language model by only using ImageNet-1K: the class names of ImageNet-1K are treated as language inputs, and are encoded through a pure text encoderWe use SimCSE as the text encoder trained on text data only. to generate the classification weights. As shown in Table. 15, our method achieves 49.5 hIoU with ResNet-101 as backbone, which is much higher than other approaches. On the other hand, we also try to integrate the CLIP with SPNet, and we find it only achieves 33.4 hIoU, which is far from our method by using the same ResNet-101 backbone. Those experiments indicate that the surpassing performance of our method does not only come from larger pre-training data, but also our two-stage framework.

5 Ablation Studies

In this section, we validate the key designs of our method. If not specified, we report the performance on the COCO Stuff dataset with the MaskFormer model of ResNet-101 and CLIP of ViT-B/16 by using the zero-shot setting.

We evaluate the performance of the mask proposal generation methods by plugging them into our pipeline. To avoid the impact of the learnable classifier trained on seen classes, we perform the comparison by directly classifying the masked regions with the CLIP model. The results are shown in Sec. 6.4, and the MaskFormer achieves better performance than the Selective Search and GPB-UCM. Note that even the other two methods are worse than the MaskFormer, they still achieve comparable performance compared with state-of-the-arts on mIoU-unseen.

5.2 Generalization of Mask Proposal Generator.

Although using MaskFormer to generate the mask proposals achieves excellent performance on zero-shot setting, it is still unknown whether Maskformer can produce good performance on the cross-dataset setting, i.e., training on one and testing on another dataset. Therefore, we evaluate the generalization ability of using MaskFormer to generate mask proposals between the COCO Stuff dataset and the ADE20K dataset.

In this experiment, we want to evaluate only the quality of the proposal without the effects of the region classifier. However, it is difficult to design a simple “recall” metric for mask proposals in semantic segmentation similar to object detection. Because a segment can consist of multiple mask proposals, this may lead low recall while the final semantic segmentation result is still correct. Therefore, we designed an “oracle” experiment to evaluate how these proposals affect the final performance of semantic segmentation. Specifically, for each mask proposal, its category is specified as the same as the ground-truth segment in which it has the largest overlap. In this case, the segmentation performance can fully reflect the proposal quality.

The results are shown in Sec. 6.4. We directly report the mIoU in this experiment because the seen class cannot be defined between different datasets. We note that the MaskFormer model trained on COCO Stuff can produce good performance on ADE20K compared to the MaskFormer model directly trained on ADE20K with acceptable performance degradation, and vice versa. That demonstrates the generalization ability to use MaskFormer as the proposal generator.

5.3 Different Strategies of Using CLIP.

We study the two different strategies of using CLIP discussed in Sec. 4.2.1: retrained vision encoder or directly using CLIP vision encoder without tuning. The results are shown in Sec. 6.5.1. The retrained vision encoder shows excellent performance on seen classes, while its performance on unseen classes is relatively low. In contrast, the CLIP vision encoder shows strong performance on unseen classes while worse than retrained vision encoder on seen classes by a large margin. By ensembling the two strategies, the performance on both seen and unseen classes is significantly improved, indicating the two strategies are complementary.

5.4 Different CLIP Variants.

CLIP provides several variants with different network architectures and model sizes. We study how these models affect the performance of our method when using them as the region classifiers. We report the results without using the learnable prompt due to the high experimental overhead. The results are shown in Sec. 6.5.1. We find that all models perform well and CLIP with ViT-B/16 achieves the best performance.

5.5 Comparison with Supervised Baseline.

We also compare our method with the supervised baseline on COCO Stuff. The supervised model is MaskFormer with ResNet-101 backbone which is trained on all classes, including seen and unseen classes. We report the results in Sec. 6.5.1. It is remarkable that while our method is worse than the supervised baseline by a large margin on mIoU-unseen and hIoU, the gap in mIoU is much close. That is because there are only 15 unseen classes in the current dataset partition. For reference, there are 156 seen classes.

To further explore the performance gap between our method and the supervised baseline, we split the unseen classes into things and stuff. The results are reported in Sec. 6.5.1. We find the performance gap of our method between things classes and the stuff classes is significantly large than the supervised baseline, and self-training can significantly reduce the gap. This observation suggests that the classification ability of CLIP models is different in things and stuff, which may be due to the bias of the pre-trained dataset used by CLIP.

Conclusion

In this work, we propose a simple and effective two-stage framework for open-vocabulary semantic segmentation with the advanced pre-trained vision-language model. We reformulate and break down the open-vocabulary semantic segmentation into two steps: 1) training a mask proposal generator to generate a set of binary masks and 2) leveraging the pre-trained CLIP to classify each mask proposal. We conduct extensive experiments to verify our approach. Notably, the proposed framework outperforms previous state-of-the-arts of zero-shot semantic segmentation on Pascal VOC 2012 and COCO Stuff by large margins. Our work reveals the potential for using pre-trained vision-language models on open-vocabulary/zero-shot semantic segmentation and provides a strong baseline for this community to facilitate future research.

Appendix 0.A Definition of hIoU

Following previous works , harmonic mean IoU (hIoU) is defined among the seen classes and unseen classes as:

Appendix 0.B Sliding Window Testing in Fully Convolutional Network

We study the different inference methods in this section for Fully Convolutional Network(FCN). For a fair comparison, we use ResNet-101 in FCN and ViT-B/16 in CLIP, same as our two-stage framework. Table. 11 shows the results. The FCN without sliding window test achieved 11.7 hIoU and 10.4 mIoU-unseen. In comparison, employing the window test improved the performance by +9.2 on hIoU and +5.6 on mIoU-unseen. This significant difference in performance is caused by the inconsistent image size between pre-training and testing of the CLIP model. Although the sliding window test can strengthen the FCN approach, it is still worse than our two-stage approach by -16.8 on hIoU and -20.3 on mIoU unseen, indicating that our two-stage framework is more suitable for the CLIP model.

Appendix 0.C Prompt Engineering for Image and Text

As the CLIP model is trained with low-resolution realistic images, given a mask proposal Mp\mathcal{M}^{p}, and the input image II, it is a problem how to extract the visual representation of the proposal with the CLIP model through the proper way, which we call image prompt engineering. The whole process is shown in Figure. 3. We crop the image with the bounding boxes of Mp\mathcal{M}^{p} and expand the bounding boxes by a ratio rr to involve more context information. And then, we fill the background pixels with values in the proposal with some patterns. We studied four choices for such patterns: a) Keep the background pixels unchanged; b) Fill the background pixels with manually designed values; c) Fill the background pixels with learnable values; d) Fill the background patches with mask token, presented in Figure. 4. The results are shown in Table. 12. The value filled in the background area can greatly affect the segmentation performance. Though our exploration to learn proper image prompts failed to achieve improvement like text , it is still an interesting problem for future research.

C.2 Prompt Engineering for Text

We compare two prompt tuning methods described in Sec. 4.2. The results are shown in Sec. 0.C.2. The learnable prompt outperforms the manually searched prompt by +9.9 hIoU, clearly showing the power of the learnable prompt. In addition, although the learnable prompt is only trained on seen classes, we notice that it achieves similar improvement on seen classes and unseen classes (+9.6 mIoU-seen and +10.2 mIoU-unseen), indicating the learnable prompt has a strong generalization ability to the unseen class.

We further study how prompt length and training data size affect the performance of learnable prompts by training on seen classes and testing on unseen classes. Sec. 0.C.2 shows that using 32 samples for each category reaches the best performance, in either prompt length of 16 or 32, and more training samples will degrade the performance. We speculate that more samples may lead to the over-fitting issue, which is also reported by other prompt learning attempts .

Appendix 0.D Detailed Study on MaskFormer and CLIP

The ablations have been studied in Table. 3 and Table. 4. We re-organize the results as shown in Table. 15. We can conclude: 1) MaskFormer outperforms FCN by +26.8 hIoU (2-th row vs 5-th row); 2) CLIP pre-training outperforms ImageNet pre-training by +24.7 hIoU (3-rd row vs 4-th row); 3) Our method outperforms SPNet by +24.4 hIoU with the same pre-training data (1-st row vs 3-rd row).

Appendix 0.E The Randomness of the Data split

In the experiments under the zero-shot setting, we use the official unseen/seen split as for a fair comparison. The thing/stuff ratio is 0.88 for seen and 0.87 for unseen classes. To study the impacts of different splits, we conduct studies on randomly generated seen/unseen splits in Table. 16. We find a more balanced unseen thing/stuff ratio yields higher hIoU.

Appendix 0.F Visualization of Results under Cross-dataset Setting

We illustrate more qualitative results in Figure. 5, 6 under the cross-dataset setting.

References