Learning to Generate Text-grounded Mask for Open-world Semantic Segmentation from Only Image-Text Pairs

Junbum Cha, Jonghwan Mun, Byungseok Roh

Introduction

Open-world semantic segmentation aims to identify the arbitrary semantic concepts in the open worldThis setting is often called both open-world and open-vocabulary. In this paper, we mainly refer to this setting as open-world for clarity.. Conventional semantic segmentation aims to learn segmentation capability for the small number of pre-defined target categories, whereas open-world semantic segmentation addresses unrestricted arbitrary categories or free-form texts. Such segmentation capability over unlimited targets drastically extends the application scope of the open-world segmentation models.

The first challenge for open-world segmentation is how to learn arbitrary concepts, beyond pre-defined categories. Inspired by the success of CLIP , previous approaches tackle this challenge by exploiting massive web-crawled image-text paired data; since the texts in web-crawled data contain a global semantic description for the paired images, the large-scale image-text pairs can provide rich knowledge for arbitrary semantic categories. However, there still remains another challenge in how to achieve precise localization of arbitrary concepts without dense annotations. There are several approaches that simply address this issue using dense annotation (segmentation masks) in addition to image-text pairs . The dense annotation helps to improve segmentation performance in a fixed benchmark dataset, but the requirements of expensive dense annotation still limit the applicable domains and scalability of the method.

In this paper, therefore, we focus on open-world semantic segmentation from only image-text pairs without any dense annotation. For this setting, the existing methods learn an image-text alignment capability during training and heavily rely on the transferability of the image-text alignment to perform region-text alignment at inference. More specifically, MaskCLIP leverages CLIP models pre-trained to learn image-text alignment. To perform region-text alignment using CLIP, MaskCLIP applies a simple heuristic modification to the CLIP image encoder. GroupViT and ViL-Seg propose to cluster region-level visual features into distinct groups and generate segmentation masks by matching the groups and texts. Note that they match the text embeddings and clustered region features in test time, but in training time, the text embeddings are aligned with global image embeddings. While the existing methods have shown impressive results even through the training with image-text alignment, they still suffer from the alignment-level discrepancy between training and testing phases as depicted in Fig. 2.

To address this train-test discrepancy, we propose the Text-grounded Contrastive Learning (TCL) framework, which allows a model to learn region-text alignment directly from the image-text pairs without any dense annotations. Our key idea is to incorporate a text grounding procedure within contrastive learning as illustrated in Fig. 2, where TCL generates a segmentation mask indicating text-grounded regions, computes grounded region embeddings using the mask, and applies contrastive learning between text and grounded region. By re-formulating the contrastive loss to be directly affected by the segmentation quality, TCL enables end-to-end training of the grounder and directly improves the quality of region-text level alignment. We also present a unified evaluation protocol using widely used 8 semantic segmentation datasets and compare existing methods in the same setting. As a result, TCL achieves state-of-the-art zero-shot segmentation performance with large margins in all datasets, as shown in Fig. 1.

Our main contributions are summarized as follows:

We introduce a novel framework for open-world segmentation, named Text-grounded Contrastive Learning (TCL), which enables learning region-text alignment directly without train-test discrepancy, thus learning to generate more precise segmentation masks through only image-text pairs.

We present a unified evaluation protocol and re-evaluate recent open-world segmentation models for a fair and direct comparison.

We achieve the new state-of-the-art zero-shot segmentation performance on 8 segmentation datasets with large margins compared to existing methods.

Related Works

Open-world scenario aims to recognize arbitrary concepts in the open world. It is also called open-vocabulary because the target vocabulary is open rather than closed. Contrastive Language-Image Pre-training (CLIP) ushered in the era of open-world image recognition using large-scale image-text pairs . CLIP learns the alignment between an image and a text in training time, then transfer it to the zero-shot classification by aligning image and texts indicating target classes at inference time. The advent of CLIP enables open-world settings in various fields such as object detection , image captioning , or semantic segmentation .

Open-world semantic segmentation with image-text pairs is addressed in two different settings. The first is a semi-supervised setting, which uses dense annotation (i.e., segmentation masks) in addition to image-text pairs . Semi-supervised approaches learn segmentation capability using dense annotation and expand the target vocabulary using image-text supervision. LSeg expands target class vocabulary using image-label datasets and CLIP text encoder . OpenSeg and OVSeg first train a mask generator using dense annotation and expand target vocabulary using image-text datasets. The use of dense annotation makes the model learn region-level alignment instead of image-level alignment, leading to high-quality segmentation masks. However, it still relies on costly dense annotation, and applicable domains are limited to the domains where dense annotation is available.

The target of this paper is an unsupervised setting, which aims to learn segmentation from only image-text pairs without any dense annotation . Since the massive image-text pairs are easily obtained by web crawling without human annotators, applicable domains of unsupervised methods become almost unlimited. In order to achieve segmentation capability using only image-text pairs, we need to learn region-text alignment instead of image-text alignment and train a text-grounded mask generator. However, the absence of dense annotation makes this approach challenging. Existing open-world semantic segmentation studies have taken a strategy to bypass this issue. Instead of learning region-level alignment directly, they transfer image-level alignment to region-level by heuristic modification or clustering . MaskCLIP proposes to obtain a dense image embedding from CLIP image encoder through heuristic modification of the last attention layer. Even though it has several limitations, such as low output resolution or noisy segmentation results, they show it is a simple yet effective way to obtain an initial segmentation map for refinement. ReCo proposes an advanced refinement method based on MaskCLIP, by retrieval and co-segmentation. Clustering-based methods learn representations using CL with image-text pairs. They compute region-level image embedding by clustering sub-region embeddings. These approaches also have shown impressive results but have several limitations: (ii) the learning objective is still image-level alignment due to lack of the region annotation, (iiii) the number of clusters is pre-defined independent of the given image, and (iiiiii) clustering sub-region image embeddings is independent of the query text. In summary, existing methods indirectly address region-level alignment problems by learning image-level alignment. To tackle this problem, we propose a novel region-level alignment objective, named Text-grounded Contrastive Learning (TCL).

2 Region-level Contrastive Learning

Learning region-level alignment instead of image-level alignment is a fundamental target objective in dense tasks, such as segmentation or object detection. There are approaches to learn region-level alignment using dense annotation in the semi-supervised setting. They first train mask or region proposal networks using dense annotation and learn alignment between the proposals and texts . For example, OpenSeg trains a class-agnostic mask generator using dense annotation. In the object detection field, RegionCLIP employs an off-the-shelf region proposal network and learns region-level alignment. In contrast to the existing region-level methods, the proposed method learns region-level alignment without any dense annotation.

Methods

Open-world semantic segmentation is a task that aims to learn a model capable of zero-shot segmentation for arbitrary visual concepts, not restricted to pre-defined ones. Our main goal is to develop an open-world segmentation algorithm using only image-text pairs. However, achieving this objective is challenging because there is no explicit supervision (i.e., pixel-level dense annotations) for text-described region segmentation. Existing methods learn models parametrized by θ\theta to maximize the mutual information between paired images and texts as follows:

where (xV,xT)(\mathbf{x}^{V},\mathbf{x}^{T}) is a random image and text pair. This objective encourages the model to learn the alignment between images and texts, however, at test time, the learned model generates segmentation masks for arbitrary concepts by computing region-text alignments. Such alignment-level discrepancy between train and test time can lead the model to a sub-optimal solution as shown in Fig. 2. With this in consideration, to bridge the gap between the objective of conventional contrastive learning (CL) and the requirement of the zero-shot segmentation, we propose Text-grounded Contrastive Learning (TCL) which incorporates a text grounding process within CL to enable learning region-text alignment directly. As a text grounding module, we introduce a grounder to generate segmentation masks for the given texts. In a nutshell, TCL learns a model to maximize mutual information between text-grounded regions and texts as follows:

where m\mathbf{m} is a text-grounded mask of random variable indicating the text-described region. Compared to contrastive learning that implicitly learns a grounding capability, TCL has a clear advantage of explicitly learning the grounding capability through the end-to-end trainable grounder.

In the rest of this section, we first explain the text-grounded mask generation procedure by the grounder. Then, we describe how we define losses using the generated mask to train our open-world grounder with text-grounded contrastive learning. Lastly, we explain how our model performs zero-shot inference for arbitrary concepts.

2 Grounder

Fig. 3 illustrates our overall training pipeline. For an input batch of paired texts XT\mathbf{X}^{T} and images XV\mathbf{X}^{V}, TCL first performs a grounding process to identify text-grounded regions for a text via a grounder. The grounder consists of three components: (ii) image encoder EvE_{v} is in charge of providing a single (L2-normalized) global feature as well as dense patch-level features, (iiii) text encoder EtE_{t} provides a (L2-normalized) text embedding feature, and (iiiiii) grounding decoder DgD_{g} converts dense features from image encoder into finer pixel-level embeddings for alignment with text. In practice, we adopt a pre-trained CLIP model to initialize two encoders and freeze them to preserve and exploit the rich knowledge of CLIP learned during large-scale pre-training. The text-grounded masks are computed by the position-wise dot product between text embedding and dense pixel-level embedding. The overall process of grounder is summarized as follows:

The generated text-grounded masks are used to extract text-grounded image embedding. By replacing the global image embedding with text-grounded image embedding in the contrastive learning framework, TCL enables the model to learn region-text alignment in an end-to-end manner. In the following section, we describe how the generated mask M\mathbf{M} is used for text-grounded contrastive learning.

3 Text-grounded Contrastive Learning

Recall that the main idea of TCL is to use text-grounded images instead of whole images, unlike conventional CL. For this purpose, we define TCL losses in three different levels—image-level, feature-level, and area-level—using the generated masks M\mathbf{M} for all pairs of images and texts in a batch; the detailed pseudo code to compute TCL losses is given in Appendix A. We also employ smooth regularization to further improve the quality of generated masks.

Feature-level TCL loss.

Note that this feature-level embedding is computed using negative masks Mi,j (i≠j)\mathbf{M}_{i,j~{}(i\neq j)}, different from the image-level TCL loss. We then compute the cosine similarity Si,jf=vi,jf⊤tjS^{f}_{i,j}={\mathbf{v}^{f}_{i,j}}^{\top}\mathbf{t}_{j} between all pairs of text embeddings and feature-level text-grounded image embeddings in the batch. The feature-level TCL loss is defined as follows:

Area TCL loss.

The image-level and feature-level TCL losses focus on generating a mask to capture the text-described region in the image. However, the model can collapse into a trivial solution with only these losses—generating a mask for the entire image instead of the desired region. To prevent this collapse, we introduce an additional objective to our TCL framework, named area TCL loss, which incorporates priors on the mask area to ensure capturing only the text-described region. To be specific, for the positive masks (masks from positive pairs) M+\mathbf{M}^{+} and the negative masks (masks from negative pair) M−\mathbf{M}^{-}, we denote the area of positive and negative masks by M+‾\overline{\mathbf{M}^{+}} and M−‾\overline{\mathbf{M}^{-}}, respectively. The area TCL loss is defined by L1-distance between the area priors and the expected area of each mask:

where p+p^{+} and p−p^{-} are positive and negative area priors. For the negative area prior p−p^{-}, intuitively, we can expect the area of the negative masks to be 0.0. We set the positive area prior p+p^{+} to 0.4, which is the average text-described region area measured by MaskCLIP in the CC3M dataset .

Smooth regularization.

In the image-text dataset, a text usually describes the salient object or concept in the paired image. We observe that the regions described by the text are generally smooth rather than noisy. We employ total variation (TV) regularization loss to incorporate this smoothness observation in the objective. The TV loss is applied to both mask and pixel-level dense embedding:

where ∥⋅∥TV\|\cdot\|_{\text{TV}} is the anisotropic TV norm.

Final loss.

where LTCL=LTCLv+LTCLf\mathcal{L}_{\text{TCL}}=\mathcal{L}_{\text{TCL}_{v}}+\mathcal{L}_{\text{TCL}_{f}}, and λTCL\lambda_{\text{TCL}}, λarea\lambda_{\text{area}}, λtv\lambda_{\text{tv}} are hyperparameters to balance three losses.

4 Inference Pipeline

Prompt templates such as “a photo of a {label}.” are used to generate text embeddings as in CLIP .

Experiments

In open-world semantic segmentation, a standard evaluation protocol is not yet established. Previous studies conduct an evaluation using their own protocols such as different data processing strategies on different datasets ; surprisingly, even for the same dataset, the target classes are sometimes different across studies. For a fair comparison, we present a unified evaluation protocol following the open-world scenario where prior access to the target data before evaluation is not allowed. Under this scenario, the proposed protocol prohibits dataset-specific hyperparameters or tricks, e.g., class name expansion or rephrasing, leading to performance overestimation. For example, we observe that TCL can get significant performance gains by expanding the target class of “person” to its sub-concepts (e.g., man, woman, worker, rider, etc.), but the kinds of class name-based tricks are not allowed in our unified evaluation protocol because the expansion depends on the target class names. With this consideration, we evaluate models using unified class names from the default version of MMSegmentation without class name-based tricks. Dense CRF is not used identically due to its expensive computational cost. All other evaluation settings follow GroupViT , where the input image is resized to have a shorter side of 448. We employ mean intersection-over-union (mIoU) as a performance metric, which is a standard metric in semantic segmentation. While we aim to provide a fair comparison, defining fair conditions can be subjective. Thus, we provide further results and discussion on this topic in Appendix C, especially regarding dataset scale and refinement methods.

Benchmark datasets and comparison methods.

We provide an extensive evaluation on widely used 8 benchmarks, categorized into two groups: (ii) with background class (PASCAL VOC , PASCAL Context , and COCO-Object ), and (iiii) without background class (PASCAL VOC20 , PASCAL Context59 , COCO-Stuff , Cityscapes , and ADE20K ). Note that open-world segmentation methods rely on the textual description of class names, which may require additional considerations for the background class, such as probability thresholding instead of using the “background” description as is. The datasets with background class evaluate this aspect. We compare TCL with all existing open-sourced methods, including GroupViT , MaskCLIP , and ReCo under the unified protocol. We also include their variants in comparison baselines for an extensive comparison. Additional details and comparisons are given in Appendix E.

Implementation details.

For the grounder, we use the CLIP ViT-B/16 model where the size of input images is 224×224224\times 224 and the patch size is 16×1616\times 16. Following MaskCLIP , we modify the last attention layer of the CLIP image encoder to acquire the dense embedding representing local semantics. The grounding decoder consists of four gated convolution blocks with two upsampling interpolations, and we use pixel-adaptive mask refinement (PAMR) for mask refinement. Further details on the model architecture are provided in Appendix B. We use CC 3M and 12M datasets for training. The loss weights of λTCL=0.1,λarea=0.4,λtv=1.0\lambda_{\text{TCL}}=0.1,\lambda_{\text{area}}=0.4,\lambda_{\text{tv}}=1.0 are used. We train the model with a batch size of 10241024 and a learning rate of 7.5×10−57.5\times 10^{-5} for total 50,00050,000 iterations with 15,00015,000 warmup steps and cosine schedule. AdamW optimizer is used with a weight decay of 0.050.05.

2 Zero-shot Transfer to Semantic Segmentation

We extensively compare existing open-world semantic segmentation methods in Table 1 using the proposed unified protocol, including two checkpoints of GroupViT and two variants of MaskCLIP . Between the existing methods, GroupViT achieves the best average performance, particularly on object-oriented datasets such as VOC, VOC20, and COCO-Object. However, its performance tends to decrease when the target dataset is dominated by stuff classes. On the other hand, MaskCLIP performs the best on stuff-oriented datasets such as Context, Context59, and COCO-Stuff, benefiting from the large-scale pre-trained CLIP model. We conjecture it benefits from leveraging a large-scale pre-trained CLIP model. The refinement techniques proposed in MaskCLIP improve the average performance but significantly degrade it on Cityscapes (21.6→12.621.6\rightarrow 12.6), suggesting the limitation of the heuristic refinement methods. The significant performance degradation of MaskCLIP and ReCo between VOC20 and VOC may imply the need for consideration for background class.

TCL remarkably outperforms existing methods.

Although the performances of existing methods vary depending on the characteristics of the evaluation datasets, TCL outperforms all the other methods by large margins across all datasets as shown in Table 1. These results demonstrate that our TCL framework successfully addresses the alignment-level train-test discrepancy that exists in the previous methods by learning the region-level alignment. In addition, region-level alignment learning of TCL allows our model to learn the capability to distinguish the background region in a data-driven manner, thus, our method can address the background class without any heuristic post-processing that the previous methods typically rely on.

3 Qualitative Results

Fig. 4 illustrates the impact of the learned grounding decoder. Since we follow MaskCLIP modification, the results in “w/o DgD_{g}” rows can be regarded as the initial results of MaskCLIP before refinement. Despite the vast pre-training scale and remarkable zero-shot classification performance of CLIP , its grounding capability is limited because the learning objective targets image-level alignment (See “w/o DgD_{g}” rows). In contrast, the grounding decoder (DgD_{g}) learns the region-level alignment by TCL, resulting in more precise, finer, and less noisy generated masks (See “w/ DgD_{g}” rows).

Qualitative comparison.

We qualitatively compare the proposed method in Fig. 5. On the PASCAL VOC dataset (Fig. 5(a)), we observe various types of errors in each comparison method. The grouping procedure of GroupViT makes the segmentation results less noisy, but it also causes an incorrect segmentation of a large group. ReCo struggles with the segmentation of background regions due to the lack of consideration about the background class. MaskCLIP does not take this into account as well, but its refinement methods make the results less noisy. In addition, we present examples in the wild to show open-world segmentation capability in Fig. 5(b). We collect test samples containing visual concepts not included in conventional segmentation datasets (e.g., moon, sunset) or free-form texts (e.g., “two women and one man with a smiling snowman”). GroupViT tends to focus on the main object of the image and regard the other objects as background, which is consistent with its good performance in object-oriented datasets. Interestingly, in this qualitative comparison in the wild, we observe ReCo consistently outperforms MaskCLIP contrary to the quantitative results. We conjecture that this is because the refinement approach of ReCo is data-driven, while the refinement approach of MaskCLIP depends on heuristic post-processing, which may not guarantee general improvement. Compared to the baselines, TCL consistently generates more precise segmentation masks. These results demonstrate that our proposed method, which learns region-level alignment, improves the segmentation quality both in the evaluation dataset and in web images in the wild. Additional qualitative results are provided in Appendix H.

Additional analysis on failure cases and model behavior

are provided in Appendices F and G, respectively.

4 Ablation Studies

We investigate the impact of individual components of the proposed framework by ablation studies on the training split of the PASCAL VOC20 dataset. We use a short learning schedule with a batch size of 512 for total 40,00040,000 iterations including 10,00010,000 warmup steps.

LABEL:table:ablations-baseline-to-tcl presents cumulative ablation studies on the grounding decoder and the TCL losses. Our initial model before training based on MaskCLIP is referred to as the baseline (a), which modifies the last attention layer of the CLIP image encoder. When we add only the grounding decoder to the baseline without TCL loss (b), there is no improvement in performance. This suggests that training the decoder with the same CL loss as the pre-training (CLIP) does not enhance the localization capabilities. As shown in (c), the proposed framework becomes complete with TCL loss.

Impact of individual TCL losses.

The influence of each component of the proposed TCL loss and its effect on the segmentation performance compared to the conventional CL loss are shown in LABEL:table:ablations-gcl. Smooth regularization is used for all experiments in this table. The CL loss (d) is computed by applying attention pooling to the dense image embedding Vs\mathbf{V}^{s}. When comparing (d) and (c), the proposed TCL loss remarkably improves the segmentation performance (61.1→77.461.1\rightarrow 77.4). Image-level or feature-level TCL loss (e, f) solely improves the performance significantly, and using both losses together provides further performance gain. Using CL in addition to TCL (g) does not improve performance, and it is essential to use area TCL loss in TCL framework to prevent model collapse (h), as described in Sec. 3.3. The difference between (b) and (d) is the use of smooth regularization.

Hyperparameters.

LABEL:table:ablations-TCL, LABEL:table:ablations-area and LABEL:table:ablations-tv shows the performance changes according to the variation of the loss weight hyperparameters (HPs). The first rows show the importance of each loss (λ=0.0\lambda=0.0 cases). The absence of area TCL loss causes a significant performance drop (LABEL:table:ablations-area), as mentioned above. Smooth regularization also significantly contributes to the final performance (LABEL:table:ablations-tv), supporting our assumption that the text-described region is smooth rather than noisy. Note that the sensitivity on HPs is about loss balancing, not about the target dataset. As an open-world segmentation method, TCL does not require any tuning with the target dataset, including model fine-tuning and inference HPs tuning. Once a TCL model is trained, we evaluate the model for every benchmark without any fine-tuning.

Conclusion

We propose a novel framework for open-world semantic segmentation with only image-text pairs, addressing the alignment-level discrepancy between training (image-text) and testing (region-text) in existing methods. In the proposed framework, we incorporate the grounding process within contrastive learning, thus allowing explicitly learning alignment between text and text-grounded regions (i.e., segmentation mask). We also present a unified evaluation protocol for a fair comparison of existing methods, where TCL achieves state-of-the-art zero-shot segmentation performance on all 8 benchmarks, remarkably surpassing previous methods. We hope that this study encourages a new research direction of explicitly learning region-text alignment for open-world semantic segmentation.

References

Appendix A Pseudo-code

For clarity, we present the pseudo-code for the core implementation of TCL in Fig. 6. As described in Sec. 3, we first generate text-grounded masks via grounder and use them to compute text-grounded images and embeddings (L10-L17). Furthermore, this pseudo-code also demonstrates our efficiency-aware design of TCL. We compute the B×BB\times B mask, where BB is the batch size, using a single einsum operation and scalar projection (L13). When computing TCLv loss, we need to perform CLIP image encoder inference again for the grounded images (L17). To reduce computational complexity, we only use the positive mask, which is the diagonal of the quadratic mask, for linear inference with a size of BB instead of quadratic (L15). When computing TCLf loss, we use the entire quadratic mask because a single einsum operation can efficiently compute the grounded embeddings without requiring additional encoder inference (L22). A more detailed discussion on the efficiency is provided in the later section (Appendix D).

Appendix B Architecture Details

Our core design principle is to preserve and exploit the diverse knowledge of pre-trained CLIP . Therefore, we freeze the pre-trained CLIP encodersAfter 30,00030,000 iterations, we unfreeze only the last block of the image encoder for richer model capability. and train the grounding decoder for the adaptation from the image-text alignment to the region-text alignment. We also considered the other techniques to preserve knowledge , but a simple freezing strategy worked the best. We follow the simple modification of MaskCLIP to the CLIP image encoder. They modify the last attention of the CLIP image encoder to acquire the dense features representing local semantics. These dense image features Vd\mathbf{V}^{d} are fed to the grounding decoder. As shown in Fig. 7, the grounding decoder consists of four gated convolution blocks, where the output of a convolution is gated by a learned gating parameter and added to the skip connection. Concretely, the process of gated convolution can be written as:

where x\mathbf{x} is input feature and gg is a learned gating parameter. The upsamplers increase the feature map resolution for the high-resolution segmentation capability. The first two upsamplers use the nearest neighbor interpolation and the last upsampler adapts the resulting embedding into the pixel-level embedding by the bilinear interpolation. In addition, as shown in Fig. 7, we employ two branches strategy: the main grounding decoder and knowledge preservation (KP) branches. In this KP branch, the CLIP dense features Vd\mathbf{V}^{d} are reshaped spatially and upsampled to pixel-level resolution by bilinear interpolation, and then we compute the text-grounded mask MKP\mathbf{M}^{\text{KP}} by Eq. 5. There are no learnable parameters in this branch and the output masks are just mixed with the output from the grounding decoder branch as follows:

where MKP\mathbf{M}^{\text{KP}} is the generated masks from knowledge preservation branch, M′\mathbf{M}^{\prime} is the final output mask, and wkpw_{kp} is a mixing hyperparameter. We use the wkpw_{kp} of 0.3. To fully leverage the massive pre-trained knowledge of CLIP , this branch is only used in the inference stage. It also can be regarded as a cost-free ensemble.

Appendix C Fair Comparison

In this paper, we present a unified evaluation protocol to facilitate a fair and rigorous comparison. However, the condition of fair evaluation protocol can be controversial. Thus, in this section, we provide additional comparisons to enhance fairness, paving the way for future research on fair comparison in open-world segmentation.

In our unified evaluation protocol, we do not unify refinement methods, as we believe that each method’s approach to refining the model output is a design choice. However, some may argue that a fair comparison protocol should unify the refinement methods as well. To address this concern, we also provide the comparison with and without PAMR , which is our refinement method used in TCL. As shown in Table 3, even using the same refinement method across all methods, TCL still achieves state-of-the-art performance with a significant margin in both settings, demonstrating the effectiveness of its underlying approach. It is also worth noting that the performance gains resulting from PAMR are specific to each method. For example, TCL and ReCo demonstrate significant performance improvements of +3.3+3.3 and +2.1+2.1 mIoU, respectively, while GroupViT only shows a marginal gain of +0.6+0.6 mIoU, and MaskCLIP actually leads to a decrease in performance of −1.5-1.5 mIoU. We believe that this is because each comparison method was designed without considering refinement by PAMR. Thus, because the effectiveness of the refinement method is closely tied to the main method, we have not unified the use of refinement methods in our proposed evaluation protocol.

Fair comparison in dataset scale.

Except for GroupViT, all comparison methods (TCL, ReCo, and MaskCLIP) utilize CLIP pre-trained models. GroupViT proposes a new encoder architecture and is therefore unable to leverage CLIP pre-trained models directly, which is one of its limitations. Hence, a fair comparison would be to evaluate GroupViT as it is. Nevertheless, it could be argued that comparing GroupViT and other methods at the same training dataset scale would be a fair comparison. To address this concern, we provide additional scale-up experiments. As mentioned, GroupViT is unable to leverage CLIP pre-trained models directly. Thus, for a fair comparison, we train GroupViT using the publicly available Coyo700M dataset of large-scale image-text pairs , which is larger than the non-public CLIP training dataset (WIT400M). However, we observe that a simple scale-up of the training dataset does not guarantee improvement in performance. As shown in Table 4, the correlation between dataset size and performance is not clear. The impact of larger datasets varies between the datasets (e.g., “CC15M+RedCaps12M” performs best on VOC20, but “CC15M+Coyo100M” performs best on Context59), and TCL outperforms all variants of GroupViT despite the fair dataset scale.

Appendix D Efficiency Analysis

While efficiency is not the primary objective of this study, it is also considered one of our core design principles, especially regarding inference efficiency for practical applications. This section provides an analysis of the inference throughput and our design choices for efficiency.

We benchmark the inference speeds and FPS using 448×448448\times 448 images and 2121 target classes on a single NVIDIA V100 GPU. Our benchmark setting follows an open-world scenario that addresses an arbitrary class, meaning that the text embeddings are computed for every inference. As shown in Table 5, TCL has slightly lower FPS compared to GroupViT or MaskCLIP, due to the relatively high resolution of segmentation masks. This trade-off between mask resolution and FPS can be controlled by the design of the grounding decoder, but we do not investigate this further, as it is outside the scope of this study. ReCo shows very low FPS due to its retrieval process, which involves retrieving similar images from the ImageNet dataset and co-segmenting them. Note that there are many methods to reduce the computational cost of TCL, e.g., utilizing depthwise convolution instead of standard convolution , but it is not the focus of this study and is left for future work.

Efficiency-aware designs in TCL.

Appendix E Additional Details and Experiments

TCL implicitly assumes that CLIP can address masked images robustly. While CLIP is widely known for its strong robustness , it is not yet clear how well CLIP can handle masked images. As such, we investigate CLIP scores between masked images and positive texts in Fig. 8. The results indicate that CLIP can address masked images. Interestingly, the scores tend to be improved, especially for the images with complex context (columns 1, 3, and 4 of Fig. 8). The masks remove noisy context and help to focus on the target object, leading to improved scores.

Details on the comparison methods.

In the quantitative evaluation, we include the variants of the comparison baselines for an extensive comparison; the GroupViT variants from the YFCC and RedCaps checkpoints and the MaskCLIP variants by the refinement process (key smoothing and prompt denoising). For the backbone of MaskCLIP, we use ViT-B/16 since its reported performance is better than ResNet-50 . For the qualitative comparison, we choose the quantitatively best variant for each method.

Comparisons with zero-shot semantic segmentation methods.

We provide an extensive and unified comparison in Table 1, but the comparison does not include non-open-sourced methods. To the best of our knowledge, ViL-Seg is the only non-open-sourced method for open-world semantic segmentation. We compare the zero-shot segmentation performance following the evaluation protocol of ViL-Seg. In this evaluation protocol for zero-shot semantic segmentation, only partial classes are used: 5 classes (potted plant, sheep, sofa, train, tv-monitor) for PASCAL VOC20, 4 classes (cow, motorbike, sofa, cat) for PASCAL Context59, and 15 classes (frisbee, skateboard, cardboard, carrot, scissors, suitcase, giraffe, cow, road, wall concrete, tree, grass, river, clouds, playingfield) for the COCO-Stuff dataset. Therefore, we additionally provide the comparison results under the partial classes protocol. As shown in Table 6, TCL achieves state-of-the-art performance with a large margin in every dataset again.

Appendix F Case Study on Failures

We investigate the failure cases of TCL via various qualitative examples. First, TCL undergoes difficulty in capturing segment boundaries accurately. For example, in Fig. 9(b), the predicted segment of the “mountain” class includes part of the sky region, and the “cell phone” segment contains the right hand and arm regions. This is a fundamental challenge of unsupervised open-world segmentation; the absence of dense annotation makes precisely capturing a segment boundary extremely difficult. Although the proposed method remarkably improves the segmentation performance compared with previous methods, this case study reveals that there are still many areas to be improved. Furthermore, despite the help of the smooth prior loss, the predictions still tend to be noisy, e.g., “sea” or “hair drier” in Fig. 9(b).

Ambiguity in benchmarks.

On the other hand, we also find crucial issues in the current benchmark datasets: ambiguities in the class label set and scene semantics. In particular, there are lots of class labels with similar semantics, especially in the datasets with a large vocabulary, e.g., COCO-Stuff (171 classes) or ADE20K (150 classes) . Mostly the labels have different semantics in detail, but the distinction between the labels can be ambiguous depending on how the image captures the scene. For example, it is hard to distinguish “clouds” and “fog” in Fig. 9(a) and “hill” and “mountain” in Fig. 9(b). Also, there are labels with superset-subset relations. For instance, the COCO-Stuff dataset has “broccoli”, “vegetable”, and “food-other” classes. In the supervised setting, a model can address this issue by training only if there is labeling consistency between images. However, in the open-world scenario, such superset-subset relations cause significant ambiguity. Furthermore, a more frequent ambiguity raises when a segment has multiple semantics. More proper descriptions of the “clouds” and “grass” segments in Fig. 9(a) are “foggy or cloudy mountain” and “bushes on the grass”, respectively. However, ground truth (GT) labels represent only part of the entire semantics. As with the superset-subset relation case, benchmarks for the open-world scenario require additional consideration to address such ambiguities. In this study, we propose a unified evaluation protocol to compare the existing methods fairly, but it only unifies the evaluation protocol and simply employs existing benchmark datasets. This analysis suggests the need for further advanced benchmarks dedicated to open-world scenarios in the future.

Appendix G Analysis on Model Behavior

In this section, we investigate how the learned TCL model generates different segmentation masks depending on the input text prompts. As shown in Fig. 10, the model tends to capture the intended region better when the input text prompt is more specific. Although this characteristic can cause performance degradation when evaluating the model on a fixed benchmark, it also improves the controllability of the model. We can exploit this controllability to maximize the benchmark performance, e.g., class name expansion. However, we do not employ these dataset-dependent tricks to prevent the overestimation of the model performance, as described in Sec. 4.1.

Appendix H Additional Qualitative Results

Qualitative comparison on PASCAL VOC in Sec. 4.3 visualizes the performance difference between comparison methods. However, the VOC dataset tends to be object-oriented and its images are generally composed of one or two segments. In this section, we qualitatively compare open-world segmentation methods in complicated scenes. Figs. 11 and 12 show examples including at least 3 segments from the Cityscapes and COCO-Stuff datasets. As shown in the figures, GroupViT and MaskCLIP tend to generate a small number of segments. It makes the results less noisy but causes a large error. For example, in Fig. 11, MaskCLIP fails to segment “building” regions and GroupViT misidentifies “road” as “traffic light”. In contrast, ReCo suffers from noisy prediction. Our TCL also generates partially noisy results, but it is relatively cleaner and better than the comparison baselines. It is also worth noting that the image resolution of the unified evaluation protocol is relatively smaller than the widely used protocols for Cityscapes. We resize a shorter side of an image to 448448 with keeping the aspect ratio, resulting in 448×896448\times 896. In contrast, 1024×20481024\times 2048 is the widely used resolution for the Cityscapes dataset . Increasing the resolution can help recognize small objects, e.g., the persons in Fig. 11.

H.2 Additional Qualitative Examples in the Wild

Fig. 13 shows additional qualitative examples from web images in the wild. In this experiment, we investigate the discrimination capability of the model in various aspects: proper nouns (Frodo, Gandalf, Pyramid, Sphinx, Samwise, Gollum, Taj Mahal, Batman, Superman), colors with the same object (red, green, yellow bananas), letters (MMU, Turkish, Fighter), and subclasses (Corgi, Shepherd). The results show that our model can recognize and segment various concepts in the wild. For the baseline models, the results show similar tendencies with the in-the-wild examples in Sec. 4.3. Contrary to the quantitative evaluation in fixed benchmarks, ReCo generates a relatively plausible segmentation map compared with GroupViT and MaskCLIP .