Learning Open-vocabulary Semantic Segmentation Models From Natural Language Supervision

Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, Weidi Xie

Introduction

Semantic segmentation considers the problem of assigning semantic labels to each pixel in the image. It plays central roles in a wide range of real-world scenarios, including autonomous driving, computer-aided diagnosis and satellite image analysis, to name a few. Generally speaking, two lines of research dominate semantic segmentation, one idea is to cluster the pixels into different groups and assign a semantic label to each group; the other idea treats segmentation as pixel-wise classification, casting each of the pixels into one category. Despite tremendous progress, the scalability of existing approaches that rely on supervised training has been fundamentally limited: (1) costly annotation procedure. Extensive manual pixel-wise annotations are required for training segmentation models; (2) closed-set segmentation. The model is restricted to segmenting objects from a closed-set of categories. Whenever a new dataset comes, the model requires fully-supervised re-training.

In this paper, our goal is to train an open-vocabulary semantic segmentation (OVS) model, by exploiting the freely available image-caption pairs on Internet, as illustrated in Fig. 1. The recent CLIP and ALIGN papers have demonstrated that a combination of large-scale image-caption pairs, with simple noise contrastive estimation can be used to learn powerful image-text embeddings from scratch, and show strong “zero-shot” generalization abilities for open-vocabulary classification. Recent works, such as GroupViT , initially extend the idea towards semantic segmentation by training a segmentation model with text supervision only. They perform hierarchical grouping of visual tokens, which are then aligned to the corresponding text embeddings via a contrastive loss. However, the following issues remain challenging and unsolved: First, the captions only provide coarse, image-level descriptions, which are insufficient for training semantic segmentation models where fine-grained, pixel-wise supervision is usually needed. Second, the diversity of web-collected data is large, that requires the model to learn visual invariance on objects of interest, with only weak supervision provided. For instance, the visual appearance of two images with similar captions can be drastically different.

To tackle the above challenges, (i) we propose a transformer-based model for open-vocabulary semantic segmentation, dubbed as OVSegmentor, that can segment objects of arbitrary categories via zero-shot transfer, with only image-caption pairs for pre-training. Specifically, we introduce learnable group tokens to cluster image patches via a slot-attention based binding module, and align the group tokens with the corresponding caption embedding. Note that our model neither needs ground-truth masks for training nor requires additional re-training on target segmentation datasets, substantially alleviating the annotation efforts and boosting transfer efficiency; (ii) As for training on the image-caption dataset, we propose two proxy tasks, namely masked entity completion and cross-image mask consistency, based on an observation that entities in the captions matter in matching specific visual objects. The former trains the model to infer all the masked entities in the sentence given the group tokens, and the latter enforces consistent mask prediction for images with the common entity. Both tasks have shown to be beneficial in learning entity-specific, fine-grained and visually invariant group semantics; (iii) We construct an image-caption dataset, termed as CC4M, by designing an automatic approach to filter CC12M with frequently appeared informative entities, significantly improving the training efficiency.

We pre-train the proposed OVSegmentor on our filtered image-caption dataset (CC4M), and manual segmentation masks are not used whatsoever. The model is evaluated on three segmentation benchmarks, PASCAL VOC 2012 , PASCAL Context and COCO in a zero-shot manner, i.e., the model is directly evaluated on downstream benchmarks without any finetuning. Extensive experiments demonstrate that our model surpasses the model with supervised finetuning and outperforms state-of-the-art methods on PASCAL VOC by using only 3% data (4M vs 134M) for pre-training, significantly improving the training efficiency.

Related Work

Vision-Language Pre-training. Vision-language pre-training (VLP) aims to learn joint visual-textual representations for a variety of multimodal downstream tasks. Existing works either learn unimodal encoders by distinguishing the positive pair(s) from the unpaired samples or focus on one multimodal encoder for joint feature learning with masked image/language modeling and image-text matching losses . Additionally, some approaches seek fine-grained supervision for cross-modal interaction . For example, GLIP proposed to align the bounding boxes with corresponding phrases in the text. However, they still rely on ground-truth grounding annotations. In contrast, our work explores fine-grained information with only weak supervision provided. Despite remarkable performance on multimodal downstream tasks, few of these vision-language models have been designed for fundamental vision tasks (e.g., semantic segmentation).

Zero-shot/Open-vocabulary Semantic Segmentation. The goal of zero-shot semantic segmentation is to segment objects of interest that are not seen in the training set. Prior works mainly transfer the knowledge from the training set (seen) to the testing set (unseen) via visual-semantic mapping. Inspired by the open-vocabulary nature of language, current approaches exploit vision-language models (e.g., CLIP ) pre-trained on large-scale image-caption pairs. They need either finetuning or self-training on the target segmentation dataset, which is less efficient and flexible than zero-shot transfer. VLP for open-vocabulary segmentation combines the merits of both zero-shot transfer and open-vocabulary recognition. GroupViT designed a grouping vision transformer and learned the alignment between groups and text via the contrastive loss. ViL-Seg combined image-text contrastive loss with online pixel clustering for segmentation. A concurrent work CLIPpy explored different aggregation operations for training spatial-aware vision-language models for segmentation. Beyond the global image-text matching, we further design masked entity completion and cross-image mask consistency to enrich the group semantics.

Fully-/Weakly-/Semi-Supervised Semantic Segmentation. Fully-supervised semantic segmentation emerges from per-pixel classification to mask classification . To relieve the laborious annotations, extensive efforts have been made to address semantic segmentation with less supervision. In the family of weakly-supervised object localization and semantic segmentation , only class labels are available for supervision. Generally, the class activation maps derived from the classification network serve as the initial segmentation results. Another line of research focuses on semi-supervised semantic segmentation where a few samples have dense per-pixel labels and the remaining samples are unlabeled. These works mainly perform supervised training on labeled samples with additional consistency regularizations posed on unlabeled samples. Despite promising results, these approaches are still limited to closed-set object categories.

Methodology

We present OVSegmentor, a vision-language pre-training framework for OVS. It assembles the image patches into groups and aligns the group to human-understandable categories, by only exploiting weak supervisions from web-collected image-caption pairs. Conceptually, the architecture consists of two stages, namely, a pixel-to-group binding that assigns all pixels with same semantics into one group, and group-to-category alignment that computes matching scores between each of these group tokens with semantic categories, the segmentation mask mm can be computed as:

The visual encoder consists of two components, namely, Transformer encoders and binding modules. Specifically, the image tokens and learnable group tokens are concatenated, and iteratively processed by the Transformer encoders, and a binding module with slot-attention being adopted for grouping. The visual encoder is defined as:

Transformer Encoder. Both Φenc1\Phi_{\text{enc}}^{1} and Φenc2\Phi_{\text{enc}}^{2} consist of 6 Transformer encoder layers , where each layer is composed of a multi-head self-attention (MHSA) layer followed by layer normalisation (LN) and a feed-forward network (FFN). Φenc1\Phi_{\text{enc}}^{1} takes the concatenation of the image patches and the randomly initialised group tokens as input, and outputs intermediate encoded group and image tokens (G′G^{\prime} and I′I^{\prime}); Φenc2\Phi_{\text{enc}}^{2} processes the output from the binding module.

Binding Module. The binding module uses slot-attention to cluster image tokens into groups in a data-dependent manner, i.e., image patches with similar appearance and semantics are encouraged to be grouped together. Formally, the binding module accepts the output from the Transformer encoder Φenc1\Phi_{\text{enc}}^{1}, and transforms them into query, key and values with linear transformations:

In contrast to the standard cross attention in Transformer Decoders , slot-attention performs normalisation over queries, encouraging each image token to be claimed by one of the group tokens. The binding process can be defined as:

where WobindW_{o}^{\text{bind}} is a linear transformation. Now we have obtained the correspondence between each pixel and group tokens, next we describe the procedure for encoding captions.

1.2 Text Encoder

Till this end, we start by filtering all the captions and only keep the ones with informative entities, followed by exploiting three variants of the caption encoding, namely, the entire caption embedding, the masked caption embedding and the prompted entity embedding. In all cases we use a pre-trained BERT as the text encoder Φtext\Phi_{\text{text}}.

Constructing Entity Set. We adopt the nltk toolkit to extract entities from all the captions, and construct an entity set Ω=Φentity({T1,…,TN})\Omega=\Phi_{\text{entity}}(\{T_{1},\dots,T_{N}\}) that only maintains the frequently appeared entities (e.g., people, cat, shirt, etc.), and exclude the abstract nouns (e.g., art, view, etc) as they usually do not correspond to any specific region in the image. For each image-caption pair, we can thus obtain an image-caption-entity triplet (I,T,E)(I,T,E), where E={e∣e∈T∩e∈Ω}E=\{e|e\in T\enspace\cap\enspace e\in\Omega\} includes all frequent entities in the caption.

2 Training

As for training, we aim to learn the alignment between group tokens and caption embeddings via three proxy tasks, namely, image-caption alignment, masked entity completion, and cross-image mask consistency.

Image-caption Alignment. For each image-text pair, the objective is to align their visual and textual embeddings. The visual embedding zIz^{\text{I}} is the average of group tokens, and the textual embedding zTz^{\text{T}} is obtained by taking the [EOT] token feature of the caption embedding Tcap\mathcal{T}^{\text{cap}}, both projected to a 256-d joint feature space followed by normalisation. The image-caption contrastive loss Lcontrast\mathcal{L}_{\text{contrast}} is formulated as:

Here, we omit the temperature parameter for simplicity.

Masked Entity Completion. The goal of masked entity completion is to infer all the masked entities in the sentence given the group tokens. In specific, we adopt a Transformer decoder layer, where a projection of the masked caption embedding is treated as query, and two linear transformations of group tokens are treated as key and values, respectively.

Intuitively, the entity completion task enables better alignment between the groups and entities.

where σ\sigma is the sigmoid activation; both M1={mk1}k=1K′\mathcal{M}_{1}=\{m_{k}^{1}\}_{k=1}^{K^{\prime}} and M^1={m^k1}k=1K′\hat{\mathcal{M}}_{1}=\{\hat{m}_{k}^{1}\}_{k=1}^{K^{\prime}} consist of K′K^{\prime} unordered masks.

To align M^1\hat{\mathcal{M}}_{1} with M1\mathcal{M}_{1}, we first adopt the bipartite matching to find the optimal permutation p1∗p^{*}_{1} over K′K^{\prime} subgroups with the lowest matching cost as:

where P\mathcal{P} is the full permutation and cos⁡(⋅)\cos(\cdot) denotes the cosine similarity. Eq. 11 is solved via the efficient Hungarian algorithm . In this way, the symmetric cross-image mask consistency loss Lmask\mathcal{L}_{mask} is defined as:

where sg(⋅\cdot) denotes the stop gradient operation; the target mask mk\textbf{m}_{k} is achieved by binarizing mkm_{k} with a threshold δ\delta; D(m,m^)=1−2∣m⋂m^∣/(∣m∣+∣m^∣)\text{D}(\textbf{m},\hat{m})=1-2|\textbf{m}\bigcap\hat{m}|/(|\textbf{m}|+|\hat{m}|) stands for the standard Dice loss. To guarantee the quality of pseudo mask targets, the masks are generated by an extra momentum model, which is updated by the exponential-moving-average (EMA) of the online model.

Training Objective. We adopt a combination of three different loss functions:

where λ\lambda is the weight for balancing the mask consistency.

3 Discussion

One work that is closely related to ours is GroupViT , that improved the image-text alignment by exploiting nouns in the caption. In specific, they extracted multiple nouns and prompted each noun to a sentence to serve as extra matched captions for the image. The model is thus supervised by a multi-label contrastive loss. In contrast, our paper differs from GroupViT from three critical aspects: (1) Entities vs nouns. Rather than using all nouns, we leverage the entities that match to visual objects, enabling high-quality image-caption correspondence. (2) Network architecture. Beyond separate visual and text encoders in GroupViT, we further devise a (very) lightweight decoder to model the fine-grained, token-wise group-word correlation. (3) Proxy tasks for training. We propose two different proxy tasks, i.e., masked entity completion and cross-image mask consistency to improve the entity-specific group semantics and further encourage visual invariance. The superiority of masked entity completion over multi-label contrastive loss is verified in Sec. 4.3.

Experiments

Pre-training Dataset. Following , we use Conceptual Captions 12M for training, which is originally constructed with over 12M image-text pairs collected from the Internet. However, due to some links have been expired, we have downloaded about 10M image-text pairs. The constructed entity set in Sec. 3.1.2 includes a total number of 100 frequently appeared entities while abstract nouns (e.g., art, view) are discarded. After filtering CC12M, we obtain 4.3 million image-text pairs for pre-training, which is termed as CC4M. Examples of entities include people, car, cup, chair, T-shirt, house, bed, cat, ball, pizza, etc. Please refer to the supplementary material for the full entity set.

Downstream Evaluation Datasets. We evaluate our model on three benchmarks, namely, PASCAL VOC 2012 , PASCAL Context and COCO Object with 20, 59 and 80 foreground classes, respectively. An extra background class is considered in all three datasets. We ignore their training sets and directly evaluate our method on the validation sets without any finetuning, including 1449, 5105, and 5000 images, respectively. In general, we report the mean Intersection-over-Union (mIoU) on all the classes.

Implementation Details. In our model, the self-attention layers in the visual encoder are initialised with DINO pre-trained on ImageNet. The text encoder is initialised with BERT model pre-trained on BookCorpus and English Wikipedia. Our decoder with one randomly initialised Transformer Decoder layer performs reasonably well. The input image is randomly cropped to 224×\times224 at training time, and the batch size is set to 2048 with an initial learning rate 3.2×10−43.2\times 10^{-4}. We train our model for 40 epochs using the Adam optimizer with weight decay set to 0.5. The coefficient for updating the momentum model is 0.99. As the generated masks are unreliable in early epochs, we set the mask consistency coefficient λ\lambda=0 for the first 30 epochs and λ\lambda=0.1 for the remaining epochs. The group selection ratio rr is 0.5. As for the threshold in mask consistency loss, we use δ\delta=0.65. At inference time, the image is resized with a shorter length of 448. We follow to set a threshold for the background class, which is 0.90.9, 0.50.5 and 0.90.9 on PASCAL VOC, PASCAL Context and COCO Object, respectively.

2 Comparison with Existing Methods

In Table 1, we compare our model with the existing models that have been trained with (1) fully-supervised finetuning transfer and (2) zero-shot transfer. In Table 2, zero-shot segmentation (ZSS) approaches are listed for comparison.

Comparison with Finetuning Transfer. We compare our method with DeiT , MoCo and DINO , which are pre-trained on ImageNet or CC12M+YFCC15M datasets with class labels or self-supervision , and finetuned on the training set from downstream benchmarks, with a randomly initialised convolution head appended on the backbone network. As shown in Table 1, our model achieves competitive performance on PASCAL Context, and outperforms the self-supervised methods by over 10% on PASCAL VOC, with zero-shot transfer.

Comparison with Zero-shot Transfer. Here, we compare with existing works under the zero-shot transfer scenario, including GroupViT , ViL-Seg and CLIPpy , with the pre-training data ranging from CC12M to 134M in-house dataset HQITP-134M . For fair comparison, we re-train GroupViT with their official codebase on our CC4M and CC12M datasets, with the same pre-trained weights as ours being adopted on CC4M. However, we observe no further performance gain of either applying pre-trained weights to GroupViT or using ViT-B as the backbone on CC12M. The results on CC12M match the reported ones (40.2 vs 41.1). Under the same pre-trained ViT-B backbone and CC4M dataset, our method surpasses the original GroupViT by 35.7%. Additionally, by only using 4M pre-training data, our model yields the best segmentation performance on PASCAL VOC, even outperforming CLIPpy, which is a concurrent work to ours, and pre-trained on 134M data, indicating the effectiveness and training efficiency of our proposed model.

Comparison with Zero-shot Segmentation Methods. In this line of research , the idea is to train the model with full mask labels (obtained either from manual groundtruth or pseudo-labelling ) on the seen classes and transfer the model to unseen classes, and the task is thus dubbed as zero-shot semantic segmentation (ZSS). To be specific, 5 classes (potted plant, sheep, sofa, train and tv-monitor) in PASCAL VOC and 4 classes (cow, motorbike, sofa and cat) in PASCAL Context are considered unseen while the remaining classes belong to seen. For comparison, we also report the zero-shot transfer performance of our model on these unseen classes, however, note that, we do not use any manual mask annotations for training. As observed in Table 2, our model surpasses majority of the models trained under the ZSS scenario, except for that adopted CLIP model pre-trained on 400M data. In terms of transfer efficiency, our model excels at zero-shot transfer ability without the need of training on seen classes.

3 Ablation Study

In this section, we conduct thorough ablation studies to validate the necessity of each proposed component.

Ablation Study on Proxy Tasks. Here, we aim to understand the effects of our proposed proxy tasks, i.e., masked entity completion and cross-image mask consistency. As shown in Table 3, the baseline model uses the image-text contrastive loss Lcontrast\mathcal{L}_{\text{contrast}} only, while adding the entity completion task, Lentity\mathcal{L}_{\text{entity}}, we observe a significant improvement by 8.4% and 4.8% on PASCAL VOC and PASCAL Context, respectively. The performance gain is due to the ability of better aligning the pixel groups with the visual entities. Additionally, the mask consistency also brings improvements, and combing both leads to the best performance. Qualitative results can be seen in Fig. 3. We refer the readers for more visualisations in the supplementary material.

On the Choice of Masking Objectives. We compare our masked entity completion with a series of variants as shown in Table 4. Our proposed objective of masking all entities is listed in the first row. (1) All entities vs one entity: the masked language modeling (MLM) in prior works normally choose 15% of the token positions in the sentence for prediction, which results in one entity in most of our cases. We observe that masking all entities is 2.9% mIoU better than single entity masking, as it forces the network to infer all possible object categories in the image, that is potentially beneficial for the group to category alignment. (2) Entities vs nouns: masking and predicting noun phrases in the sentence is one feasible option to learn fine-grained vision-text matching. However, noun masking leads to 3.5% lower mIoU than our entity masking strategy. This is because not all of the nouns in the sentence are visually corresponding to the objects in the image (e.g., illustration, night, etc), thus the group tokens are not expected to align with these nouns. Our method avoids this issue by only masking the visual entities. (3) Masked entity completion vs multi-label contrastive loss: comparing with the multi-label contrastive loss used in GroupViT , our proposed strategy shows superior performance. (4) Masked entity completion vs masked language modeling: MLM originally predicts the masked token over the entire vocabulary via a cross-entropy loss. Here, we restrict the vocabulary to our constructed entity set for fair comparison. Our masking strategy surpasses both MLM variants by a clear margin, which we conjecture is because: (1) MLM classifies each masked token individually, and the model can easily refer to the context words without relating to groups. (2) MLM focuses on word-level representations, which is not in accordance with the sentence-level representation of the class embeddings we used during inference. Fig. 4 shows the effect of the masked entity completion in improving visual grouping (left) and group-text alignment (right).

On Cross-image Mask Consistency. Here, we analyze two key factors in cross-image consistency: (1) group numbers and selection ratios. As shown in Table 6, 8 groups perform comparably well for all selection ratios, and we pick 0.5 as the default. Smaller ratios miss the entities encoded in remaining groups while larger ratios introduce entity-irrelevant information (e.g., background). However, while increasing the group number to 16, it brings over-segmentation for large objects, deteriorating the performance. (2) the objective for cross-image consistency. We study another variant of cross-image consistency by directly aligning two sets of group tokens encoded with shared entity. We adopt the contrastive loss (NCE in Table 6) to pull two sets of group tokens closer, while other group tokens in each mini-batch are pushed farther. Table 6 reveals the superiority of our proposed mask consistency over group consistency, which we believe is because mask consistency involves the image content to realize visual invariance.

Performance on Unseen Entities. One might question whether the performance gain is mainly attributed to the selected (seen) entities in the entity set. We measure the mean IoU for 65 objects within entity set (e.g., person, bus) and 16 objects out of entity set (e.g., frisbee, stop sign) on COCO. As shown in Table 7, the mIoU of unseen entities is comparable to that of seen entities, indicating our model retains strong open-vocabulary segmentation ability without being affected by the choice of entities in pre-training. The text encoder pre-trained on large text corpus remains its ability to encode the semantic concept of objects out of entity set.

Conclusion

In this paper, we present OVSegmentor, a transformer-based model for open-vocabulary semantic segmentation. The model exploits web-collected image-caption pairs for pre-training without any mask annotations, and transfers to target benchmark segmentation datasets (including PASCAL VOC, PASCAL Context and COCO Object) in a zero-shot manner. The model clusters the image pixels into learnable group tokens, which are then aligned with the corresponding caption embeddings. We further devise two proxy tasks, namely masked entity completion and cross-image mask consistency, to learn entity-specific, fine-grained and visually invariant group semantics. OVSegmentor outperforms the state-of-the-art method on PASCAL VOC by using only 3% (4M vs 134M) for pre-training, indicating the effectiveness and training efficiency of our model.

References

Appendix A Additional Experiments

The constructed entity set contains 100 frequently appeared entities, including: people, man, men, woman, women, girl, boy, lady, kid, child, children, baby, student, bride, groom, couple, prince, princess, car, bus, truck, motorcycle, train, bicycle, boat, aeroplane, airplane, motorbike, bike, cup, bottle, bowl, knife, spoon, glass, fork, chair, table, bench, clock, laptop, light, vase, plant, remote, microwave, toaster, oven, mouse, keyboard, sofa, monitor, desk, tv, TV, couch, flower, refrigerator, house, building, hotel, handbag, umbrella, book, backpack, phone, shirt, tie, suitcase, T-shirt, bag, box, sink, bed, toilet, cat, dog, horse, bird, cow, sheep, elephant, bear, zebra, giraffe, ball, racket, skateboard, skis, snowboard, surfboard, kite, pizza, cake, apple, banana, sandwich, orange, carrot, donut. Note that, we exclude the word “person” in the entity set as CC12M claimed that they performed person-name substitutions to protect the privacy of the individuals in the images, specifically, all named entities of type Person (e.g., the name of the artist) detected by the natural language APIs are replaced with “person”.

A.2 Additional Ablation Studies

Effect of the Pre-trained Backbones. We show the effect of applying different unimodal/multimodal pre-trained weights for visual and textual encoders in Table 8, with Lcontrast\mathcal{L}_{\text{contrast}} being adopted only. Training both encoders from scratch only achieves 28.8 mIoU on PASCAL VOC. Initialization from CLIP visual and text encoders (including the visual/textual projection heads) brings significant improvement. However, it requires 400M image-text pairs for pre-training. Besides, a potential drawback of applying CLIP pre-trained weights is that the model can easily learn the visual-text alignment while ignoring the visual grouping. In comparison, initializing the model from single-modality sources, i.e. DINO and BERT, yields better performance. This design choice requires no manual annotation as both DINO and BERT use self-supervised training.

On the Choice of Mask Threshold. Here, we study the influence of different mask thresholds δ\delta as mentioned in Sec.3.2 in the manuscript. As observed in Table 9, our model reaches a decent mIoU of 53.6 on PASCAL VOC when δ\delta is 0.6, while smaller thresholds lead to false-positive pixels of the objects.

Effect of the Momentum Model. Our proposed OVSegmentor adopts a momentum model for encoding the cross-image, which is updated by the exponential-moving-average (EMA) of the online model. Table 10 reveals that applying the momentum model brings about 2% mIoU gain on PASCAL VOC and COCO Object. We attribute this to the improved quality of the pseudo targets generated by the momentum model. In Fig. 5, we also show the object masks generated by our online model M^1,M^2\hat{M}_{1},\hat{M}_{2} and momentum model M1,M2M_{1},M_{2} for both the input image I1I_{1} and the sampled cross-image I2I_{2} with the shared entity.

Mask Probing. Following DINO and GroupViT , we evaluate the quality of the generated masks regardless of the class predictions, termed as mask probing. Mask probing directly reflects the effect of the pixel-to-group assignment in our proposed model. For ViT-based methods that adopt finetuning transfer, i.e., DeiT , MoCo and DINO , the self-attention maps in the last ViT block are probed.

Per-class Segmentation Performance. We compare the mIoU over total 20 object categories in PASCAL VOC, as shown in Table 12. Our proposed OVSegmentor surpasses VIL-Seg on all the categories, while significantly outperforming GroupViT on categories such as aeroplane, car, motorbike. OVSegmentor achieves inferior results on the “person” class, owing to its large variation of visual appearance in web-collected images, posing additional challenges for our proposed cross-image mask consistency to learn visual invariance.

Appendix B More Visualization Results

Additional qualitative results on PASCAL VOC, PASCAL Context, and COCO Object can be found in Fig. 6, Fig. 7, and Fig. 8, respectively. Generally, our proposed OVSegmentor successfully groups semantically related pixels together and aligns the group to the correct category. On PASCAL VOC, OVSegmentor successfully segments objects with various scales (e.g. small aeroplanes and distant cars in the 2nd{}^{\text{nd}}, 5th{}^{\text{th}} and 7th{}^{\text{th}} rows) and multiple objects of the same class (4th{}^{\text{th}}, 6th{}^{\text{th}} and 8th{}^{\text{th}} rows). In terms of PASCAL Context where objects of more categories are annotated, our model manages to segment the salient objects while failing to recognize stuff classes that usually appear as the background in web-collected data (e.g. grass, floor, wall, etc.). On COCO Object, we observe that our model can not separate co-occurring objects from different classes into distinct groups very well (e.g. laptop and mouse), which we conjecture is because the captions sourced from the Internet usually lack fine-grained descriptions to cover the full image content.