A Simple Framework for Open-Vocabulary Segmentation and Detection

Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianfeng Gao, Jianwei Yang, Lei Zhang

Introduction

Developing vision systems that can be transferable to novel concepts or domains has emerged as an important research topic in the community. In the light of strong zero-shot transferability demonstrated in the seminal work CLIP , a number of researchers have attempted to build advanced open-vocabulary models by leveraging large-scale image-text pairs for fine-grained vision tasks like detection and segmentation .

Arguably, core vision tasks like detection and segmentation are fairly distinct in their vocabulary sizes and spatial granularities of supervision, as illustrated in Fig. 2 (a). For example, the commonly used public detection dataset Objects365 contains box annotations for 365 concepts in around 1.7M images, while mask annotations in COCO cover merely 133 categories in 0.1M images. Previous works have explored different ways of leveraging a large amount of image-text data for open-vocabulary detection or segmentation, such as distilling the visual-semantic representations from multi-modal foundation models , designing fine-grained or augmented contrastive learning methods or utilizing pseudo-labeling techniques . To the best of our knowledge, most (if not all) of them focused on how to improve the performance for either detection or segmentation. Moreover, transferring weak image-level supervision to fine-grained tasks usually requires sophisticated designs to mitigate the huge granularity gap and is vulnerable to noises in image-text pairs. This leads to a natural question: can we bridge detection and segmentation that are cleaner and have a closer gap to attain a good open-vocabulary model for both?

Taking one step back, marrying detection and segmentation had been previously explored in two main ways. On one hand, Mask R-CNN is one of the first works that proposed to jointly learn detection and instance segmentation on COCO. On the other hand, it is shown that detection models pre-trained on Objects365 can be feasibly transferred for COCO panoptic segmentation . However, as depicted in Fig. 2 (b), the former method requires the model to be trained on the same dataset containing aligned box and mask annotations, while the latter method follows pre-train-then-fine-tune protocol, leading to two separate closed-set models. In this work, we are the first to propose jointly learning from detection and segmentation data, and more importantly serving an open-vocabulary model for both tasks (Fig. 2 (b) bottom). Achieving this goal requires answering two critical questions: ii) how to transfer the semantic knowledge across detection and segmentation data; iiii) how to bridge the gap between box and mask supervision. First, the vocabulary shares commons but also bear substantial differences between the two tasks. We need to accommodate the two vocabularies and further go beyond towards open vocabulary. Second, semantic and panoptic segmentation tasks require segmenting not only foreground objects (things like “dog” and “cat”.) but also background concepts (stuff like “sky” and “building”), while detection task solely cares about foreground objects. Third, box supervision by nature is coarser than mask supervision. We can convert masks into boxes but hardly vice versa.

To the end, we propose OpenSeeD, a simple encoder-decoder framework to reconcile the two tasks by mitigating the aforementioned problems. Concretely, we first exploit a single text encoder to encode all concepts occurring in the data and train our model to align the visual tokens with the semantics in a common space. Second, we explicitly divide the object queries in the decoder into two sub-types: foreground and background queries, where the first group is responsible for foreground objects from both segmentation and detection while the second group is only for background stuffs in segmentation. Third, we introduce conditioned mask decoding which learns to decode masks from ground-truth boxes from segmentation data and generates the mask assistant for detection data. As a result, our OpenSeeD is able to learn from separate detection and segmentation data seamlessly and achieves outstanding or competitive zero-shot and transfer performance across various tasks/datasets. Fig 1 shows a visualization of our model on instance, panoptic and semantic segmentation tasks. It also shows the segmentation results on datasets that largely differ from our training data such as the SeginW datasets and demonstrates the conditioned segmentation ability of OpenSeeD. Given the encouraging results, we hope our work can contribute as the first strong baseline for developing a single open-vocabulary model for both tasks.

Contributions. To summarize, our main contributions are:

We are the first to present a strong baseline model that can jointly learn from detection and segmentation data towards an open-vocabulary model for both tasks.

We locate the discrepancies in two tasks/datasets and propose separate techniques including shared semantic space, decoupled decoding, and conditioned mask assistance to mitigate the issues.

By jointly training our model on segmentation and detection data, we achieve new state-of-the-art segmentation performance for zero-shot and task transfer across a variety of datasets, and competitive performance for zero-shot object detection.

Related Work

Generic Segmentation and Detection. Detection and segmentation have been long-standing problems in the vision community . Both tasks require understanding what and where the visual concepts are but with different spatial granularities. Generic segmentation mainly includes instance, semantic and panoptic segmentation , with respect to different semantics. Recently, Detection Transformer (DETR) that is based on Transformer has achieved significant progress in many detection and segmentation models . However, all these methods are constrained to a limited vocabulary size. Open-Vocabulary Segmentation. Many open-vocabulary segmentation models leverages large pretrained vision-language models (e.g., CLIP or ALIGN) to distill or transfer the visual-semantic knowledge. Apart from using foundation models, DenseCLIP and GroupViT show that fine-tuning from a foundation model or training from scratch can also yield superior zero-shot performance. Recently, X-Decoder proposes to unify all types of segmentation tasks and several vision-language tasks for open-vocabulary segmentation. In ODISE , the authors study a new way of using a text-to-image diffusion model as the backbone for open-vocabulary segmentation. Unlike the previous works, our model instead explores connecting segmentation and detection which have cleaner data and closer gap between each other. Open-Vocabulary Detection. Similarly, some open-vocabulary detection models directly leverage foundation models for distillation or transfer like OV-DETR and VILD . Recently, GLIP formulates detection as a special grounding problem to unify detection and phrase grounding tasks. These grounding data help improve the alignment between phrases and regions for open detection. RegionCLIP and DetCLIP generate pseudo box labels from image-text pairs for more generalized detection. Weakly-Supervised Segmentation. Weakly-supervised segmentation typically only uses box annotation as supervision to generate segmentation. Prominent methods design teacher models or weak supervision loss, like BoxInst , Box2Mask , DiscoBox and Mask Auto-Labelers . All these models are with closed-set and usually inferior to models with segmentation supervision. In contrast, we attempt to leverage as much supervision as possible from both segmentation and detection for an open-vocabulary model. Learning from Box and Mask. There are primarily two ways to learn from both box and mask. The first one is to train on a single dataset with both box and mask annotations. Prominent methods include Mask R-CNN and HTC . However, they are constrained to foreground instances. The second way is to pretrain with only box supervision and then transfer to segmentation. For example, HTC and Mask DINO can both learn from large-scale detection data and then be fine-tuned to a specific segmentation dataset. However, such a pretrain-and-finetune protocol leads to two separate models that are only capable of either detection or segmentation. Moreover, both models are closed-set and thus not transferable to novel concepts.

Method

Given segmentation and detection datasets, OpenSeeD is aimed at learning an open-vocabulary model for both tasks. Formally, let Dm={Ii,(ci,mi)}i=1M\mathcal{D}_{m}=\{{I}_{i},(\mathbf{c}_{i},\mathbf{m}_{i})\}_{i=1}^{M} denote the segmentation dataset of size MM and Db={Ij,(cj,bj)}j=1N\mathcal{D}_{b}=\{I_{j},(\mathbf{c}_{j},\mathbf{b}_{j})\}_{j=1}^{N} the detection dataset of size NN, where c\mathbf{c} are the visual concepts in an image, and m\mathbf{m} and b\mathbf{b} the corresponding masks and boxes, respectively. Suppose V={c1,...cK}\mathcal{V}=\{c_{1},...c_{K}\} be the vocabulary of unique KK visual concepts appearing in Dm\mathcal{D}_{m} and Db\mathcal{D}_{b}. The goal of OpenSeeD is learning to detect and segment visual concepts in V\mathcal{V} and beyond.

To achieve the goal, we exploit a general encoder-decoder design and employ a text encoder for our OpenSeeD, as shown in Fig. 3. Our model takes as input an image II and the vocabulary V\mathcal{V} and output a set of predictions including masks Pm\mathbf{P^{m}}, boxes Pb\mathbf{P^{b}}, and classification scores Pc\mathbf{P^{c}}. As a whole, ⟨Pm,Pb,Pc⟩=OpenSeeD(I,V)\langle\mathbf{P^{m}},\mathbf{P^{b}},\mathbf{P^{c}}\rangle=\mathsf{OpenSeeD}(I,\mathcal{V}). More specifically, our model consists of one image encoder EncI\mathsf{Enc_{I}}, one text encoder EncT\mathsf{Enc_{T}}, and one decoder Dec\mathsf{Dec}. Given an image II and the vocabulary V\mathcal{V}, we first encode them by EncI\mathsf{Enc_{I}} and EncT\mathsf{Enc_{T}}, respectively:

where the image features O∈RH×W×C\mathbf{O}\in\mathcal{R}^{H\times W\times C}, and the text features T={t1,t2,...,tK}\mathbf{T}=\{t_{1},t_{2},...,t_{K}\}. Afterward, the decoder takes LL queries Q∈RL×C\mathbf{Q}\in\mathcal{R}^{L\times C} as inputs and cross-attends the image features to get outputs:

where Ps\mathbf{P^{s}} is the decoded semantics. The visual-semantic matching scores Pc\mathbf{P^{c}} is derived from Sim(Ps,T)\mathbf{Sim}(\mathbf{P^{s}},\mathbf{T}) by calculating the similarity scores between Ps\mathbf{P^{s}} and T\mathbf{T}, which is used to compute the loss during training and predict the category during inference.

In this basic formula, we attempt to reconcile the two tasks by promoting a shared visual-semantic space without touching other issues. For multiple tasks and datasets, our loss function can be written as follows.

For clarity, we omit the weight for each loss term. Note that for the segmentation task, we can derive accurate boxes b^\hat{\mathbf{b}} from masks m\mathbf{m} and use them to compute the box loss as in term Lb(Pb,b^)\mathcal{L}_{b}(\mathbf{P^{b}},\hat{\mathbf{b}}). By summing over all the terms, our model can achieve a reasonably good open-vocabulary performance. Furthermore, it can be pre-trained end-to-end with detection and segmentation data, allowing it to perform open-vocabulary segmentation and detection using a single set of weights.

Despite building a strong baseline, we must consider the intrinsic discrepancies between the two tasks, as previously discussed. Semantic and panoptic segmentation require the recognition of both foreground and background, while detection focuses solely on localizing foreground objects. As a result, using the same queries for both tasks creates conflicts that can significantly degrade performance. Additionally, good box predictions are typically indicative of good masks, and vice versa. Separately training the box and mask head on detection and segmentation data obstructs the synergy of spatial supervision from both datasets.

To address the aforementioned discrepancies, we introduce a new decoder design for our OpenSeeD. We divide the queries Q\mathbf{Q} into three types: LfL_{f} foreground queries Qf\mathbf{Q_{f}}, LbL_{b} background queries Qb\mathbf{Q_{b}} and LdL_{d} conditioned queries Qd\mathbf{Q_{d}}, and propose query-specific computations for each type. In the following, we will describe how we decouple the foreground and background decoding to address the task discrepancy in Sec. 3.2, and employ the conditioned mask decoding to tackle the data discrepancy in Sec. 3.3.

2 Bridge Task Gap: Decoupled Foreground and Background Decoding

Without loss of generality, we have defined the visual concepts that appear in instance segmentation and detection as foreground, while the stuff categories in panoptic segmentation are considered background. To mitigate the task discrepancy, we perform foreground and background decoding with foreground queries Qf\mathbf{Q_{f}} and background queries Qb\mathbf{Q_{b}}, respectively. Specifically, for these two query types, our decoder predicts two sets of outputs: ⟨Pfm,Pfb,Pfc⟩\langle\mathbf{P}^{m}_{f},\mathbf{P}^{b}_{f},\mathbf{P}^{c}_{f}\rangle and ⟨Pbm,Pbb,Pbc⟩\langle\mathbf{P}^{m}_{b},\mathbf{P}^{b}_{b},\mathbf{P}^{c}_{b}\rangle. We also divide the ground truths in segmentation dataset into two groups: (cf,mf)(\mathbf{c}_{f},\mathbf{m}_{f}) and (cb,mb)(\mathbf{c}_{b},\mathbf{m}_{b}), and then perform two independent Hungarian Matching processes for these two sets correspondingly, as shown in Fig.4 (a). Consequently, both foreground and background decoding are used for segmentation, while only foreground decoding is used for detection. As a result, our basic loss function in Eq. (3) is reformulated to:

where b^f\hat{\mathbf{b}}_{f} and b^b\hat{\mathbf{b}}_{b} are derived from mf{\mathbf{m}}_{f} and mb{\mathbf{m}}_{b}, respectively. Based on such explicit decoupling, our model maximizes the cooperation of foreground supervision from both detection and segmentation datasets and significantly reduces the interference between foreground and background categories. Though decoupled, we note that these two types of queries share the same decoder and interact with each other with self-attention, as shown in Fig. 4 (b). Below we explain how the foreground and background queries are determined.

Language-guided foreground query selection. Open-vocabulary setting differs from the conventional closed-set setting in that a model is required to localize a large number of foreground objects far beyond the training vocabulary. However, the fact is that our decoder contains a limited number of foreground queries (a few hundred typically), making it hardly handle all possible concepts in the image.

To address this issue, we propose a method called language-guided foreground query selection to adaptively select queries with respect to given text concepts as shown in Fig. 3 left part. Given the image features O\mathbf{O} and text features T\mathbf{T}, we employ a lightweight module to predict the box and score for each feature:

where Head\mathsf{Head} is the box head. Then we select LfL_{f} top-ranked entries from Eb\mathbf{E}^{b} and O\mathbf{O} according to the scores in Ec\mathbf{E}^{c}. These selected LfL_{f} image features and boxes are then fed to the decoder as the foreground queries (blue squares in Fig. 3). By selecting only the text-related tokens as decoder queries, we mitigate the problem of decoding irrelevant semantics and provide better query initialization. Such an adaptive way of proposing foreground queries enables our model to effectively transfer to novel vocabulary during test scenarios.

Learnable background queries. Different from foreground queries, we use learnable query embeddings for our background queries for two reasons. Firstly, query selection does not work well because the selected reference points often extend beyond large and non-convex background regions, leading to suboptimal results. Secondly, background stuff has a relatively smaller number of categories than the foreground, and a single image typically contains a few different stuffs (e.g., “sky”, “building”). As a result, using learnable queries for our model can sufficiently and effectively handle background stuff categories and generalize well to open-vocabulary settings. The background queries are marked by green squares in Fig. 3.

Comparison with previous works. In Fig. 5, we show a comparison between our approach and others on handling foreground and background. Mask2Former and MaskDINO treat foreground and background equally when conducting panoptic segmentation, resulting in suboptimal mask average precision (AP) for foreground objects compared to the same model solely trained on foreground classes (instance segmentation). Panoptic Segformer separates foreground and background queries, but their background queries have fixed semantics, with each query corresponding to a pre-defined background category, limiting their ability to handle open-vocabulary categories. In contrast, our approach proposes foreground queries through a language-guided selection mechanism, and our background queries are fully learnable, eliminating the restrictions of a predefined vocabulary.

3 Bridge Data Gap: Conditioned Mask Decoding

Our ultimate goal is to bridge the data gap by using a single loss function to train multiple tasks, resulting in the following loss function:

Here, D\mathcal{D} represents the union of segmentation and detection datasets. However, the loss function requires mask annotations for detection data and box annotations for segmentation data, leading to a discrepancy in the granularity of spatial supervision between the two tasks. As we discussed earlier, we can easily convert an object mask mm to a box b^\hat{b}, which augments the original segmentation data Dm={Ii,(ci,mi)}i=1M\mathcal{D}_{m}=\{I_{i},(\mathbf{c}_{i},\mathbf{m}_{i})\}_{i=1}^{M} into D^m={Ii,(ci,mi,b^i)}i=1M\hat{\mathcal{D}}_{m}=\{I_{i},(\mathbf{c}_{i},\mathbf{m}_{i},\hat{\mathbf{b}}_{i})\}_{i=1}^{M}. For detection data Db\mathcal{D}_{b}, however, we are only given coarse location (box) and category. Then an interesting question comes – can we obtain its mask given these priors?

To address this problem, we resort to the segmentation data which contains rich mappings from label&box to mask, i.e., (c,b)→m(c,b)\rightarrow m and propose conditioned mask decoding to learn the mappings as shown in Fig. 3 right-most part. Given the ground-truth concepts and boxes, (c,b)(\mathbf{c},\mathbf{b}), we employ the decoder to decode the mask:

where tt is the text features extracted for the concepts. Based on Eq. (7), the question becomes, can we learn from segmentation data a good mapping which generalizes well to detection data with different categories? Mapping Hypothesis Verification. To answer the question, we conduct a pilot study. We train a model which learns to decode masks conditioned on GT concepts and boxes on COCO , and then evaluate the conditioned decoding performance on ADE20K . The results are shown in Table 1. Comparing the top two rows, we can find mask decoding conditioned on the GT concept and box significantly improves the quality (mask AP from 8.6 to 46.4), which even reaches a similar level to COCO (46.4 v.s. 53.2). These results indicate that our learned mask decoding generalizes well to a new dataset with novel categories. To further verify, we visualize decoded masks in Fig. 6.

Interactive Segmentation. The above study implies a new interface of image segmentation. Apart from segmenting an image from scratch, users can give a hint about the object location by drawing a box (click four points), and our OpenSeeD can generate its mask with fairly high quality. This capacity can potentially help accelerate the annotation of segmentation data, especially for those with boxes. We leave a comprehensive study on this as future work. Conditioned Mask Decoding Training. Based on the verified hypothesis, we add all the GT boxes and labels as the conditioned queries to simultaneously learn foreground/background decoding and conditioned mask decoding, as shown in Fig. 3. It unifies all our tasks and enables OpenSeeD to learn more generalized conditioned decoding in the joint semantic space. Based on this, we can literally derive the pseudo masks m^\hat{\mathbf{m}} for object detection data and obtain augmented D^b={Ii,(ci,mi^,bi)}i=1N\hat{\mathcal{D}}_{b}=\{I_{i},(\mathbf{c}_{i},\hat{\mathbf{m}_{i}},{\mathbf{b}}_{i})\}_{i=1}^{N}. Below we elaborate on how these pseudo masks are used for training.

Conditioned Mask Generation to Guide Detection Data. The trained conditioned mask decoding component can also be used to assist detection data as segmentation guidance. We propose two methods to utilize the generated mask to guide our model training, Online Mask Assistance and Offline Mask Assistance. For Online Assistance, we only train one model and generate the masks on the fly. Instead of directly using the generated masks as mask supervision, we use the masks to assist in matching predictions and GT instances because the mask quality is not strong enough for supervision. especially in the early stage (shown in Tab. 1 third row). As for Offline Assistance, we train our model with conditioned mask decoding until convergence and generate mask annotations for detection data. The annotated dataset can be used to train a segmentation model. Considering detection data only has instance-level annotations, the generated masks are expected to improve instance segmentation in both cases. More details about these two methods are discussed in the Appendix. Comparison with Denoising Training. Compared with models using denoising training (DN), conditioned mask decoding differs in two aspects. First, their design choices are different. DN adds noise to the GT boxes and labels for reconstruction, but our model learns to generate masks conditioned on the GT priors. Second, their design purposes are different. DN is designed to accelerate training convergence (understanding), while our method aims to generate masks for detection data (generation).

Experiment

Datasets and Settings. In our experiments, we jointly pre-train on two types of data, including panoptic segmentation and object detection. For panoptic segmentation, we use COCO2017 with segmentation annotations (around 110k images). For object detection, we use Objects365 (660k images for v1 and 1700k images for v2). We use Objects365v1 for training and ablating our tiny model and Objects365v2 only for training our large model. We evaluate our models on all tasks covered by pretraining, including semantic, instance, panoptic segmentation, and object detection. In particular, we benchmark on more than 60 datasets covering a wide range of domains on zero-shot segmentation and detection.

Implementation Details. We build on Mask DINO to implement our model. Mask DINO is a unified detection and segmentation framework which simultaneously predicts box and mask. We follow to use 300 latent queries and nine decoder layers for thing categories in instance segmentation and add 100 panoptic queries for stuff categories. For the visual backbone, we adopt pretrained Swin-T/L by default. We also use Focal-T in our ablation studies following . For the language backbone, we adopt the pretrained base model in UniCL . Particularly, our model only uses these pretrained backbones and does not use other image-text pairs or grounding data for pretraining . During pretraining, we set a minibatch for segmentation to 3232 and detection to 6464, and the image resolution is 1024×10241024\times 1024 for both segmentation and detection. During fine-tuning, we use 512×1024512\times 1024 for Cityscapes and 640×640640\times 640 for ADE20K by default. Following the balanced sampling strategy in , the segmentation data are always sampled for a consistent number of epochs, regardless of the total number of detection data. We use AdamW as the optimizer. We pre-train our model on the joint dataset for 30 epochs. The learning rate is set to 0.00010.0001, which is decayed at 0.9 and 0.95 fractions of the total number of steps by 10. Unless otherwise specified, we use online mask assistance during our pretraining by default.

2 Open-Vocabulary Benchmarking

After pretraining our OpenSeeD on COCO and Objects365, we evaluate it on a wide range of datasets in a zero-shot manner. Following , we cover six commonly used segmentation datasets, including indoor scenes (ADE20K ), outdoor scenes (Cityscapes ), and driving scenes (BDD100K ). In addition, we evaluate both segmentation and detection performance on LVIS . We report PQ, mask AP, and mIoU for panoptic, instance, and semantic segmentation, respectively, and use box AP for detection. The results are shown in Table 2.

We first compare with previous works on segmentation tasks. Overall, our model achieves significantly better performance on instance segmentation and comparable performance for panoptic and semantic segmentation. Compared with state-of-the-art methods ODISE and X-Decoder , OpenSeeD achieves 1.1 and 1.9 mask AP improvements on ADE20K, respectively. This gap is even larger on Cityscapes and LVIS. Our OpenSeeD outperforms X-Decoder by 10.2 and 8.3 mask AP with a tiny and large model on Cityscapes, respectively. On LVIS, we evaluate the mask AP with the released X-Decoder tiny model and the comparison shows 9.8 mask AP improvement with our OpenSeeD tiny model. These results indicate that the proposed joint learning method can effectively transfer the instance-level knowledge in detection data for instance segmentation. Compared with instance segmentation, both panoptic and semantic segmentation requires the segmentation of background stuff, which is fully absent in the detection data. Despite that, our OpenSeeD still outperforms X-Decoder for panoptic segmentation on 3 out of 4 datasets (except for ADE20K), and achieves comparable semantic segmentation performance. The results on three segmentation tasks suggest that detection data significantly benefit the instance-level understanding while image-text pairs mainly augment the semantic understanding for semantic segmentation. In addition to segmentation, OpenSeeD also produces reasonably good detection performance. Compared with GLIP, our OpenSeeD (T) outperforms GLIP (T) (setting A) for zero-shot detection on LVIS (21.8 v.s. 18.5), where both only use Objects365 as the pretraining detection dataset. At last, we highlight that our model is the first one that can be pretrained with segmentation and detection data jointly and perform zero-shot transfer to both tasks.

3 Direct and Task-Specific Transfer

After pretraining, our model can be directly transferred to downstream segmentation and detection tasks. In Table 3, we compare our model with both closed-set and open-vocabulary methods. Remarkably, our model achieves SOTA performance 59.5 PQ for COCO panoptic segmentation without any further fine-tuning. After data-specific fine-tuning, OpenSeeD establishes a new SOTA on ADE20K panoptic (53.7 PQ) and instance segmentation (42.6 AP) when trained with 1280×12801280\times 1280 image size. In addition, we also achieve a new SOTA on Cityscapes instance segmentation (48.5 AP). These results indicate the jointly pretrained open-vocabulary model can also be well transferred to closed-set detection and segmentation.

4 Segmentation and Detection in the Wild

To investigate the generalization ability of OpenSeeD for segmentation, we evaluate our model on more domain-specific datasets. We conduct a zero-shot evaluation of our model on the Segmentation in the Wild (SeginW) benchmark , which includes 25 datasets. As this benchmark focuses on instance segmentation, we report the average and median mAP of all the datasets following the common practice. The results in Table 4 indicate that combining detection supervision significantly improves the segmentation performance by more than 10 AP under the same setting.

To further study the object detection ability of OpenSeeD, we follow GLIP to evaluate detection performance on Object Detection in the Wild (ODinW) benchmark. It collects over 35 datasets and is closer to real-world scenarios. We report the average and median mAP of all 35 datasets. With the jointly pretrained model weights, we directly evaluate this challenging benchmark in a zero-shot manner. Under the same setting that only uses Objects365 as the detection data for training, our tiny model outperforms GLIP-T (setting A) by 2.82.8 AP in average.

5 Ablation

Ablation on our basic components. We first ablate on our basic model to verify whether each basic component works as expected. We evaluate our model on the COCO closed-set panoptic segmentation task. The first two rows in Table 6 show that our model can achieve a comparable closed-set panoptic segmentation performance as MaskDINO using the same ResNet-50 backbone. It is also shown in the last four rows that our framework can significantly improve the segmentation performance by combining tasks and utilizing detection data across different backbones. Ablation on the offline mask supervision. As we discussed earlier, the proposed conditioned mask decoding can be used either in online or offline manner. Here, to investigate the effectiveness of offline setting, we generate pseudo annotations with our large models and then use them to tune the tiny models. Concretely, we first generate pseudo masks on Objects365 conditioned on boxes with our OpenSeeD (L) model. Given the masks, we conduct experiments on two models, including the open-vocabulary model (OpenSeeD) and the closed-set model (MaskDINO) by using the pseudo masks for supervision. As shown in Table 7, when evaluating our models, we find that the mask AP and box AP are significantly improved. When conducting experiments on MaskDINO, we perform object detection for the setting without pseudo annotation and instance segmentation for the setting with pseudo annotations during pretraining. After pretraining, we do a zero-shot evaluation on COCO. We also use the pretrained model for fine-tuning. The results in Table 8 indicate that in both settings, pseudo-annotations improve performance. It is also shown that the box AP can also be improved accordingly with extra mask annotations. In addition, we are the first to report zero-shot segmentation performance on COCO which is comparable with the full-shot performance of other methods. Ablation on decoupled decoding and online mask assistance. In Table 9, we conduct experiments to show the effectiveness of our proposed components by removing them one at a time. In the second row, after removing online mask assistance, the instance mask and box performance on the open-segmentation dataset ADE20K drops by 0.6 AP and 0.5 AP, respectively. In the third row, when we remove the decoupled decoding, the performance of the mask and box is further impacted by a large margin (-1.1 AP and -2.3 AP for ADE mask and box). Both results suggest that the proposed techniques help to mitigate the gap between segmentation and detection data.

Conclusion

We have presented OpenSeeD, a simple open-vocabulary segmentation and detection framework, which jointly learns from different segmentation and detection datasets with a single model. To bridge the task gap between foreground objects and background stuff, we propose a decoupled decoding method with language-guided foreground query selection. We also jointly train a conditioned mask decoding task, which provides an interactive segmentation interface during inference and helps bridge data gap for detection data during training. The result indicates our unified model significantly improves open-segmentation while keeping a reasonable detection performance. The jointly pre-trained model can also be seamlessly transferred to improve close-vocabulary performance. Limitations. In this work, we aim at exploring the potential of training an open-vocabulary model for both segmentation and detection. OpenSeeD does not utilize either referring/grounding data or large-scale image-text pairs to further enrich our training data and semantic coverage. We leave a grander joint training to future work.

References