Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, Shalini De Mello
Introduction
Humans look at the world and can recognize limitless categories. Given the scene presented in Fig. 1, besides identifying every vehicle as a “truck”, we immediately understand that one of them is a pickup truck requiring a trailer to move another truck. To reproduce an intelligence with such a fine-grained and unbounded understanding, the problem of open-vocabulary recognition has recently attracted a lot of attention in computer vision. However, very few works are able to provide a unified framework that parses all object instances and scene semantics at the same time, i.e., panoptic segmentation.
Most current approaches for open-vocabulary recognition rely on the excellent generalization ability of text-image discriminative models trained with Internet-scale data. While such pre-trained models are good at classifying individual object proposals or pixels, they are not necessarily optimal for performing scene-level structural understanding. Indeed, it has been shown that CLIP often confuses the spatial relations between objects . We hypothesize that the lack of spatial and relational understanding in text-image discriminative models is a bottleneck for open-vocabulary panoptic segmentation.
On the other hand, text-to-image generation using diffusion models trained on Internet-scale data has recently revolutionized the field of image synthesis. It offers unprecedented image quality, generalizability, composition-ability and, semantic control via the input text. An interesting observation is that to condition the image generation process on the provided text, diffusion models compute cross-attention between the text’s embedding and their internal visual representation. This design implies the plausibility of the internal representation of diffusion models being well-differentiated and correlated to high/mid-level semantic concepts that can be described by language. As a proof-of-concept, in Fig.1 (center), we visualize the results of clustering a diffusion model’s internal features for the image on the left. While not perfect, the discovered groups are indeed semantically distinct and localized. Motivated by this finding, we ask the question of whether Internet-scale text-to-image diffusion models can be exploited to create universal open-vocabulary panoptic segmentation learner for any concept in the wild?
To this end, we propose ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation (pronounced o-di-see), a model that leverages both large-scale text-image diffusion and discriminative models to perform state-of-the-art panoptic segmentation of any category in the wild. An overview of our approach is illustrated in Fig. 2. At a high-level it contains a pre-trained frozen text-to-image diffusion model into which we input an image and its caption and extract the diffusion model’s internal features for them. With these features as input, our mask generator produces panoptic masks of all possible concepts in the image. We train the mask generator with annotated masks available from a training set. A mask classification module then categorizes each mask into one of many open-vocabulary categories by associating each predicted mask’s diffusion features with text embeddings of several object category names. We train this classification module with either mask category labels or image-level captions from the training dataset. Once trained, we perform open-vocabulary panoptic inference with both the text-image diffusion and discriminative models to classify a predicted mask. On many different benchmark datasets and across several open-vocabulary recognition tasks, ODISE achieves state-of-the-art accuracy outperforming the existing baselines by large margins.
To the best of our knowledge, ODISE is the first work to explore large-scale text-to-image diffusion models for open-vocabulary segmentation tasks.
We propose a novel pipeline to effectively leverage both text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation.
We significantly advance the field forward by outperforming all existing baselines on many open-vocabulary recognition tasks, and thus establish a new state of the art in this space.
Related Work
Panoptic Segmentation. Panoptic segmentation is a fundamental vision task that encompasses both instance and semantic segmentation. However, previous works follow a closed closed-vocabulary assumption and only recognize categories present in the training set. They are hence limited in segmenting things/stuff present in finite-sized vocabularies, which are much smaller than the typical vocabularies that we use to describe the real world.
Open-Vocabulary Segmentation. Most prior works on open-vocabulary segmentation either perform object detection with instance segmentation alone or open-vocabulary semantic segmentation alone . In contrast, we propose a novel unified framework for both open-vocabulary instance and semantic segmentation. Another distinction is that prior works only use large-scale models pre-trained for image discriminative tasks, e.g., image classification or image-text contrastive learning . The concurrent work MaskCLIP also uses CLIP . However, such discriminative models’ internal representations are sub-optimal for performing segmentation tasks versus those derived from image-to-text diffusion models as shown in our experiments.
Generative Models for Segmentation. There exist prior works, which are similar in spirit to ours in their use of image generative models, including GANs or diffusion models to perform semantic segmentation . They first train generative models on small-vocabulary datasets, e.g., cats , human faces or ImageNet and then with the help of few-shot hand-annotated examples per category, learn to classify the internal representations of the generative models into semantic regions. They either synthesize many images and their mask labels to train a separate segmentation network ; or directly use the generative model to perform segmentation . Among them, DDPMSeg shows the state-of-the-art accuracy. These prior works introduce the key idea that the internal representations of generative models may be sufficiently differentiated and correlated to mid/high-level visual semantic concepts and could be used for semantic segmentation. Our work is inspired by them, but it is also different in many respects. While previous works primarily focus on label-efficient semantic segmentation of small closed vocabularies, we, on the other hand, tackle open-vocabulary panoptic segmentation of many more and unseen categories in the wild.
Method
Following , we train a model with a set of base training categories , which may be different from the test categories, , i.e., . may contain novel categories not seen during training. We assume that during training, the binary panoptic mask annotation for each category in an image is provided. Additionally, we also assume that either the category label of each mask or a text caption for the image is available. During testing, neither the category label nor the caption is available for any image, and only the names of the test categories are provided.
2 Method Overview
An overview of our method ODISE, for open-vocabulary panoptic segmentation of any category in the wild is shown in Fig. 2. At a high-level, it contains a text-to-image diffusion model into which we input an image and its caption and extract the diffusion model’s internal features for them (Sec 3.3). With these extracted features as input, and the provided training mask annotations, we train a mask generator to generate panoptic masks of all possible categories in the image (Sec 3.4). Using the provided training images’ category labels or text captions, we also train an open-vocabulary mask classification module. It uses each predicted mask’s diffusion features along with a text encoder’s embeddings of the training category names to classify a mask (Sec 3.5). Once trained, we perform open-vocabulary panoptic inference with both the text-image diffusion and discriminative models (Sec 3.6 and Fig. 3). In the following sections, we describe each of these components.
3 Text-to-Image Diffusion Model
We first provide a brief overview of text-to-image diffusion models and then describe how we extract features from them for panoptic segmentation.
A text-to-image diffusion model can generate high-quality images from provided input text prompts. It is trained with millions of image-text pairs crawled from the Internet . The text is encoded into a text embedding with a pre-trained text encoder, e.g., T5 or CLIP . Before being input into the diffusion network, an image is distorted by adding some level of Gaussian noise to it. The diffusion network is trained to undo the distortion given the noisy input and its paired text embedding. During inference, the model takes image-shaped pure Gaussian noise and the text embedding of a user provided description as input, and progressively de-noises it to a realistic image via several iterations of inference.
Visual Representation Extraction
The prevalent diffusion-based text-to-image generative models typically use a UNet architecture to learn the denoising process. As shown in the blue block in Fig. 2, the UNet consists of convolution blocks, upsampling and downsampling blocks, skip connections and attention blocks, which perform cross-attention between a text embedding and UNet features. At every step of the de-noising process, diffusion models use the text input to infer the de-noising direction of the noisy input image. Since the text is injected into the model via cross attention layers, it encourages visual features to be correlated to rich semantically meaningful text descriptions. Thus the feature maps output by the UNet blocks can be regarded as rich and dense features for panoptic segmentation.
Our method only requires a single forward pass of an input image through the diffusion model to extract its visual representation, as opposed to going through the entire multi-step generative diffusion process. Formally, given an input image-text pair , we first sample a noisy image at time step as:
where is the diffusion step we use, represent a pre-defined noise schedule where , as defined in . We encode the caption with a pre-trained text encoder and extract the text-to-image diffusion UNet’s internal features for the pair by feeding it into the UNet
It is worth noting that the diffusion model’s visual representation for is dependent on its paired caption . It can be extracted correctly when paired image-text data is available, e.g., during pre-training of the text-to-image diffusion model. However, it becomes problematic when we want to extract the visual representation of images without paired captions available, which is the common use case for our application. For an image without a caption, we could use an empty text as its caption input, but that is clearly suboptimal, which we also show in our experiments. In what follows, we introduce a novel Implicit Captioner that we design to overcome the need for explicitly captioned image data. It also yields optimal downstream task performance.
Implicit Captioner
Instead of using an off-the-shelf captioning network to generate captions, we train a network to generate an implicit text embedding from the input image itself. We then input this text embedding into the diffusion model directly. We name this module an implicit captioner. The red block in Fig. 2 shows the architecture of the implicit captioner. Specifically, to derive the implicit text embedding for an image, we leverage a pre-trained frozen image encoder , e.g., from CLIP to encode the input image into its embedding space. We further use a learned MLP to project the image embedding into an implicit text embedding, which we input into text-to-image diffusion UNet. During open-vocabulary panoptic segmentation training, the parameters of the image encoder and of the UNet are unchanged and we only fine-tune the parameters of the MLP.
Finally, the text-to-image diffusion model’s UNet along with the implicit captioner, together form ODISE’s feature extractor that computes the visual representation for an input image . Formally, we compute the visual representation as:
4 Mask Generator
The mask generator takes the visual representation as input and outputs class-agnostic binary masks and their corresponding mask embedding features . The architecture of the mask generator is not restricted to a specific one. It can be any panoptic segmentation network capable of generating mask predictions of the whole image. We can instantiate our method with both bounding box-based and direct segmentation mask-based methods. While using bounding box-based methods like , we can pool the ROI-Aligned features of each predicted mask’s region to compute its mask embedding features. For segmentation mask-based methods like , we can directly perform masked pooling on the final feature maps to compute the mask embedding features. Since our representation focuses on dense pixel-wise predictions, we use a direct segmentation-based architecture. Following , we supervise the predicted class-agnostic binary masks via a pixel-wise binary cross entropy loss along with their corresponding ground truth masks (treated as class-agnostic ones as well). Next, we describe how we classify each mask, represented by its mask embedding feature, into an open vocabulary.
5 Mask Classification
To assign each predicted binary mask a category label from an open vocabulary, we employ text-image discriminative models. These models , trained on Internet-scale image-text pairs, have shown strong open-vocabulary classification capabilities. They consist of an image encoder and a text encoder . Following prior work , while training, we employ two commonly used supervision signals to learn to predict the category label of each predicted mask. Next, we describe how we unify these two training approaches in ODISE.
Here, we assume that during training we have access to each mask’s ground truth category label. Thus, the training procedure is similar to that of traditional closed-vocabulary training. Suppose that there are categories in the training set. For each mask embedding feature , we dub its corresponding known ground truth category as . We encode the names of all the categories in with the frozen text encoder , and define the set of embeddings of all the training categories’ names as
where the category name . Then we compute the probability of the mask embedding feature belonging to one of the classes via a classification loss as:
where is a learnable temperature parameter.
Image Caption Supervision
Here, we assume that we do not have any category labels associated with each annotated mask during training. Instead, we have access to a natural language caption for each image, and the model learns to classify the predicted mask embedding features using the image caption alone. To do so, we extract the nouns from each caption and treat them as the grounding category labels for their corresponding paired image. Following , we employ a grounding loss to supervise the prediction of the masks’ category labels. Specifically, given the image-caption pair , suppose that there are nouns extracted from , denoted as . Suppose further that we sample image-caption pairs to form a batch. To compute the grounding loss, we compute the similarity between each image-caption pair as
where and are vectors of the same dimension and is the -th element of the vector defined in Eq. 6 after Softmax. This similarity function encourages each noun to be grounded by one or a few masked regions of the image and avoids penalizing the regions that are not grounded by any word at all. Similar to the image-text contrastive loss in , the grounding loss is defined by
where is a learnable temperature parameter. Finally, note that we train the entire ODISE model with either or , together with the class-agnostic binary mask loss. In our experiments, we explicitly state which of these two supervision signals (label or caption) we use for training ODISE when comparing to the relevant prior works.
6 Open-Vocabulary Inference
During inference (Fig. 3), the set of names of the test categories is available, The test categories may be different from the training ones. Additionally, no caption/labels are available for a test image. Hence we pass it through the implicit captioner to obtain its implicit caption; input the two into the diffusion model to obtain the UNet’s features; and use the mask generator to predict all possible binary masks of semantic categories in the image. To classify each predicted mask into one of the test categories, we compute defined in Eq. 6 using ODISE and finally predict the category with the maximum probability.
In our experiments, we found that the internal representation of the diffusion model is spatially well-differentiated to produce many plausible masks for objects instances. However, its object classification ability can be further enhanced by combining it once again with a text-image discriminative model, e.g., CLIP , especially for open-vocabularies. To this end, here we leverage a text-image discriminative model’s image encoder to further classify each predicted masked region of the original input image into one of the test categories. Specifically, as Fig. 3 illustrates, given an input image , we first encode it into a feature map with the image encoder of a text-image discriminative model. Then for a mask , predicted by ODISE for image , we pool all the features at the output of the image encoder that fall inside the predicted mask to compute a mask pooled image feature for it
We use from Eq.6 to compute the final classification probabilities from the text-image discriminative model. Finally, we take the geometric mean of the category predictions from the diffusion and discriminative models as the final classification prediction,
where is a fixed balancing factor. We find that pooling the masked features is more efficient and yet as effective as the alternative approach proposed in , which crops each of the predicted masked region’s bounding box from the original image and encodes it separately with the image encoder (see details in the supplement).
Experiments
We first introduce our implementation details. Then we compare our results against the state of the art on open-vocabulary panoptic and semantic segmentation. Lastly, we present ablation studies to demonstrate the effectiveness of the components of our method.
We use the stable diffusion model pre-trained on a subset of the LAION dataset as our text-to-image diffusion model. We extract feature maps from every three of its UNet blocks and, like FPN , resize them to create a feature pyramid. We set the time step used for the diffusion process to , by default. We use CLIP as our text-image discriminative model and its corresponding image and text encoders everywhere. We choose Mask2Former as the architecture of our mask generator, and generate binary mask predictions.
Training Details
We train ODISE for 90k iterations with images of size and use large scale jittering . Our batch size is 64. For caption-supervised training, we set . We use the AdamW optimizer with a learning rate and a weight decay of . We use the COCO dataset as our training set. We utilize its provided panoptic mask annotations as the supervision signal for the binary mask loss. For training with image captions, for each image we randomly select one caption from the COCO dataset’s caption annotations.
Inference and Evaluation
We evaluate ODISE on ADE20K for open-vocabulary panoptic, instance and semantic segmentation; and the Pascal datasets for semantic segmentation. We also provide the results ODISE for open-vocabulary object detection and open-world instance segmentation in the supplement. We use only a single checkpoint of ODISE for mask prediction on all tasks on all datasets. For panoptic segmentation, we report the panoptic quality (PQ) , mean average precision (mAP) on the “thing” categories, and the mean intersection over union (mIoU) metrics (additional SQ and RQ metrics are in the supplement). In panoptic segmentation annotations , the “thing” classes are countable objects like people, animals, etc. and the “stuff” classes are amorphous regions like sky, grass, etc. Since we train ODISE with panoptic mask annotations, we can directly infer both instance and semantic segmentation labels with it. When evaluating for panoptic segmentation, we use the panoptic test categories as , and directly classify each predicted mask into the test category with the highest probability. For semantic segmentation, we merge all masks assigned to the same “thing” category into a single one and output it as the predicted mask.
Speed and Model Size
ODISE has 28.1M trainable parameters (only 1.8% of the full model) and 1,493.8M frozen parameters. It performs inference for an image () at 1.26 FPS on an NVIDIA V100 GPU and uses 11.9 GB memory.
2 Comparison with State of the Art
For open-vocabulary panoptic segmentation, we train ODISE on COCO and test on ADE20K . We report results in Table 1. ODISE outperforms the concurrent work MaskCLIP by 8.3 PQ on ADE20K. Besides the PQ metric, our approach also surpasses MaskCLIP at open-vocabulary instance segmentation on ADE20K, with 8.4 gains in the mAP metric. The qualitative results can be found in Fig. 4 and more in the supplement.
Open-Vocabulary Semantic Segmentation
We show a comparison of ODISE to previous work on open-vocabulary semantic segmentation in Table 2. Following the experiment in , we evaluate mIoU on 5 semantic segmentation datasets: (a) A-150 with 150 common classes and (b) A-847 with all the 847 classes of ADE20K , (c) PC-59 with 59 common classes and (d) PC-459 with full 459 classes of Pascal Context , and (e) the classic Pascal VOC dataset with 20 foreground classes and 1 background class (PAS-21). For a fair comparison to prior work, we train ODISE with either category or image caption labels. ODISE outperforms the existing state-of-the-art methods on open-vocabulary semantic segmentation by a large margin: by 7.6 mIoU on A-150, 4.7 mIoU on A-847, 4.8 mIoU on PC-459 with caption supervision; and by 6.2 mIoU on A-150, 4.5 mIoU on PC-459 with category label supervision, versus the next best method. Notably, it achieves this despite using supervision from panoptic mask annotations, which is noted to be suboptimal for semantic segmentation .
We provide comparisons to the state of the art for additional open-vocabulary tasks of object detection and discovery in the supplement.
3 Ablation Study
To demonstrate the contribution of each component of our method, we conduct an extensive ablation study. For faster experimentation, we train ODISE with resolution images and use image caption supervision everywhere.
We compare the internal representation of text-to-image diffusion models to those of other state-of-the-art pre-trained discriminative and generative models. We evaluate various discriminative models trained with full label, text or self-supervision. In all experiments we freeze the weights of the pre-trained models and use exactly the same training hyperparameters and mask generator as in our method. For each supervision category we select the best-performing and largest publicly available discriminative models. We observe from Table 3 that ODISE outperforms all other models in terms of PQ on both datasets. To offset any potential bias arising from the larger size of the LAION dataset (2B image-caption pairs) with which the stable diffusion model is trained, versus the smaller datasets used to train the discriminative models, we also compare to CLIP(H) . It is trained on an equal-sized LAION dataset. Despite both models being trained on the same data, our diffusion-based method outperforms CLIP(H) by a large margin on all metrics. This demonstrates that the diffusion model’s internal representation is indeed superior for open-vocabulary segmentation that that of discriminative pre-trained models.
The recent DDPMSeg model is somewhat related to our model. Besides us, it is the only prior work that uses diffusion models and obtains state-of-the-art performance on label-efficient segmentation learning. Since DDPMSeg relies on category specific diffusion models it is not designed for open-world panoptic segmentation. Hence its direct comparison against our approach is not feasible. As an alternative, we compare against the internal representations of a class-conditioned generative model trained on more categories from ImageNet (LDM row in Table 3). Not surprisingly, we find that despite both generative models being diffusion-based, our approach of using a model trained on Internet-scale data is more effective at generalizing to open-vocabulary categories.
Captioning Generators
As discussed in Sec. 3.3, the internal features of a text-to-image diffusion model are dependent on the embedding of the input caption. To derive the optimal set of features for our downstream task, we introduce a novel implicit captioning module to directly generate implicit text embeddings from an image. This module also facilitates inference on images sans paired captions at test time. Here, we construct several baselines to show the effectiveness of our implicit captioning module. The results are shown in Table 4. The various alternatives that we compare are: providing an empty string to the text encoder for any given image, such that the text embedding for all images is fixed (row (a)); employing two different off-the-shelf image captioning networks to generate an explicit caption for each image on-the-fly (rows (b) and (c)), where (c) is trained on the COCO caption dataset, while (b) is not; and our proposed implicit captioning module (row (d)). Overall, we find that using an explicit/implicit caption is better than using empty text. Furthermore, (c) improves over (b) on COCO but has similar PQ on ADE20K. It may be because the pre-trained BLIP model does not see ADE20K’s image distribution during training and hence it cannot output high-quality captions for it. Lastly, since our implicit captioning module derives its caption from a text-image discriminative model trained on Internet-scale data, it is able to generalize best among all variants compared.
Diffusion Time Steps
We also study which diffusion step(s) are most effective for extracting features from, similarly to DDPMSeg . The noise process is defined in Eq.1. The larger the value is, the larger the noise distortion added to the input image is. In stable diffusion there are a 1000 total time steps. From Table 5, all metrics decrease as increases and the best results are for (our final value). Concatenating 3 time steps, 0, 100, 200, yields a similar accuracy to only, but is slower. We also train our model with as a learnable parameter, and find that many random training runs all converge to a value close to zero, further validating our optimal choice of .
Mask Classifiers
For final open-vocabulary classifciation (Fig. 3), we fuse class prediction from the diffusion and discriminative models. We report their individual performance in Table 6. Individually the diffusion approach performs better on both datasets than the discriminative only approach. Nevertheless, fusing both together results in higher values on both ADE20K and COCO. Finally, note that even without fusion, our diffusion-only method already surpasses existing methods (see Tables 1, 2).
Conclusion
We take the first step in leveraging the frozen internal representation of large-scale text-to-image diffusion models for downstream recognition tasks. ODISE shows the great potential of text-to-image generation models in open-vocabulary segmentation tasks and establishes a new state of the art. This work demonstrates that text-to-image diffusion models are not only capable of generating plausible image but also of learning rich semantic representations. It opens up a new direction for how to effectively leverage the internal representation of text-to-image models for other tasks as well in the future.
Acknowledgements. We thank Golnaz Ghiasi for providing the prompt engineering labels for evaluation. Prof. Xiaolong Wang’s laboratory was supported, in part, by NSF CCF-2112665 (TILOS), NSF CAREER Award IIS-2240014, DARPA LwLL, Amazon Research Award, Adobe Data Science Research Award, and Qualcomm Innovation Fellowship.
References
Appendix A Implementation Details
We open-source our code and models at https://github.com/NVlabs/ODISE.
We train ODISE for 90k iterations with images of size and use large scale jittering with random scales between as data augmentation. We use 32 NVIDIA V100 GPUs with 2 images per GPU with an effective batch size is 64. We use the AdamW optimizer with a learning rate and a weight decay of . We use a step learning rate schedule and reduce the learning rate by a factor of at 81k and 86k iterations. We set the balancing factor between the diffusion and discriminative models to for all tasks. Following , we use Hungarian matching to match the predicted masks to the ground-truth ones. We compute the training losses between the matched pairs.
Open-Vocabulary Inference
An object can often be described by more than one possible description, e.g., the dog category could be described by “dog” or “puppy”. We use the same prompt engineering strategy as in to create an ensemble of text prompts for each test category and predict the category with the maximum probability.
Speed and Model Size
It takes 5.3 days to train ODISE for 90k iterations on the COCO dataset. It has 28.1M trainable parameters (only 1.8% of the full model) and 1,493.8M frozen parameters (including Stable Diffusion and CLIP). It performs single image inference at 1.26 FPS on an NVIDIA V100 GPU and uses 11.9 GB memory with an image of size . We also replace the bounding box cropping proposed in that runs at 0.38 FPS, with mask feature pooling described in Section 3.6 of the main paper. Mask pooling yields a 3x speedup, while maintaining similar PQ on ADE20K: 23.4 for mask pooling versus 23.7 for bounding box cropping.
Appendix B Experiments
Besides panoptic quality (PQ), we additionally report the detailed metrics of segmentation quality (SQ) and recognition quality (RQ) for ODISE and MaskCLIP on both the thing (Th) and stuff (St) categories of the ADE20K dataset in Table B.1. Here, all models were trained on COCO. ODISE outperforms MaskCLIP w.r.t. all metrics.
We also evaluate ODISE trained on COCO on the Cityscapes and Mapillary Vistas datasets in Table B.2. Since the source code for MaskCLIP is not publicly available, we regard ODISE’s implementation with CLIP(H) features (from Table 3 of the main paper) as a close proxy to MaskCLIP and compare against it (Table B.2). Here too, ODISE, which is based on diffusion features, outperforms its CLIP(H) variants by large margins. Note that in this experiment, we use the original text labels provided with the respective test datasets and didn’t carefully select the category names for computing the text embedding. Hence, the results could be further improved if categories like “terrain” are converted into more detailed descriptions.
Finally, to additionally verify the effectiveness of ODISE, we also swap the training and evaluation datasets, i.e., we train on ADE20K and evaluated on COCO, and report the results in Table B.3. Here too, we regard the variant of ODISE with CLIP(H) features as a proxy to MaskCLIP and compare against it. ODISE outperforms its CLIP(H) variant by a large margin.
Open-Vocabulary Object Detection
We also evaluate ODISE for the task of open-vocabulary object detection on the LVIS dataset (Table B.4). By regarding all categories to belong to “things”, we directly evaluate on LVIS’s object detection labels, which contain annotations for 1203 fine-grained categories for COCO images. For this task, we measure , which denotes the mAP score on 337 rare categories only. We evaluate ODISE trained with both types of supervision: mask category labels or image captions. ODISE outperforms MaskCLIP by a large margin w.r.t. both mAP and . Note that the validation split of LVIS has overlapping images with COCO’s training split, but the category labels of LVIS are only available during inference.
Open-World Instance Segmentation
The task of open-world instance segmentation aims at discovering at test time, all plausible instance masks that may be present in an image in a class-agnostic manner. We also evaluate ODISE in for this task. Following , we report the average recall of 100 mask proposals (AR@100) on the UVO and ADE20K datasets. As reported in Table B.5, here too we outperform the existing state of the art by 14.3 points on UVO and 9.3 points on ADE20K. It demonstrates that with the internal representation of pre-trained text-to-image diffusion models it is plausible to discover open-world instances.
B.2 Ablation Study
In Fig. B.1 we show k-means clustering of the text-to-image diffusion model’s and CLIP’s frozen internal features; diffusion features are much more semantically differentiated. Quantitative comparisons of ODISE and its CLIP(H) variant in Table 3 of the main paper and Table B.2 and Table B.3 further substantiate diffusion features’ superiority over those of CLIP’s.
Appendix C Qualitative Results
To demonstrate the open-vocabulary recognition capabilities of ODISE, we merge the category names from LVIS , COCO , ADE20K together and perform open-vocabulary inference with test classes. We only train ODISE on COCO’s training dataset and evaluate open-vocabulary panoptic inference on ADE20K and Ego4D. The qualitative results on COCO’s validation dataset, ADE20K and Ego4D are shown in Fig. B.2, Fig. B.3 and Fig. B.4, respectively. Most categories, e.g., “police cruiser”, “flag”, “conveyor belt”, “chandelier”, “aquarium”, “grocery bag”, “power shovel”, etc., are novel categories from LVIS or ADE20K that are not annotated in COCO . It is worth noting that Ego4D is a video dataset, which consists of diverse ego-centric videos. Despite the large domain gap between the testing dataset Ego4D and our training dataset COCO , ODISE still outputs good-quality plausible panoptic segmentation results on Ego4D’s novel categories.
Appendix D Limitations and Future Work
In the current datasets, the category definitions are sometimes ambiguous and non-exclusive, e.g., in ADE20K, “tower” is often mis-classified as “building”. Although this could be mitigated by prompt and ensemble engineering, how category definitions affect evaluation accuracy, would be interesting to analyze in the future.
Appendix E Ethics Concerns
The text-to-image diffusion model that we use is pre-trained with web-crawled image-text pairs collected by previous works. Despite applying filtering, there may still be potential bias in its internal representation.