Towards Open Vocabulary Learning: A Survey
Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan, Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang, Yunhai Tong, Xudong Jiang, Bernard Ghanem, Dacheng Tao
Introduction
Deep neural networks have revolutionized scene understanding tasks, including object detection, segmentation, and tracking . Nonetheless, using conventional approaches for real-world applications can be challenging due to limitations such as inadequate class annotations, close-set class definitions, and costly labeling expenses. These limitations can increase the difficulty and cost of implementing a deep model on new scenes, particularly when the number of categories or concepts in the scene is much larger than what is included in the training dataset.
For example, object detection is a core computer vision task involving scene understanding. It requires human annotations for each category and each object location, which can be costly and time-consuming. For instance, the COCO dataset , widely used for benchmarking object detection algorithms, only includes 80 categories. But in reality, natural scene images often have more than 80 different types of objects. We would need to incur significant annotation costs to extend object detectors to cover all these categories. Current research focuses on developing methods to train more flexible object detectors on the subset of COCO with base classes and let them identify new or unfamiliar objects without requiring additional annotations.
Several previous solutions adopt zero-shot learning (ZSL) . These approaches extend a detector to generalize from annotated (seen) object classes to other (unseen) categories. The annotations of seen object classes are used during the training, while the annotations of unseen classes are strictly unavailable during training. Most approaches adopt word embedding projection to constitute the classifier for unseen class classification.
However, these approaches come with several limitations. ZSL, in particular, is highly restrictive. Typically, these methods lack examples of unseen objects and treat these objects as background objects during training. As a result, during inference, the model identifies novel classes solely based on their pre-defined word embeddings , thereby limiting exploration of the visual information and relationships of those unseen classes. This is why ZSL approaches have been shown to yield unsatisfying results in novel classes.
Open vocabulary learning is proposed to handle the above issue and make the detector extendable. It has been successfully applied to multiple tasks, e.g., segmentation, detection, video understanding, scene understanding, etc. In particular, open vocabulary object detection is first proposed. Also, open vocabulary segmentation is proposed . Similarly, the model trains on base classes and inferences on both base and novel classes. The critical difference between zero-shot learning and open vocabulary is that one can use visual-related language vocabulary data like image captions as auxiliary supervision in open vocabulary settings. The motivations to use language data as auxiliary weak supervision are: 1) Language data requires less labeling effort and thus is more cost-friendly. The visual-related language data, like image captions, is widely available and more affordable than box or mask annotations. 2) Language data provides a more extensive vocabulary size and thus is more extendable and general. For example, words in captions are not limited to the pre-defined base categories. It may contain novel class names, attributes, and motions of objects. Incorporating captions in training has been proven to be extremely useful in helping improve the models’ scalability.
Moreover, recently, visual language models (VLMs) , which pre-train themselves on large-scale image-text pairs, show remarkable zero-shot performance on various vision tasks. The VLMs align images and language vocabularies into the same feature space, fulfilling the visual and language data gap. Many open vocabulary methods effectively eliminate the distinction between close-set and open-set scenarios by utilizing the alignment learned in VLMs, making them highly suitable for practical applications. For example, An open vocabulary object detector can be easily extended to other domains according to the demand without the need to gather relevant data or incur additional labeling expenses.
As open vocabulary models continue to advance rapidly and demonstrate impressive results, it is worthwhile to track and compare recent research on open vocabulary learning. Several surveys work on low-shot learning , including few-shot learning and zero-shot learning. There are also several surveys on multi-modal learning , including using transformer for vision language tasks and vision language pre-training . However, these works focus on learning with few examples, multi-modal fusion, or pre-training for better feature representation. As far as we know, there haven’t been any surveys that thoroughly summarize the latest developments in open vocabulary learning, including methods, settings, benchmarks, and the use of vision foundation models. We aim to fulfill the blank with this work.
Contribution. In this survey, we systematically track and summarize recent literature on open vocabulary learning, including object detection, segmentation, video understanding, and 3D scene understanding. The survey covers the most representative works in each domain by extracting the common technical details. It also contains the background of open vocabulary learning and related concepts comparison, including zero-shot learning (ZSL) , open set recognition (OSR) , and out-of-distribution detection (OOD) . It also includes large-scale visual language models and representative detection and segmentation works, which makes the survey self-contained. In addition, we present a comprehensive analysis and comparison of benchmarks and settings for each specific domain. As far as we know, we are the first to concentrate on the specific area. Finally, since the field of open vocabulary learning is rapidly evolving, we may not be able to keep up with all the latest developments. We welcome researchers to contact us and share their new findings in this area to keep us updated. Those new works will be included and discussed in the revised version.
Survey pipeline. In Sec. 2, we will cover the background knowledge, including definition, datasets, metrics, and related research domains. Then, in Sec. 3, we conduct main reviews on various methods according to different tasks. In particular, we will first include the preliminary knowledge of close-set detection and segmentation methods. Then, we present the method details of each direction, including detection, segmentation, video understanding, and 3D scene understanding. Next, we point out challenges and future directions in Sec. 4 and conclude Sec. 5. Finally, the appendix compares the results for different tasks and benchmarks.
Background
Overview. In this section, we first present the concept definition of open vocabulary and related concepts comparison. Then, we present a historical review of open vocabulary learning and point out several representatives. Next, we present the standard datasets and metrics. We also present the unified notations for open vocabulary object detection and segmentation tasks. Finally, we review the related research domains.
We take the classification task for concept illustration. The supervised methods assume that the training data and testing data share the same closed-set label space. However, the model trained under this assumption cannot be extended to new categories. To address this issue, researchers have introduced several new concepts like open-set learning and zero-shot learning, ultimately leading to open vocabulary learning.
These concepts have similar settings and notations. In particular, the examples are classified into base and novel classes (or called out-of-distribution examples). The base classes can be accessed for training, while the novel classes are not. We denote the base classes in the label set space and the novel classes in the label set space .
Due to the low performance on novel classes in earlier works for open-set, open world, and OOD tasks and the easier acquisition of image-text pairs, recent works propose the open vocabulary setting. It allows using additional low-cost training data or pre-trained vision language models like CLIP , which have much larger language vocabularies . It may contain the concepts of both and , but its main goal is to enable the model to generalize across more classes in the open domain. In mathematical terms and data view, open vocabulary learning can be defined as follows:
Training Data (): The training dataset is a collection of data-label pairs, where each pair consists of an input and its associated label . The input can be an image or a video according to the task type, and in scene understanding tasks such as detection and segmentation, the label also includes visual labels, such as bounding boxes or masks, in addition to the class labels. Alongside the standard input-label pairs, open vocabulary learning includes vision-aware language vocabulary data, represented as . The language vocabulary data can be image-caption data or vision-aware class name embeddings in VLMs like CLIP . The language augmented training set can be denoted as , where is the training data length, is the visual data, is the label from base classes , and is the associated language data from a large vocabulary space . Note that is not strictly required to contain or , as the language vocabulary may not cover all the class names in the vision data. On the contrary, may also have words out of the pre-defined novel categories, which can further extend models’ generalizability.
Evaluation Data (): The evaluation dataset is similarly a collection of data pairs. However, the labels for evaluation data include both base classes and novel classes. This can be represented as , where is the evaluation data length, belongs to either or . During the evaluation, open vocabulary methods need to predict given in the realm .
Here, we emphasize the concept differences between open vocabulary learning and analogous concepts as follows:
Open-Set Learning. Open-set learning aims to classify known classes and reject unknown classes during testing . Concretely, the training data is , where is from base classes . These are also called known classes in open-set learning . The evaluation data is , where . is a single class that represents the ‘unknown’ class. Open-set learning tasks do not require further classifying the unknown classes.
Open World Learning. Open world learning addresses real-world environments’ dynamic and ever-evolving nature by recognizing and learning new categories incrementally over time, without the need for complete system retraining . This process includes classifying known objects, identifying unknowns, labeling the unknowns by humans, and incrementally learning new categories as they are labeled and added to the system. Suppose an open world learning process contains learning steps. In the -th learning step, where , the training data is , where is in the labels annotated in the -th step , and is the input data corresponding to the newly added labels. The evaluation data is , where . is a single class that represents the ‘unknown’ class. is the known class label set for timestamp . It is an accumulation of labeled classes in the current and previous steps. . Note that evaluation input keeps the same across learning steps.
Out-of-Distribution Detection. Out-of-distribution (OOD) detection focuses on the ability to detect data that is different in some way from the data used during training . The training data is and is supposed to be sampled from a probability distribution . During testing, the model encounters data that are sampled from a different distribution , which is not represented in the training dataset. The goal is to identify or appropriately handle these OOD samples. Metrics for OOD detection often involve measuring the model’s certainty or confidence in its predictions, using scores like softmax probability, and checking whether it correctly identifies OOD samples .
Zero-Shot Learning. Zero-shot learning aims to recognize objects or concepts not seen during training. The training data is , where y is from base classes , which is usually called seen classes in zero-shot learning . The evaluation data is , where . is novel classes or unseen classes. Zero-shot learning tasks require models to make clear classifications among the new unseen classes.
We briefly compare these concepts in Fig. 1, including open vocabulary, open-set, open world, OOD, and zero-shot.
2 History and Roadmap
Before introducing the open vocabulary setting, reviewing the progress of open vocabulary learning is necessary. In Fig. 2, we summarize the timeline of open vocabulary learning. Localization and classification of arbitrary objects in the wild have been challenging problems due to the limitations of existing datasets. The concept of open vocabulary in scene understanding comes from the work , where the authors build a joint image pixel and word concept embedding framework. The concept is hierarchically divided. Then, multi-modal pre-training was well studied with the rise of BERT in NLP. Motivated by the process of vision language pre-training, OVR-CNN proposed the concept of open vocabulary object detection, where the caption data are used for connecting novel classes semantics and visual region. Later on, CLIP was presented and open-sourced. After that, VilD is the first work that uses the knowledge of CLIP to build open vocabulary object detection. Meanwhile, LSeg first explored the CLIP knowledge of language-driven segmentation tasks. After these works, recently, there have been more and more works on improving the performance of open vocabulary detectors or building new benchmarks for various settings. SAM is proposed to build the segmentation foundation model, which is trained by billion-level masks. Combined with CLIP, SAM can also achieve good zero-shot segmentation without fine-tuning. Recently, with the rapid process of large language model (LLM) , open vocabulary learning has become a more promising direction since more language knowledge can be embedded in multi-modal architecture. As shown in Fig. 3(a), the number of research works in open vocabulary learning has increased significantly since 2021. We also summarize the statistics of different directions in Fig. 3(b) and (c). The details of these directions can be found in Sec. 3.
3 Tasks, Datasets, and Metrics
Tasks. Open vocabulary learning has included a wide range of computer vision tasks, including object detection , segmentation , video understanding, and 3D scene understanding. The center goal of these tasks is similar, recognizing the novel classes with the aid of large vocabulary knowledge for their corresponding tasks. In this survey, we mainly focus on the methods of scene understanding tasks, including object detection, instance segmentation, semantic segmentation, and object tracking. Nonetheless, we also consider other closely related tasks, such as open vocabulary attribution prediction, video classification, and point cloud classification.
Datasets. For object detection, the common datasets are COCO and LVIS . Recently, a more challenging dataset, v3Det with more than 10,000 categories, is proposed. For image segmentation, the most commonly used datasets are COCO , ADE20k , PASCAL-VOC 2012 , PASCAL-Context , and Cityscapes . For video segmentation and tracking, the frequently used datasets are VSPW , Youtube-VIS , LV-VIS , MOSE , and TAO .
Metrics. For detection tasks, the commonly used metrics are mean average precision (mAP) and mean average recall (mAP) for both base and novel classes. Among segmentation tasks, the commonly used metrics are mean intersection over union (mIoU) for semantic segmentation, mask-based mAP for instance segmentation, and panoptic quality (PQ) for panoptic segmentation.
4 Related Research Domains
Open-Set Recognition. The concept of Open-Set Recognition (OSR) addresses the challenge of identifying unknown classes during classification tasks. In traditional classification systems, models are trained to recognize a finite set of known classes, but in real-world scenarios, they may encounter data that doesn’t belong to these predefined categories. During testing, OSR aims to classify known classes seen during the training and reject unknown classes that are unseen . The main methods used for OSR can be broadly categorized into two types: discriminative models and generative models . The discriminative models focus on differentiating between known classes and identifying unknown classes by enhancing the boundary or margin between these classes. The generative models are either instance generation-based or non-instance generation-based . They emphasize generating new instances or features to improve the ability of the system to recognize new, unseen classes during testing. Each category employs specific techniques and approaches to address the challenges inherent in open-set recognition, which involves dealing with classes not seen during the training phase. extends the OSR task to require further identifying novel classes.
Open World Learning. Open world learning involves identifying and labeling new, unknown categories (novel unknowns) . This process includes incrementally learning new categories as they are labeled and added to the system. The goal is to create a system that remains robust to unknown categories and adapts continually to include new information, balancing the risks associated with open spaces in the learning model. first propose open word recognition. Recent works also explore the open world object detection task.
Out-of-Distribution Detection. Out-of-distribution (OOD) detection methods can be structured into several categories . Classification-Based Methods : These include output-based methods , label space redesign , and OOD data generation techniques . Density-Based Methods : This category involves methods that detect OOD by modeling data density. Distance-Based Methods : These methods use distance metrics, typically in the feature space, to identify OOD instances. Reconstruction-Based Methods : This approach achieves OOD detection that features reconstruction capabilities.
Zero-Shot Detection and Segmentation: This task aims to segment classes that have not been encountered during training. Two streams of work have emerged: discriminative methods and generative methods . Representative works in this field include SPNet and ZS3Net . SPNet maps each pixel to a semantic word embedding space and projects pixel features onto class probabilities using a fixed semantic word embedding projection matrix. On the other hand, ZS3Net first trains a generative model to produce pixel-wise features for unseen classes based on word embeddings. With these synthetic features, the model can be trained in a supervised manner. Both of these works treat zero-shot detection and segmentation as a pixel-level zero-shot classification problem. However, this formulation is not robust for zero-shot learning, as text embeddings are typically used to describe objects/segments rather than individual pixels. Subsequent works follow this formulation to address different challenges in zero-shot learning. In a weaker assumption where unlabeled pixels from unseen classes are available in the training images, self-training is commonly employed. Despite promising results, self-training often requires model retraining whenever a new class appears. ZSI also employs region-level classification for bounding boxes but focuses on instance segmentation rather than semantic segmentation. More recently, PADing proposes a unified framework to tackle zero-shot semantic segmentation, zero-shot instance segmentation, and zero-shot panoptic segmentation.
Most approaches in open vocabulary learning are based on zero-shot learning settings, such as replacing the fixed classifier with language embeddings. However, these methods struggle to generalize well to novel classes due to the absence of novel class knowledge. Consequently, their performance is limited, and they are not practical for real-world applications.
Long-tail Object Detection and Instance Segmentation. This task addresses the challenge of class imbalance in instance segmentation. Many approaches tackle this issue through techniques such as data re-sampling , loss re-weighting , and decoupled training . Specifically, some studies employ image-level re-sampling. However, these methods tend to introduce bias in instance co-occurrence. To address this issue, other works focus on more refined re-sampling techniques at the instance or feature level. Regarding loss re-weighting, most studies rebalance the ratio of positive and negative samples during training. Additionally, decoupled training methods introduce different calibration frameworks to enhance classification results. Resolving long-tail object detection can lead to improved accuracy of rare classes. However, these methods currently cannot be applied to the detection of novel classes.
Few-Shot Detection and Segmentation: Few-shot object detection aims to expand the detection capabilities of a model using only a few labeled samples. Several approaches have been proposed to advance this field. Notably, TFA introduces a simple two-phase fine-tuning method, while DeFRCN separates the training of RPN features and RoI classification. Moreover, SRR-FSD combines multi-modal inputs while LVC proposes a pipeline to generate additional examples for novel object detection and train a more robust detector. Few-shot segmentation comprises Few-Shot Semantic Segmentation (FSSS) and Few-Shot Instance Segmentation (FSIS). FSSS involves performing pixel-level classification on query images. Previous approaches typically build category prototypes from support images and segment the query image by computing the similarity distance between each prototype and query features. On the other hand, FSIS aims to detect and segment objects with only a few examples. FSIS methods can be categorized into single-branch and dual-branch methods. The former primarily focuses on designing the classification head, while the latter introduce an additional support branch to compute class prototypes or re-weighting vectors of support images. This support branch assists the segmenter in identifying target category features through feature aggregation. For example, Meta R-CNN performs channel-wise multiplication on RoI features, and FGN aggregates channel-wise features at three stages: RPN, detection head, and mask head. However, few-shot learning still requires examples of novel classes during training, and such data may not be available.
Methods: A Survey
Overview. In this section, we first review preliminary knowledge and vision language modeling by extending close-set detectors into open vocabulary detectors via VLMs in Sec. 3.1 and Sec. 3.2. Then we sequentially survey the six subsidiary tasks, including object detection (Sec. 3.3), segmentation (Sec. 3.4), video understanding (Sec. 3.5), 3D scene understanding (Sec. 3.6), and closely related tasks (Sec. 3.7). Note that we only record and compare the most representative works. We list numerous works in Tab. I and Tab. IV. Moreover, since there are several similar tasks (Sec. 3.7), we also survey and compare other related tasks, including class agnostic detection and segmentation, open world object detection, and open-set panoptic segmentation.
Method taxonomy and relation of each subsection. We summarize the most commonly used methods among these different directions in Fig. 4. However, we cannot include all methods due to the wide range of tasks. In each subsection, we summarize different methods according to different problems, segmentation, detection, video understanding, and closely related topics. We cluster methods according to common practice, such as knowledge distillation and region text pre-training. We start with object detection since it was first proposed. For segmentation tasks, we ignore the specific setting in Sec. 3.4. We argue that despite these two directions, they may share similar ideas. However, most meta-architectures and tasks are different. Thus, we still review each direction individually. For video, 3D understanding, and other closely related topics, due to different task definitions and extra input cues, such as temporal information and multi-view inputs, we also survey two directions individually.
Motivation of survey organization. We argue that using the unified symbolic system to summarize all methods is hard. There are several reasons. (1), The task definitions and settings are different. In addition to different input and output formats, taking the open vocabulary segmentation as an example, there are at least three different settings for open vocabulary segmentation. Thus, the design principles are different. For example, OpenSeg uses supervised pretraining and contrastive losses while GroupViT adopt unsupervised setting without mask annotations. (2), Different methods use different detectors or baseline models. Several works use R-CNN framework while several methods adopt query-based approaches. Several methods adopt the RPN’s features to develop their methods, while query-based methods cannot perform these designs. (3), Different methods may also use different datasets during training. Several methods mainly explore the dataset effect on open-vocabulary detection or segmentation. Since the supervision signals are different, it will hard to put all methods in one unified symbolic system. In summary, our survey focus on more comprehensive review to extract the common features of various domains in open vocabulary learning.
Pixel-based Object Detection and Segmentation. In general, it can be divided into two aspects: semantic-level tasks and instance-level tasks, where the difference lies in whether to distinguish each instance. For the former, we take semantic segmentation as an example. It was typically approached as a dense pixel classification problem, as initially proposed by FCN . Then, the following works are all based on the FCN framework. These methods can be divided into the following categories, including better encoder-decoder frameworks , larger kernels , multiscale pooling , multiscale feature fusion , non-local modeling , and better boundary delineation . After the transformer was proposed, with the goal of global context modeling, several works propose the variants of self-attention operators to replace the CNN prediction heads .
For the latter, we take object detection and instance segmentation for illustration. Object detection aims to detect each instance box and classify each instance. It mainly has two different categories: two-stage approaches and one-stage approaches. Two-stage approaches rely on an extra region proposal network (RPN) to recall foreground objects at the first stage. Then, the proposals (Region of Interests, RoI) with high scores are sent to the second-stage detection heads for further refinement. One-stage approaches directly output each box and label in a per-pixel manner. In particular, with the help of focal loss and feature pyramid networks, one-stage approaches can surpass the two-stage methods in several datasets. Both approaches need anchors for the regression of the bounding box. The anchors are the pixel locations where the objects are possibly emerging. Since two-stage approaches explicitly explore the foreground objects, where they have better recall than single-stage approaches. Two-stage approaches are more commonly used in open vocabulary object detection tasks.
Instance segmentation segmenting each object goes beyond object detection. Most instance segmentation approaches focus on how to represent instance masks beyond object detection, which can be divided into two categories: top-down approaches and bottom-up approaches . The former extends the object detector with an extra mask head. The designs of mask heads are various, including FCN heads , diverse mask encodings , and dynamic kernels . The latter performs instance clustering from semantic segmentation maps to form instance masks. The performance of top-down approaches is closely related to the choice of detectors , while bottom-up approaches depend on both semantic segmentation results and clustering methods . Besides, several approaches use gird representation to learn instance masks directly.
Query-based Object Detection and Segmentation. With the rise of vision transformers , recent works mainly use transformer-based approaches in segmentation, detection, and video understanding. Compared with previous pixel-based approaches, transformer-based approaches have more advantages in cases of flexibility, simplicity, and uniformity . One representative work is detection transformer (DETR) . It contains a CNN backbone, a standard transformer encoder, and a standard transformer decoder. It also introduces the concepts of object query to replace the anchor design in pixel-based approaches. Object query is usually combined with bipartite matching during training, uniquely assigning predictions with ground truth. This means each object query builds the one-to-one matching during training. Such matching is based on the matching cost between ground truth and predictions. The matching cost is defined as the distance between prediction and ground truth, including labels, boxes, and masks. By minimizing the cost with the Hungarian algorithm , each object query is assigned by its corresponding ground truth. For object detection, each object query is trained with classification and box regression loss . For instance-wised segmentation, each object query is trained with classification loss and segmentation loss. The output masks are obtained via the inner product between object query and decoder features. Recently, mask transformers further removed the box head for segmentation tasks. Max-Deeplab is the first to remove the box head and design a pure-mask-based segmenter. It combines a CNN-transformer hybrid encoder and a query-based decoder as an extra path. Max-Deeplab still needs many auxiliary loss functions. Later, K-Net uses mask pooling to group the mask features and designs a dynamic convolution to update the corresponding query. Meanwhile, MaskFormer extends the original DETR by removing the box head and transferring the object query into the mask query via MLPs. It proves simple mask classification can work well enough for all three segmentation tasks. Then, Mask2Former proposes masked cross-attention and replaces the cross-attention in MaskFormer. Masked cross-attention forces the object query only attends to the object area, guided by the mask outputs from previous stages. Mask2Former also adopts a stronger Deformable FPN backbone , stronger data augmentation , and multiscale mask decoding. In summary, query-based approaches are stronger and simpler. They cannot directly detect novel classes but are widely used as base detectors and segmenters in open vocabulary settings.
2 Vision Language Modeling
Large Scale Visual Language Pre-training. Better visual language pre-training can lead to a better understanding of semantics on given visual inputs . Previous works focus on cross-modality research via learning the visual and sentence-dense connection between different modalities. Some works utilize two-stream neural networks based on the vision transformer model, while several works adopt the single-stream neural network, where the text embeddings are frozen. The two-stream neural networks process visual and language information and fuse them afterward by another transformer module. Recently, most approaches adopt pure transformer-based visual language pre-training. Both CLIP and Align are concurrent works that explore extremely large-scale pre-training on image text pairs. In particular, with such large-scale training, CLIP demonstrates that the simple pre-training task of predicting which caption goes with which image can already lead to stronger generalizable models. Several following works aim to improve the CLIP training via mask image modeling, scaling up the training data and model size. These VLMs are the foundation of open vocabulary learning in different tasks, which means that the open vocabulary approaches aim to distill or utilize knowledge of VLMs into their corresponding tasks.
Moreover, in addition to achieving better zero-shot recognition, several works focus on designing better vision language models for language-related tasks, including visual question answering (VQA). Several works explore how to better align caption loss and contrastive loss during the image-text pertaining. In particular, CoCa adopts cross attention to connect caption generation part and contrastive learning. Recently, BLIP-2 bootstraps vision-language pre-training from off-the-shelf frozen pre-trained image encoders and frozen large language models via a lightweight Querying Transformer.
Visual Grounding Tasks. Visual grounding tasks aim to localize the specific objects according to given text descriptions. Referring segmentation tasks segment the specific object. Referring expression comprehension localizes the bounding boxes of given object texts. Recent research is mainly based on two-stream networks: one for visual encoder and the other for text encoder . These works focus on how to better match text features and visual features via different architectures. For referring segmentation, previous works adopt “decoder-fusion”, where they design a separate decoder at the end of the two-stream networks. Recent works explore the “encoder-fusion” to directly fuse the text features into a visual backbone before mask prediction. For referring expression comprehension, similar approaches are adopted such as encoder-fusion. Recently, several works aim to unify visual grounding and object detection. In particular, Grounding DINO unifies detection and grounding in one framework. It uses stronger detectors with multiple dataset pre-training, which achieves strong results for many downstream tasks. Compared with open vocabulary learning, visual grounding tasks require visual text matching with specific text descriptions. Open vocabulary learning tasks require the model to automatically detect, segment and recognize new objects without the given text information, such as class names, which is more challenging.
Turning Close-set Detector and Segmenter Into Open Vocabulary Setting. A common way towards the open vocabulary setting is to replace the fixed classifier weights with the text embeddings from a VLM model. In Fig. 5, we present a meta-architecture. In particular, the vision model generates a visual embedding for each box/mask proposal and computes similarity scores by computing a dot product with the text embeddings from both base and novel classes. The classification scores are computed as follows:
where is the -th vision embedding and is the -th text embedding, and is the prediction score of the -th vision proposal predicted to the -th class. The represents the dot product operation. The denominator is added with one because the background text embedding is set to all 0 or make it learnable. and are base and novel classes, respectively. Finally, the class with the highest prediction score is chosen as the prediction. In practice, the vision embeddings are trained with base annotations to fit the text embeddings. Therefore, the model combines the knowledge of VLM and the learned visual features, and the detector/segmenter can detect/segment novel classes via semantically related text embeddings.
3 Open Vocabulary Object Detection
In this section, we divide the methods into five categories: knowledge distillation, region text pre-training, training with more balanced data, prompting modeling, and region text alignment. Finally, we summarize the common features in Tab. II.
Knowledge Distillation. These techniques aim to distill the knowledge of Vision-and-Language Models (VLMs) into close-set detectors . Since the knowledge of VLMs is much larger than close-set detectors, distilling novel classes into based classes trained detectors is a straightforward idea. Knowledge distillation aims to distill visual knowledge directly into close-set detectors since the visual features are aligned with text features during the VLM pre-training stage. One earlier method is ViLD , a two-stage detection approach that utilizes instance-level visual-to-visual knowledge distillation. ViLD consists of two branches: the ViLD-text branch and the ViLD-image branch. In the former branch, fixed text embeddings obtained from VLMs’ text encoder output are treated as classifiers. Meanwhile, in the latter branch, pre-computed proposals are fed into a detector to obtain region embeddings using the RoIAlign function. The cropped image is then sent to the VLMs’ image encoder to generate image embeddings. Subsequently, ViLD proposes distilling this information onto each Region-of-Interest (RoI) via Loss. LP-OVOD extends the ViLD framework by making two main modifications. Firstly, it replaces the softmax cross entropy loss with the sigmoid focal loss . Secondly, LP-OVOD introduces a new classification branch that is supervised by pseudo labels. HierKD uses a single-stage detector and introduces a global-level language-to-visual knowledge distillation module. The module aims to narrow down performance gaps between one-stage and two-stage methods. This technique employs global-level knowledge distillation modules (GKD), which align global-level image representations with caption embeddings through contrastive loss. Both ViLD and HierKD use pixel-based detectors. However, Rasheed et al. leverages query-based detector Deformable DETR for its detection process. To ensure consistency between their detection region representations and CLIP’s region representations, the authors employed inter-embedding relationship matching loss (IRM). Furthermore, the authors adopt mixed datasets pre-training to enhance the ability of novel class discovery. OADP thinks previous works only distill object-level information from VLMs to downstream detectors and ignore the relation between different objects. To tackle this problem, OADP employs object-level distillation as well as global and block distillation methods. These supplementary techniques aim to compensate for the lack of relational information in object distillation by optimizing the L1 distance between the CLIP visual encoder and the detector backbone’s global features or block features. Rather than using simple novel class names for text distillation, several works also utilize more fine-grained information, including attributes, captions, and relationships of objects. PCL adopts an image captioning model to generate more comprehensive captions that describe object instances. OVRNet simultaneously detects objects and their visual attributes in open vocabulary scenarios. By exploring the COCO attributes dataset via a joint co-training strategy, the authors find that the recognition of fine-grained attributes works complementary for OVD. In summary, knowledge distillation is a common design, which effectively transfer VLM’s knowledge into close set detectors. However, the recognition ability is still within the scope of teacher VLMs.
Region Text Pre-training. Another assumption of open vocabulary learning is the availability of large-scale image text pairs, which can be easily obtained in daily life. Since these pairs contain large enough knowledge to cover the most novel or unseen datasets for detection and segmentation. Most approaches adopt web-scale caption data for pre-training, which contains millions of image text pairs. The learning of region text alignment maps the novel classes of visual features and text features into an aligned feature space. Once trained for the alignment, it is nature to generalize the detector for novel class classification. OVR-CNN first introduces the concept of open vocabulary object detection by using caption data for novel class detection. The model first trains a ResNet and vision to language (V2L) layer using image-caption pairs via grounding, masked language modeling, and image-text matching. Since captions are not constrained in language space, the V2L layer learns to map features from visual space into semantic space without limiting to closed-label space. During the next stage of training, Faster R-CNN is used as a detection method with pre-trained ResNet as its backbone. OVR-CNN replaces only learnable classifiers with fixed text embeddings from pre-trained language models . For classification purposes, region visual features obtained from RoI-Align are sent into the V2L layer and mapped into semantic space. Attribute-Sensitive OVR-CNN proposes a different approach from OVR-CNN . Instead of grounding vision regions to the input word embeddings of BERT , Attribute-Sensitive OVR-CNN aligns vision regions with contextualized word embeddings that are output from BERT . Additionally, Attribute-Sensitive OVR-CNN proposes using an adjective-noun negative caption sampling strategy to enhance the model’s sensitivity to adjectives, verb phrases, and prepositional phrases other than object nouns in the caption. GLIP series unify the object detection and phrase grounding for pre-training. In particular, it leverages massive image-text pairs by generating grounding boxes in a self-training fashion, which sets strong results for both detection and grounding. RegionCLIP learns visual region representation by matching image regions to region-level descriptions. It creates pseudo labels by CLIP for region-text pairs and then uses contrastive loss to match them before fine-tuning the visual encoder using human-annotated detection datasets. OWL-ViT removes the final token pooling layer of pre-trained VLMs’ image encoder and attaches a lightweight classification head and box regression head to each transformer output token before fine-tuning it on standard detection datasets using bipartite matching loss. Meanwhile, MaMMUT presents a simple text decoder and visual encoder for multimodal pre-training. Designing a two-pass text decoder combines both contrastive and generative learning in one framework. The former is for grounding text visual entities, while the latter learns to generate. DITO presents a new image-level pretraining strategy to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pertaining phase, DITO replaces the classification architecture used in CLIP with the detector architecture, which better serves the region-level recognition needs of detection by enabling the detector heads to learn from noisy image-text pairs. In summary, adopting more text-image pair can improve the performance on rare and novel classes. However, such process needs more computation cost for extra dataset training.
Training with More Balanced Data. Rare and unseen data are common in image classification datasets. Joint training can be used to address this issue. The core idea of these approaches is to leverage more balanced data, including image classification datasets, pseudo labels from image-text data, extra-related detection data, or even data generated by generation models. Detic improves long-tail detection performance with image-level supervision. The classification head of Detic is trained using image-level data from ImageNet21K . During training, the max area proposal from RPN is chosen as the RoI. Detic’s basic insight is that most classification data are object-centric. Therefore, the maximum area proposal may completely cover an object represented by an image-level class. The mm-ovod method improves upon Detic by utilizing multi-modal text embeddings as the classifier. This approach employs a large language model to create a description of each class to generate text-based embeddings. In addition, mm-ovod also uses vision-based embeddings from image exemplars. By fusing the text-based and vision-based embeddings, the multi-modal text embeddings can significantly enhance Detic’s performance. To use more data, several methods generate pseudo bounding box annotations from large-scale image-caption pairs. PB-OVD uses Grad-CAM and generates the pseudo-bounding boxes combined with RPN’s proposals. The activation map of Grad-CAM is obtained from the alignment between region embeddings and word embeddings that come from pre-trained VLMs. Then, the boxes are generated from the activation map and jointly trained with existing box annotations. Meanwhile, several works leverage the rich semantics available in recent vision and language models to localize and classify objects in unlabeled images. VL-PLM proposes to train Faster R-CNN as a two-stage class-agnostic proposal generator using a detection dataset without category information. LocOV uses class-agnostic proposals in RPN to train Faster R-CNN by matching the region features and word embeddings from image and caption, respectively. From the data generation view, several works adopt the diffusion model to generate the on-target data for effective training. In particular, X-Paste generates the rare class data to improve the classification ability for existing approaches. It uses an extra segmentation model to provide the foreground object masks and adopt simple copy and paste to augment training data. Recently, several works have explored the self-training approaches to generate large pseudo labels for more balanced learning. OWLv2 presents a self-training pipeline. It generates huge pseudo-box annotations on WebLI dataset and pre-trains a model on such generated dataset. Finally, it fine-tunes this model on a specific OVD dataset. Since the WebLI contains a large vocabulary size, OWLv2 achieves significant gains on LVIS and COCO datasets. In summary, these approaches are more effective in unbalanced data setting. Designing a more efficient data augmentation still have room to explore.
Prompting Modeling. Prompt modeling is an effective technique for adapting foundation models to various domains, such as language modeling and image classification . By incorporating learned prompts into the foundation model, the model can transfer its knowledge to downstream tasks more easily. To generate text embeddings of category names, prompts are fed to the text encoder of pre-trained VLMs. However, negative proposals do not belong to any specific category. To address this issue, DetPro forces the negative proposal to be equally dissimilar to any object class instead of using a background class. PromptDet introduces category descriptions into the prompt and explores the position of the category in the prompt. It also proposes to use cached web data to enhance the novel classes during training. Based on the DETR framework, CORA proposes region prompting and anchor pre-matching. The former reduces the gap between the whole image and region distributions by prompting the region features of the CLIP-based region classifier, while the latter learns generalizable object localization via a class-aware matching mechanism. Prompt-OVD follows the pipeline of OV-DETR . It presents RoI-based masked attention and RoI pruning techniques by utilizing CLIP visual features to improve the novel object classification. Despite the effectiveness, without more data training, the performance of prompting is limited compared with other directions.
Region Text Alignment. Using language as supervision instead of a ground truth bounding box is an attractive alternative for open vocabulary object detection. However, obtaining enough object-language annotations is difficult and costly. Compared with region text pre-training, region text alignment aims at a better matching between region visual features and text features during the base class training without introducing extra data. OV-DETR introduces a transformer-based detector for open vocabulary object detection by replacing the bipartite matching method with a conditional binary matching mechanism. VLDet converts the image into a set of regions and the caption into a set of words, and solves the object-language alignment problem using a set matching method. It proposes a simple matching strategy to align caption and vision features. DetCLIPv2 uses ATSS as an object detector and trains it with three datasets: a standard detection dataset, a grounding dataset, and an image-text pairs dataset for word-text alignment. Recently, BARON proposes to align the embedding in bags of different regions rather than only individual regions. It first groups contextually interrelated regions as a bag and treats each region in the bag as a word in a sentence. Then, it sends the bag of regions into the text encoder to get bag-of-regions embeddings. These bag-of-regions embeddings will align with cropped region embeddings from the image encoder of VLMs. CoDet reformulates the region-word alignment as a co-occurring object discovery problem and aligns the co-occurring objects with the shared concept. F-VLM finds that the origin CLIP features already have grouping effects. It is a two-branch method similar to ViLD-text . F-VLM uses a CLIP vision encoder as the backbone and applies the VLM feature pooler on the region features from the backbone to get VLM predictions. The final result of F-VLM combines the detection scores and the VLM predictions. RO-ViT builds upon F-VLM , which believes that the difference in position embeddings between image-level and region-level is responsible for the gap between vision-language pre-training and open vocabulary object detection. To address this issue, RO-ViT proposes a cropped positional embedding module in VLM that bridges the gap between vision-language pre-training and downstream open vocabulary object detection tasks. In summary, region-text alignment is a core research topic, and it still has room to explore when considering recent large language models .
4 Open Vocabulary Segmentation
Although segmentation problems can be defined in various ways, such as semantic segmentation or panoptic segmentation, we categorize the methods based on their technical aspects.
Utilizing VLMs to Leverage Recognition Capabilities. VLMs have shown remarkable performance in image classification by learning rich visual and linguistic representations. Therefore, it is natural to extend VLMs to semantic segmentation, which can be seen as a dense classification task. For instance, LSeg aligns the text embeddings of category labels from a VLM language encoder with the dense embeddings of the input image. This enables LSeg to leverage the generalization ability of VLMs and segment objects that are not predefined but depend on the input texts. Following LSeg, several works propose methods to utilize the VLMs for open vocabulary segmentation tasks. Fusioner utilizes self-attention operations to combine visual and language features in a transformer-based framework at an early stage. ZegFormer decouples the problem into a class-agnostic segmentation task and a mask classification task. It uses the label embeddings from a VLM to classify the proposal masks and applies the CLIP-vision encoder to obtain language-aligned visual features for them. After that, more approaches are proposed to extract the rich knowledge in VLMs. SAN attaches a lightweight side network to the pre-trained VLM to predict mask proposals and classification outputs. CAT-Seg jointly aggregates the image and text embeddings of CLIP by fine-tuning the image encoder. Han et al. develop an efficient framework that does not rely on the extra computational burden of the CLIP model. OPSNet proposes several modulation modules to enhance the information exchange between the segmentation model and the VLMs. The modulation modules fuse the knowledge of the learned vision model and VLM, which improves the zero-shot performance in novel classes. MaskCLIP inserts Relative Mask Attention (RMA) modules into a pre-trained CLIP model. It can utilize the CLIP features more efficiently. TagCLIP proposes a trusty token module to explicitly predict pixels that contain objects (both base and novel) before classifying each pixel, which avoids the problem that models tend to misidentify pixels into novel categories. Recently, several works directly fuse the frozen CLIP visual encoder. They fuse the learned CLIP score and prediction score to achieve better close-set and open-set recognition ability trade-offs. Same as the open vocabulary detection, there are still several improve space when adopting more advanced VLMs.
Learning from Caption Data. Image caption data can provide weak supervision for identifying novel classes, as they expose potential novel category names. Like such approaches in open vocabulary object detection , image caption data is also explored in open vocabulary segmentation tasks also explore image caption data. OpenSeg applies a region-word grounding loss to directly ground objects and nouns in the caption data. CGG combines caption grounding and caption generation losses to fully exploit the knowledge in caption data. It leverages the role of object nouns in visual grounding and the mutual benefits of words in caption generation.
Generating Pseudo Labels. Intuitively, providing the model with more data on novel categories can improve classification performance. MaskCLIP+ modifies the CLIP-vision model by replacing the last pooling layer with a convolution layer, which produces dense feature maps. These feature maps are then used to generate pseudo labels for training a segmentation model. OVSeg matches the proposed image regions with nouns in captions using CLIP to generate pseudo annotations. It also proposes a mask prompt tuning module to help CLIP adapt to masked images without changing their weights. XPM follows a similar pseudo data generation procedure by aligning words in captions with regions in images to generate pseudo instance mask labels. However, these approaches are limited by the data annotations, since both mask and region caption are hard to collection. Maybe more automatic data generation can be used in the future.
Training without Pixel-Level Annotations. Despite the absence of annotations on novel objects, most methods still require base mask annotations during training, which are costly and labor-intensive to obtain. To reduce the annotation burden, many recent works explore training segmentation models with merely weak supervision, such as image caption. GroupViT proposes a semantic segmentation framework that leverages a grouping mechanism to automatically merge image patches with the same semantics. It trains the model with a contrastive loss on image-caption pairs and does not need pixel-level annotations. PACL enhances the contrastive loss with a patch alignment objective that aligns the image patches and the CLS token of the captions. ViL-Seg combines both contrastive loss and clustering loss. SegCLIP further introduces a reconstruction loss and a superpixel-based KL loss. To compute the superpixel-based KL loss, it uses an unsupervised graph-based segmentation model to generate pseudo labels. TCL proposes a finer-grained contrastive loss, namely text-grounded contrastive loss, which explicitly aligns captions and regions. The model can segment the region that corresponds to a given text expression during inference. OVSegmentor introduces masked entity completion and cross-image mask constituency tasks to improve the training efficiency. Mask-free OVIS generates pseudo mask annotations with purely image-text pair data to assist the segmentation models. Since these methods are trained without mask supervision, the quality of segmentation results is low which make it hard to use in real application.
Jointly Learning Several Tasks. Open vocabulary segmentation encompasses different tasks, including OVSS, OVIS, and OVPS. How to jointly learn multiple segmentation tasks in one model becomes a practical problem. X-Decoder proposes a framework that can handle various tasks, including open vocabulary semantic segmentation, open vocabulary instance segmentation, and open vocabulary panoptic segmentation. It utilized a query-based segmentation architecture, pre-trains the model on a mixture of segmentation data and image-text pairs, and then fine-tuned or applied in zero-shot settings for downstream tasks. FreeSeg also proposes a generic framework to tackle the three tasks in a unified manner. It designs an adaptive task prompt module and performs test time tuning on the learnable prompts to capture the task-specific features. POMP first trains class prompts at a large vocabulary dataset (Imagenet-21K) and then transfers the prompts into multiple open vocabulary tasks. Moreover, if the model can learn from multiple tasks, it raises the question of whether it can benefit from multi-sourced data, e.g., detection and segmentation data. To handle this question, OpenSeeD and OpenSD jointly learn from segmentation and detection datasets. To bridge the task gap, they propose decoupled decoding frameworks, which decode foreground and background masks separately and generate masks for the bounding box proposals. One shortcoming of these approaches is extra computation costs that brought other tasks.
Adopting Denoising Diffusion Models. Recently, diffusion-based generative models have achieved remarkable success in text-based image generation suggests. There are mainly two ways to let diffusion models enhance open vocabulary tasks. First, The reality and diversity of the images generated by diffusion models suggest that the intermediate representations in the diffusion models may be highly aligned with natural language vocabularies. Inspired by this, ODISE proposes a framework that leverages the middle representation of the diffusion model. The middle representations contain rich semantic information for both base and novel classes. Thus, ODISE can perform open vocabulary segmentation by training a decoder head on the representations. follows ODISE to use the intermediate representations of diffusion models. Another way to incorporate diffusion models into the open vocabulary setting is to take advantage of the image and mask generation ability . OVDiff proposes a prototype-based method to tackle open vocabulary semantic segmentation. It uses the diffusion model to generate images for various categories and treat them as prototypes. The method does not need training. During testing, the input images are compared with these generated prototypes, and the best-matching one is the predicted class. Meanwhile, several works use the diffusion model to generate images and masks for the rare classes to augment data. Thus, the segmenter can be trained in a more balanced manner or even totally using generated data. Despite the inference time of these methods is quite limited, it still has room to explore joint generation and segmentation in one framework.
5 Open Vocabulary Video Understanding
We also review several video open vocabulary tasks. Most works focus on designing VLM models’ temporal fusion or association in various settings, including action recognition and tracking.
Video Classification. Traditional video classification methods usually require large datasets specific to video (e.g., Kinetics ). However, annotating video datasets requires very high costs. Using semantic information of label texts from web data may alleviate this deficiency. Building on the success of CLIP in image open vocabulary recognition, ActionCLIP uses knowledge from image-based vision-language pre-training. It adds a temporal fusion layer for the zero-shot capability in video action recognition. A concurrent work I-VL follows a similar paradigm by training the transformer layer on top of a frozen CLIP image encoder and also has a strong zero-shot capability. Another concurrent work, EVL , also employs the frozen CLIP for efficient video action recognition. X-CLIP uses a video encoder that leverages the temporal information in the encoder for video recognition. It aligns the feature from the video encoder with the text encoder that is pre-trained on the image-based vision-language pairs (e.g., CLIP). To take advantage of other modalities, MOV further fuses the audio information and the pre-trained CLIP model to build a multi-modal open vocabulary video classification model. Recently, Open-VCLIP formulates the CLIP-to-video knowledge transfer as a continual learning problem and proposes Interpolated Weight Optimization to address the issue. AIM tackles the open vocabulary video classification problem by adding adapt layers on top of the CLIP image encoder. ViFi-CLIP reveals that a simple fine-tuning baseline (image-level feature extraction with CLIP visual encoder following temporal pooling) instead of advanced fusion layers can have a strong performance and further investigates the influence of prompt tuning. ASU explores the use of fine-grained language features extracted by semantic units (e.g., head, arms, balloon, knee bend posture of exercise in the video) to guide the training of video classification. Recent advancements in open vocabulary video classification have turned our attention to the language part of VLM modeling. VicTR introduces video-conditioned text representations to optimize the visual and text information jointly. MAXI leverages LLMs to build a text bag for video without annotation by text expansion. The verbs, which are lacking in the realm of image, have the opportunity to participate in video-language modeling.
Object Tracking and Video Instance Segmentation. The object tracking and video instance segmentation can also enjoy the rich knowledge of VLMs to build an open vocabulary tracker based on a close-set tracker. Going beyond the large vocabulary object tracking , to tackle the real-world multiple object tracking (MOT), OVTrack first introduces large VLMs to the object tracking and tackles their proposed open vocabulary MOT (OV-MOT) task. Specifically, OVTrack extracts RoIs via an RPN and uses CLIP for knowledge distillation. A separate tracking head is used for tracking and supervised by pseudo-LVIS videos. As for video instance segmentation (VIS), to make the VIS model capable of generalizing to novel classes in the real world, MindVLT adopts a frozen CLIP backbone and proposes an end-to-end method with an open vocabulary classifier for segmenting and tracking unseen categories. It collects a large-vocabulary VIS dataset and tests their MindVLT on it. A concurrent work, OpenVIS , also tackles the open vocabulary video instance segmentation task but in a different manner. It generates the class-agnostic mask of instances and leverages the masks to crop raw images to feed into the CLIP visual encoder for calculating class scores. Beyond the open vocabulary video instance segmentation, DVIS++ proposes the first open vocabulary universal video segmentation scheme supporting video semantic segmentation, video instance segmentation, and video panoptic segmentation. DVIS++ leverages the knowledge from the CLIP backbone and adopts mask pooling on the CLIP backbone for enabling open vocabulary capability.
6 Open Vocabulary 3D Understanding
In this section, we review the open vocabulary approaches used in 3D scene understanding tasks. We mainly survey the works for point cloud classification and segmentation, which explores the knowledge of 2D VLMs into 3D.
3D Recognition. The VLM has shown great success in the zero-shot or few-shot learning of 2D images by leveraging a huge amount of image data. However, such internet-scale data is not available for the 3D point cloud. To this extent, extending the open vocabulary mechanism into 3D perception is not trivial since it is hard to collect enough data for point-language contrastive training like what CLIP does in 2D perception. To mitigate the gap, PointCLIP takes the first step to 3D open vocabulary perception. The principle insight behind PointCLIP is that the 3D point cloud can be converted to CLIP-recognizable images. By projecting the 3D point cloud to a 2D plane and extracting visual features from the projected depth map via the CLIP visual encoder, the point cloud feature can be naturally aligned with the language feature extracted by the language encoder. Then, the whole framework is with 3D point cloud zero-shot capability, just like 2D images. However, simply projecting the point cloud to depth maps may yield inferior performance. To further improve the performance, rather than directly using the CLIP visual encoder for visual feature extraction of depth map, CLIP2Point aligns the RGB image feature from the CLIP visual encoder and depth feature from a depth encoder through contrastive learning. Specifically, they collect image-depth pairs to align the from-scratch depth map encoder with the CLIP image encoder at training time and only use the depth map encoder at inference time. Then, the depth feature can be aligned with the language embedding, facilitating the 3D point cloud zero-shot capability with CLIP. Also aiming at improving the performance of PointCLIP , PointCLIPV2 focuses on different perspectives and proposes a realistic shape projection scheme along with an LLM-based 3D prompting generation scheme. Although without extra annotations, PointCLIPV2 achieves a significant improvement over PointCLIP .
Though projecting the point cloud to depth maps can make it easy to leverage the existing 2D visual encoders pre-trained in the CLIP, it may not exhaustively use the 3D information in the point cloud. Accordingly, beyond the depth map-based methods, ULIP first collects multi-modalities triplets (point cloud, image, and text) for training the 3D backbone. Since CLIP already aligns the language and image encoders, it only needs to align the 3D backbone to the image-language feature space. With only small-scale data, ULIP aligns the 3D feature extracted by the 3D backbone to the CLIP-aligned visual and text feature to enable the zero-shot capability and enhance the standard 3D recognition capability. The advantage of ULIP also includes unifying the three modalities, which may bring more downstream applications, such as image-to-3D retrieval. Though ULIP aligns the 3D and 2D and thus enables the native 3D open vocabulary capability, it trains on a small-scale dataset, which limits its performance. CLIP2 resorts to finding training samples from the real world and proposes Triplet Proxies Collection scheme for finding instance-level 3D point cloud, 2D image crop, and text triplets for 3D recognition. OpenShape scales up both the dataset and backbone. For the dataset, they build a text-3D shape dataset containing 876k training shapes with over 1k categories. They also explore adopting larger 3D backbones and their performance. As a result, OpenShape drastically enhances the performance on 3D zero-shot capability. LidarCLIP also aligns the 3D and 2D features with a similar approach, but LidarCLIP focuses on the driving scenes. Besides the 3D zero-shot capability, LidarCLIP also shows point cloud captioning and lidar-to-image generation capability with off-the-shelf foundation models.
3D Object Detection. In 3D object detection, the training data is often hard to get, and existing 3D detectors are often trained on very limited classes, which means they are hard to generalize to novel classes. Inspired by the success of 2D open vocabulary detection , aiming at taking full use of both image and point cloud modalities for 3D detection, OV-3DETIC and OV-3DET proposes to decouple the localization and recognition in point cloud object detection into localization and recognition. For localization, 2D pre-trained detectors can be adapted to train the 3D detector by back-projecting the 2D bounding box and further optimizing it by the point cloud. For recognition, they propose aligning the regional features of 3D and 2D encoders with the corresponding text feature. Therefore, during the inference, 3DETIC and OV-3DET can generalize to unseen classes with the knowledge in the CLIP. While OV-3DET achieves significant performance on the Open-Vocabulary 3D detection, it relies on the localization capability of pre-trained 2D detectors, which may limit the 3D object discovery capability. To solve this problem, CODA proposes to discover objects with both 2D semantic prior and 3D geometry prior, and further proposes a cross-modal alignment module to align 3D and 2D features. Also trying to get rid of the 2D image detector, concurrent work Object2Scene proposes to augment the existing 3D object detection datasets with large-scale 3D object datasets. It also introduces a new framework, L3Det , for 3D object-text alignment. As a result, both CODA and L3Det outperforms OV-3DET . As driving scenarios may yield new challenges, the above-mentioned methods may have an inferior performance. OpenSight focuses on the outdoor scenes and proposes to leverage the knowledge from 2D grounding DINO to train the 3D detectors. In particular, it uses the size prior (e.g., cars are 1.96m wide) generated by the LLM to re-calibrate the output of the model to make the model more suitable for the driving scene.
3D Scene Understanding. In the understanding of the 3D scene, the same obstacle, lacking training data, also limits the capability of generalization. Sharing a similar motivation that there is a lack of 3D-text pairs for point-language contrastive training, PLA proposes to extract features from multi-view images sampled from a scene to generate descriptions with pre-trained language models. The descriptions can then be used to extract language features to train 3D backbone-language alignment. With the language-aligned 3D backbone, it is natural to conduct open vocabulary segmentation like with the 2D image. OpenScene proposes to link the 3D features for each point with CLIP features for each pixel. By back-projecting the 2D pixels to 3D space, each point in the point cloud can be ensembled with features of several pixels in different views. The 3D-text co-embedding makes it possible for open vocabulary semantic segmentation. Inspired by MaskCLIP , a concurrent work CLIP-FO3D also focuses on 3D semantic segmentation and has a similar approach. It leverages the feature map extracted by the CLIP visual encoder and directly transfers CLIP’s knowledge without any extra annotation. A recent work, RegionPLC on 3D open vocabulary semantic segmentation, considers the benefits of both PLA and OpenScene . It proposes to use the dense caption of any random region to extract more information in CLIP for 3D backbone distillation. Fine-grained dense supervision in RegionPLC beyond view-level or instance-level further improves performance compared to PLA . Different from training 3D backbones like in , PartSLIP and SATR directly projecting the 3D point cloud and mesh to 2D plane and uses GLIP to localize and segment. Meanwhile, OpenMask3D mainly focuses on the 3D instance segmentation. It first generates class-agnostic instance mask proposals with the 3D point cloud. Then, OpenMask3D selects views and projects 3D masks into 2D images. The 2D masks are further refined by SAM . Finally, the 2D masks can be fed into the CLIP visual encoder to generate label prediction with the help of the CLIP language encoder. While OpenMask3D has a strong performance, it requires point cloud-2D image pairs for training. On the contrary, OpenIns3D does not require 2D images but only needs point clouds with colors. OpenIns3D proposes a ”Mask-Snap-Lookup” pipeline that first learns class-agnostic mask proposals, then generates synthetic scene-level images, and finally assigns categories for each proposal. Another work trying to improve OpenMask3D is Open3DIS . Open3DIS improves the 3D mask proposal quality by employing a 3D instance network and a 2D-guide-3D Instance Proposal Module. After the mask proposal, it aggregates CLIP features in a multi-scale, multi-view manner. Open3DIS achieves a stronger performance over OpenMask3D on several datasets. In the driving scene, to mitigate the heavy dependence on the point cloud data annotation and fast generalization to the new scenes, CLIP2Scene proposes to train a 3D network via semantic-driven cross-modal contrastive learning for 3D segmentation. CLIP2Scene also considers spatial-temporal semantic consistency regularization to facilitate training.
7 Closely Related Tasks
Class Agnostic Detection and Segmentation. The goal of class agnostic detection and segmentation is to learn a general region proposal system that can be used in different scenes. For detection, OLN replaces the binary classification head in RPN by predicting the IoU score of foreground objects, which proves the generalization ability on cross-dataset testing. For example, the model trained on COCO shows the localization ability on the LVIS dataset. Open world instance segmentation aims to correctly detect and segment instances whether their categories are in training taxonomy. GGN leverages the bottom-up grouping idea, combining a local pixel affinity measure with instance-level mask supervision, producing a training regimen designed to make the model more generalizable. Meanwhile, from the data augmentation view, MViT combines the Deformable DETR and CLIP text encoder for the open world class-agnostic detection, where the authors build a large dataset by mixing existing detection datasets. For the segmentation domain, Entity segmentation aims to segment all visual entities without predicting their semantic labels. The goal of such segmentation is to obtain high-quality and generalized segmentation results. Fined-grained entity presents a two-view image crop ensemble method named CropFormer to enhance the fine-grained details. Recently, SAM proposes more generalized prompting methods, including masks, points, boxes, and texts. In particular, the authors build a larger dataset with 1 billion masks following the spirits of CLIP. The SAM achieves zero-shot testing in various segmentation datasets, which are highly generalizable.
Open World Object Detection. Open world recognition requires the model to identify novel categories and label them as “unknown.” Then, the novel categories are progressively annotated. The model learns incrementally with new data and recognizes the newly-annotated categories. Inspired by this setting, open world object detection (OWOD) expands the recognition task to object detection. The authors propose a novel method that utilizes contrastive clustering and an energy-based unknown identification module. Following them, OW-DETR introduces an end-to-end transformer-based framework. It consists of three dedicated components, namely, attention-driven pseudo-labeling, novelty classification, and objectness scoring, to address the OWOD challenge explicitly. Meanwhile, open world DETR proposes a two-stage training approach based on Deformable DETR. It focuses on alleviating catastrophic forgetting when the annotations of the unknown classes become available incrementally using knowledge distillation and exemplar replay technologies. PROB proposes a probabilistic objectness head into the OW-DETR to better mine unknown background classes. Recently, researchers propose unknown-classified open world object detection (UC-OWOD), which aims to detect and classify unknown instances into different unknown classes. This task is close to the open vocabulary object detection but requires incremental learning.
Open-Set Panoptic Segmentation. Similar to other open-set tasks, open-set panoptic segmentation (OSPS) requires the model to identify novel categories as ‘unknown’ in a panoptic segmentation task. The ’un classes are chosen from the thing classes (foreground objects). The authors apply exemplar theory for novel class discovery. After that, Dual proposes a divide-and-conquer scheme to develop a dual decision process for OSPS. The results indicate that by properly combining a known class discriminator with an additional class-agnostic object prediction head, the OSPS performance can be significantly improved.
CHALLENGES AND OUTLOOK
Base Classes Over-fitting Issues. Most approaches detect and segment novel objects by learning proposals from base class annotations. Thus, there are natural gaps in shapes and semantic information between novel and base objects. VLM models can bridge such gaps via pre-trained visual-text knowledge. However, most detectors still easily overfit the base classes when the novel classes have similar shapes and semantics, since these classes are trained with higher confidence scores. More fine-grained feature discriminative modeling , including parts or attributes, is required to handle these problems.
Training Costs. Most state-of-the-art methods need huge data for pre-training to achieve good performances. However, the costs are expensive and even unavailable for many research groups to follow. Thus, with the aid of VLM, designing more efficient data learning pipelines or learning methods is more practical and affordable. One simple solution can be adopting a frozen backbone. However, this may limit the representation capacity.
Across Dataset Generation and Evaluation. As shown in the benchmark section, current state-of-the-art methods design specific models for each benchmark. Designing one shared model across different datasets on open vocabulary detection and segmentation follows the origin spirits of open vocabulary learning. However, there are still performance gaps between unified models and dataset-specific models in several OVD benchmarks, such as LVIS datasets.
Better Benchmarks and Metrics. Since several classes contain overlapping concepts (for example, building and tower, person and woman), designing more new metrics is also needed to better measure the open vocabulary methods. Moreover, current datasets are still small. To realize real open vocabulary settings, more datasets like SAM-1B are needed.
2 Future Work
Explore Temporal Information. In practical applications, video data is readily available and used more frequently. Accurately segmenting and tracking objects that are not predetermined requires a great deal of attention, which is necessary for a wide range of real-world scenarios, such as short video clips and autonomous vehicles. However, there are only a few works exploring open vocabulary learning on detection and tracking in video. Moreover, the input scenes are simple. For example, several clips only contain a few instances, which makes the current tracking solution more trivial. Thus, a more dynamic, challenging video dataset is needed to fully explore the potential of vision language models for open vocabulary learning.
3D Open Vocabulary Scene Understanding. Compared with image and video, point cloud data are more expensive to annotate, in particular for dense prediction tasks. Thus, research on 3D open vocabulary scene understanding is more urgent. Current solutions for 3D open vocabulary scene understanding focus on designing projection functions for better usage of 2D VLMs. More new solutions for aligning 2D models knowledge into 3D models will be a future direction.
Explore Foundation Models With Specific Adapter For Custom Tasks. Vision foundation models can achieve good zero-shot performance on several standard classification and segmentation datasets. However, for several custom tasks, such as medical image analysis, and aerial images, there are still many corner cases. Thus, designing a task-specific adapter is needed for these custom tasks. Such adapters can fully utilize the knowledge of pre-trained foundation models. One possible solution is to explore the in-context learning to fully explore or connect the knowledge of VLMs and LLMs.
Combining with Incremental Learning. In real scenarios, the data annotations are usually open-world and non-stationary, where novel classes may occur continuously and incrementally. However, directly turning into incremental learning may lead to catastrophic forgetting problems. On the other hand, current open world object detection only focuses on novel class localization, rather than classification. How to handle both catastrophic forgetting problems and novel class detection in one framework is worth exploring in the future.
Combining with Large Language Models. Compared with VLMs, most LLMs contain more text concepts, which naturally have a broader scope than various dataset taxonomies, even larger than the recent V3Det dataset. Thus, how to better align the LLMs knowledge with visual detectors or segmenters to achieve stronger zero-shot results still needs exploration.
Conclusion
This survey offers a detailed examination of the latest developments in open vocabulary learning in computer vision, which appears to be a first of its kind. We provide an overview of the necessary background knowledge, which includes fundamental concepts and introductory knowledge of detection, segmentation, and vision language pre-training. Following that, we summarize more than 50 different models used for various scene understanding tasks. For each task, we categorize the methods based on their technical viewpoint. Additionally, we provide information regarding several closely related domains. In the experiment section, we provide a detailed description of the settings and compare results fairly. Finally, we summarize several challenges and also point out several future research directions for open vocabulary learning.
Acknowledgement. This work is supported by the National Key Research and Development Program of China (No. 2023YFC3807600) and the interdisciplinary doctoral grants (iDoc 2021-360) from the Personalized Health and Related Technologies (PHRT) of the ETH domain.
Appendix A Benchmark Results
This section systematically compares the different settings and methods in each task. For each setting, first, we introduce the datasets used in the evaluation. Next, we present the details of each setting separately. Then, we compare the results in detailed tables. We list the detailed results and settings for reference.
Settings. Open vocabulary object detection methods, such as OVR-CNN and ViLD , have primarily focused on the COCO , LVIS and V3Det datasets. The COCO dataset comprises 80 classes, with 48 considered base classes and 17 treated as novel classes. Any annotations not labeled with base classes are removed from the training data. Therefore, for open vocabulary object detection using the COCO dataset, there are 107,761 training images containing 665,387 bounding box annotations of base classes and 4,836 test images with 28,538 bounding box annotations of both base and novel categories. On the other hand, the LVIS dataset is designed for long-tail object detection tasks and includes 1203 classes. 866 frequent and common ones serve as base categories, while the remaining rare ones (377) act as novel categories. The LVIS dataset is specifically designed for long-tail object detection tasks and consists of 1203 classes. Among these, there are 866 frequent and common categories that serve as base categories, while the remaining 377 rare ones serve as novel categories. The V3Det dataset is a vast vocabulary visual detection dataset that contains 13,204 categories, 243k images, and 1,753k bounding box annotations. In the open vocabulary setting, there are 6,709 base categories and 6,495 novel categories.
Evaluation Metrics. When evaluating open vocabulary object detection, we calculate the box mAP using an IoU threshold of 0.5 for the COCO dataset and mask mAP for the LVIS dataset. To assess a method’s ability to detect novel classes, we split the metric into two categories: and . Our main focus is on measuring .
Results Comparison. In Table IX, X, and XI, we present the performance of current open vocabulary object detection methods on COCO, LVIS, and V3Det. BARON outperforms other models on both datasets when using Faster R-CNN detector with the ResNet-50 backbone, achieving a high score of 42.7 on COCO and 22.6 on LVIS. FVLM achieves 28.0 on the COCO dataset and 18.6 on LVIS using only frozen backbone and detection data. When utilizing the ResNet50x64 backbone, F-VLM achieves a remarkable performance of 32.8 on the LVIS dataset, surpassing other methods that employ the Swin-B backbone and additional data like CC3M and ImageNet. And for one-stage detector methods such as HierKD and GridCLIP , there is still a significant gap compared to two-stage detectors.
A.2 Open Vocabulary Semantic Segmentation
Settings. There are two main settings in open vocabulary semantic segmentation, according to different train/test data choices. 1) Self-evaluation setting. Methods like ZegFormer and MaskCLIP+ train and test their models on the train/test split from the same dataset. For example, ZegFormer trains on base annotations on the COCO-Stuff training set and tests on the COCO-Stuff validation set with base and novel categories. 2) Cross-evaluation setting. Methods like OpenSeg and OVSeg train their models on COCO-Stuff and test on other datasets, such as ADE20K, Pascal VOC, and Pascal Context. There is also another setting that splits a dataset into multiple folds and performs n-fold evaluations. We do not compare the results under this setting considering only few works adopt it .
The most commonly used datasets for the cross-evaluation setting are ADE20K, Pascal Context, and Pascal VOC. The ADE20K dataset has 2k validation images that comprise various indoor and outdoor semantic categories. The full dataset contains 2,693 foreground and background classes. Open vocabulary works often evaluate their models with the 847 classes split, named A-847, and the 150 classes split, named A-150. The Pascal Context dataset has 5k validation images with 459 categories (PC-459). Some works also evaluate their models on the dataset with the 59 most frequent classes (PC-59). The Pascal VOC dataset has 20 classes used for evaluation, denoted as PAS-20. OpenSeg proposes only assigning pixels with a Pascal Context category to background classes. We denote this setting as PAS-20b.
Evaluation Metrics. We adopt mean Intersection-over-Union (mIoU) to evaluate a model’s performance on open vocabulary semantic segmentation. Specifically, for the cross-evaluation setting, the IoU averaged on all the categories in the evaluation datasets is listed because all the classes can be seen as novel. For the self-evaluation setting, the testing dataset usually contains both base and novel classes. For a fair comparison, we only report mIoUs on novel classes while highlighting that the number is in a different setting.
Results under the Self-evaluation Setting. Tab. XII shows the open vocabulary semantic segmentation results under the self-evaluation setting. MaskCLIP+ achieves the highest novel mIoU on COCO-Stuff, PASCAL-VOC, and PASCAL-Context, while in general, FreeSeg has the best scores on base and harmonic mIoU. PADing∗ means including complicated crop-mask image preprocess (CLIP-Image encoder).
Results under the Cross-evaluation Setting. Tab. XIII shows the open vocabulary semantic segmentation results under the cross-evaluation setting. SCAN achieves the highest mIoU of 14.0 on the A-847 dataset. SED achieves 22.6 on the PC-459 dataset and 35.2 on the A-150 dataset. X-Decoder achieves 65.1 and 97.9 on PC-59 and PAS-20 datasets. For the PAS-20b dataset, ODISE achieves the highest mIoU of 84.6. Note that the works at the bottom blank use large-scale extra data for training. For example, X-Decoder uses Conceptual Captions , SBU Captions , Visual Genome , and COCO Captions as its training set. It contains 4M image-text pairs. Among works training without pixel-level annotations, PACL achieves a much higher score than others. It uses the biggest extra training data that contains CC3M, CC12M, and YFCC.
A.3 Open Vocabulary Instance Segmentation
Settings. Following previous works , we adopt two settings in evaluating open vocabulary instance segmentation. 1) Constrained setting. The base and novel categories are evaluated separately. 2) Generalized setting. The base and novel categories are evaluated together. Concretely, in the generalized setting, all the category names are input to the model. The model needs to not only classify novel classes but also to distinguish novel classes from base ones. The second setting is harder than the first one.
Evaluation Metrics We adopt mask mean Average Precision (mAP) as the evaluation metric. For the constrained setting, we report and separately. For the generalized setting, we report mAP for the base, novel, and all classes.
Results on COCO. Tab. XIV shows the open vocabulary instance segmentation results on the COCO dataset. The CGG method achieves the best results on both constrained and generalized settings while not using pre-trained VLMs or any extra data. And Mask-free OVIS gets a relatively high score on novel classes without any mask labels.
A.4 Open Vocabulary Panoptic Segmentation
Settings. The models are trained using COCO-Panoptic, and tested on other datasets in a zero-shot manner. We report both results evaluated on the ADE20K and COCO datasets.
Evaluation Metrics Following previous works , we mainly adopt Panoptic Quality (PQ), Segmentation Quality (SQ), and Recognition Quality (RQ) as the evaluation metrics.
Results. Tab. XV shows the results on the ADE20K dataset. ODISE-cap achieves the best PQ score of 23.4. It surpasses the second-best score by 0.8. Tab. XVI shows the results on the COCO dataset. PADing achieves a better PQ of 41.5 for seen classes, while Freeseg achieves the highest PQ score of 29.8 for unseen classes.
A.5 Open Vocabulary Video Recognition
Settings. In video classification, existing methods usually test the zero-shot capability on downstream datasets (e.g., UCF and HMDB) pre-trained on Kinetics-400 datasets. In most of the existing methods such as , the ground truth of the Kinetics-400 is used for training, while recent method MAXI does not require the ground truth when training on Kinetics-400.
Results Comparison. The comparison results are in Tab.V. Among the existing methods, Open-VCLIP performs best on all three downstream datasets. It is also worth noting that the MAXI performs well even without annotations from the Kinetics-400.
A.6 Open Vocabulary Video Instance Segmentation
Settings. To test the performance of open vocabulary video instance segmentation methods, MindVLT collects LV-VIS datasets inheriting the categories of LVIS and splits the categories into 659 base categories (frequent and common) and 553 novel categories. In the evaluation protocol, the LV-VIS is not used for training but for evaluation only. The mean Average Precision (mAP) is reported for comparison. and refer to mAP for base and novel categories, respectively.
Results Comparison. The comparison results are in Tab.VI. Among the methods, MindVLT achieves the best results on both base and novel categories. MindVLT shows stronger improvement in the novel categories.
A.7 Open Vocabulary 3D Recognition
Settings. For open vocabulary 3D recognition, the zero-shot classification results on three different scale datasets, including ModelNet40 , ScanObjectNN , and Objaverse-LVIS are reported. The first two datasets have 40 and 15 common categories, and the Objaverse-LVIS has 1156 LVIS categories. The methods are trained on different datasets but tested on the three datasets for fair comparison.
Results Comparison. As in Tab.VII, using the large-scale OpenShape datasets and larger backbones can significantly boost the performance on the downstream testing datasets.
A.8 Open Vocabulary 3D semantic segmentation
Settings. To test the open vocabulary 3D semantic segmentation methods, two datasets, ScanNet and nuScenes , are adopted for covering broad application scenarios. The ScanNet and nuScenes have 19 and 15 categories, respectively. The categories are split into base and novel categories for testing. For example, B15/N4 refers to 15 base categories for training and four novel categories that are missing in the training set. The mIoUB (base category mIoU), mIoUN (novel category mIoU), and hIoU (harmonic mIoU) are reported for comparison.
Results Comparison. As in Tab. VIII, RegionPLC achieves remarkable performance on novel categories and thus indicates that it has a good capability to generalize to unseen categories.