DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, Hang Xu

Introduction

Most state-of-the-art object detection methods can only recognize and localize a pre-defined number of categories. Their detection performance greatly relies on sufficient training data for each category, which requires expensive and time-consuming human annotations, especially for the rare classes that can only be distinguished from the expert. Even though a great effort has been made, existing publicly available detection datasets only have a limited number of object categories, for instance, 80 for COCO , 365 for Object365 , and 1203 for LVIS . However, versatile open-world object detection is still out of reach mainly for two reasons: a) Due to the long-tail problem , finding sufficient examples for rare categories is surprisingly challenging; b) Developing a larger detection dataset requires extremely costly and labor-intensive manual annotation.

On the other hand, image-text pair data are cheap and abundant on the Internet. Recent vision-language (VL) pre-training methods (e.g., CLIP , ALIGN ) utilize those data to extend their open-domain capacity and have shown good zero-shot ability on various downstream classification tasks. Intuitively, some works try to extend a two-stage detector to an open-world detector by distilling the learned image embeddings of the cropped proposal regions from a pre-trained VL model. However, this paradigm requires costly feature extraction from cropped images. There also exists discrepancy between instance-level features and conventional image-level features extracted from VL models. proposes a self-training pipeline with pseudo labels (i.e., noun phrases in caption) generated by a pre-trained VL model while the phrases in caption are always too limited to cover all the objects in an image. Moreover, their open-domain ability is determined by the performance of the pre-trained VL models.

Another line of work, GLIP further proposes a grounding formulation of open-world detection by directly utilizing both detection and grounding data for pre-training. Specially for detection data, GLIP takes an image and a text prompt sentence that concatenates all the category names as input (i.e., sequential formulation). A text encoder will output features for this sentence and then GLIP aligns them with the region features extracted from the image encoder. However, restricted by the max length of the input token size of the text encoder, it is difficult for GLIP to operate on a large number of categories or extend to more detailed description of the categories. Furthermore, full attention matrix upon all the categories has to be learned in the text encoder, which is unnecessary and low-efficient for detection especially when the size of input categories increases.

To alleviate the above problems and further improve the open-domain ability, we present DetCLIP, a dictionary-enriched visual-concept paralleled pre-training method for open-world detection. Note that the “concepts” denotes the category names in detection data, and the phrases in grounding and image text pair data. Specifically, we first design a novel paralleled concept formulation to improve learning efficiency. Instead of feeding the whole prompt text sentence into the text encoder like GLIP, DetCLIP extracts each concept separately and parallelly feeds them into the text encoder (i.e., paralleled formulation). This paralleled formulation allows the model to avoid unnecessary interaction between uncorrelated categories/phrases, and produce a longer description for each concept. By converting detection data, grounding data, and image-text pair data into paralleled formulation, DetCLIP can be pre-trained under different types of supervision to support both localization and open-domain capability.

During pre-training, existing datasets often have a large domain gap and difference in their labeling space, e.g., the same object concept with different names or hierarchical/inclusive structures of different concepts. However, it is difficult to obtain these implicit relationships among concepts simply by their short names. In order to form a more unified concept space and provide prior knowledge (i.e., implicit relationships) for each input concept, we propose a novel concept dictionary to enrich our prompt text concepts during joint pre-training, as shown in Fig.2. Firstly, we construct the dictionary with concepts extracted from online resources and existing large-scale detection datasets by considering both commonality and coherence. Based on this concept dictionary, DetCLIP can automatically enrich the current concept with the concepts and their descriptions existing in the dictionary to facilitate open-domain learning. To alleviate the partial label problem for the grounding and image-text pair data, DetCLIP further randomly samples concepts from the dictionary as negative samples to efficiently pre-train with the alignment loss. Besides, all concepts in the dictionary are served as the additional category inputs for label completion on the image-text pair data.

The proposed DetCLIP outperforms the state-of-the-art GLIP by large margins on large-scale open-world detection benchmark LVIS with 1203 classes. Particularly, DetCLIP-T achieves around 9.9% mAP improvement over GLIP-T on LVIS without utilizing any images in LVIS during pre-training. Compared to the baseline ATSS model (with the same swin-T backbone) trained on LVIS, our DetCLIP-T achieves 2.3% mAP improvement in the total performance and 13.5% mAP improvement on the rare classes.

Related Work

Vision-Language Pre-training (VLP). Current Vision-Language Pre-training is a natural extension and development of the successful pre-train-and-fine-tune scheme in the domains of natural language processing (NLP) and computer vision community. Dual-stream methods such as CLIP and ALIGN have shown great zero-shot classification ability by performing cross-modal contrastive learning on large-scale image-text pairs from the Internet. Single-stream approaches directly model the interaction between the visual and textual embeddings together by a single transformer-based model, which can perform tasks such as image caption and VQA. Recent approaches such as VLMo and BLIP further explore a mixed architecture of single-stream and dual-stream models to enable a unified way of vision-language understanding and generation. However, those approaches usually focus on whole-image representation learning and the pre-trained models are usually designed for retrieval/generation tasks. Current vision-language pre-training approaches can not directly be applied to object detection task, i.e., a core computer vision task.

General Object Detection. Object detection is a core problem in computer vision. CNN-based object detection methods (one-stage detectors: YOLO , SSD and ATSS and two-stage detectors: Faster R-CNN and R-FCN ) usually use a classifier to map ROI (Region Of Interest) features into categories. Further improvement of the CNN-based detectors such as K-Net and Panoptic-FCN further introduce dynamic kernels and replace the static kernels in the convolution layers to improve the flexibility of models. Recent transformer-based methods such as DETR and Deformable DETR try to formulate the object detection as a set prediction problem that can eliminate post-process NMS. Those methods are constrained to predefined categories, while our method tries to endow the detector with open-domain recognition ability which can detect any categories by learning a wide range of concepts.

Zero-shot Object Detection/ Open-vocabulary Object Detection. In the early setting of this area, zero-shot object detection aims to generalize the detector from known categories (training) to unknown categories (inference). Under this setting, various works try to find relationships between existing categories and unknown categories through pre-trained semantic/text features knowledge graphs and so on. However, the evaluation under this setting is not general enough since people usually simply split the class name into known/unknown categories in the single dataset and the transfer learning is still under similar domain. On the other hand, inspired by the success of vision-language~(VL) pre-training methods (e.g., CLIP ) and their good zero-shot ability, several methods attempt to perform zero-shot detection on a wider range of domains by leveraging a pre-trained VL model. For example, tries to distill the learned image embeddings of the cropped proposal regions from CLIP } to a student detector. proposes a self-training pipeline which utilizes Grad-CAM and ALBEF . However, those methods are very slow because the feature extraction is repeated and the image-level representation may be sub-optimal for the instance-wise tasks. Recently, GLIP and X-DETR try to align region and language features using a dot-product operation and can be trained end-to-end on both grounding data and detection data. Our method aims to design an open-domain detector that learns new concepts efficiently and expands domain coverage from low-cost data from the Internet.

The Proposed Approach

This paper aims to develop a vision-language pre-training pipeline to enhance new concept learning for open-world detection. To achieve this goal, our DetCLIP leverages a new paralleled vision-concept per-training pipeline (Sec.3.1) for efficient training and a concept dictionary (Sec.3.2) to provide external knowledge to automatically enrich the current input concept and alleviate the partial-labeling problem of the grounding and image-text pair data. Our DetCLIP is pre-trained under hybrid supervision from detection data, grounding data and image-text pair data (Sec.3.3).

A robust open-world detector is required to be trained with sufficient data covering enough vision concepts. Current available detection datasets lack enough concepts and have limited capacity due to the costly annotation. Leveraging data from other sources, e.g., grounding data and image-text pair data, is a feasible solution to augment the semantic coverage. To enable training with heterogeneous supervisions, we are required to find a unified formulation for different formats of data.

Fig.3 illustrates a comparison of the different concept formulation strategies that are used for pre-training with different types of data. As shown in Fig.3(a)-(b), traditional detection and grounding training utilize different input formulations. Detection treats each category as a fixed label and aligns region features to the pre-defined label space, while the grounding training leverages the whole caption sentence as input, models the attention between words, and adopts each token embedding in the output embeddings for alignment. To utilize both detection and grounding datasets for pre-training, GLIP converts the detection data to phrase grounding data, by replacing the object classification logits with the word-region alignment scores, which are associated with a sentence that concatenates all the category names (see Fig.3(c)), i.e., the text input is [“person, bicycle, car, … , toothbrush”].

We argue that this sequential form is not an effective formulation to model open-world object detection as a vision-language task because it (1) leads to unnecessary interaction between category names in the attention module; and (2) constraints the number of negative samples in contrastive learning due to the limited context length of text input. Ablation studies conducted for GLIP show that randomly shuffling the word order in the grounding training data can even bring a slight improvement for the downstream detection task, indicating that the noun phrase is more critical for the detection task, compared to the context information.

To address the above problem, our DetCLIP introduces a paralleled concept formulation to train with different data sources. As shown in Fig.3(d), we extract the concept noun phrase for each bounding box and then feed them into the text encoder individually to obtain the corresponding text embedding. In our paralleled design, context information is removed and model directly learns language features from each separate concepts, which is a more straightforward modeling for the detection task and improves the learning efficiency (see Fig.4). Furthermore, our paralleled design can enable easy expansion of the augment of category descriptions (see Sec.3.2.2). Detailed paralleled formulation for different types of data are as follows:

Detection data: Suppose there are kk positive categories in an image, we first pad the category number to NN by randomly sampling the negative categories, where NN is a predefined number of the concepts for constructing the alignment loss. More specifically, the text input PP can be formulated as {pn}n=1N\{p_{n}\}_{n=1}^{N}, where pnp_{n} represents for the nn-th category name, e.g.,

P=P= [“person”, “bicycle”, “car”, … , “toothbrush”].

Then, we feed NN category names as separate sentences into the text encoder and use the embedding of [end of sentence] token as the text embedding for each category. The NN text embeddings are then concatenated and matched with ground truth bounding boxes to construct the alignment loss.

Grounding data: We extract positive phrases for bounding boxes (provided by the grounding annotations), and drop other words in the caption. To align the input format with detection data, we also pad the category number to NN by sampling negative categories from our proposed concept dictionary (see Sec.3.2.2). An example of input PP for grounding data can be:

P=P= [“a woman”, “a herding dog”, “three cattle”, “neg1neg_{1}”, … , “negmneg_{m}”],

where neg1neg_{1} to negmneg_{m} are sampled negative categories from the constructed concept dictionary. Then similar steps with detection data can be conducted to construct the alignment loss.

Image-text pair data: Image-text pair data only contains images and the corresponding captions, while without any annotated bounding boxes. To obtain object-level dense labels, we first use a pre-trained class-agnostic region proposal network (RPN) to extract object proposals. Then a powerful pre-trained vision-language model like CLIP or FILIP is utilized to assign pseudo labels for the proposals (see Appendix for more details.). To tackle the concept-missing problem in the captions of image-text pair data, instead of using noun phrases in the captions as the candidate category names following , we propose to assign open-domain categories to the object proposals via constructing a large-scale concept dictionary (see Sec.3.2.2). After obtaining object-level dense pseudo labels, we convert the image-text pair data to the same format as the detection and grounding data, and then a similar training procedure can be applied .

2 Concept Dictionary

In this work, we try to build an open-world object detector that can cover a wide range of concepts and can be applicable to different types of datasets. However, the existing detection/grounding/image-text-pair datasets have a very large domain gap and difference in their labeling space. For example, a boy can be annotated as “man”, “child” or “people” in the different datasets. Moreover, there often exists hierarchical/inclusive relationships between different concepts. Knowing these implicit relationships can effectively facilitate the pre-training, however it is also clearly challenging to discover these relationships with only a limited set of concept names.

Therefore, we propose to build a large-scale concept dictionary to form a unified concept space for different data sources and explicitly provide the useful relationships between various concepts through definitions. For example, the car is defined as “a motor vehicle with four wheels usually propelled by an internal combustion engine”, and the motorcycle is defined as “a motor vehicle with two wheels and a strong frame”. By aggregating these definitions, we can conclude that the car and the motorcycle are both motor vehicles with a difference in the number of wheels.

To construct a unified concept dictionary OO, we collect concepts from multiple sources: (1) noun phrases extracted from the large-scale image-text pair dataset, i.e., YFCC100m; (2) category names of existing public detection datasets (e.g., Objects365 , OpenImages ); (3) object names from the manually-collected concept database, i.e., Things ). For object names from the detection datasets and Things, we directly add them into the dictionary after deduplication. For noun phrases extracted from the YFCC, we filter out the concepts that appear less than a frequency of 100100 or without definition in WordNet to ensure the commonality. Our concept dictionary OO is then constructed by concatenating each word olo_{l} with its definition defldef_{l} in WordNet: O={ol:defl}l=1L,O=\{o_{l}:def_{l}\}_{l=1}^{L}, where LL is the number of concepts in our dictionary. The constructed dictionary covers about 14k concepts along with their definitions, and can be updated by directly adding new concepts and corresponding definitions. Based on the constructed concept dictionary, we then propose 2 techniques utilizing it to boost the pre-training in the following section.

2.2 Knowledge Enrichment with Concept Dictionary

P∗=P^{*}= [“person, a human being.”, “bicycles, a wheeled vehicle that has two wheels and is moved by foot pedals.”, … , “toothbrush, small brush has long handle used to clean teeth.”]

Partial Annotation Enrichment. In the grounding or image-text pair data, only main objects that people care about are labeled in the caption, which is known as the partial labeling problem. Compared with standard detection datasets which have sufficient positive and negative classes for each image, pre-training with grounding and image-text pair datasets encounters two severe issues: 1) lack of annotations of negative concepts for learning discriminative concept embeddings; 2) lack of annotations of partial positive concepts to efficiently train the model. For the first problem, DetCLIP randomly samples the concepts in the constructed dictionary OO as the negative concepts to construct the alignment loss, instead of directly padding empty inputs (Fig.5(b)). Note that since the number of concepts in the dictionary OO is large (i.e., about 14k), the probability that the sampled concepts are indeed in the image is extremely small. For the second problem, to perform label completion on image-text pair data during pseudo labeling, we add all the concepts in dictionary {ol}l=1L\{o_{l}\}_{l=1}^{L} as the additional category inputs, instead of using the original noun phrase in the caption to calculate the similarity matrix. Therefore, the concepts shown in the image while not in the caption can also be labeled and then get pre-trained. An illustration is also shown in Fig.7 to qualitatively verify the effectiveness of label completion.

3 Model Architecture/Training Objective

where LALI\mathcal{L}_{ALI}, LCEN\mathcal{L}_{CEN}, and LREG\mathcal{L}_{REG} denote the alignment loss, the centerness loss, and the regression loss, respectively. α\alpha and β\beta represent the weight factor for LCEN\mathcal{L}_{CEN} and LREG\mathcal{L}_{REG}, respectively. Following GLIP , we adopt the ATSS detector as our image encoder. We use the sigmoid focal loss for LALIL_{ALI}, the sigmoid loss for LCENL_{CEN}, and the GIoU loss for LREGL_{REG}.

Experimental Results

Implementation Details. We pre-train all the models based on Swin-Transformer backbones with 32 GPUs. AdamW optimizer is adopted and batch size is set to 128. The learning rate is set to 2.8x10-4 for the parameters of the visual backbone and detection head, and 2.8x10-5 for the language backbone. Without otherwise specified, all models are trained with 12 epochs and the learning rate is decayed with a factor of 0.1 at the 8-th and the 11-th epoch. The max token length for each input sentence is set to 48. The number of the concepts NN in text input PP is set to 150 and the number of region features MM is determined by the feature map size and the number of pre-defined anchors. The loss weight factors α\alpha and β\beta are both set to 1.0. The training of DetCLIP-T with Swin-T backbone tasks 63 hours when 32 GPUs are used. MMDetection code-base is used.

Training Data. Our model is trained with a hybrid supervision from different kinds of data, i.e., detection data, grounding data, and image-text pair data. More specifically, for detection data, we use a sampled Objects365 V2 dataset (denoted as O365 in the following sections) with 0.66M training images. Here we do not use the whole dataset since training with sampled data is more efficient and suffices to demonstrate the effectiveness of our method. For grounding data, we use gold grounding data (denoted as GoldG) introduced by MDETR . Moreover, following GLIP , we remove the training samples contained in LVIS dataset for fair zero-transfer evaluation, which results in 0.77M training data. For image-text pair data, we perform object-level dense pseudo labeling on YFCC100m dataset with a pre-trained CLIP model, and sample a subset of 1M training images from the results containing objects with similarities scores above a given threshold. Finally, our training set contains a total of 2.43M images. Compared to GLIP’s 27M training data, DetCLIP only uses less than 10% data, but achieves better results. More details are in Appendix.

Benchmark Settings. We evaluate our method mainly on LVIS which contains 1203 categories. Following GLIP and MDETR , we evaluate on the 5k minival subset and report the zero-shot fixed AP for a fair comparison. We do not focus the performance on COCO since it only contains 80 common categories that are fully covered by the training dataset Objects365 , which may not sufficient to reflect generalization ability of a model in the open-domain detection setting. To further study the generalization ability of our method, following GLIP , we also evaluate the averaged AP on other 13 downstream detection datasets published on Roboflowhttps://public.roboflow.com/object-detection.

We train our DetCLIP with two backbones, i.e., swin-T and swin-L. Distinct from GLIP , which introduces additional heavy modules like Dynamic Head and cross-modal fusion, we directly adopt the vanilla ATSS as our vision encoder to keep our architecture as neat as possible. Following GLIP, for Swin-T backbone, we also train 3 versions of models, i.e., DetCLIP-T(A), DetCLIP-T(B) and a complete version DetCLIP-T, which differ in using different training data.

Table 1 reports the results on LVIS . Results with LVIS as pre-training data stand for the fully-supervised models trained with annotated data. With our proposed paralleled formulation, introducing more training data from different sources can consistently improve the performance. I.e., comparing DetCLIP-T(A) trained with only detection data with DetCLIP-T trained with additional grounding and image-text pair data, we can observe a considerable performance gain (28.8% AP v.s. 35.9% AP). Besides, befitting from the proposed effective framework, our DetCLIP models outperform their GLIP counterparts by a large margin, i.e., DetCLIP-T(A) (resp. DetCLIP-T) surpasses GLIP-T(A) (resp. GLIP-T) by 10.3% (resp. 9.9%). Besides, DetCLIP-T also significantly outperforms GLIPv2-T by 6.9%. Note that our models are more lightweight (without heavy DyHead and cross-modal fusion) and trained with fewer epochs (our 12 vs. GLIP’s 24) and much less data. In addition, our DetCLIP-T’s zero-shot performance even beats the fully-supervised model with the same backbone by utilizing weak-annotated data like image-text pair data. Results on LVIS full validation dataset can refer to the Appendix.

Efficiency Comparison. To demonstrate the efficiency of our proposed DetCLIP, we directly compare the training and inference speed of DetCLIP-T with the GLIP-T in Table 2.

With the same setting of training with 32 V100 GPUs, the total training time for GLIP-T is about 10.7K GPU hours (5X than us) due to its heavy backbone and more image-text pair training data. On the other hand, DetCLIP-T achieves the 2.3 FPS (0.43 s/image) on a single V100 when performing inference on LVIS , while GLIP-T can only achieve 0.12 FPS (8.6 s/image). With much better training and inference efficiency, DetCLIP-T can still outperform GLIP-T 9.9% on LVIS.

Qualitative Visualizations. We illustrate the bounding box predictions on LVIS dataset from DetCLIP-T and GLIP-T model in Fig.6. We can observe that by adopting the paralleled concept formulation and the external knowledge from the concept dictionary, our DetCLIP can outperform GLIP both on the accuracy and completeness of the predicted labels, especially on the rare classes.

2 Ablation Studies

Table 3 studies the effectiveness of two core components of DetCLIP, i.e., the paralleled formulation and the knowledge enrichment with concept dictionary. The first row stands for our implementation of a GLIP-A model, which is modeled with sequential formulation and uses only detection data for training. Due to the implementation discrepancy, our version can achieve 23.7% zero-shot AP on LVIS , which is higher than the official’s 18.5% and serves as a stronger baseline.

First, applying the paralleled formulation can bring significant improvements (row 2), boosting the performance to 27.8%. This indicates that paralleled formulation is much more effective than sequential formulation for modeling object detection as a vision-language task. However, directly applying the same approach to a larger scale dataset with heterogeneous supervisions, e.g., detection plus grounding, can hurt the performance (row 4). We speculate this is because paralleled formulation weakens the interaction between text concepts, resulting in the model’s inability to effectively construct the connections between semantic-related concepts. Therefore, we introduce word definitions to the class names to help bridge relationships between different concepts, which boosts the performance to 32.2% (row 5). Sampling negative categories from the concept dictionary also helps better utilize grounding data, improving the performance to 34.4% (row 7). Further introducing image-text pair datasets like YFCC can bring substantial improvement for rare categories (row 8), while utilizing concept dictionary for label completion during pseudo labeling finally improves the overall AP to 35.9% (row 9). A similar performance pattern is also observed on 13 downstream detection datasets.

Impact of Concept Dictionary’s Size. To study the impact of the scale of the proposed concept dictionary, we build three concept dictionaries with different sizes by using: (1) class names from Objects365 ; (2) class names from Objects365 + Things ; (3) class names in (2) plus noun phrases extracted from YFCC100m . We equip our DetCLIP-T(B) with these three concept dictionaries and compare their performance in Table 4. It can be seen that using a small size dictionary, e.g., Objects365 + Things, can even bring a performance drop compared to without the dictionary, while scaling up the dictionary with nouns from YFCC can significantly improve the performance. We speculate this is because a large dictionary can provide rich negative concepts for grounding data, encouraging the model to learn more discriminative features.

Pseudo labels with concept dictionary. Fig.7 visualizes the results of label completion via concept dictionary. Multiple effectiveness can be observed. E.g., cases (a) and (b) demonstrate concept dictionary contributes to labeling bbox with the category names not shown in the captions. In (c), a finer-grained pseudo label can be produced, i.e., ‘antheraea polyphemus’ v.s. the original ‘brown butterfly’, which helps the learning of rare categories. In (d), label noise in the caption is alleviated.

Importance of concept enrichment. To directly illustrate how the category definition helps the model achieve better detection performance, we conduct the experiments to infer with different text inputs with DetCLIP-T utilized. We compare the three different cases, i.e., no definition, true definition and false definition (use definition of another category), and report the alignment score for different objects in Fig.8. Observations are: (1) True definition helps the model to better determine the category (i.e., increases the confidence score); (2) Wrong definition may confuse the model and bring the false positive samples; (3) Both the category name and definition matter to the model.

Conclusion

In this paper, we propose a novel open-world detection pre-training framework named DetCLIP, aiming at improving the open-domain ability and learning efficiency of the open-world detector. By unifying three kinds of supervision via paralleled concept formulation, our DetCLIP can learn from different domains and enhance the training efficiency. We further propose a concept dictionary module to improve the discovery and coverage of the novel knowledge by importing external knowledge. The proposed usages of the concept dictionary achieves better open-world detection result in terms of both common and rare categories on LVIS. Experiments on multiple downstream detection datasets suggest that our DetCLIP is more powerful than current SOTA open-world detectors such as GLIP .

Acknowledgements We gratefully acknowledge the support of MindSporehttps://www.mindspore.cn/, CANN (Compute Architecture for Neural Networks) and Ascend AI Processor used for this research.

Appendix for DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

Appendix A Negative Impacts and Limitations

Potential Negative Social Impact. Our method has no ethical risk on dataset usage and privacy violation since all the benchmarks are publicly available and transparent.

Limitations and Future Works. The localization ability of the region proposals is still limited by the annotation of the bounding box. More weakly-supervision can be included to learn from image-text pairs. Furthermore, although we prove the effectiveness of our method on the web-collected dataset YFCC , we expect to extend our method to larger image-text pair datasets from the Internet.

Appendix B Dataset Details

In this section, we provide more details of training datasets used in our experiments, which include (1) the approach of generating pseudo detection labels for image-text pair dataset; and (2) a dataset comparison with GLIP .

Pseudo Labeling on Image-Text Pair Data. We use image-text pair data from the web-collected dataset YFCC100m . To generate pseudo detection labels for image-text pair data, we first use a Region Proposal Network (RPN) pre-trained on Objects365 to extract object proposals. To ensure the quality of proposals, we filter result bounding boxes with objectness scores below a threshold of 0.3 or region area smaller than 6000. This operation also helps significantly reduce the number of proposal candidates and accelerates the pseudo-labeling process.Then a powerful pre-trained CLIP model (ViT-L) is used to predict pseudo class labels for each retained bounding box. To alleviate the partial-label problem, we use concept names from our proposed concept dictionary (Sec.3.2) instead of the raw caption as the text input. Following CLIP, the prompt "a photo of a category." is used to pad a category name into the sentence. Since the proposed dictionary consists of a large number of concepts (i.e., 14k), to accelerate the inference, we pre-compute the text embeddings of all concepts and store them for the later computation. For each proposal bounding box, we first crop it from the raw image, then resize it to 224×224224\times 224 and feed it into the visual encoder to obtain the visual embedding. We use the cosine similarity as the classification score, which is computed as

where fθIf_{\theta}^{I} and gϕTg_{\phi}^{T} stand for the image and text encoder of the CLIP model; RiR_{i} and cjc_{j} are ii-th cropped proposal and jj-th category, respectively. Both embeddings are L2-normalized before the similarity calculation. After category prediction, a second-stage filtering is adopted to drop proposals with a classification score below 0.240.24. Finally, we sample 1M images from the results to form our final training image-text pair data.

Training Data Comparison (with GLIP ). Table 5 compares the training data used by DetCLIP and GLIP. The explanation of each dataset can be found in the table caption. Our DetCLIP-T uses less than half training data compared to GLIP-T, while DetCLIP-L uses less than 10% training data compared to GLIP-L.

Appendix C More Results on LVIS and 13 Detection Datasets

More results under GLIP-protocal. We study the generalization ability of our models by zero-shot transferring them to LVIS full validation set. Following GLIP , we use manually designed prompts for some downstream datasets, as illustrated in Table 6. AP for LVIS full validation set is reported in Table 7. Despite using much less training data, DetCLIP models can dominate their GLIP’s counterparts in most cases (except APfAP_{f} on LVIS for DetCLIP-L). Notably, compared to GLIP, DetCLIP considerably boosts the performance for rare categories, which is an important indicator reflecting models’ generalization ability for the open-world detection task. More AP performances for 13 downstream detection datasets can refer to Table 8.

More results under VILD-protocal. To make a more comprehensive evaluation of our method, we also perform experiments under the VILD protocol, i.e., the method is trained on base categories and then evaluated on novel categories using the original LVIS AP metric. We replace the Objects365 part in our training data with LVIS-base, and GoldG and YFCC1M are still included. Including additional data will lead to somehow unfair comparison with VILD but it is necessary since this is the core component in our method to enable zero-shot capability, which differs from VILD that distills knowledge from a pre-trained CLIP model. Note that we implement DetCLIP using the same training/testing setting as in the paper, and do not use techniques such as large-scale jittering and prompt ensemble which is adopted by VILD to boost the performance. The results are shown in the Table 9. Our method (27.3 mAP) outperforms VILD (22.5 mAP) by 4.8% mAP.

Appendix D Ablation Studies

Sequential Formulation with Shuffled Grounding Data (Sec 3.1). The sequential formulation (e.g., GLIP ) is not effective for modeling open-world object detection as a visual-language task since it leads to unnecessary interaction between category names in the attention module. To demonstrate the idea, we randomly shuffle the word order in the grounding training data and report the performance comparison in Table 10. It can be seen that randomly shuffling the word order in the grounding data can even bring a slight improvement (i.e., +1.4% on LVIS minival) on zero-shot transfer AP for the downstream detection task, indicating that the noun phrase is more critical for the detection task, compared to the context information. Therefore, DetCLIP drops the context information and treats each noun phrase as a paralleled text input, which avoids unnecessary attention among class names and achieves better training efficiency.

Important Role of Class Definition. In DetCLIP, we augment the class names in the detection dataset with their definitions during both training and inference stage, which is termed as concept enrichment. To verify that DetCLIP learns knowledge from class definitions, we compare the performances of including/excluding definitions in text input during the inference stage. Table 11 reports the results. It can be found that adding definitions to class names can significantly improve the zero-shot transfer performance.

Impact of pre-trained language models in Concept enrichment. During training, we use a pre-trained language model to retrieve a definition in our dictionary for concepts without a direct match in WordNet. We conduct experiments to study how the pre-trained language model in the this process affects the final performance. Three different settings are considered: 1. do not use language model, i.e., directly adopt the category name as the input for the concepts not in WordNet; 2. use a pre-trained FILIP text encoder; and 3. use a pre-trained RoBERTa as in GLIP. The results are shown in the Table 12. We can observe that: 1) the concept enrichment procedure can bring significant improvements, (e.g., +3.6% on rare categories) even without using a pre-trained language model; 2) using FILIP can further boost the AP performance from 28.3 to 28.8, while using RoBERTa achieves similar performance with no language model is used.

Other Important Training Techniques. Training a vision-language model that works for the open-world detection task is not easy. We highlight two important training techniques we found in our experiments: (1) using a small learning rate for the pre-trained language backbone, since it helps maintain the language model’s knowledge learnt in the large-scale pretraining; and (2) removing the regression loss for non-detection data, since it helps alleviate the negative impact caused by inaccurate localization annotation of grounding/image-text pair data. Table 13 provides the ablation studies of these techniques.

Appendix E Qualitative Results

More visualizations of pseudo labels with concept dictionary. Fig. 9 shows extra examples of YFCC data that pseudo labeled with the concept dictionary, as well as their comparisons with the results generated by using the original caption. Concept dictionary alleviates partial-label problem and helps CLIP model provide finer-grained and higher quality pseudo labels.

Retrieval with Concept Dictionary. In our concept enrichment, to augment a given class name with its definition, we retrieve it in the constructed concept dictionary. If there is an exact match, we directly use the corresponding definition; otherwise, we use semantic similarity computed by a pre-trained language model to find the closest one. Table 14 illustrates some example retrieval results. For class names that are not contained in the dictionary, our method can find proper synonyms.

Illustrations of Concept Dictionary. We illustrate some examples in our concept dictionary in Table 15. We observe that the concepts collected from the image-text pair data can cover more fine-grained categories (e.g., cotswold, cuniculus paca) and a wider range of classes (e.g., giant, cathedral).

References