Grounded Language-Image Pre-training

Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, Jianfeng Gao

Introduction

Visual recognition models are typically trained to predict a fixed set of pre-determined object categories, which limits their usability in real-world applications since additional labeled data are needed to generalize to new visual concepts and domains. CLIP shows that image-level visual representations can be learned effectively on large amounts of raw image-text pairs. Because the paired texts contain a boarder set of visual concepts than any pre-defined concept pool, the pre-trained CLIP model is so semantically rich that it can be easily transferred to downstream image classification and text-image retrieval tasks in zero-shot settings. However, to gain fine-grained understanding of images, as required by many tasks, such as object detection , segmentation , human pose estimation , scene understanding , action recognition , vision-language understanding , object-level visual representations are highly desired.

In this paper, we show that phrase grounding, which is a task of identifying the fine-grained correspondence between phrases in a sentence and objects (or regions) in an image, is an effective and scalable pre-training task to learn an object-level, language-aware, and semantic-rich visual representation, and propose Grounded Language-Image Pre-training (GLIP). Our approach unifies the phrase grounding and object detection tasks in that object detection can be cast as context-free phrase grounding while phrase grounding can be viewed as a contextualized object detection task. We highlight our key contributions as follows.

Unifying detection and grounding by reformulating object detection as phrase grounding. The reformulation changes the input of a detection model: it takes as input not only an image but also a text prompt that describes all the candidate categories in the detection taskDifferent from typical phrase grounding tasks, phrases in the text prompt for an object detection task may not be present in the image.. For example, the text prompt for COCO object detection is a text string that consists of 80 phrases, i.e., the 80 COCO object class names, joined by “. ”, as shown in Figure 3 (Left). Any object detection model can be converted to a grounding model by replacing the object classification logits in its box classifier with the word-region alignment scores, i.e., dot product of the region (or box) visual features and the token (or phrase) language features, as shown in Figure 3 (Right). The language features are computed using a language model, which gives the new detection (or grounding) model a dual-encoder structure. Different from CLIP that fuses vision and language only at the last dot product layer , we show that deep cross-modality fusion applied by GLIP, as shown in Figure 3 (Middle), is crucial to learn high-quality language-aware visual representations and to achieve superior transfer learning performance. The unification of detection and grounding also allows us to pre-train using both types of data and benefits both tasks. On the detection side, the pool of visual concepts is significantly enriched thanks to the grounding data. On the grounding side, detection data introduce more bounding box annotations and help train a new SoTA phrase grounding model.

Scaling up visual concepts with massive image-text data. Given a good grounding model (teacher), we can augment GLIP pre-training data by automatically generating grounding boxes for massive image-text-paired data, in which noun phrases are detected by an NLP parser . Thus, we can pre-train our (student) GLIP-Large model (GLIP-L) on 27M grounding data, including 3M human-annotated fine-grained data and 24M web-crawled image-text pairs. For the 24M image-text pairs, there are 78.1M high-confidence (>0.5>0.5) phrase-box pseudo annotations, with 58.4M unique noun phrases. We showcase two real examples of the generated boxes in Figure 3. The teacher model can accurately localize some arguably hard concepts, such as syringes, vaccine, beautiful caribbean sea turquoise, and even abstract words (the view). Training on such semantic-rich data delivers a semantic-rich student model. In contrast, prior work on scaling detection data simply cannot predict concepts out of the teacher models’ pre-defined vocabulary . In this study, we show that this simple strategy of scaling up grounding data is empirically effective, bringing large improvements to LVIS and 13 downstream detection tasks, especially on rare categories (Sections 4.2 and 5). When the pre-trained GLIP-L model is fine-tuned on COCO, it achieves 60.8 AP on COCO 2017val and 61.5 on test-dev, surpassing the current public SoTA models that scale up object detection data in various approaches.

Transfer learning with GLIP: one model for all. The grounding reformulation and semantic-rich pre-training facilitate domain transfer. GLIP can be transferred to various tasks with few or even no additional human annotations. When the GLIP-L model is directly evaluated on the COCO and LVIS datasets (without seeing any images in COCO during pre-training), it achieves 49.8 and 26.9 AP on COCO val2017 and LVIS val, respectively, surpassing many supervised baselines. When evaluated on 13 existing object detection datasets, spanning scenarios including fine-grained species detection, drone-view detection, and ego-centric detection, the setting which we term “Object Detection in the Wild” (ODinW) (Section 5.1), GLIP exhibits excellent data efficiency. For example, a zero-shot GLIP-L outperforms a 10-shot supervised baseline (Dynamic Head) pre-trained on Objects365 while a 1-shot GLIP-L rivals with a fully supervised Dynamic Head. Moreover, when task-specific annotations are available, instead of tuning the whole model, one could tune only the task-specific prompt embedding, while keeping the model parameters unchanged. Under such a prompt tuning setting (Section 5.2), one GLIP model can simultaneously perform well on all downstream tasks , reducing the fine-tuning and deployment cost.

Related Work

Standard object detection systems are trained to localize a fixed set of object classes predefined in crowd-labeled datasets, such as COCO , OpenImages (OI) , Objects365 , and Visual Genome (VG) , which contains no more than 2,000 object classes. Such human-annotated data are costly to scale up . GLIP presents an affordable solution by reformulating object detection as a phrase grounding (word-to-region matching) problem, and thus enables the use of grounding and massive image-text-paired data. Though our current implementation is built upon Dynamic Head (DyHead) , our unified formulation can be generalized to any object detection systems .

Recently, there is a trend to develop vision-and-language approaches to visual recognition problems, where vision models are trained with free-form language supervision. For example, CLIP and ALIGN perform cross-modal contrastive learning on hundreds or thousands of millions of image-text pairs and can directly perform open-vocabulary image classification. By distilling the knowledge from the CLIP/ALIGN model into a two-stage detector, ViLD is proposed to advance zero-shot object detection. Alternatively, MDETR trains an end-to-end model on existing multi-modal datasets which have explicit alignment between phrases in text and objects in image. Our GLIP inherits the semantic-rich and language-aware property of this line of research, achieves SoTA object detection performance and significantly improves the transferability to downstream detection tasks.

This paper focuses on domain transfer for object detection. The goal is to build one pre-trained model that seamlessly transfers to various tasks and domains, in a zero-shot or few-shot manner. Our setting differs from zero-shot detection , where some categories are defined as unseen/rare and not present in the training set. We expect GLIP to perform well on rare categories (Section 4.2) but we do not explicitly exclude any categories from our training set, because grounding data are so semantically rich that we expect them to cover many rare categories. This resembles the setting in open-vocabulary object detection , which expects raw image-text data to cover many rare categories. A line of work identifies building a open-world object proposal module that could propose any novel object at test time as the key challenge ; GLIP offers a new perspective: the model does not need to propose every possible novel objects from an open set; rather it only needs to propose objects mentioned in the text prompt as the detection branch is conditioned on the prompt.

Beyond performance on rare categories, we also consider the transfer cost in real-world scenarios, i.e., how to achieve the best performance with the least amount of data, training budget, and deployment cost (Section 5). In particular, we show that GLIP supports prompt tuning , which matches the performance of full fine-tuning but only tunes a fraction of the model parameters. We also present a novel finding that in object detection, prompt tuning is most effective for a model with deep vision-language fusion such as GLIP, while being far less effective for shallow-fused models. This stands in contrast to recent work that investigates prompt tuning only for shallow-fused vision-language models such as CLIP .

Grounded Language Image Pre-training

Conceptually, object detection and phrase grounding bear a great similarity. They both seek to localize objects and align them to semantic concepts. This synergy motivates us to cast the classical object detection task into a grounding problem and propose a unified formulation (Sec 3.1). We further propose to add deep fusion between image and text, making the detection model language-aware and thus a strong grounding model (Sec 3.2). With the reformulation and deep fusion, we can pre-train GLIP on scalable and semantic-rich grounding data (Sec 3.3).

Background: object detection. A typical detection model feeds an input image into a visual encoder EncI\text{Enc}_{I}, with CNN or Transformer as backbone, and extracts region/box features OO, as shown in Figure 3 (Bottom). Each region/box feature is fed into two prediction heads, i.e., a box classifier C\mathcal{C} and a box regressor R\mathcal{R}, which are trained with the classification loss Lcls\mathcal{L}_{\text{cls}} and the localization loss Lloc\mathcal{L}_{\text{loc}}, respectively:

In two-stage detectors, a separate region proposal network (RPN) with RPN loss Lrpn\mathcal{L}_{\text{rpn}} is used to distinguish foreground from background and refine anchors. Since Lrpn\mathcal{L}_{\text{rpn}} does not use semantic information of object classes, we merge it into the localization loss Lloc\mathcal{L}_{\text{loc}}. In one-stage detectors, localization loss Lloc\mathcal{L}_{\text{loc}} may also contain the centerness loss .

The box classifier C\mathcal{C} is typically a simple linear layer, and the classification loss Lcls\mathcal{L}_{\text{cls}} can be written as:

Object detection as phrase grounding. Instead of classifying each region/box into cc classes, we reformulate detection as a grounding task, by grounding/aligning each region to cc phrases in a text prompt (see Figure 3). How to design a text prompt for a detection task? Given object classes [person, bicycle, car, ..., toothbrush][\text{person, bicycle, car, ..., toothbrush}], one simple way is

in which each class name is a candidate phrase to be grounded. One could design better prompts, by providing more expressive descriptions of these classes and/or by exploiting the preference of a pre-trained language model. For example, when the pre-trained BERT model is used to initialize our language encoder EncL\text{Enc}_{L}, the prompt “person. bicycle. car. … . toothbrush” works better than the more human-friendly prompt described above. We will discuss the prompt design in Section 5.2.

In a grounding model, we compute the alignment scores SgroundS_{\text{ground}} between image regions and words in the prompt:

Equivalence between detection and grounding. With the above reformulation, we can convert any detection model into a grounding model, and the two views, i.e., detection and grounding, are theoretically equivalent for both training and inference. We also verify this empirically: the SoTA DyHead detector with Swin-Tiny backbone gives the same performance on COCO val2017 before and after our reformulation. Please refer to the appendix for discussions. With the reformulation, a pre-trained phrase grounding model can be directly applied to any object detection task, thanks to the free-form input of the language encoder. This makes it possible to transfer our GLIP model to arbitrary detection tasks in a zero-shot manner.

2 Language-Aware Deep Fusion

In (3), the image and text are encoded by separate encoders and only fused at the end to calculate the alignment scores. We call such models late-fusion models. In vision-language literature , deep fusion of visual and language features is necessary to learn a performant phrase grounding model. We introduce deep fusion between the image and language encoders, which fuses the image and text information in the last few encoding layers, as shown in Figure 3 (Middle). Concretely, when we use DyHead as the image encoder and BERT as the text encoder, the deep-fused encoder is:

where LL is the number of DyHeadModules in DyHead , BERTLayer is newly-added BERT Layers on top of the pre-trained BERT, O0O^{0} denote the visual features from the vision backbone, and P0P^{0} denote the token features from the language backbone (BERT). The cross-modality communication is achieved by the cross-modality multi-head attention module (X-MHA) (4), followed by the single modality fusion and updated in (5) & (6). Without added context vectors (Ot2iiO^{i}_{\text{t2i}} for vision modality and Pi2tiP^{i}_{\text{i2t}} for language modality), the model is reduced to a late-fusion model.

In the cross-modality multi-head attention module (X-MHA) (4), each head computes the context vectors of one modality by attending to the other modality:

where {W(symbol,I),W(symbol,L):symbol∈{q,v,out}}\{W^{(\text{symbol},I)},W^{(\text{symbol},L)}:\text{symbol}\in\{q,v,out\}\} are trainable parameters and play similar roles to those of query, value, and output linear layers in Multi-Head Self-Attention , respectively.

The deep-fused encoder (4)-(6) brings two benefits. 1) It improves the phrase grounding performance. 2) It makes the learned visual features language-aware, and thus the model’s prediction is conditioned on the text prompt. This is crucial to achieve the goal of having one model serve all downstream detection tasks (shown in Section 5.2).

3 Pre-training with Scalable Semantic-Rich Data

Considerable efforts have been devoted to collecting detection data that are rich in semantics and large in quantity. However, human annotations have been proven costy and limited . Prior work seeks to scale up in a self-training fashion . They use a teacher (a pre-trained detector) to predict boxes from raw images and generate pseudo detection labels to train a student model. But the generated data are still limited in terms of the size of the concept pool, as the teacher can only predict labels defined in the concept pool, constructed on the existing datasets. In contrast, our model can be trained on both detection and, more importantly, grounding data. We show that grounding data can provide rich semantics to facilitate localization and can be scaled up in a self-training fashion.

First, the gold grounding data cover a much larger vocabulary of visual concepts than existing detection data. The largest attempts at scaling up detection vocabulary still cover no more than 2,000 categories . With grounding data, we expand the vocabulary to cover virtually any concepts that appear in the grounded captions. For example, Flickr30K contains 44,518 unique phrases while VG Caption contains 110,689 unique phrases, orders of magnitude larger than the vocabulary of detection data. We provide an empirical study in Section 4.4 to show that 0.8M gold grounding data brings a larger improvement on detecting rare categories than additional 2M detection data.

Further, instead of scaling up detection data, we show a promising route to obtaining semantically rich data: scaling up grounding data. We use a simple approach inspired by self-training. We first pre-train a teacher GLIP with gold (human-annotated) detection and grounding data. Then we use this teacher model to predict boxes for web-collected image-text data, with noun phrases detected by an NLP parser . Finally, a student model is trained with both the gold data and the generated pseudo grounding data. As shown in Figure 3, the teacher is capable of generating accurate boxes for semantically rich entities.

Why can the student model possibly outperform the teacher model? While discussions remain active in the self-training literature , in the context of visual grounding, we posit that the teacher model is utilizing the language context and language generalization ability to accurately ground concepts that it may not inherently know. For example, in Figure 3, the teacher may not directly recognize certain concepts such as vaccine and turquoise, if they are not present in gold data. However, the rich language context such as syntactic structures can provide strong guidance for the teacher model to perform an “educated guess”. The model can localize vaccine if it can localize a small vail; it can localize turquoise if it can find caribbean sea. When we train the student model, the “educated guess” of the teacher model becomes a “supervised signal”, enabling the student model to learn the concept of vaccine and turquoise.

Transfer to Established Benchmarks

After pre-training, GLIP can be applied to grounding and detection tasks with ease. We show strong direct domain transfer performance on three established benchmarks: 1) MS-COCO object detection (COCO) containing 80 common object categories; 2) LVIS covering over 1000 objects categories; 3) Flickr30K , for phrase grounding. We train 5 variants of GLIP (Table 1) to ablate its three core techniques: 1) unified grounding loss; 2) language-aware deep fusion; 3) and pre-training with both types of data. Implementation deails are in the appendix.

GLIP-T (A) is based on a SoTA detection model, Dynamic Head , with our word-region alignment loss replacing the classification loss. It is based on the Swin-Tiny backbone and pre-trained on O365 (Objects365 ), which contains 0.66M images and 365 categories. As discussed in Section 3.1, the model can be viewed as a strong classical zero-shot detection model , relying purely on the language encoder to generalize to new concepts.

GLIP-T (B) is enhanced with language-aware deep fusion but pre-trained only on O365.

GLIP-T (C) is pre-trained on 1) O365 and 2) GoldG, 0.8M human-annotated gold grounding data curated by MDETR , including Flickr30K, VG Caption , and GQA . We have removed COCO images from the dataset. It is designed to verify the effectiveness of gold grounding data

GLIP-T is based on the Swin-Tiny backbone and pre-trained on the following data: 1) O365, 2) GoldG as in GLIP-T (C), and 3) Cap4M, 4M image-text pairs collected from the web with boxes generated by GLIP-T (C). We also experiment with existing image caption datasets: CC (Conceptual Captions with 3M data) and SBU (with 1M data) . We find that CC+SBU GLIP-T performs slightly better than Cap4M GLIP-T on COCO, but slightly worse on the other datasets. For simplicity, we report both versions on COCO but only the Cap4M model for the other tasks. We present the full results in the appendix.

GLIP-L is based on Swin-Large and trained with: 1) FourODs (2.66M data), 4 detection datasets including Objects365, OpenImages , Visual Genome (excluding COCO images) , and ImageNetBoxes ; 2) GoldG as in GLIP-T (C); and 3) CC12M+SBU, 24M image-text data collected from the web with generated boxes.

We conduct experiments on MS-COCO to evaluate models’ transfer ability to common categories. We evaluate under two settings: 1) zero-shot domain transfer, and 2) supervised transfer, where we fine-tune the pre-trained models using the standard setting. For the fine-tuning setting, we additionally test the performance of a GLIP-L model, where we include the COCO images in the pre-training data (the last row). Specifically, we add the full GoldG+ grounding data and COCO train2017 to the pre-training data. Note that part of COCO 2017val images are present in GoldG+ . Thus we only report the test-dev performance of this model. Please see more details in the appendix.

We introduce an additional baseline: DyHead pre-trained on Objects365. We find that COCO 80 categories are fully covered in Objects365. Thus we can evaluate DyHead trained on Objects365 in a “zero-shot” way: during inference, instead of predicting from 365 classes, we restrict the model to predict only from the COCO 80 classes. We list standard COCO detection models for reference. We also list two state-of-the-art models pre-trained with extra data.

Results are present in Table 2. Overall, GLIP models achieve strong zero-shot and supervised performance. Zero-shot GLIP models rival or surpass well-established supervised models. The best GLIP-T achieves 46.7 AP, surpassing Faster RCNN; GLIP-L achieves 49.8 AP, surpassing DyHead-T. Under the supervised setting, the best GLIP-T brings 5.5 AP improvement upon the standard DyHead (55.2 v.s. 49.7). With the Swin-Large backbone, GLIP-L surpasses the current SoTA on COCO, reaching 60.8 on 2017val and 61.5 on test-dev, without some bells and whistles in prior SoTA such as model EMA, mix-up, label smoothing, or soft-NMS.

We analyze the zero-shot performance of GLIP and find three contributing factors: close domain overlap between Objects365 and COCO, deep fusion, and grounding data. As Objects365 covers all categories in COCO, the O365 pre-trained DyHead-T shows strong performance, reaching 43.6 zero-shot AP; reformulating the model into a grounding model, we observe a slight performance drop (GLIP-T (A)); adding deep fusion boosts the performance by 2 AP (GLIP-T (B)); the largest contributor is the gold grounding data, with which GLIP-T (C) reaches a zero-shot AP of 46.7. While the addition of image-text data brings slight or no improvement on COCO (GLIP-T v.s. GLIP-T (C)), we find it essential in generalizing to rare classes, as we show in the LVIS experiments.

2 Zero-Shot Transfer on LVIS

We evaluate the model’s ability to recognize diverse and rare objects on LVIS in a zero-shot setting. We report on MiniVal containing 5,000 images introduced in MDETR as well as the full validation set v1.0. Please see the evaluation details in the appendix.

Results are present in Table 3. We list three supervised models trained on the annotated data of LVIS. GLIP exhibits strong zero-shot performance on all the categories. GLIP-T is on par with supervised MDETR while GLIP-L outperforms Supervised-RFS by a large margin.

The benefit of using grounding data is evident. Gold grounding data brings a 4.2-point improvement on MiniVal APr (model C v.s. model B). Adding image-text data further improves performance by 3.1 points. We conclude that the semantic richness of grounding data significantly helps the model recognize rare objects.

3 Phrase Grounding on Flickr30K Entities

We evaluate the model’s ability to ground entities in natural language on Flickr30K entities . Flickr30K is included in the gold grounding data so we directly evaluate the models after pre-training as in MDETR . We use the any-box-protocol specified in MDETR. Results are present in Table 4. We evaluate three versions of GLIP with different pre-training data. We list the performance of MDETR, the SoTA grounding model. MDETR is trained on GoldG+, containing 1.3M data (GoldG is a subset of GoldG+ excluding COCO images).

GLIP-T with GoldG (Row 3) achieves similar performance to MDETR with GoldG+, presumably due to the introduction of Swin Transformer, DyHead module, and deep fusion. More interestingly, the addition of detection data helps grounding (Row 4 v.s. 3), showing again the synergy between the two tasks and the effectiveness of our unified loss. Image-text data also helps (Row 5 v.s. 4). Lastly, scaling up (GLIP-L) can achieve 87.1 Recall@1, outperforming the previous SoTA by 2.8 points.

4 Analysis

In this section, we perform ablation study by pre-training GLIP-T on different data sources (Table 5). We answer two research questions. First, our approach assumes the use of a detection dataset to bootstraps the model. One natural question is whether grounding data brings improvement when paired with different detection data. We find that adding grounding data brings consistent improvement with different detection data (Row 1-6).

Second, we have shown the effectiveness of grounding data for both common and rare categories. One orthogonal direction is to scale up detection data by including more images and categories (Section 3.3). We intend to provide an empirical comparison between scaling up detection data and grounding data. We present GLIP trained with 4 public detection datasets (Row 8) as an extreme attempt at scaling up detection data with human annotations. The model is trained with 2.66M detection data in total, with an aligned vocabulary of over 1,500 categories. However, it still trails behind Row 6 on COCO and APr of LVIS, where Row 6 is trained with only 0.66M detection data and 0.8M gold grounding data. Adding image-text data further widens the gap on LVIS APr (20.8 versus 15.0). We conclude that grounding data are indeed more semantic-rich and a promising alternative to scaling up detection data.

Object Detection in the Wild

To evaluate GLIP’s transferability to diverse real-world tasks, we curate an “Object Detection in the Wild” (ODinW) setting. We choose 13 public datasets on Roboflowhttps://public.roboflow.com/object-detection, each requiring a different localization skill. Many of the datasets are designed with a specific application purpose to mimic real-world deployment scenarios. For example, EgoHands requires locating hands of a person; Pothole concerns detecting holes on the road; ThermalDogsandPeople involves identifying dogs and persons in infrared images. Please refer to the appendix for details.

We demonstrate that GLIP facilitates transfer to such diverse tasks. (1) GLIP brings great data efficiency, reaching the same performance with significantly less task-specific data than baselines (Section 5.1). (2) GLIP enables new domain transfer strategies: when adapting to a new task, we can simply change the text prompt and keep the entire grounding model unchanged. This greatly reduces deployment cost because it allows one centralized model to serve various downstream tasks (Section 5.2).

We vary the amount of task-specific annotated data, from zero-shot (no data provided), to XX-shot (providing at least XX examples per category ), to using all data in the training set. We fine-tune the models on the provided data and use the same hyper-parameters for all models. Each dataset comes with pre-specified category names. As GLIP is language-aware, we find it beneficial to re-write some pre-specified names with more descriptive language (see Section 5.2 for a discussion). We compare with the SoTA detector DyHead-T, pre-trained on Objects365. We test with the standard COCO-trained DyHead-T and find it giving similar performance. For simplicity, we report only the former. We also experiment with the scaled cosine similarity approach but find it slightly underperforming the vanilla approach so we report only the latter. Please refer to the appendix for full statistics, including three independent runs for XX-shot experiments.

Results are shown in Figure 4. We find that unified grounding reformulation, deep fusion, grounding data, and model scale-up all contribute to the improved data efficiency (from the bottom red line (Dyhead-T) up to the upper purple line (GLIP-L)). As a result, GLIP exhibits transformative data efficiency. A zero-shot GLIP-T outperforms 5-shot DyHead-T while a one-shot GLIP-L is competitive with a fully supervised DyHead-T.

We further plot the zero-shot performance of GLIP variants on 5 different datasets in Figure 5. We find that the introduction of grounding data brings significant improvement on certain tasks that test novel concepts, e.g., on Pothole and EgoHands, models without grounding data (A&B) performs terribly, while models with grounding data (C) outperform them with ease.

2 One Model for All Tasks

As neural models become larger, how to reduce deployment cost has drawn an growing research interest. Recent work on language models , image classification , and object detection has explored adapting a pre-trained model to a new domain but only changing the least amount of parameters. Such a setting is often denoted as linear probing , prompt tuning , or efficient task adapters . The goal is to have a single model serving various tasks, and each task adds only a few task-specific parameters or no parameters to the pre-trained model. This reduces training and storage cost. In this section, we evaluate models against the metric of deployment efficiency.

Manual prompt tuning. As GLIP performs language-aware localization, i.e., the output of GLIP is heavily conditioned on the language input, we propose an efficient way for GLIP to do task transfer: for any novel categories, the user can use expressive descriptions in the text prompt, adding attributes or language context, to inject domain knowledge and help GLIP transfer. For example, on the left hand side of Figure 6, the model fails to localize all occurrences of the novel entity “stingray”. However, by adding the attributes to the prompt, i.e., “flat and round”, the model successfully localizes all occurrences of stringrays. With this simple prompt change, we improve the AP50 on stingray from 4.6 to 9.7. This resembles the prompt design technique in GPT-3 and is practically appealing, as it requires no annotated data or model re-training. Please refer to the appendix for more details.

Prompt tuning. We further consider the setting where we have access to task-specific training data but wish to tune the least amount of parameters for easy deployment. For classical detection models, Wang et al. report the effectiveness of “linear probing” (i.e., train only the box regression and classification head). GLIP can also be “linear probed”, where we only fine-tune the box head and a projection layer between the region and prompt embeddings. Because of the language-aware deep fusion, GLIP supports a more powerful yet still efficient transfer strategy: prompt tuning . For GLIP, as each detection task has only one language prompt (e.g., the prompt for Pothole could be “Detect pothole.” for all images), we first get prompt embeddings P0P^{0} from the language backbone, then discard the language backbone and only fine-tune P0P^{0} as the task-specific input (Section 3.2).

We evaluate the models’ performance under three settings (Figure 7): linear probing, prompt tuning (only applicable for GLIP), and full-model tuning. For DyHead-T, prompt tuning is not applicable as the traditional object detection model cannot accept language input; the gap between linear probing and full-model tuning is large. GLIP-T (A) has no language-aware deep fusion; thus prompt tuning and linear tuning achieve similar performance and lag significantly behind full-model tuning. However, for GLIP-T and GLIP-L, prompt tuning almost matches the full-tuning results, without changing any of the grounding model parameters. Interestingly, as the model and data size grow larger, the gap between full-model tuning and prompt tuning becomes smaller (GLIP-L v.s. GLIP-T), echoing the findings in NLP literature .

Conclusion

GLIP unifies the object detection and phrase grounding tasks to learn an object-level, language-aware, and semantic-rich visual representation. After pre-training, GLIP showed promising results on zero-shot and fine-tuning settings on well-established benchmarks and 13 downstream tasks. We leave a detailed study of how GLIP scales with text-image data size to future work.

Acknowledgement

We thank anonymous reviewers for their comments and suggestions. We thank Xiyang Dai, Zicheng Liu, Yi-Ling Chen for help with the project. LL and KC are supported in part by DARPA MCS program under Cooperative Agreement N66001-19-2-4032.

References

Appendix

In Section A, we provide more visualizations of our model’s grounding predictions on the Conceptual Caption 12M dataset .

In Section B (referred by Section 3.1), we discuss the equivalence between detection and grounding.

In Section C.1 (referred by Section 4), we introduce the pre-training details of the models we use in Section 4.

In Section C.2 (referred by Section 4), we introduce the evaluation details of experiments on COCO, LVIS, and Flickr30K.

In Section C.3 (referred by Section 4), we discuss the difference between the public image-text data (Google Conceptual Captions,SBU) and the image-text data we collected.

In Section D, we provide a detailed analysis on the computational cost and performance effect of the language-aware deep fusion.

In Section E.1 (referred by Section 5), we introduce the 13 datasets in Object Detection in the Wild (ODinW).

In Section E.2 (referred by Section 5), we detail the manual prompt design.

In Section E.3 (referred by Section 5.1), we give the details for the data efficiency experiments.

In Section E.4 (referred by Section 5.3), we give the details for the linear probing and prompt tuning experiments.

In Section E.5, we present per-dataset results for all experiments in Section 5.

Appendix A Visualization

We provide more visualizations of the predictions from our teacher model. Even given noise image-text pairs, our model is still capable of grounding semantic-rich phrases accurately.

Appendix B Equivalence Discussion between Detection and Grounding

In Section 3.1 of the main paper, we discussed the equivalence between detection and grounding. We corroborate the discussion with empirical experiments.

We first confirm that when all categories fit into one prompt, our grounding formulation is equivalent to classical object detection. We conduct the experiments on COCO . We first choose the SoTA detection model Dynamic Head (DyHead) based on the Swin-Tiny Transformer backbone as the base object detection mode. We then transform this model into a grounding model as described in Section 3.1: we concatenate the 80 class names with “. ” into one prompt and replace DyHead’s classification loss with our grounding loss. We use BERT (base-uncased) to encode the text prompt. When concatenating the class names, we follow a fixed order.

We train the two models with the exact same hyperparameters as in : we train with the standard 2x training configurations . We train with batch size 32 and learning rate 1×10−41\times 10^{-4} (for the model with grounding reformulation, we use 1×10−51\times 10^{-5} for the BERT text encoder). We decay the learning rate at 67% and 89% of the total training steps.

The two models achieve the same performance on COCO 2017val: 49.4 AP. Their results are close to the 49.7 reported in the last row of Table 6 of Dai et al. (the small difference is presumably due to the implementation difference). Thus, we conclude that when all categories can fit into a single prompt, grounding and detection tasks are equivalent.

When not all object categories can fit into a single prompt.

The text encoder for the prompt has a limit on the input sentence length. For example, BERT can only encode sentences containing at most 512 tokens. In our implementation, to reduce computational costs, we limit the input length to 256. Thus, for certain datasets with a large vocabulary (e.g., Objects365 has 365 object categories), we cannot fit all category names into one prompt. As a practical solution, we can split the category names into multiple prompts, during both training time and inference time. We find that this incurs minor performance drop. For example, in Table 2 in the main paper, DyHead-T pre-trained on Objects365 achieves 43.6 on COCO zero-shot, while GLIP-T (A) (the grounding reformulated model of DyHead) achieves 42.9 on COCO.

Appendix C Transfer to Established Benchmarks

We introduce the implementation details of the models used in Section 4 and discuss the difference between public image-text data and the data crawled by us.

In Section 4, we introduced GLIP-T (A), GLIP-T (B), GLIP-T (C), GLIP-T, and GLIP-L. We introduce the implementation details in the following. We pre-train models based on Swin-Tiny models with 32 GPUs and a batch size of 64, and models based on Swin-Large with 64 GPUs and a batch size of 64. We use a base learning rate of 1×10−51\times 10^{-5} for the language backbone and 1×10−41\times 10^{-4} for all other parameters. The learning rate is stepped down by a factor of 0.1 at the 67% and 89% of the total training steps. We decay the learning rate when the zero-shot performance on COCO saturates. The max input length is 256 tokens for all models.

As noted in Section B, when we pre-train on datasets such as Objects365, we cannot fit all categories into one prompt. During pre-training, we randomly down-sample the categories and keep only the down-sampled categories in the prompt. We randomly shuffle the categories’ order in the prompt.

The down-sampling is done randomly on the fly for each training example and serves as data augmentation. Specifically, for an example, we denote the positive classes that appear in the image as CposC_{\text{pos}} and the rest negative classes as CnegC_{\text{neg}}. We always keep all of CposC_{\text{pos}}. With a probability of 0.50.5, we sample from CnegC_{\text{neg}} till we have 85 categories in the prompt; with a probability of 0.50.5, we uniformly choose an interger NN from [1,85−∣Cpos∣][1,85-|C_{\text{pos}}|] and put NN categories in the prompt.

Augmentation for image-text data with generated boxes.

When we pre-train the model on image-text data with generated boxes, we find it beneficial to increase the difficulty. We mix a few negative captions (that are from other examples and do not match with the image) with the positive caption (that is matched to the image) to form a longer text input. The model is trained to predict boxes and align them to the correct phrases in the positive caption. The model would need to first identify the positive caption among a few potential captions and then align the box to the correct phrases in the positive caption. This makes the grounding task more challenging and help the model learn a semantic-rich representation during pre-training. This augmentation is also done randomly on the fly. For each training example, with a probability of 0.3, we conduct such augmentation and mix in 19 negative captions; with a probability of 0.3, we mix in a random number (uniformly drawn between 1-19) of negative captions; for the rest of the time, we do not conduct such augmentation.

C.2 Evaluation Details

For fine-tuning on COCO, we use a base learning rate of 1×10−51\times 10^{-5} for pre-trained models.

For zero-shot evaluation on LVIS, since LVIS has over 1,000 categories and they cannot be fit into one text prompt, we segment them into multiple chunks, fitting 40 categories into one prompt and query the model multiple times with the different prompts. We find that models tend to overfit on LVIS during the course of pre-training so we monitor the performance on minival for all models and report the results with the best checkpoints.

For zero-shot evaluation on Flickr30K, models may also overfit during the course of pre-training so we monitor the performance on the validation set for all models and report the results with the best checkpoints.

C.3 Difference Between Public Data and Web-Crawled Data

For GLIP-T pre-trained with image-text data, as mentioned in Section 4, we train two versions, one with public data (CC3M,SBU) and another with data we crawled (Cap4M). Here we provide a comparison between the two models in Table 6.

The two models differ only slightly, with the Cap4M version better on LVIS while the CC3M+SBU version better on COCO. We conjecture that this is potentially because the public data is more extensively screened and contains more common categories and less rare concepts. Thus it performs slightly better on COCO while lags slightly on LVIS.

Appendix D Computation Cost and Performance Analysis of Deep Fusion

In this section, we provide a more detailed ablation on the computational cost and performance effect of the language-aware deep fusion proposed in Section 3.

We test the additional computational cost of the language-aware deep fusion for both GLIP-T and GLIP-L. For inference, we test on a P100 GPU with batch size 1. Note that for inference with GLIP without deep fusion, we could cache the language embeddings of the prompts; thus the inference time of GLIP without deep fusion is equivalent to that of DyHead .

For training, we test on a standard DGX-2 machine with 16 V100 GPUs (we test under the multi-GPU setting as it mimics the actual training environment): for GLIP-T models, we use 2 images per batch and for GLIP-L models, we use 1 images per batch. As the fusion module invovles multi-head attention over a large number of input elements, we turn on gradient checkpointinghttps://pytorch.org/docs/stable/checkpoint.html for the deep fusion module, which increases training time but reduces GPU memory consumption.

Table 7 shows that the language-aware deep fusion brings less than 1x additional computational cost overall.

D.2 Performance

We provide an analysis on the effect of language-aware deep fusion when different kinds of pre-training data are used. We pre-train four variants of GLIP-T and show the results In Table 8. Deep fusion is beneficial for testing on 1) common categories (i.e., COCO); 2) grounding tasks (i.e., Flickr30K), and 3) low-resource transfer to real-world downstream tasks (i.e., ODinW).

However, on LVIS, the effect of deep fusion seems unclear: when only detection data are used, deep fusion seems to degrades performance (row 1 v.s. row 2); when grounding data are present, deep fusion degrades common category performance but improves rare category performance. Our assumption is that when GLIP is only trained with detection data (e.g., O365), the language model could “overfit” to the categories in O365 and does not generalize to novel categories well (i.e., outputs out-of-distribution text representation). The deep fusion could “amplify” such overfit as the visual representation is conditioned on the language model. Thus, when tested on prompts containing novel categories (e.g., LVIS), deep fusion could degrade performance. When grounding data are used, such overfit could be mitigated.

Appendix E Object Detection in the Wild

In this section, we provide the details and additional results for the experiments in Section 5.

We use 13 datasets from Roboflowhttps://public.roboflow.com/object-detection. Roboflow hosts over 30 datasets and we exclude datasets that are too challenging (e.g., detecting different kinds of chess pieces) or impossible to solve without specific domain knowledge (e.g., understanding sign language).

We provide the details of the 13 datasets we use in Table 9. We include the PASCAL V0C 2012 dataset as a reference dataset, as public baselines have been established on this dataset. For PascalVOC, we follow the convention and report on validation set. For Pistols, there are no official validation or test sets so we split the dataset ourselves.

E.2 Manual Prompt Tuning

As discussed in Section 5, we find it beneficial to manually design some prompts to provide language guidance. We provide the prompts we use in Table 10. We design the prompts for 6 datasets. Since some prompts are sentences, we only apply these prompts for models trained with grounding data (GLIP-T (C), GLIP-T, and GLIP-L). For GLIP-T (A) and GLIP-T (B), we find it beneficial to use prompts for the Rabbits and Mushrooms datasets, as the prompts there are just single word or short phrases. Overall, using prompts improves AP without any model re-training (e.g., the AP improves from 22.1 to 50.0 for EgoHands).

E.3 Data Efficiency

We provide details for the experiments in Section 5.1. We train with batch size 4, learning rate 1×10−41\times 10^{-4} (for the model with grounding reformulation, we use 1×10−51\times 10^{-5} for the BERT text encoder), and weight decay of 0.05. We do not find that increasing batch size improves performance significantly. For computational reasons, we use a batch size of 4. Following convention, we freeze the bottom 2 layers of the backbone during fine-tuning. We monitor the performance on validation and decay the learning rate by 0.1 when the validation performance plateaus. In XX-shot settings, we randomly sample the dataset such that there are at least XX examples per category . We change the random seeds (and thus change the sampled data) and conduct 3 independent runs for each XX-shot experiment. We provide two DyHead-T variants as baselines, one trained on COCO and one trained on Objects365. We report the full zero-shot results in Table 14 and few-shot results in Table 11.

E.4 One Model for All Tasks

In Section 5.2, we conduct experiments with respect to deployment efficiency: tuning the least amount of parameters for the best performance. For all models, we experiment with the linear probing setting; for GLIP models, we also experiment with the prompt tuning setting. For linear probing, we try both the vanilla approach (simply tune the classification and localization head) and the cosine scale approach . Below we provide the implementation details.

For the vanilla linear probing, we train with a learning rate of 1×10−41\times 10^{-4}, batch size of 4, and weight decay of 0.05. For linear probing with the cosine scale, we use a scale of 20.020.0 per suggestions of Wang et al. , learning rate of 0.010.01, batch size of 4, and weight decay of 0.05. For prompt tuning, we train with a learning rate of 0.050.05, batch size of 4, and weight decay of 0.25. We have conducted preliminary searches for the hyper-parameters.

Results are present in Table 12 (linear probing) and Table 13 (prompt tuning). Comparing them with full-tuning results (Table 11), we see prompt tuning performance of GLIP is competitive, showing the deployment efficiency. Contrary to Wang et al. who report that linear probing can deliver competitive performance for classical detection models, we find that linear probing does not work well compared to full tuning. We find that the reason could be the transfer datasets (ODinW) in our case contain a lot of novel tasks and domains, while experiments in Wang et al. focus on transferring to common domains (e.g., PascalVOC and COCO). In Table 15, we report the per-dataset performance. We find that for some common tasks or domains (e.g., PascalVOC and Vehicles), linear probing of DyHead COCO performs competitively with full fine-tuning but the gap is large for some other tasks of a novel domain (e.g., AerialDrone).

E.5 All Results

We report the per-dataset performance under 0,1,3,5,10-shot and full data as well as linear probing, prompt tuning, and full-model tuning in Table 14, Table 15, and Table 16 (on the next pages).