RegionCLIP: Region-based Language-Image Pretraining

Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, Jianfeng Gao

Introduction

Recent advances in vision-language representation learning has created remarkable models like CLIP and ALIGN . Such models are trained using hundreds of millions of image-text pairs by matching images to their captions, achieving impressive results of recognizing a large set of concepts without manual labels, and capable of transferring to many visual recognition tasks. Following their success on image classification, a natural question is that whether these models can be used to reason about image regions, e.g., for tasks like object detection.

To answer this question, we construct a simple R-CNN style object detector using a pretrained CLIP model, similar to adapting a pretrained convolutional network. This detector crops candidate object regions from an input image, and applies the CLIP model for detection by matching visual features of cropped regions to text embeddings of object categories. Fig. 1(a-b) shows the results on LVIS dataset . When using object proposals as the input regions, scores from CLIP often fail to capture the localization quality (Fig. 1a). Even with ground-truth object boxes, classification accuracy using CLIP drops significantly from 60% on ImageNet to 19% on LVIS, with a similar number of classes (Fig. 1b). There is thus a major performance degradation when applying a pretrained CLIP model for object detection. How can we empower a vision-language pretrained model to reason about image regions?

We believe the main gap lies in the training of these vision-language models. Many existing vision-language models, including CLIP, are trained to match an image with its image-level text description. The training is unaware of the alignment between local image regions and text tokens. Thus, the models are unable to precisely ground a textual concept to an image region. Further, cropping image regions and matching them to text tokens largely ignore the surrounding visual context that is critical for object recognition, not to mention the high computational cost, e.g. a few seconds per image on a modern GPU.

In this paper, we explore learning region representations for object detection via vision-language pretraining. Our key idea is to explicitly align image regions and text tokens during pretraining. However, two key challenges arise. First, the fine-grained alignment between image regions and text tokens is not available in image-text pairs. Second, the text description of its paired image is often incomplete, i.e. many image regions are not described by the text. To address these challenges, we propose to bootstrap from a pretrained vision-language model to align image regions and text tokens, and to fill in the missing region descriptions, as illustrated in Fig. 1c.

Specifically, our method starts with a pool of object concepts parsed from text corpus, and synthesizes region descriptions by filling these concepts into pre-defined templates. Given an input image and its candidate regions from either object proposals or dense sliding windows, a pretrained CLIP model is used to align the region descriptions and the image regions, creating “pseudo” labels for region-text alignment. Further, we use both “pseudo” region-text pairs and ground-truth image-text pairs to pretrain our vision-language model via contrastive learning and knowledge distillation. Although the “pseudo” region-text pairs are noisy, they still provide useful information for learning region representations and thus bridge the gap to object detection, as validated by our experiments.

We pretrain our models on captioning datasets (e.g., Conceptual Caption) and mainly evaluate models on the benchmarks of open-vocabulary object detection (COCO and LVIS datasets). When transferred to open-vocabulary object detection, our pretrained model establishes new state-of-the-art (SoTA) results on COCO and LVIS. For instance, our model achieves a relative gain of 37.7% over published SoTA in AP50 for novel categories on COCO. Moreover, our model supports zero-shot inference and outperforms baselines by a clear margin.

Our contributions are summarized as follows: (1) We propose a novel method that aligns image regions and their descriptions without manual annotation, thereby enabling vision-language pretraining for learning visual region representations. (2) A key technical innovation that facilitates our pretraining is a scalable approach for generating region descriptions, neither relying on human annotations nor limited to the text paired with an image. (3) Our pretrained model presents strong results when transferred to open-vocabulary object detection, and demonstrates promising capability on zero-shot inference for object detection.

Related Work

Visual representation learning for images. Early works on visual representation learning focused on learning from intensive human labels by training image classifiers . These classifiers can be further used to label un-annotated images for training student models in semi-supervised learning . To reduce the annotation burden, self-supervised learning was proposed to match the visual representation of different views from the same image. The most relevant work is learning from natural language, such as image tags and text descriptions . Coupled with millions of image-text pairs collected from the Internet, recent vision-language pretraining learned to match images with image descriptions and demonstrated impressive performance on zero-shot inference and transfer learning for image classification. However, these works focus on image representation and target at image classification. In this paper, we propose to learn visual representation for image regions which supports zero-shot inference and transfer learning for region reasoning tasks (e.g., object detection).

Visual representation learning for image regions. By leveraging human annotations contributed by , major progress has been made to reason about image regions, such as object detection . With the object detectors trained on these human annotations as teacher models, semi-supervised learning creates pseudo labels for image regions in return for training student detectors. Beyond object labels, the region representation learned from additional labels of object attributes demonstrated noticeable improvement on vision-language tasks . However, these works heavily rely on expensive human annotation and are limited to predefined categories. To reduce annotation cost, the idea of self-supervised learning is extended to region representation learning by maximizing the representation similarity among augmented views of image regions. Different from these works, we propose to learn region representation via vision-language pretraining, inspired by CLIP . The learned region representation supports recognizing image regions with a large vocabulary.

Zero-shot and open-vocabulary object detection. Zero-shot object detection aims at detecting novel object classes which are not seen during detector training . Bansal et al. learned to match the visual features of cropped image regions to word embeddings using max-margin loss. Rahman et al. proposed polarity loss to model background category and to cluster categories with similar semantics. Zhu et al. explored improving localization performance for novel categories by synthesizing visual features with a generative model. These zero-shot object detectors usually rely on the semantic space of pretrained word embeddings . Recently, Zareian et al. proposed OVR for open-vocabulary object detection, where a visual encoder was first pretrained on image-text pairs to learn broad object concepts and then transferred to zero-shot object detection setting. Another close work is ViLD that focused on the training of zero-shot object detectors by distilling visual features from a pretrained CLIP model . Similar to OVR and ViLD, our detector also leverages the visual-semantic space learned from vision-language pretraining. Different from OVR, we propose to learn visual region representation from our “pseudo” region-text pairs given by another pretrained CLIP model. Our method is thus not restricted to particular text that pairs with an image. Unlike ViLD, our method focuses on pretraining and the resulting regional representations support both zero-shot inference and transfer learning.

Method

Our goal is to learn a regional visual-semantic space which covers rich object concepts so that it can be used for open-vocabulary object detection. Consider a text description tt that describes the content of region rr in an image II. In the visual-semantic space, the visual region representation V(I,r)\mathcal{V}(I,r) extracted from rr should be matched to text representation L(t)\mathcal{L}(t). V\mathcal{V} is a visual encoder that takes image II and a region location rr, and outputs a visual representation for this region. L\mathcal{L} is a language encoder that converts a text in natural language to a semantic representation.

Disentanglement of recognition and localization. There are two key components for image region understanding: localization and recognition. Inspired by , we disentangle these two components, use existing region localizers, and focus on region recognition by learning regional visual-semantic space without heavy human annotation.

Method overview. As shown in Fig. 2, we denote Vt\mathcal{V}_{t} and L\mathcal{L} as visual and language encoders pretrained to match images to their descriptions, such as CLIP. Our goal is to train a visual encoder V\mathcal{V} so that it can encode image regions and match them to region descriptions encoded by language encoder L\mathcal{L}. To address the challenge of lacking large-scale region descriptions, as shown at the bottom of Fig. 2, we construct a pool of object concepts, create the region descriptions by filling concepts into prompts, and leverage teacher encoder Vt\mathcal{V}_{t} to align these text descriptions with the image regions proposed by an image region localizer. Given the created region-text pairs, our visual encoder V\mathcal{V} learns to match these pairs via contrastive learning and concept distillation. Once pretrained, our model supports zero-shot inference for region recognition and can be transferred to object detector when the human annotation is available.

2 Region-based Language-Image Pretraining

We introduce how we obtain region-level visual and semantic representation, and then describe how we build the alignment between image regions and region descriptions.

Visual region representation. Image regions can be proposed by either off-the-shelf object localizers (e.g., RPN ) or dense sliding windows (e.g., random regions). By default, we use an RPN which is pretrained on human-annotated object bounding boxes without object labels. We use RPN to propose image regions for all images in a batch and finally obtain NN image regions in total. The set of image regions denotes as {ri}i=1,...,N\{r_{i}\}_{i=1,...,N}. Given the proposed regions, the visual representation viv_{i} of region rir_{i} is extracted from our visual encoder V\mathcal{V} with a feature pooling method, such as RoIAlign . RoIAlign pools regional visual features from the feature map of full image by using interpolation. Specially, we note that our visual encoder V\mathcal{V} is initialized by the teacher Vt\mathcal{V}_{t} so that it can have a good starting point in visual-semantic space.

Semantic region representation. A single image usually contains rich semantics, covering one or more objects out of thousands of categories. It is costly to annotate all these categories in the large-scale image-text datasets. To this end, we first build a large pool of concepts to exhaustively cover regional concepts, regardless of individual full images. As shown at the bottom of Fig. 2, we create a pool of object concepts which are parsed from text corpus (e.g., the image descriptions collected from the Internet), by using off-the-shelf language parsers . Given the concept pool, the semantic representations for regions are created by two steps: (1) We create a short sentence for each concept by filling it to prompt templates (e.g., prompts of CLIP ). For example, the “kite” concept will be converted to “A photo of a kite”. (2) We encode the created text descriptions into semantic representations by using the pretrained language encoder L\mathcal{L}. Finally, all regional concepts are represented by their semantic embeddings {lj}j=1,...,C\{l_{j}\}_{j=1,...,C} and CC denotes the size of concept pool.

While our region descriptions are built based on the image descriptions, our method is not constrained by the particular text description that pairs with an image. More importantly, in light of the powerful language encoder L\mathcal{L} which has seen many words in natural language, we can easily customize our concept pool and scale it up, which is difficult to achieve from human annotations. Similarly, in vision modality, the disentanglement of visual recognition and localization makes our method flexible to adopt different ways of extracting candidate regions.

2.2 Visual-Semantic Alignment for Regions

Alignment of region-text pairs. We leverage a teacher visual encoder Vt\mathcal{V}_{t} to create the correspondence between image regions and our created texts (represented as semantic embeddings). Again, visual representation vitv^{t}_{i} of region rir_{i} is extracted from teacher encoder Vt\mathcal{V}_{t} by pooling features from the loca image region with RoIAlign. Then we compute the matching score between vitv^{t}_{i} and each concept embedding ljl_{j}. The matching score S(v,l)S(v,l) is given by

The object concept lml_{m} that has highest matching score is selected and linked to region rir_{i}. Finally, we obtain the pseudo labels for each region, namely the pairs of {vi,lm}\{v_{i},l_{m}\}.

Our pretraining scheme. Our pretraining leverages both created region-text pairs and the image-text pairs from the Internet. Given the aligned region-text pairs (represented by {vi,lm}\{v_{i},l_{m}\}), we pretrain our visual encoder with contrastive loss and distillation loss based on the image regions across different images. The contrastive loss is computed as

τ\tau is a predefined temperature. Nri\mathcal{N}_{r_{i}} represents a set of negative textual samples for region rir_{i}, i.e., the object concepts that are not matched to region rir_{i} but matched to other regions in the batch. Beyond contrastive learning over positive and negative region-text pairs, we also consider knowledge distillation for each image region over all object concepts. The distillation loss is defined as

where LKLL_{KL} is KL divergence loss; both qitq^{t}_{i} and qiq_{i} are probabilities over all object concepts. qitq^{t}_{i} is a soft target from teacher model computed as softmax(S(vit,l1)/τ,...,S(vit,lC)/τ)softmax(S(v^{t}_{i},l_{1})/\tau,...,S(v^{t}_{i},l_{C})/\tau). qiq_{i} is computed as softmax(S(vi,l1)/τ,...,S(vi,lC)/τ)softmax(S(v_{i},l_{1})/\tau,...,S(v_{i},l_{C})/\tau) coming from our student model.

Given image-text pairs collected from the Internet, our region-level contrastive loss LcntrstL_{cntrst} can naturally extend to image-level contrastive loss Lcntrst−imgL_{cntrst-img}. It can be considered as a special case where (1) the visual representation is extracted for single global box that covers the whole image, (2) text descriptions are collected from the Internet, and (3) negative samples are the text descriptions that come with other images. Finally, our overall loss function is given by

Zero-shot inference. Once pretrained, our visual encoder can be directly applied to region reasoning tasks. For example, given image region proposals from RPN, region representation extracted from our visual encoder are matched to the embeddings of target object concepts, thereby predicting the most likely category. Inspired by , we fuse RPN objectness scores and category confidence scores by geometry mean. Empirically, we observe that RPN scores significant improve zero-shot inference.

3 Transfer Learning for Object Detection

In pretraining, our visual encoder learns from region-text alignment which is created by teacher model. Such alignment does not require human efforts but it is inevitably noisy and weak. When strong supervision for image regions is available (e.g., the human-annotated detection labels), our visual encoder can be further fine-tuned by simply replacing the region descriptions, as shown in Panel 3 of Fig. 2.

Specifically, we transfer our pretrained visual encoder to object detectors by initializing their visual backbones. To detect image objects, same as our pretraining, we use an off-the-shelf RPN to localize object regions and recognize these regions by matching their visual region representation and the semantic embeddings of target object classes (e.g., the object classes in detection dataset).

Training for open-vocabulary object detection . In this setting, the detectors are trained by the annotation of base categories while expected to detect novel categories never seen in detector training. Specially, we apply class-wise weighted cross-entropy loss to train our detectors. (1) For base categories, inspired by focal loss , we apply focal scaling and calculate the weight for a base category as (1−pb)γ(1-p^{b})^{\gamma}, where pbp^{b} is probability after softmax for this base category and γ\gamma is a hyperparameter. Empirically, focal scaling is effective to alleviate the forgetting of previously learned object concepts in pretraining, especially when there are very few base categories in dataset (e.g., COCO). We conjecture that the detector might overfit to the small set of base categories, thereby hurting the generalization on novel categories. (2) For background category, we use a fixed all-zero embedding and apply a predefined weight to background regions following .

Experiments

Our models are primarily evaluated on transfer learning for open-vocabulary object detection. We also present results of zero-shot inference for object detection. Finally, we present ablation study on different model components.

Datasets. For pretraining, we use the image-text pairs from Conceptual Caption dataset (CC3M) which collects 3 millions of image-text pairs from the web. We also consider a smaller dataset COCO Caption (COCO Cap) to pretrain our model when conducting ablation study. COCO Cap contains 118k images with each image annotated by human for 5 captions. We parsed object concepts from COCO/CC3M dataset and filtered the concepts whose frequency is lower than 100, resulting in 4764/6790 concepts.

For transfer learning of open-vocabulary object detection, we train detectors with base categories of COCO detection dataset and LVIS dataset (v1) , respectively. On COCO, We follow the data split of with 48 base categories and 17 novel categories which are subsets of COCO object classes. We use the processed data from with 107,761 training images and 4,836 test images. On LVIS, following , we use the training/validation images for training/evaluation and adopt the category split with 866 base categories (common and frequent objects) and 337 novel categories (rare objects).

We evaluate object detection performance on COCO and LVIS for both transfer learning and zero-shot inference.

Evaluation protocol and metrics. We adopt the standard object detection metrics: mean Average Precision (AP) and AP50 (AP at an intersection over union of 0.5). We evaluate our models on two benchmarks for open-vocabulary object detection, including COCO and LVIS. On COCO, we report AP50 and follow the evaluation settings in : (1) only predicting and evaluating novel categories (Novel), (2) only predicting and evaluating base categories (Base), (3) a generalized setting that predicts and evaluates all categories (Generalized). On LVIS, we follow the benchmark of where the rare objects are defined as novel categories. We report AP for novel categories (APr), base categories (APc, APf) and all categories (mAP), respectively.

Implementation details. During pretraining, the default student model and teacher model were both ResNet50 of pretrained CLIP. The RPN used in pretraining was trained with the base categories of LVIS dataset. Our default model was pretrained on CC3M dataset with the concepts parsed from COCO Cap. SGD was used with the image batch of 96, initial learning rate of 0.002, maximum iteration of 600k, and 100 regions per image. For transfer learning of object detection, our detectors were developed on Detectron2 using Faster RCNN with ResNet50-C4 architecture. The RPN used in transfer learning was trained by the base categories of target dataset (e.g., the transfer learning on COCO used the RPN trained on COCO). SGD was used with image batch of 16, initial learning rate 0.002, and 1x schedule. The weight of background category was set to 0.2/0.8 on COCO/LVIS. Focal scaling was particularly applied to COCO training with γ\gamma as 0.5. For zero-shot inference of object detection, RPN was the same as pretraining stage and NMS threshold was set to 0.9. For all experiments, the temperature τ\tau was 0.01.

We present the results of transfer learning for open-vocabulary object detection on COCO and LVIS datasets. Additionally, we report results for fully supervised setting where all categories are used during training.

Setup. The detectors are trained by base categories while evaluated on base and novel categories (e.g., 48/866 base categories and 17/337 novel categories on COCO/LVIS). To compare with ViLD , all experiments on LVIS additionally use mask annotation to train detector.

Baselines. We consider several baselines as follows:

Zero-shot object detectors (SB , DELO , PL ): Zero-shot object detection is the closest area to open-vocabulary object detection. These detectors usually rely on the pretrained word embeddings of object classes for generalization to novel categories.

Open-vocabulary object detectors (OVR , ViLD ): These detectors leverage pretrained vision-language models that have learned a large vocabulary from image-text pairs. OVR is our close competitor in the sense that we both pretrain visual encoders and use them as the detector initialization. ViLD is a recent unpublished work that focuses on detector training by distilling visual features of a pretrained model from CLIP. ViLD specially uses the data augmentation of copy-paste with 16x training schedule.

Fully supervised detectors: On COCO, we include the supervised baseline from OVR which is a Faster RCNN trained by the base categories with 1x schedule. On LVIS, we include the supervised baseline from ViLD which is a Mask RCNN trained by base and novel categories with special data augmentation as ViLD. We additionally report a Mask RCNN trained in standard 1x schedule from Detectron2 .

Our detector variants: We consider initializing our detector with different pretrained visual encoders, including CLIP and our model pretrained on COCO Cap.

Results. Table 1 and Table 2 show the results on COCO and LVIS datasets, respectively.

On COCO dataset, initialized by our pretrained backbone, our detector significantly outperforms previous published SoTA method OVR on all metrics (e.g., 31.4 vs. 22.8 on novel categories). Compared with the CLIP backbone from which we start our region-based pretraining, our model brings a remarkable gain across all metrics, particularly +17.2 AP50 on novel categories. When compared with ViLD, an unpublished SoTA method with sophisticated training strategy, our model is still comparable on Base and All, while substantially better on Novel (e.g., 31.4 vs. 27.6) which is the main focus in open-vocabulary detection. On LVIS dataset, with comparable backbone size (RN50x4-C4 of ours: 83.4M, RN152-FPN of ViLD: 84.1M), our detector outperforms ViLD by a large margin (e.g., +2.2 APr and +3.6 mAP). Note that these superior detection results on COCO and LVIS are achieved by using a single pretrained backbone, with standard data augmentation and 1x training schedule. These results suggest that our region-based vision-language pretraining has learned better alignment between image regions and object concepts, and thus facilitates open-vocabulary object detection.

1.2 Fully Supervised Object Detection

Setup. Detection annotation of all object categories are used during training and evaluation. Again, all experiments on LVIS additionally use mask annotation to train detector.

Baselines. We consider the following baselines: (1) Faster RCNN intialized by ImageNet pretrained backbone: This is a common object detector in the community. (2) Our detector initialized by pretrained CLIP. This baseline is to validate our proposed pretraining method.

Results. In Table 3, the detector initialized by our pretrained visual backbone largely outperforms the baselines that are initialized by ImageNet and CLIP backbones (e.g., +2.4 mAP on COCO and +2.8 mAP on LVIS). These results suggest that our proposed pretraining method helps the fully supervised detector converge faster and achieves better performance at 1x schedule. Again, when using RN50x4 as the backbone for both teacher model and student model, the performance is significantly improved (eg, 38.8 vs. 42.7 mAP on COCO, 29.0 vs. 32.5 on LVIS).

2 Zero-shot Inference for Object Detection

Setup. Without finetuning on the detection annotation, the pretrained vision-language models are directly used to recognize the proposed regions. We use the same evaluation datasets and metrics as the experiments in transfer learning. We consider two types of region proposals: (1) The ground-truth bounding boxes are used as region proposals. This setting aims at evaluating the recognition performance by eliminating the localization error. (2) The region proposals come from a RPN which is also used in pretraining. The performance in this setting is dependent on both the quality of RPN and recognition ability.

Baselines. We consider two baselines: (1) OVR pretrains visual backbone on image-text pairs of COCO Cap which has close object concepts as COCO detection dataset. We evaluate the pretrained model provided in their code base. (2) CLIP is pretrained on 400M image-text pairs. Both OVR and CLIP pretrain model on the image-text pairs while our pretraining leverages the created region-text pairs for learning visual region representation.

Results. Table 4 summarizes the results. With ideal region proposals, our pretrained model outperforms CLIP baseline by a clear margin across datasets (e.g., 61.4 vs. 58.3 All AP50 on COCO, 44.4 vs. 42.2 mAP on LVIS). When compared with OVR, our model demonstrates a much larger margin (e.g., 61.4 vs. 44.5 All AP50 on COCO), not to mention that OVR is pretrained on the same dataset as evaluation. Even if using RPN proposals, our model still clearly outperforms CLIP and OVR (e.g., 26.8 vs. 19.6 & 25.5 on COCO, 9.6 vs. 9.2 on LVIS). These promising results suggest that our pretraining method with region-text alignment improves the visual recognition ability for image regions. With RN50x4 architecture as the backbones of teacher and student models, the zero-shot inference performance is further improved across datasets and different types of region proposals (e.g., +6.3 mAP on LVIS with GT boxes, +2.8 All on COCO with RPN boxes).

3 Ablation Study

The evaluation in this section uses COCO dataset and the same metrics as zero-shot inference and transfer learning.

Pretraining supervision. Table 5 studies the effect of different pretraining supervisions. Accordingly, though using the region-text pairs already attains plausible results, the additional supervision from image-text pairs can further improve the performance (e.g., +2.4 AP50 with GT boxes on zero-shot inference, +5.4 Novel AP50 on transfer learning). We suspect that image-text pairs provide extra contextual information from global image description which compensates our created region descriptions.

Types of image regions. Table 6 studies the effects of region proposal quality during pretraining. We replace the RPN proposals by sampling the same number of image regions with random location and random aspect ratio. Random boxes hurt zero-shot inference (-2.0 AP50 with GT boxes) while reserve comparable performance in transfer learning (46.9 vs. 47.5 All AP50). These results indicate that our pretraining is robust to the quality of region proposals. Zero-shot inference benefits from higher quality of proposals but the gap becomes smaller when human supervision is available to finetune the model.

Pretraining dataset and concept pool. In Table 7, using COCO Cap dataset or using the COCO concepts achieves better zero-shot inference performance (62.8 vs. 61.4 vs. 60.8 AP50 with GT boxes). We hypothesize that COCO Cap has a smaller domain gap to COCO detection dataset. However, the model pretrained on CC3M achieves significant boost on transfer learning (47.5 vs. 50.4 All AP50). We conjecture that the model learns more generic visual representation from a larger number of images in CC3M.

Pretraining losses. Table 8 studies the effects of different losses. With both contrastive loss and distillation loss, the model achieves close results as distillation-only model on zero-shot inference (e.g., 62.8 vs. 63.1 AP50 with GT boxes) while the best performance on transfer learning (e.g., 26.8 Novel AP50). These results suggest that two losses play different roles. Distillation loss helps to inherit the visual-semantic knowledge from the teacher model, while contrastive loss enforces more discriminative representations for transfer learning.

Teacher model and student model. Table 9 studies the effects of using different teacher and student models. Compared with the default setting at first row, using ResNet50x4 as the teacher model can largely improve the zero-shot inference performance (+4.2 AP50 with GT boxes). However, in the transfer learning setting, the performance using a stronger teacher remains roughly the same (both are 50.4 AP50 for All). When we further replace the student model with ResNet50x4, the transfer learning performance is significantly boosted (+5.3 AP50 for All), but the zero-shot inference performance remains (29.6 vs. 29.3 AP50 with RPN boxes). Based on these results, we conjecture that zero-shot inference performance relies on the teacher model that guides the region-text alignment, while transfer learning is more likely constrained by the capacity of student model.

Focal scaling. Table 10 studies the effects of focal scaling during transfer learning. With focal scaling, the finetuned detector achieves a better balance between novel categories and base categories on COCO dataset. We conjecture that the detector overfits to the small set of base categories in COCO (e.g., 48 base categories), which hurts the generalization on novel categories. Focal scaling effectively alleviates the potential overfitting.

4 Discussion

Visualization. Fig. 3 visualizes the results of zero-shot inference with ground-truth boxes and 65 categories from COCO dataset. Our model predicts more reasonable categories than CLIP (e.g., the blue regions in 1st and 2nd columns are correctly predicted as “umbrella” and “person” by our model). These results suggest that our proposed region-based vision-language pretraining can help to recognize image regions precisely.

Further, the pretrained models can predict the customized object concepts by simply replacing the language embeddings of target categories. Fig. 4 visualizes results of zero-shot inference with ground-truth boxes and 1203 categories from LVIS dataset, instead of the small set of 65 categories from COCO dataset. We show the top-3 predictions for each region with their confidence scores.

As shown by the successful cases in Fig. 4, our pretrained model can correctly recognize the image regions while the CLIP model often fails to predict the correct labels (e.g., “teddy bear” is predicted by our model with a high confidence score 99.5%). Interestingly, other than the most-confident category, our model can also predict reasonable categories with top-3 scores (e.g., “bear” in 1st example and “truffle chocolate” in 2nd example). Even in the failure case where both CLIP and our model fail to recognize the dog as most-confident category, our model can still recognize the image region as visually similar concepts (e.g., “ferret” and “cub”) or a fine-grained type of dog (e.g., “shepherd dog”). On the contrary, CLIP predicts less visually similar concepts, such as “grizzly” and “gorilla”.

Limitations. Our work has several limitations that can be further investigated. (1) We focus on learning the object concepts without explicitly attending to other information in natural language, such as object attributes and object relationships, which are beneficial to some vision tasks (e.g., visual grounding). Learning comprehensive region representations can be a future work. (2) Our method relies on CLIP’s visual-semantic space and has not updated the language encoder. When given similar scale of data as CLIP, unfreezing the language encoder may bring more gain in our region-based language-image pretraining.

Conclusion

In this paper, we proposed a novel region-based vision-language pretraining method that learned to match image regions and their descriptions. Our key innovation is a scalable approach to associate region-text pairs beyond the tokens presented in the paired text data without using human annotation. Learning from such region-level alignment, our pretrained model established new state of the art when transferred to open-vocabulary object detection on COCO and LVIS datasets. Moreover, our pretrained model demonstrated promising results on zero-shot inference for object detection. We hope that our work can shed light on vision-language pretraining for visual region understanding.

References