Position-guided Text Prompt for Vision-Language Pre-training
Alex Jinpeng Wang, Pan Zhou, Mike Zheng Shou, Shuicheng Yan
Introduction
The vision-and-language pre-training (VLP) models like CLIP , ALIGN and CoCa have greatly advanced the state-of-the-art performance of many cross-modal learning tasks, e.g., visual question answering , reasoning , and image captioning . Typically, a generic cross-modal model is first pre-trained on large-scale image-caption data in a self-supervised fashion to see sufficient data for better generalization ability, and then fine-tuned on downstream tasks for adaptation. With remarkable effectiveness, this pre-training-then-fine-tuning paradigm of VLP models has dominated the multi-modality field.
In VLP, visual grounding is critical for many tasks as observed in previous research . To model the position information, traditional VLP models (the top of Fig. 1 (a)) employ a faster-rcnn pre-trained on the 1600 classes Visual Genome to extract salient region features and bounding boxes. Then these models use both the bounding box and object feature as input. In this way, these models not only learn what objects are contained in the salient region and where are these objects. However, when using region features as input, the model pays attention to the items inside the bounding boxes and ignores the contextual data outside of them . More seriously, on downstream task, these methods still need to use detectors to extract objects, giving very slow inference speed.
To get rid of region feature for higher efficiency, recent works (the middle of Fig. 1 (a)) adopt raw-pixel image as input instead of region features, and train the model with Image Text Matching and Masked Language Modeling loss end-to-end. Despite their faster speed, these models cannot well learn the object positions and also their relations. As shown in Fig. 1 (b), we observe that a well-trained ViLT model well know what objects are in an image. But this model does not learn the object positions accurately. For example, it wrongly predicts “the dog is on the right of this image”. However, during fine-tuning, downstream tasks actually require the object position information to comprehensively understand the image. Such a gap largely impairs the performance on downstream tasks.
In this work, we aims to ease the position missing problem for these end-to-end models, and keep fast inference time for downstream tasks at the same time. Inspired by the recently prompt learning methods , we propose a novel and effective Position-guided Text Prompt (PTP) paradigm (the bottom of Fig. 1 (a)) for cross-modality model pre-training. The key insight is that by adding position-based co-referential markers in both image and text, visual grounding can be reformulated into a fill-in-the-blank problem, maximally simplify the learning of object information. To ground natural language expressions in image data, PTP contains two components: (1) block tag generation to divide image into blocks and to identify object in each block, and (2) text prompt generation that puts the query text into a position-based text query template.
By bringing the position information into pre-training, our PTP enables strong visual grounding capabilities of VLP models. At the same time, as we do not used object detector for downstream tasks, we keep fast inference time. Experimental results show that our method outperforms their counterparts by a large margin especially for zero-shot setting. For example, our PTP-BLIP achieves 3.4% absolute accuracy gain over CoCa in zero-shot retrieval Recall@1 on coco dataset with much less training data (4M vs. 3B) and a much smaller model (220M vs. 2.1B). In addition to the zero-shot task, we show that PTP can achieve strong performance for object position guided visual reasoning and the other common VLP tasks such as visual question answering, and image captioning.
Related Work
Existing VLP models can be roughly grouped into three categories according to their architectures: one-stream models, dual-stream models and dual-stream + fusion encoder model. All three architectures are introduced below:
1) One-stream Model (e.g., UNITER , ViLT ) in Fig. 2 (a) operates on a concatenation of image and text inputs. 2) Dual-stream Model (e.g., CLIP ) in Fig. 2 (b) uses separate but equally expensive transformer encoders for each modality. The two modalities are not concatenated at the input level and interaction between the pooled image vector and text vector at shallow layer. 3) Dual-stream with Fusion Model (e.g., BLIP ) Fig. 2 (c) is a combination of one-stream and dual-stream model.
In this work, without loss of generality, we focus on prompting all these three kinds of VLP models due to their prevalence and adaptability to different downstream tasks.
2 Prompt Learning for Computer Vision
Prompt learning is originally designed for probing knowledge in pre-trained language models to specific downstream tasks . Recent years have seen a rise in the study of prompt tuning on vision taks, e.g. multi-modal learning and image understanding. The pioneer Color Prompt adds color prompt on image and text color description for visual grounding. Most related to our work is Multi-modality Prompt which presents multi-modality prompt tuning for VLPT models, achieving promising results on some vision-language tasks.
However, these efforts, like earlier NLP research, concentrate on prompt engineering in fine-tuning while leaving the pre-training phase unaffected. The goal of using the prompt design in this work, in contrast, is to provide the model the ability to understand semantic concepts at a finer level while it is still in the pre-training stage.
3 Learn Position Information in VLP
The grounding ability has shown to be essential for multiple cross-modality tasks . To introduce this ability into VLP models, bottom-up and top-down and its follow-up works concatenate region feature and bounding box vector together. But object extraction is time-consuming in inference for downstream task. Recently, some works propose train the VLP models with additional object localization loss or word patch alignment loss which, however, are hard to extend because they are specifically designed for particular frameworks. In contrast, we aim to propose a general framework for learning position information. To this end, we propose a simple text prompt that can be plug into existing frameworks easily.
Position-guided Text Prompt
In this section, we first elaborate on our proposed Position-guided Text Prompt paradigm (PTP for short). Then we introduce how to incorporate it with current vision-language pre-training (VLP) frameworks for boosting their visual grounding capabilities by taking the classical and popular VILT , CLIP and BLIP as examples.
To enhance the visual grounding ability of cross-modal models trained by VLP, we propose a novel and effective Position-guided Text Prompt (PTP) that helps a cross-modal model perceive objects, and also align these objects with pertinent text. PTP differs from the conventional vision language alignment methods, e.g. , that concatenate object feature and bounding box together as input to learn the alignment between objects and pertinent text, and thus paves an alternative way which indeed enjoys some advantages as shown and discussed in Sec. 3.2. As illustrated in Fig. 3, PTP has two steps: 1) block tag generation which divides an input image into several blocks and also identifies the objects in each block; and 2) text prompt generation that reformulates the visual grounding task into a fill-in-the-blank problem according to the object position information in step 1). Based on these steps, one can easily plug PTP into a VLP model by solving fill-in-the-blank problem in PTP. We will introduce these two steps below.
As shown in Fig. 3, for each image-text pair in the training phase, we evenly divide the input image into blocks. Then we identify the object in each block by one of the following two ways:
(1) Object Detector. We first adopt a strong Faster-rcnn used in VinVL to extract all objects for each image. This Faster-rcnn version is based on ResNeXt152 and is trained on 1600-classes Visual Genome . Then we select top- objects denoted by with highest prediction confidence, where denotes an object with 4-dimensional region position vector and object category . For each block, we select the objects whose region center are in that block. At last, the final block tag for this block is of these selected objects. In this work, we generate object tag with object detector as default.
(2) CLIP Model. Instead of heavy object detector, some recent works also try to generate region supervision based on CLIP because of its efficiency and effectiveness. Inspired by these works, PTP can also generate block-wise object supervision via CLIP (ViT-B) model https://huggingface.co/openai/clip-vit-base-patch16. First, we extract (3000 in default) key words/phrases that are most frequent on the whole text corpus Extract key word/phrase with NLTK (/https://github.com/nltk/nltk). These key words/phrases are regarded as our vocabulary . Then we extract the text feature of all these key words/phrases embedding via CLIP text encoder.
Additionally, we take the image embedding from each block and compute the similarity across every text feature. The keyword/phrase with the highest similarity score is selected as the final object tag for this particular block. Formally, the index of object tag per block is computed as
where is the visual feature embedding of selected block. Comparing with object detector, the CLIP model have two advantages. Firstly, as opposed to pre-defined object categories, more diverse object tags are produced. Secondly, the generation of block tag is much faster than object detector, e.g. 40 faster than Faster-RCNN (ResNeXt152) model. Please refer to Sec. 4.3 for comparison.
1.2 Text Prompt Generation
For the input image of each training pair, Sec. 3.1.1 already generate the object tags and positions which allows us to design a simple text prompt as follows:
where denotes the index of selected block and is used to denote the object position; denotes the object tag generated for the block . Note, we explore more prompt design choices in Section 4.3. For a certain , we may have various options for because the block may contain multiple objects. For such situation, we select one at random for each time. In this way, each sentence in our PTP incorporates fine-grained object position and language into a model, and thus provides a new way to align the objects and pertinent text.
2 Pre-training with PTP
In this work, we integrate our PTP into mainstream VLP frameworks, leading to PTP-ViLT , PTP-CLIP and PTP-BLIP . Following receipt of the PTP, we have two options for training these models:
Integrate into existing tasks. The simplest method for using text prompt is to change the text input. As shown in Fig. 3, the prompted text and original caption were simply padded together. Formally, the input caption of our method is represented as:
where is text and is our generated text prompt. Then we train the VLP models end-to-end with conventional objectives. Following , we employ Language Modeling (LM) loss, Image-text Matching (ITM), and Image-text Contrastive (ITC) loss for our PTP-BLIP; we use ITM and Masked Language Modeling (MLM) loss to train our PTP-ViLT; we only use ITC loss to train our PTP-CLIP. We use this method as default for all experiments because of its good performance.
As a new pretext task. Alternatively, we explore the position prediction as an additional language modeling task. Formally, if is the pretraining data and is a training token sequence of our generated text prompt , then at the timestep , we devise our model to predict a probability distribution . Then we regressively try to maximize the probability of being the correct token. The object prediction loss is computed as follow:
where is the trainable parameters of the model. In this way, the model is asked to predict which block has objects and what object is in this block.
Discussion. Notably, our method does not need to modify the base network and can be applied to any VLP models without bells and whistles. The model is designed to learn position information from raw-pixel image. Note that only during the pre-training stage, we would require the object’s position information; yet on downstream tasks, we evaluate model in normal end-to-end ways without object information to get rid of the heavy object feature extraction.
Experiments
In this section, we empirically evaluate PTP on multiple downstream tasks and present a comprehensive study.
We first describe the pre-training experimental conditions, including the datasets, training configurations, evaluation procedures, and baseline models used in our studies.
Datasets. As in earlier studies , we begin by using a 4M setup made up of four popular pre-training datasets (COCO , VG , SBU and CC3M ). Following recent work , we also explore 14M setting, which includes additional CC12M (actually only 10M image urls available) dataset besides 4M datasets. We refer readers to supplementary material for more dataset details.
Training Settings. Our models are implemented in PyTorch and pre-trained on 8 NVIDIA A100 GPUs. For the optimizer and training hyperparameter, we follow the original implementation in baseline works for fair comparison. For image augmentation, we explore RandAugment and use all the original policies except for color inversion since color information is important. We augment the bounding box in same way as image for affine transformation like rotation. We take random image crops of resolution during pre-training, and increase the image resolution to for finetuning.
Baselines. We evaluate three variants of pre-training frameworks, including one-stream ViLT , dual-encoder CLIP , and fusion-encoder BLIP , for their superior performance. For fair comparisons, we adopt the ViT-B/16 as base vision encoder and use same dataset.
2 Main Results
In this section, we integrated our PTP into existing networks and compare to existing VLP methods on a wide range of vision-language downstream tasks. Then we introduce each task and finetuning strategy. More details can be found in the supplementary material.
We evaluate PTP for both image-to-text retrieval (TR) and text-to-image retrieval (IR) on COCO and Flickr30K benchmarks. For PTP-BLIP, following original implementation, we adopt an additional re-ranking strategy.
We first report zero-shot retrieval result on both image-to-text and text-to-image setting in Tab. 1. We find PTP significantly improves baselines on all metrics. For example, for ViLT baseline, PTP leads to 13.8 % absolute improvement (from 41.3 % to 55.1 %) over Recall@1 of image to text retrieval on MSCOCO. In addition, based on strong BLIP , our PTP-BLIP even outperforms CoCa on most recalls of MSCOCO with much less data.
A summary comparison about fine-tuned setting between different models appears in Tab. 2, from which we observe that: (1) PTP outperforms the BLIP and ViLT baselines by a large margin in both datasets. For example, PTP-ViLT achieves an impressive 5.3% improvement on R@1 of TR in MSCOCO. (2) With strong BLIP as baseline, PTP-BLIP leads to state-of-the-art performance at same scale. Notice that the training cost remains the same BLIP baseline, because we train PTP with the same settings as the baseline and do not increase the maximum input text token. We can even reduce the gap between 4M setting and ALBEF (14M data), with similar framework.
From all these results above, we point out UNITER , OSCAR , VinVL , ImageBERT all use faster-rcnn as we used. However, our PTP leads to much better results than these related works. Besides, we only use object detector in pre-training stage. This indicates object detector is not the secret for success and how to leverage the position information is essential important for VLP models.
2.2 Image Captioning
This task asks the model to describe the input image. We consider two datasets for image captioning: No-Caps and COCO , both evaluated using the model finetuned on COCO with the LM loss. Similar to BLIP, we start each caption with the phrase “a picture of,” which yields marginally better results. We do not pre-train using the COCO dataset to avoid information leakage. For No-Caps dataset, following BLIP, we adopts a zero-shot setting (evaluate directly with the captioning model trained on CoCo dataset).
As shown in Tab. 3, related works utilizing a comparable quantity of pre-training data perform significantly worse than PTP-BLIP. The results of our method are closed to the VinVL with fewer training samples and smaller image. Finally, with 14M setting, our method leads to close result with LEMON, which trained on billions data and requires two times higher resolution image.
2.3 Visual Question Answering
VQA requires the model to predict an answer given an image and a question. For PTP-ViLT, we formulating VQA as a multi-answer classification task. For PTP-BLIP, we follow and consider it as an answer generation task that allows open-vocabulary VQA for better result.
The results are reported in Tab. 4. Compared to ViLT baseline, PTP brings 1.8% gains on both dev split. With 14M setting, PTP-BLIP achieves better performance than SimVLM , which uses 1.8B training samples and a ViT-Large based vision backbone.
2.4 Visual Reasoning
Natural Language Visual Reasoning (NLVR2) task is a binary classification task given triplets of two images and a question in natural language. This task relies on position information heavily. As shown in Tab. 4, SimVLM is outperformed by PTP-BLIP, which has a reasonable model size and was pretrained on fewer instances. Meanwhile, our method is also closed to VinVLlarge model that adopt larger model and use object feature from strong object detector instead of raw-pixel image as input.
2.5 Video-Language Tasks
We analyze the generalization ability of our method to video-language tasks in this experiment. Specifically, we perform zero-shot transfer to text-to-video retrieval in Tab. 5, where we directly evaluate the models trained on COCO-retrieval. We just uniformly sample 8 frames each video in order to process video input, then concatenate the frame features into a single sequence. Our method leads to better result than OA-Trans that focus on retrieval task, which showcase the generality capability of PTP.
3 Ablation & Design Choices
In this section, we first evaluate our method on retrieval task over three well-known baselines under 4M setting for comparison. Then we train a BLIP model on CC3M as baseline and perform various ablations.
We experiment with three distinct kind baselines: ViLT, CLIP, and BLIP in order to explore the impact of PTP. Tab. 6 reports the performance on the COCO 5K test set. Comparing the outcomes of these baseline experiments, we find that PTP greatly improves the i2t and t2i performance. This suggests that PTP has good generality.
In addition, we also compare the running time. Since we do not use object detector or prompt in downstream task, the computation cost keep consistent with baseline models but 20 times faster than object feature based VinVL .
3.2 Text Prompt vs. Additional Pretext Task
We examine the effects of regarding PTP as a new pretext task. In this way, the pretext task does not influence the other pre-training objectives, such as ITM and ITC, but it does add to the cost of computation. Contrarily, the prompt design simply modifies the text input, therefore it will have an impact on all pre-training objectives.
We report the result in Tab. 7. We observe both Pretext and Prompt design improved the baseline over all four tasks. However, prompting is far preferable to pretext, particularly for COCO captioning CIDER (127.2 vs 123.5). In this work, we use prompt as default due to its efficiency.
3.3 Other Types of Text Prompt
In this experiment, we explore six different kind of prompts: i. The [O] is in block [P]. ii. The block [P] looks like [O]. iii. The [O] is in which block? In [P]. iv. The [O] is located in block [P]. v. (, , , ) has a [O]. is the top left point and are the width and height for bounding box. vi. The block [P] has a [O]. vii. The block [NP] has a [O]. NP means we use nouns to represent the block position. e.g, from upper left to bottom right. More variations can be found in the supplementary.
We report the result in Tab. 8 and observe precise position does not produce superior results to block, the reason maybe precise position is hard to learn. In addition, we find use block ID (like 0) or nouns (like upper left) remain similar results. In the end, we discover that the hybrid version does not produce the best outcomes.
3.4 The Importance of Position in Text Prompt
In this experiment, we examine the efficacy of prompting our PTP for information at various granularities, such as without Positional. We simply use [P] has [O] when remove prompt. We list the results in Tab. 9. We observe: i. It’s interesting to see that each component is crucial. Without any one component, the downstream performance to get progressively poorer. ii. Although OSCAR discovered that using object tags as a supplementary input improved results when area features were used as input, we have shown that object tags are ineffective when raw pixel images are used. This serves as an illustration of the need to create a workable prompt for understanding the alignment between object tags and image region.
3.5 Number of Blocks
We explore if more fine-grained position information helps in our PTP. In Fig. 4, we varying the number of blocks from (remove position information in PTP) to and report the relative performance based on both BLIP and ViLT models. As can be seen, the results for both backbones are improved when the number of blocks is more than 1. However, once there are 16 blocks, all downstream activities experience a relative drop in performance. The reason may be that the predicted bounding box deviates from the localization of the real object, resulting in a mesh that is too small and may not contain the selected object. We hence recommend using blocks, as it enjoys accurateness.
3.6 Is Object Detector Necessary?
In this work, a part of predicted bounding box information is coming from Faster-rcnn . In order to verify the expressive power of object, we also consider two variations: i. Pure clip similarity. This design choice is adapted mainly for efficiency reasons, where utilizing object detector is time consuming and not easy to access sometimes. ii. In addition to the powerful ResNext152-based object detector, we also use a smaller Faster-rcnn network that utilizes ResNet101 as backbone.
The results are reported in Tab. 10. We also report the overall feature extracting time on 8 NVIDIA V100 GPUs. As can be seen from the table, we found that using stronger detector leads to better result, but bring huge computation cost at the same time. Moreover, we observe the result of CLIP embedding is very closed to Faster-rcnn (ResNeXt152). In addition, it takes only around 2.3% time of Faster-rcnn (ResNeXt152) version to extract pseudo label for each grid. We came to the conclusion that a clip model is a good alternative of object detector in PTP.
4 Visualization
To explore whether model training with the PTP framework does indeed learn position information, we design a fill-in-the-blank evaluation experiment in this section. Follow ViLT , we masked some key words and asked the model to predict the masked words and show its corresponding heatmap. We design two text prompts, given the noun to predict the localization and given the localization to predict the missing noun. We show top-3 predictions and more visualization results can be found in supplementary.
The results are shown in Fig. 5. On the one hand, we find that the PTP-ViLT can make correct object prediction based on the block position information and its visual concepts. On the other hand, when only masked the position information, we witness a high predicted probability value for corrected block. For example, in the bottom of Fig. 5, our model find all patches looks like “man” correctly. Based on these experiments and Fig. 1, we conclude that the PTP can help the base VLP model learn position information very well based on our simple text prompt.
Furthermore, we cluster the token-level features with K-Means algorithm for ViLT and PTP-ViLT. Intuitively, the token with similar semantic should be clustered together. We show the visualization result in Fig. 6. Comparing with ViLT baseline, we observe that our method can cluster similar patches more accurate. This illustrate our PTP have fairly accurate learns semantic information.
Limitations and Conclusion
We first try to leverage the position information from existing object detector/trained model to VLP models with simple prompt. We provide a success practice cross-modal prompt settings to aid prompt engineering. Through rigorous experiments, we showed that PTP could serve as a general-purpose pipeline and improve the learning of position information without much extra computation cost. However, at this time, PTP does not take into account how to deal with the wrong object tag. Additionally, this work does not adequately explore more complicated prompts. Future research will also examine how well PTP performs on additional vision-language tasks.
References
Appendix
Appendix A Pre-training and Fine-tuning Details
In this work, we explore both 4M and 14M setting. The 14M setting is a combination of 4M setting and CC-12M. We report the data statistics in Tab. 11. Since these URLs are from Interent and a part of them already invalid, we only download 2.8M data of CC3M and 10.2M data of CC12M dataset, correspondingly. Notice that BLIP baseline use 3M data for CC3M, which is slightly more than our version. The amount of images that containing bounding box for CC3M is 2.69M, and for CC12M is 7M. These bounding boxs are used in our PTP. For quick evaluation, we pre-train the BLIP model for 50K steps rather than the 200K steps used in earlier works .
A.2 Hyper-parameters for Downstream Tasks
We first report the hyper-parameters of BLIP baseline in Tab. 12. The final decoder outputs from the encoder-decoder model BLIP can be used for multimodal understanding and generation. Thus, we evaluate on popular vision-language benchmarks. We mainly follow the same setup introduced in BLIP . The optimizer for all task is AdamW . We only train the retrieval task for 6 epochs in order to increase efficiency, and we think that more epochs will produce better results.
For ViLT baseline, we evaluate mainly on three tasks: vision-question answering, image-text retrieval and natural language visual reasoning. The hyper-parameters for ViLT on downstream tasks are reported in Tab. 12. For the CLIP baseline, we use the same hyper-parameters setting as BLIP baseline.
Appendix B More Ablation Study
We also explore multiple other prompt design choices in this section. The model is trained on CC3M and we evaluate on three downstream tasks. Specifically, we exploring the following ways: i. Multiple Tags. We observe that a block may contain many objects for many cases. We try to refine the text prompt as The block [P] has objects [O1], [O2] and [O3]. Keep in mind that each block has a different object number. ii. Multiple Position. We created a multiple position setup taking into account that one object could appear in numerous blocks. In practical, we refine the prompt as a question-answer pairs. iii. Synonymous Substitution We replace “block” with “region” and “is” with “looks like”.
The results is reported in Tab. 14. We observe that the multiple objects or multiple position not helps the model’s performance on downstream tasks too much, we also observe that the language modeling loss is higher than baseline. This shows that the assignment is too challenging for the model to learn. We also see that the outcome of simple synonymous replacement maintains consistency with the outcome of the original text prompt. We find that modeling location information only requires a straightforward prompt.
B.2 How many Objects do we need?
To generate object tag, we use Faster-RCNN as default and detect at least 10 objects from each image. In this experiment, we varied the number of items from 5 to 30, exploring the effects of various object counts.
The result is shown in Fig. 8. We observe there exist a slightly rising trend at the beginning of BLIP baseline. This demonstrates how crucial data diversity is for activities that come afterwards. The findings, however, are poorer when there are more objects. The cause is because a large number of objects with low confidence simultaneously produce false predictions. In this work, we set the object number as 10 as default.
B.3 Part Bounding Box Annotation
As some urls for CC dataset already invalid and some images have wrong fromat, we extract objects from 2.7M data of CC3M and 7M data of CC12M. In this way, only 10M of pre-training sample have objects available. We also report the result with 14M setting. Specifically, we only use original text without text prompt if we do not have object available.
The result is shown in Tab. 15. We observe 68.6% result in 134.6 CiDER value on COCO Captioning and 82.9 on NLVR accuracy. This illustrates more annotated samples leads to better result. This also encourage PTP is suitable for large-scale pre-training.
Appendix C More Visualization
In this section, we show object detection result with our generated text prompt. Specifically, we random select one object from and then we visualize the original image and its bounding box’s mask. Notice we augment these bounding box as the same as original image for affine transformation.
We random select some samples from the overall dataset and the result is reported in Fig. 7. We also observe the bounding box maybe very large and cross multiple blocks in some examples (e.g. the first case in the third row). Since we use RandAugment in this work, some object may be outside of the broder of input image. For such situation, we just replace the specific position with [X], the final PTP is The block [X] has a [O]. We also find that some masks may be no square. For example, the last example in the third row.
C.2 Case Analysis
In this experiment, we show some cases about position information in Fig. 9. The position information is important for various downstream tasks on the top.
Then, since a large of samples in VQA tasks include position information usually. We ask our model to do the vqa tasks and select some representative samples. Specifically, we show the prediction probability and predicted nouns in the bottom of this figure. We observe that the PTP give accurate prediction on most cases, which illustrates our PTP learns position information better.