A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models
Woojeong Jin, Yu Cheng, Yelong Shen, Weizhu Chen, Xiang Ren
Introduction
Fine-tuning large pre-trained language models (PLMs) have led to strong results in various domains including vision-language tasks Devlin et al. (2019); Raffel et al. (2020); Brown et al. (2020); Radford et al. (2021). Such large PLMs can learn a new task with a few examples or generalize to a new task without fine-tuning on any training examples, i.e., few-shot and zero-shot learning Brown et al. (2020); Radford et al. (2021); Tsimpoukelli et al. (2021). Few-shot learning overcomes the challenges of data-hungry supervised learning, where collecting human-labeled data is costly and slow. However, recent few-shot models such as GPT3 Brown et al. (2020), Frozen Tsimpoukelli et al. (2021), and PICa Yang et al. (2021) are too large to deploy in small or moderate computing machines due to their gigantic model sizes
In this paper, we study low-resource learning of VL tasks with our proposed method, FewVLM, a moderate-sized vision-language model, in which we fine-tune the model with no or a handful of training examples. For FewVLM, we pre-train a sequence-to-sequence transformer model Cho et al. (2021); Raffel et al. (2020) with prefix language modeling (PrefixLM) and masked language modeling (MaskedLM). This setup is more practical in that training and inference can be run economically using standard computing hardware and it is expensive to obtain a large number of quality training examples in the real world. In such a few-shot setting, task-specific prompts or task descriptions are important and have shown effectiveness in few-shot NLP tasks Gao et al. (2021); Radford et al. (2021); Schick and Schütze (2021a, b); Brown et al. (2020).
To extend the success to VL tasks, we aim to answer the following questions for prompt-based low-resource VL learning. Q1) How does prompt design affect zero/few-shot learning on new tasks? Q2) Does prompt design still matter given larger training? Q3) How do different pre-training objectives affect zero/few-shot learning? To answer these questions, we explore various prompt formats including hand-crafted and noisy prompts on zero/few-shot VL learning datasets. In addition, we study pre-training objectives on few-shot tasks inspired by Raffel et al. (2020): prefix language modeling (PrefixLM) inspired by Raffel et al. (2020) and masked language modeling (MaskedLM). To this end, we investigate the model’s performance on few-shot VL tasks including visual question answering Goyal et al. (2017); Marino et al. (2019); Hudson and Manning (2019), captioning Agrawal et al. (2019); Young et al. (2014) (Fig. 1), and miniImageNet Vinyals et al. (2016).
In our empirical analysis, our FewVLM with prompt-based learning outperforms Frozen Tsimpoukelli et al. (2021) which is 31 larger than FewVLM by 18.2% point on zero-shot VQAv2 and achieves comparable results to a 246 larger model, PICa Yang et al. (2021). Furthermore, we observe that (1) prompts significantly affect zero-shot performance but marginally affect few-shot performance on new tasks (§6.2 and §6.3), (2) models with noisy prompts learn as quickly as hand-crafted prompts given larger training data (§6.5), and (3) MaskedLM helps few-shot VQA tasks while PrefixLM boosts captioning performance (§6.6).
Related Work
Vision-language few-shot learning. Recently, several few-shot learners on vision-language tasks were proposed including GPT Radford et al. (2019); Brown et al. (2020), Frozen Tsimpoukelli et al. (2021), PICa Yang et al. (2021), and SimVLM Wang et al. (2021). Frozen Tsimpoukelli et al. (2021) is a large language model based on GPT-2 Radford et al. (2019), and is transformed into a multimodal few-shot learner by extending the soft prompting to incorporate a set of images and text. Their approach shows the few-shot capability on visual question answering and image classification tasks. Similarly, PICa Yang et al. (2021) uses GPT-3 Brown et al. (2020) to solve VQA tasks in a few-shot manner by providing a few in-context VQA examples. It converts images into textual descriptions so that GPT-3 can understand the images. SimVLM Wang et al. (2021) is trained with prefix language modeling on weakly-supervised datasets. It demonstrates its effectiveness on a zero-shot captioning task. While these models achieve improvement on few-shot tasks, they are impractical to use in real-world applications due to their model sizes.
Language model prompting. Providing prompts or task descriptions play an vital role in improving pre-trained language models in many tasks Gao et al. (2021); Radford et al. (2021); Schick and Schütze (2021a, b); Brown et al. (2020). Among them, GPT models Radford et al. (2019); Brown et al. (2020) achieved great success in prompting or task demonstrations in NLP tasks. In light of this direction, prompt-based approaches improve small pre-trained models in few-shot text classification tasks Gao et al. (2021); Schick and Schütze (2021a, b). CLIP Radford et al. (2021) also explores prompt templates for image classification which affect zero-shot performance. We follow these core ideas so we aim to improve zero-shot and few-shot performance using prompts in vision-language tasks.
Analysis Setup
In this work, we study the zero-shot and few-shot performance of vision-language models . We introduce our analysis setup: problem formulation, analysis questions, downstream tasks and datasets, evaluation metrics, and baselines.
For zero-shot tasks, a pre-trained VL model have no access to training set and development set , and directly makes inference on the test instances . For few-shot tasks, we compose a dev set from training data and ensure that following Perez et al. (2021); Gao et al. (2021) to tune the hyper-parameters and select the model. We limit the sizes of training and development sets to meet the goal of learning from limited data. The size of and are small — i.e., we set the size of both to 16 in our study.
2 Analysis Questions
We aim to answer the following questions in this study through experiments on multiple VL datasets.
Q1) How does prompt design affect zero/few-shot learning on new tasks? Providing a pre-trained language model with task-specific prompts or significantly improves zero-shot and few-shot performance on NLP domains Gao et al. (2021); Schick and Schütze (2021a, b); Brown et al. (2020). For this question, we test several ad-hoc prompts on vision-language tasks and analyze how large zero-shot and few-shot performance is affected by different prompts, hand-crafted and noisy prompts, in Sec. 6.5.
Q2) Does prompt design still matter given larger training data? As we will see in our experiments, prompts affect the zero/few-shot performance. However, prompts may have different effects when models are given different sizes of training data. To answer this question, we train models with different sizes of training data and various prompts, and compare the performance between different prompts.
Q3) How do different pre-training objectives affect zero/few-shot performance? We study two different pre-training objectives on few-shot performance: prefix language modeling (PrefixLM) inspired by Raffel et al. (2020) and masked language modeling (MaskedLM). In this setup, we pre-train our model with different objectives and test the model on zero-shot and few-shot tasks in Sec. 6.6.
3 Downstream Tasks and Datasets
In this work, we mainly focus on three tasks: visual question answering, captioning, and categorical learning. The visual question answering task requires models to answer a question to a given context image. We convert the visual question answering task into a generation task so that the model can generate answers in the zero-shot setting. The captioning task requires a model to generate descriptions for a given context image. The categorical learning requires a model to choose the correct category or class. We evaluate our model in an open-ended fashion to quantify fast learning of categories, in which it must generate correct labels unlike other classification methods.
We include VQAv2 Goyal et al. (2017), OK-VQA Marino et al. (2019), and GQA Hudson and Manning (2019) for visual question answering tasks, and NoCaps Agrawal et al. (2019), and Flickr30k Young et al. (2014) for image captioning.We include COCO captioning results on Sec. B of Appendix. We use Karpathy split Karpathy and Li (2015) for Flickr30k, which re-splits train and val images into 29,000 / 1,014 / 1,000 for train / validation / test. For categorical learning, we include miniImageNet Vinyals et al. (2016), a meta learning dataset. Following Tsimpoukelli et al. (2021), we use only meta test data to evaluate FewVLM in a few-shot manner and test on 5-way -shot setup, where 5 classes and examples per class are given.For VQA and captioning, we include samples in total, not per class.
4 Evaluation Metrics
To evaluate few-shot performance, we randomly sample 5 different training and dev splits and measure average performance on the 5 splits. We fine-tune the vision-language models with 200 epochs for the few-shot setup and choose the best checkpoint on the dev set. For NoCaps task, it does not have training data. Thus we use the training data from COCO captioning in the experiments following Wang et al. (2021). We evaluate on the VQAv2 validation set, GQA test-dev, OK-VQA test set, test set of Karpathy split for Flickr30k captioning, and NoCaps validation set. We adopt accuracy for VQA datasets and miniImageNet, and CIDEr Vedantam et al. (2015) and SPICE Anderson et al. (2016) as evaluation metrics for captioning.
5 Baselines
We evaluate strong zero/few-shot vision-language learners for comparison: Frozen Tsimpoukelli et al. (2021), PICa Yang et al. (2021) for VQA datasets and SimVLM Wang et al. (2021) for captioning datasets. We include Unified VLP Zhou et al. (2020) for few-shot VQAv2 and Flickr30k. Also, we compare them with fully fine-tuned models as upper bounds of few-shot models for each task; these models are fine-tuned on the entire datasets while few-shot models can access a small amount of data. For fully fine-tuned models , we borrow numbers from Uniterlarge Chen et al. (2019) for VQAv2, Oscar Li et al. (2020b) for GQA, SimVLM Wang et al. (2021) and VinVL Zhang et al. (2021) for NoCaps CIDER and SPICE respectively, and Unified VLP Zhou et al. (2020) for Flickr30k captioning. We include VL-T5 as a baseline which is pre-trained without visual question answering datasets Cho et al. (2021). For miniImageNet, we include Frozen and AFHN Li et al. (2020a). Frozen is designed for few-shot learning while AFHN is for meta learning, which is smaller and faster.
Method
Before diving into the analysis, we introduce our model, FewVLM, to do zero/few-shot learning on VL tasks and answer the analysis questions we raised. We introduce FewVLM architecture and pre-training objectives.
We adopt an encoder-decoder architecture Cho et al. (2021); Vaswani et al. (2017), to encode visual and text inputs and generate target text. We represent an input image with 36 object regions from a Faster R-CNN Ren et al. (2015) trained on Visual Genome Krishna et al. (2017). The sets of region representations are fed into the encoder by appending them to the text Cho et al. (2021). We train the model parameters by minimizing the negative log-likelihood of target text tokens given input text and image :
The model is not task-specific, so it is a good option for zero/few-shot settings.
2 Pre-training Objectives
We pre-train the models with both prefix language modeling (PrefixLM) and masked language modeling (MaskedLM). Fig. 3 illustrates the PrefixLM and MaskedLM.
Prefix language modeling. We include prefix language modeling (PrefixLM) following Raffel et al. (2020). Given an image and a span of text, this objective randomly splits the text into two separate components; the former component with the given image is used as inputs to the encoder and the latter component is used as target text to be generated by the decoder.
Masked language modeling. We follow Cho et al. (2021) to do masked language modeling. This objective is to replace random spans with numbered sentinel tokens, e.g.,
Pre-training data. To pre-train FewVLM, we collect image-caption data from MS COCO Lin et al. (2014); Chen et al. (2015) and Visual Genome (VG) Krishna et al. (2017). The pre-training datasets contains 9.18M image-text pairs and 180K distinct images.
Low-resource Adaptation
In downstream tasks, we train our model with few-shot examples. Fig. 2 shows an illustration of FewVLM in inference time. Given a prompt template , we first get input text and target text using the template . Then we train model parameters by minimizing the negative log-likelihood in Eq. (1). In inference, we use the same prompt and the model generates the label text. Here we obtain the final label by removing the target prompt template.
Prompts affect the performance of the vision-language model Cho et al. (2021); we study the effect of different prompts on the zero-shot and few-shot performance on downstream tasks. Tables 1 and 11 show prompts we used in our experiments.
The visual question answering tasks (VQA, OK-VQA, and GQA) require models to answer a question to a given context image. Recent approaches Chen et al. (2019); Tan and Bansal (2019); Su et al. (2020); Li et al. (2019, 2020b) tackle visual question answering tasks as multi-label classification over a predefined set of answer candidates. Instead, we approach the visual question answering tasks as a generation task so that the model can produce the answers without introducing any task-specific heads. In this setup, prompts act as constraints to guide the models to generate proper formats of answers; models might generate a sentence for VQA, which is not the correct format, without prompts.
Therefore, we study several prompts for input and output as shown in Tables 1 and 11; we explore hand-crafted prompts (Table 1) and noisy prompts for ablation study (Table 11).
Hand-crafted prompts. For input prompts, we explore three different templates: “question: [Q] answer:” and with the
Noisy prompts. To understand the effect of noisy prompts in zero/few-shot learning, we include irrelevant prompts, noisy tokens, and random sentences as in Table 11. Irrelevant prompts are random questions or instructions that mislead models to answer wrong questions or follow irrelevant instructions. Noisy tokens are randomly selected from T5’s vocabulary, so we test how robust our model is to random tokens. Finally, random sentences are captions from MS COCO and this gives false information to models.
1.2 Captioning.
In NoCaps and Flickr30k, we explore three hand-crafted input prompts: “a picture of”, “a photo of”, and “an image of”. We study the effect of different word choices in this captioning task. While the three different words have similar meanings, they show different performance in zero-shot and few-shot tasks as we will see in our experiments.. For target prompts, we just train the model with the original caption without any additional prompts.
1.3 MiniImageNet
In miniImageNet, we train our model with a hand-crafted input prompt, “This is
Results and Discussion
In this section, we first discuss our main results on zero-shot and few-shot tasks and then answer the questions we raised: does prompt design matter in zero/few-shot learning?
For pre-training, we set batch size 1,280 and 800 for FewVLMbase and FewVLMlarge, respectively and pre-train them with 30 epochs. We use learning rate 1e-4 with 5% linear warmup. For few-shot learning, we train models with 200 epochs, learning rate 5e-5 and 5% linear warmup and choose the best checkpoint on the dev set. For FewVLM, we use “question: [Q] answer
2 Performance on Zero-shot Learning
We evaluate the existing models in a zero-shot manner, in which models do not have access to any training data. Tables 3 and 5 show the results on VQA and captioning datasets, respectively. First, FewVLM with the hand-crafted prompt (P3) achieves better performance than other baselines on VQA datasets. In particular, our FewVLMbase significantly outperforms Frozen which is about 31 larger than ours. Also, PICa based on GPT3 Brown et al. (2020) shows the best performance on OK-VQA. It is noticeable that our FewVLMlarge, the 246 smaller model, achieves the comparable result to PICa. Compared to VL-T5 which is the same architecture as ours, FewVLMbase improves VQAv2 performance by about 30% point. As we will see in the later section, our pre-training objectives and the prompts boost the VQA performance. On NoCaps, SimVLMhuge shows the best performance. Our FewVLMbase significantly improves the performance compared to VL-T5. As we will see in the later section, our pre-training objectives and the prompts boost the VQA and captioning performance.
3 Performance on Few-shot Learning
Tables 3 and 5 show the few-shot performance on VQA and captioning datasets. Sizes of training and validation sets are 16 for FewVLM, VL-T5, and Unified VLP; and Frozen and PICa use 4 and 16 in-context demonstration examples, respectively.
On VQAv2 and OK-VQA, PICa shows the best performance while our FewVLMlarge achieves the comparable result on VQAv2. OK-VQA requires external knowledge to answer unlike other VQA datasets, so larger models and large pre-training data (prior knowledge) are necessary to improve. Interestingly, FewVLM, which is trained with 4 training examples, outperforms Frozen. On captioning data, FewVLMbase notably outperforms VL-T5 by 31.1% point on NoCaps CIDEr.
Unified VLP slightly underperforms FewVLM on Flickr30k captioning task. We conjecture that their architecture is based on a encoder-decoder transfomer and it is pre-trained with a captioning task Zhou et al. (2020).
4 MiniImageNet
Table 6 shows results on miniImageNet, where models must choose the correct class for each image. We train and evaluate FewVLM in an generative manner; the model must generate correct label text to get the credit. FewVLM significantly outperforms Frozen in all shots. Note that we train FewVLM with a few training samples while Frozen uses them as in-context demonstration. Interestingly, FewVLM with a hand-crafted prompt improves performance a lot on the 1-shot case, while it marginally improves on the 5-shot case.
5 Study of Prompt Design
Here we examine the effect of different prompts on FewVLMbase in Table 7 and Figs. 6, 5, and 4. We test the model on VQAv2 and Flickr30k datasets.
Table 7 shows the zero-shot performance on VQAv2 and Flickr30k. We observe that zero-shot results are remarkably affected by input prompts on both datasets. For input prompts,
On Flickr30k, we examine different word choices of prompts: “a picture of” (Q1), “a photo of” (Q2), and “an image of” (Q3). For instance, using “an image of” outperforms using no prompt by 21.4 point. It is noticeable that different word choices significantly affect the zero-shot results.
5.2 Few-shot Predictions
We study various input prompts including irrelevant prompts, noisy tokens, and random sentences on VQAv2 (Fig. 4). First, noisy prompts and no prompt achieve near 0 accuracy on the zero-shot setting. In few-shot predictions, FewVLM with noisy prompts learns as quickly as hand-crafted prompts given larger data. For example, our model with noisy prompts achieves comparable results to the best hand-crafted prompt. Among all different types of noisy prompts, random sentences deteriorate performance the most. This is because the random sentences come from captions in MS COCO, so the model might choose the answer from wrong captions not from images. Interestingly, no prompt outperforms the other noisy prompts and even shows similar to or better than the hand-crafted prompt with larger training data. We also observe a similar phenomenon on Flickr30k; no prompt performs similar to hand-crafted prompts in Fig. 5.
In addition, we explore two different target prompts, “ We investigate how pre-training objectives affect different tasks. We pre-train FewVLM with different pre-training objectives: masked language modeling (MaskedLM) and prefix language modeling (PrefixLM). In Table 8, we observe that MaskedLM helps VQA tasks while PrefixLM helps captioning tasks in zero-shot and few-shot settings. We conjecture that MaskedLM is to predict spans, which is analogous to predict correct answers to questions, and PrefixLM is to generate the rest of the given prefix, which is similar to captioning tasks. In other words, if the pre-training task is similar to the downstream tasks, then it will help performance further. When pre-training with both objectives, they create a synergetic effect and thus improve cross-task generalization. In this work, we present FewVLM, a few-shot prompt-based learner on vision-language tasks. On diverse datasets, FewVLM outperforms baselines and shows comparable results to PICa which is 246 larger than ours. We observe that prompts are vital in zero-shot and few-shot tasks and each pre-training objective helps different few-shot tasks. Also, we find out that models with larger training data are not significantly affected by noisy prompts. Future work includes exploring automatic prompt generation and diverse formats of few-shot tasks such as multiple-choice VQA. Finding optimal prompts require exhaustive engineering to achieve the best performance and leads to impressive results. We leave the exploration of these directions to future investigations. Table 9 shows model parameters in our model, FewVLM. FewVLMbase and FewVLMlarge is based on VL-T5 Cho et al. (2021) and T5 Raffel et al. (2020), respectively. We evaluate our model with COCO captioning data. We use Karpathy split Karpathy and Li (2015) for MS COCO captioning, which re-splits train and val images into 113,287 / 5000 / 5000 for train / validation / test. Table 10 shows the results on COCO. Tables 7, 8, and 9 show the results of each prompt on VQAv2 and Flickr30k with various training sizes. We pre-train our model with different datasets: MS COCO and Visual Genome (VG), and Conceptual Captions (CC). We investigate which pre-training dataset helps the downstream tasks in a few-shot manner. In Table 12, we observe that MS COCO and VG datasets are more helpful to the downstream tasks than CC.6 Pre-training Objectives
Conclusion
References
Appendix A Model Architectures
Appendix B COCO Captioning
Appendix C Prompt Study
Appendix D Effect of Pre-training Data