Unified Vision-Language Pre-Training for Image Captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, Jianfeng Gao
Introduction
Inspired by the recent success of pre-trained language models such as BERT (?) and GPT (?; ?), there is a growing interest in extending these models to learning cross-modal representations like image-text (?; ?) and video-text (?; ?), for various vision-language tasks such as Visual Question Answering (VQA) and video captioning, where traditionally tedious task-specific feature designs and fine-tuning are required.
Our vision-language Transformer network, which unifies the Transformer encoder and decoder into a single model, is depicted in Fig. 2 (left). The model input consists of the class-aware region embedding, word embedding and three special tokens. The region embedding is defined as:
The word embeddings are similarly defined as in (?), adding up with positional embeddings and segment embeddings, which is again overloaded as . We define three special tokens [CLS], [SEP], [STOP], where [CLS] indicates the start of the visual input, [SEP] marks the boundary between the visual input and the sentence input, and [STOP] determines the end of the sentence. The [MASK] tokens indicate the masked words which will be explained in the next section.
Pre-training Objectives
In the BERT masked language modeling objective, 15% of the input text tokens are first replaced with either a special [MASK] token, a random token or the original token, at random with chances equal to 80%, 10%, and 10%, respectively. Then, at the model output, the hidden state from the last Transformer block is projected to word likelihoods where the masked tokens are predicted in the form of a classification problem. Through this reconstruction, the model learns the dependencies in the context and forms a language model. We follow the same scheme and consider two specific objectives: the bidirectional objective (bidirectional) as in BERT and the sequence to sequence objective (seq2seq), inspired by (?).
For simplicity, we assume a single attention head in the self-attention module. Then, the self-attention output on can be formulated as:
where , , and are the embedding weights (the bias terms are omitted). The intermediate variables , , and indicate values, queries and keys, respectively, as in the self-attention module (?). is further encoded by a feed-forward layer with a residual connection to form the output . During the pre-training, we alternate per-batch between the two objectives and the proportions of seq2seq and bidirectional are determined by hyper-parameters and , respectively.
It is worth noting that in our experiments we find that incorporating the region class probabilities () into region feature () leads to better performance than having a masked region classification pretext as in (?; ?). Therefore, differing from existing works where masked region prediction tasks are used to refine the visual representation, we indirectly refine the visual representation by utilizing it for masked language reconstruction. We also choose not to use the Next Sentence Prediction task as in BERT, or in our context predicting the correspondence between image and text, because the task is not only weaker than seq2seq or bidirectional but also computationally expensive. This coincidentally agrees with a concurrent work of RoBERTa (?).
Sequence-to-sequence inference. Similar to the way seq2seq training is performed, we can directly apply VLP to sequence-to-sequence inference, in the form of beam search. More details follow next in the Image Captioning section.
We fine-tune the pre-trained VLP model on the target dataset using the seq2seq objective. During inference, we first encode the image regions along with the special [CLS] and [SEP] tokens and then start the generation by feeding in a [MASK] token and sampling a word from the word likelihood output (e.g., greedy sampling). Then, the [MASK] token in the previous input sequence is replaced by the sampled word and a new [MASK] token is appended to the input sequence to trigger the next prediction. The generation terminates when the [STOP] token is chosen. Other inference approaches like beam search could apply as well.
Visual Question Answering
We frame VQA as a multi-label classification problem. In this work we focus on open domain VQA where top most frequent answers are selected as answer vocabulary and used as class labels. Following (?) we set to .
During the fine-tuning, a multi-layer Perceptron (Linear+ReLU+Linear+Sigmoid) on top of the element-wise product of the last hidden states of [CLS] and [SEP] is learned, similar to (?). We optimize the model output scores with respect to the soft answer labels using cross-entropy loss. Note that unlike (?) where the task-specific objective (i.e., VQA) is exploited during pre-training by using the target datasets (from intensive human annotations), our pre-training does not have this requirement and is therefore more general.
Data preparation. We conduct pre-training on the Conceptual Captions (CC) dataset (?) which has around 3 million web-accessible images with associated captions. The datasets for downstream tasks include COCO Captions (?), VQA 2.0 (?) and Flickr30k (?). For COCO Captions and Flickr30k, we follow Karpathy’s splitcs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip, which gives 113.2k/5k/5k and 29.8k/1k/1k images for train/val/test splits respectively. For VQA 2.0, we split the dataset with the official partition, i.e., 443.8k questions from 82.8k images for training, 214.4k questions from 40.5k images for validation and report the results on Test-Standard set through the official evaluation server. We trim long sentences and pad short sentences to 20 words and all the words are tokenized and numericalized as in BERT (?).
Implementation details. Our Transformer backbone is the same as BERT-base (?). The input of the network consists of image (regions) and the associated/target caption. We represent each input image as 100 object regions extracted from a variant of Faster R-CNN (?) pre-trained on Visual Genome (?; ?). We take the model output from fc6 layer as the region feature () and the class likelihood on the 1600 object categories as region object labels (). Note that if not specified, the weights in our BERT model are initialized from UniLM (?) pre-trained on text corpora only. For caption inference, we use greedy search on the validation set and beam search with beam size 5 on the test set. We perform light model hyper-parameter search with the configurations presented in Appendix. is set to 0.75 for CC pre-training from light model validation (out of ), and set to 1 for image captioning (i.e., full seq2seq) and 0 for VQA (i.e., full bidirectional).
Model variants and metrics. To demonstrate the effectiveness of our vision-language pre-training, we first include a baseline model without this pre-training. We then include two extreme settings of our model with (seq2seq pre-training only) and (bidirectional pre-training only) to study how each objective individually works with different downstream tasks. Our full model conducts joint training on the two objectives. The fine-tuning procedure is performed the same regardless of the pre-training configurations. Regarding evaluation metrics, we use standard language metrics for image captioning, including Bleu@4, METEOR, CIDEr, and SPICE and the official measurement on accuracy for VQA, over different answer types including Yes/No, Number, and Other.
Comparisons against SotAs. Results comparing our methods and SotA methods on the test set are in Tab. 2. We include state-of-the-art published works (upper part of Tab. 2), unpublished works that are currently in submission (middle part), and our methods (lower part). All the image captioning methods are single models, with cross-entropy optimization only for a fair comparison. Our full model (Unified VLP) outperforms SotA methods on three out of four metrics on COCO, overall accuracy on VQA 2.0, and all four metrics on Flickr30k. The improvements are particularly sound on Flickr30k, where we get 5.1% absolute gain on CIDEr metric and 2.8% on BLEU@4.
We further perform CIDEr optimization on COCO Captions through Self-Critical Sequence Training (SCST) (?), as in most of the recent image captioning literatures. The results are in Tab. 3 where our full model sets new SotA on all the metrics.
Boost from pre-training. Our full model leads our baseline model by a large margin on most of the metrics thanks to our pre-training. Some noticeable improvements include over 10% absolute gain on CIDEr metric on Flickr30k, and over 2% gain on CIDEr on COCO and B@4, METEOR on Flickr30k. Small datasets (i.e., Flickr30k) benefit the most as vision-language pre-training alleviates overfitting issues. Our model variants under the two extreme settings work well as expected on their “favorable” tasks, i.e., seq2seq pre-training alone improves downstream captioning tasks significantly and bidirectional pre-training benefits understanding tasks (i.e., VQA), but not the opposite. They set new SotAs on all metrics except the “Number” accuracy on VQA 2.0. The joint training organically combines the representations learned from the two rather different objectives and yields slightly compromised but decent accuracy on all the downstream tasks. That said, from an engineering perspective, if we can afford having separate pre-training models for generation task or understanding task, we will get the optimal model performance. If we value model architecture and parameter sharing, the joint model is a good trade-off.
Impact of pre-training types. Depending on how the base model Transformer is initialized, we define four “degrees” of pre-training from weakest to strongest as i) without any pre-training, i.e., base model is trained from scratch, ii) bidirectional language pre-training, i.e., base model is initialized from BERT weights (?), iii) seq2seq and bidirectional language pre-training, i.e., base model is initialized from UniLM weights (?) which is our baseline setting, and iv) our full Vision-Language Pre-training. The corresponding fine-tuning results on downstream tasks are presented in Fig. 1 on the val set (full results see Appendix) and Tab. 4 on the test set. As shown from the figure, our vision-language pre-training significantly accelerates the learning process of downstream tasks and contributes to better overall accuracy. It is worth noting that the learning process of VQA is greatly shortened despite that the hidden states associated with tokens [CLS] and [SEP] are not learned during the pre-training. This indicates that the contextualized vision-language representations can generalize to unseen domains and work reasonable well as a warm-start for new tasks.
We also study how the pre-training types 1-3 influence our vision-language pre-training in terms of caption generation. The results on Conceptual Captions val set at epoch 20 are shown in Tab. 5. All the models are trained based on the unified VLP objective () for a fair comparison. We observe that initializing base model with weights transferred from pure language pre-training benefits vision-language pre-training. The training objectives of UniLM are closer to our seq2seq and bidirectional objectives than the ones in BERT and hence we hypothesize that this counts for the slightly larger improvement. Note that our intention here is to demonstrate how different weight initializations can influence pre-training performance rather than pursuing possibly high quantitative scores (with full seq2seq training, CIDEr could climb to 77.2 after training for 30 epochs).
Region object labels as pretext. Existing works (?; ?) regard region object labels (probabilities) () as an important auxiliary to enrich image region features and here we follow a similar design. We can also instead use these labels for a masked region classification pretext as in (?). Here we have a comparison over the two design choices. “region label probability as input” is equivalent to our full model Unified VLP and “region label as pretext” is the implementation from (?). As shown in the results, predicting class labels as a pretext has a negative impact on the pre-training, in terms of captioning performance. We hypothesize that this is because the class labels from the off-the-shelf object detector might be noisy which compromises the learned feature representation. In contrast, our model refines the visual representation through a more reliable masked language modeling and could correct the errors exist in the class labels.
Qualitative results and analyses. Qualitative examples on COCO Captions and VQA 2.0 are shown in Fig. 3. In the first two examples, our full model with vision-language pre-training captures more details in the image, such as “umbrellas” and “a blue wall” than the baseline methods. It also answers questions correctly. In the third example, all the methods dis-identify the gondola as a train due to their visual similarity. When it comes to the question answering, our methods all give correct answers while the GT answer is incorrect (note that there is a person in the gondola). In the fourth example, all the models mistakenly classify the activity as surfing while the correct one is kayaking/boating. This is consistent across both the caption model and the VQA model, which implies that the feature representations are indeed shared across tasks.
This paper presents a unified Vision-Language Pre-training (VLP) model that can be fine-tuned for both vision-language generation and understanding tasks. The model is pre-trained on large amounts of image-text pairs based on two objectives: bidirectional and seq2seq vision-language prediction. The two disparate objectives are fulfilled under the same architecture with parameter sharing, avoiding the necessity of having separate pre-trained models for different types of downstream tasks (i.e., generation-based or understanding-based). In our comprehensive experiments on image captioning and VQA tasks, we demonstrate that the large-scale unsupervised pre-training can significantly speed up the learning on downstream tasks and improve model accuracy. Besides, compared to having separate pre-trained models, our unified model combines the representations learned from different objectives and yields slightly compromised but decent (SotA) accuracy on all the downstream tasks. In our future work, we would like to apply VLP to more downstream tasks, such as text-image grounding and visual dialogue. Methodology-wise, we would want to see how multi-task fine-tuning can be applied to our framework to alleviate interference between different objectives.
Acknowledgement. The technical work was performed during Luowei’s summer internship at Microsoft Research. Luowei Zhou and Jason Corso were partly supported by DARPA FA8750-17-2-0125 and NSF IIS 1522904 as part of their affiliation with University of Michigan. This article solely reflects the opinions and conclusions of its authors but not the DARPA or NSF. We thank Li Dong and Furu Wei for generously sharing us their UniLM source code. We thank Kezhen Chen for his helpful discussions.
We include the validation results on fine-tuning tasks in Tab. 7. Note that for VQA 2.0, all the methods here are only trained on the training set while for the results reported on the test set (Tab. 3 and Tab. 4 in the main paper), all the models are trained on both training set and validation set following the practice from early works.
Implementation Details
Region proposal and feature. We use a variant of Faster RCNN model (?) with ResNeXt-101 FPN backbone (?) for region proposal and feature extraction. The Faster RCNN model is pre-trained on the Visual Genome dataset (?), following the same procedure in (?) for joint object detection (1600 classes) and attribute classification. We set the number of regions per image to exact 100 as suggested in (?). We take the output of the fc6 layer as the feature representation for each region, and fine-tune the fc7 layer.
Model hyper-parameters. The model hyper-parameters on pre-training and fine-tuning are in Tab. 8. The SCST training on COCO is performed after the VLP pre-training and COCO fine-tuning.
Training details. We use the same training optimizer as in BERT (?) and other training hyper-parameters are in Tab. 8. Our VQA models are trained on 2x V100 GPUs, COCO Captions SCST training on 4x Titan Xp GPUs, and all others are on 8x V100 GPUs.