Scaling Up Vision-Language Pre-training for Image Captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, Lijuan Wang
Introduction
Recent advances in image captioning can be largely attributed to vision-language pre-training (VLP) , the current prevailing training paradigm for vision-language (VL) research. VLP is usually conducted on a combined image-text dataset comprising of several or tens of millions images in total, e.g., Visual Genome , SBU and Conceptual Captions . While previous studies have analyzed various choices of pre-training objectives and model architectures, it remains unclear to what extent the pre-training dataset would impact the performance, and how it correlates with different model settings. Along the journey of pushing the limit of VLP, it becomes increasingly important to answer this question.
Scale is believed to be an important ingredient in attaining excellent performance . Recent work has investigated the Pareto frontier of training transformer models, often referred to as the neural scaling law, in the domains of natural language processing and computer vision , via unsupervised or weakly-supervised learning methods. These studies have observed consistent benefits of increasing the model size to billions of parameters, given billion magnitude of pre-training data available.
More recently, contrastive image-text pre-training has also been scaled up to 400 million and 1.8 billion data sizes for image representation learning and image-text retrieval. Both CLIP and ALIGN employ two individual networks to encode the image and the text separately for alignment, which well fits the image-text retrieval task, but little is known about the scaling properties when it comes to image captioning.
To study the characteristics of this scaling trend on the captioning task, we first construct a large-scale image-text dataset (dubbled as ALT200M), consisting of up to 200 million image-text pairs from web based on the alt attribute of the images. Then, we conduct extensive experiments to scale VLP for image captioning from both the data and model perspectives, and name our model as LEMON , short for a LargE-scale iMage captiONer. To simulate the process of data scaling, we create multiple subsets of ALT200M, ranging from to million. In terms of model, we use the state-of-the-art image captioning model VinVL as our reference model, composed of an image feature extractor and a transformer model. We adapt the pre-training task to be consistent with the captioning task, and then scale the width and depth of the transformer model with the number of parameters ranging from (i.e., tiny) to (i.e., huge) millions. Combining different models and pre-training data sizes, we summarize our results in Figure 1 and 2, which characterize the linear-logarithmic scaling trend. Larger models tend to benefit more when we have more than million data for pre-training. However, with only million data, the performance starts to saturate early as the model size increases. Moreover, we also investigate other design choices of VLP, e.g., model architectures and training objectives.
Our contributions are summarized as follows.
We present the VLP scaling rule for image captioning. Not only does this prove the effectiveness of learning from large-scale noisy data, but it also sheds lights on how performance can be efficiently improved by increasing the model and pre-training data sizes together to avoid a saturation plateau.
We achieve new state-of-the-art results for image captioning across several major benchmarks, including COCO Caption, nocaps, and Conceptual Captions.
Related Work
Since the birth of ViLBERT and LXMERT , we have witnessed a boom of methods for vision-language pre-training . Prominent examples include UNITER , VL-BERT , OSCAR , UNIMO , and VinVL . Along the journey of VLP, researchers have investigated different training strategies , robustness , compression , probing analysis , and the extension to video-text modeling . More recently, instead of using object detectors for image feature extraction, end-to-end VLP based on convolution networks and transformers are becoming popular .
However, as another important factor in achieving superior performance, the scaling behavior of VLP is less studied. While most works pre-train transformer of base/large sizes on no more than M images, we train models from tiny to huge, on up to M images. CLIP and ALIGN scaled up contrastive pre-training to 400M and 1.8B images, and SimVLM further use 1.8B images for prefix language modeling pre-training. However, CLIP and ALIGN focus on image-text retrieval, while SimVLM did not study its scaling behavior w.r.t. pre-training data sizes. Compared with them, we focus on image captioning, provide a more comprehensive study on the scaling behavior via altering data and model sizes, and show that by using 200M images, we can outperform SimVLM on image captioning.
Scaling Law.
With the success of large-scale pre-trained models in both the language and vision domains, there has been a surging research interest in discovering the empirical scaling law of these models. presented that the language model performance scales as power-law across many orders of magnitude with dataset size, model size, and computation used in training. further studied the scaling of autoregressive generative modeling. Aside from the model size, showed that the model shape also matters for efficient transfer from upstream pre-training to downstream finetuning. In the vision domain, scaled a series of vision transformer models evaluated on image classification tasks. While the scaling protocols have been investigated for many NLP and vision tasks, we are the first to study the scaling behavior of VLP for image captioning, and push multimodal transformer pre-training to a much larger scale.
In Appendix, we also provide a detailed related work review on non-pretraining-based image captioning methods.
Method
In this section, we present the pre-training dataset in Section 3.1, the model structure in Section 3.2, and training objective in Section 3.3.
We construct a data collection pipeline to crawl the images from the Internet and the associated alt attribute, which usually provides the description of the image content. In order to scale up easily, we follow the natural distribution of images without re-balancing, and apply only minimal rule-based filtering. We keep images with the longer side more than pixels and aspect ratio smaller than . As some alt-texts are too long, we split them up by punctuation marks, such as period and exclamation mark, and select the longest part. To filter out some rare or misspelled words, we build a vocabulary of unigrams with English Wikipedia titles and body text. We remove unigrams that are present less than times, resulting in approximately million unique unigrams. We remove the alt-text if any of its unigrams cannot be found in the vocabulary. Afterwards, we count the frequency of all the remaining sentences, and filter out some boilerplate sentences that are too generic, e.g., stock image, 3D illustration, vector photo. For the sake of privacy, we use a Named Entity Recognition model spaCyhttps://github.com/explosion/spaCy to identify person and location names, and replace them with special tokens ⟨PERSON⟩, ⟨LOC⟩, respectively. At last, we perform duplication check on all the collected images to ensure that they do not overlap with existing test sets, such as COCO, nocaps, and Conceptual Captions.
The final dataset, named as ALT200M, contains more than million images, each corresponding to one alt-text. The word cloud of most frequent words is visualized in Figure 3. As shown in Table 1, compared to CC12M, ALT200M has nearly more images. The vocabulary is almost doubled. We observe that of unigrams sum up to only of total occurrences, characterizing an extremely long tail of rarely occurring unigrams. The average length of the captions is , more than that of the COCO caption dataset (). We also observe that our dataset contains much more shorter captions with only or unigrams. This indicates a shift in the distribution of captions from pre-training to finetuning.
Besides CC12M, there also exist some other large-scale image-text datasets, such as WIT , WenLan , LAION-400M , and the datasets used in CLIP and ALIGN . More detailed discussions on them are provided in Appendix.
2 VLP Model for Captioning
We use the pre-trained Faster R-CNN detector from to extract image region features, which are concatenated with scaled bounding boxes as position encoding. Following , we also add the detected object tags as input. The text input, including the caption and objects tags, are tokenized by WordPiece, with a vocabulary of tokens. A multi-layer transformer model is used for multimodal fusion, which consists of a stack of encoder layers, each of which has a multi-head self-attention (MSA) layer followed by a feed-forward layer. To enable text generation with the encoder layers, we use the sequence-to-sequence attention mask in each self-attention layer for the captioning module. Specifically, the input consists of image embeddings , object tag embeddings , and token embeddings for the caption , where are the number of image regions, tags, and caption tokens, respectively. The corresponding outputs are:
where is the MSA layer with mapped to query, and mapped to key/value. means concatenation of matrices, and the index of denotes the position corresponding to . The output representation is fed into the next layer, or used for prediction at the end. In this way, during inference, the model can decode the token from left to right in an auto-regressive manner. To study the scaling trend, we experiment with model configurations, ranging from “tiny” of M parameters to “huge” of M parameters, detailed in Table 2.
3 Training Objective
Experiments
In this section, we first present our experimental setup in Section 4.1, and then detail our results in Section 4.2, followed by comprehensive analysis in Section 4.3.
To measure the progress brought about by large-scale pre-training, we aim to evaluate the model’s capability of describing varieties of (long-tail) visual concepts, which is essential for captioning in the wild. For this purpose, we choose nocaps as the evaluation benchmark, which is developed to evaluate object captioning at scale. The dataset consists of images from Open Images, and covers more than object categories, of which nearly of them are unseen from the training set in COCO . Based on whether the image contains novel objects unseen in the COCO training set, the nocaps images are divided into three domains: “in”, “near”, and “out”. None of the objects in the out-domain are seen in COCO. This discrepancy raises the importance of learning from external resources for recognizing novel objects, rather than relying on the clean and fully annotated captioning training data. As the external training resources may vary for different methods, in Table 3, we only compare our model with other models that also use extra image-caption pairs, and take the pre-training dataset size into account.
Implementation details.
To study the scaling trend, we experiment with model configurations and pre-training data sizes. We train all the models from scratch if not otherwise specified. In the pre-training, we do not include COCO or Visual Genome data, to exclude the possible impact of data quality when plotting the scaling trend, as these datasets are manually annotated instead of web collected. To create pre-training dataset of different sizes, we randomly sample from ALT200M at different data scales. Note that the larger dataset is a superset of the smaller ones.
We use AdamW optimizer with linearly decaying learning rate. During pre-training, the batch size is . The initial learning rate is set to for the base and large model, and to for the huge model. The models are trained for epochs. The maximum length of image regions, tags and caption tokens are , , , respectively. During finetuning, the model is trained for epochs with batch size . The initial learning rate is , , and for the base, large, and huge models, respectively. During inference, the caption is generated with beam search and the beam size is . The generation ends when the token is predicted, or the maximum length of tokens is reached. More training details are provided in Appendix.
2 Captioning Results
Results on nocaps validation and test sets are shown in Table 3. By leveraging large-scale pre-training on the automatically collected alt-texts, LEMON has achieved remarkable improvement, especially for out-of-domain images. Compared to the baseline trained on COCO only (row 8), after pre-training on ALT200M (row 12), the CIDEr score is improved by for the in-domain part, and for the out-of-domain part. This evidences that large-scale pre-training improves the model’s ability to recognize a wide range of long-tailed visual objects. We also present results of models pre-trained on CC3M and CC12M. Compared to the best reported results on these datasets (row 1, 2), our evaluated CIDEr scores (row 9, 10) are increased by and , respectively. This demonstrates the performance improvement in our captioning results brought about by the proposed training scheme when the pre-training dataset is the same. On the leaderboardhttps://eval.ai/web/challenges/challenge-page/355/leaderboard/1011 test set, our large and huge models (row 19, 20) both surpassed the top-ranking model (row 18) that is pre-trained on B image-text pairs, creating the new state-of-the-art of in CIDEr. We also achieve the state of the art on other image captioning benchmarks, including COCO Caption and Conceptual Captions, as summarized in Table 4 and 5.
Large-scale pre-training not only benefits VL representation learning, but also equips the model with the capability to zero-shot generalization. We use the pre-trained model to generate captions directly without further finetuning. The prefix “a picture of” is added as prompt to improve the quality of generated captions. Some examples are illustrated in Figure 5. The pre-trained model demonstrates strong ability in recognizing various long-tail visual concepts. Compared to the model trained only on small clean set, it shows the knowledge of many fine-grained categories (e.g., “metal instrument” vs. “tuba”), which are learned from the large-scale noisy supervision of alt-texts from web. We also notice that our pre-trained model tends to generate very short descriptions when used in a zero-shot manner, but this is mitigated after finetuning on COCO. We posit that the reason for this is the relatively large proportion of short alt-texts in our pre-training datasets.
3 Ablation and Analysis
We conduct comprehensive experiments to understand how much gain can be obtained in the downstream tasks by scaling up pre-training. Figure 2 shows the relationship between the number of images used in pre-training and the CIDEr scores evaluated in the downstream captioning tasks. All the models are pre-trained from scratch, and then finetuned on COCO. While all the models can be improved after pre-training with more data, the improvement is obviously less significant for the smaller models than for the larger models. On COCO, the gap between “small” and “large” models is negligible at M scale, but it becomes large as the data size increases. Moreover, when evaluating on nocaps, the gap in out-of-domain set is consistently larger than that in in-domain. This implies the advantage of large models in transferring knowledge from pre-training to downstream tasks, especially when the finetuning data are too limited to cover all test scenarios.
Besides, we observe that the model capacity becomes the performance bottleneck as the amount of available data increases. Figure 1 plots the scaling trend w.r.t. the number of model parameters. When pre-training with M data, the “base” size appears to be sufficient, and there is no significant benefit to using larger models. However, with more than M data, the larger models start to outperform the smaller ones by a significant margin. When the data magnitude reaches hundreds of millions, and if the observed trend from “base” to “huge” can be kept, there is promise in training an even larger model to push the limits of VLP for captioning tasks.
At last, to have a better understanding of the data quality, we perform pre-training with the same settings on CC12M and the M subset of ALT200M. With the only difference in pre-training data source, the models yield fairly similar results ( to differences in CIDEr) on COCO and nocaps. This indicates that our data quality is comparable to that of CC12M. The observed performance improvement should be attributed to the pre-training scale.
Sample efficiency.
We examine the improvement of learned representations along with the progress of pre-training. Progress is measured quantitatively by the number of image-text paired samples seen in pre-training, i.e., the effective batch size multiplied by the training steps. In Figure 6, we report the results on COCO Caption after finetuning intermediate pre-trained checkpoints. We also evaluate the finetuned models on nocaps, indicating the ability of generalization under domain shift. We present two models in the figure, one with “base” size, the other with “huge” size. Both models are pre-trained on ALT200M.
We observe that both models continue to improve after seeing more samples in pre-training, but the larger model learns much “faster”. To achieve similar results in the downstream COCO captioning task, the base model must see more than to times more samples in pre-training. This factor is even greater when evaluating on the nocaps out-of-domain images. The result of the “base” model seeing 19 billion samples is still slightly worse than that of the “huge” model seeing 0.8 billion samples. This demonstrates the efficiency of large models in learning from large-scale data, as well as the robustness in generalization.
Further ablation.
We compare with other common model structures and training objectives, such as the encoder-decoder transformer model and unidirectional language modeling (LM). Experiments are conducted with models of “base” size as specified in Table 2. For the encoder-decoder structure, we use encoder layers (with self-attention) followed by decoder layers (with cross-attention after self-attention), while other model configurations remain unchanged. The training objectives are illustrated in Figure 4. For each experiment setting, we sweep the hyperparameters, e.g., pre-training epochs from to , finetuning epochs from to , and learning rates from to . The results of the best hyperparameters are reported.
We train the models under different settings on COCO and CC3M, respectively. Results are summarized in Table 6. On COCO, the differences among the settings are small ( relative change in CIDEr), with the worst being from encoder+LM, and the best being from encoder-decoder+LM. In contrast, on CC3M, the difference is much larger ( relative change in CIDEr). The worst is from encoder-decoder+LM, while the best is from encoder+MLM. As CC3M is collected over the Internet and contains much more noise, we assume that the model that tends to overfit is prone to error when data quality is low, even though it performs well given well-annotated data.
Moreover, to compare training objectives, we first pre-train the models on CC12M, using s2s-MLM or LM, then finetune the intermediate checkpoints on COCO. As shown in Figure 7, we observe that although the model trained with LM converges faster at the beginning, it enters saturation early, and does not achieve scores as high as the model using s2s-MLM. We also find that training with LM is very sensitive to learning rates. Given the above results, we choose the s2s-MLM model and the encoder structure to scale up with the noisy pre-training data.
Conclusions
In this paper, we study the scaling behavior of VLP models for image captioning, and construct our own large-scale dataset ALT200M. Our experiments show that scaling up pre-training leads to remarkable improvement for the downstream captioning tasks. Our model LEMON has achieved new state-of-the-arts on multiple benchmarks, including COCO Caption, nocaps, and Conceptual Captions. LEMON also has impressive capability of recognizing a wide range of long-tail visual objects, even in the zero-shot manner. Moreover, our study on large transformer models indicates that with orders of magnitude larger training data available, the model capacity tends to be the bottleneck. It is a promising direction to train a substantially larger model to take more advantage from the large amounts of alt-text data widely circulated on the Internet.
References
Appendix A Related Work on Image Captioning
There is a rich literature on image captioning studying different model structures and learning approaches. Recent works have proposed enhanced attention-based models to improve the performance, such as ORT , AoANet , M2 Transformer , X-LAN , and RSTNet . Besides, researchers have explored to leverage semantic attributes , scene graphs , and graph convolutional networks for captioning. While those methods focus on learning from well-annotated captions in relatively small datasets, such as COCO Caption, we investigate the impact of data scales with a much larger and noisier dataset. This allows the model to learn more diverse visual concepts, and makes one step further towards in-the-wild captioning.
Appendix B Comparison of Image-Text Datasets
Besides CC3M and CC12M , we also compare with other web-crawled image-text datasets as described in the following.
The dataset used in CLIP has million image-text pairs. Unlike our dataset crawled from web without re-balancing, their dataset is built upon a set of queries. The queries include all the words occurring at least times in the English version of Wikipedia and are augmented with bi-grams. The image-text pairs are searched such that the text includes one of the queries. The final results are also balanced to include up to image-text pairs per query.
The dataset used in ALIGN has billion image-text pairs. A later work SimVLM also uses this dataset. The data collection pipeline is similar to that used in Conceptual Captions , but relaxes most cleaning steps. Only some rule-based filters are applied, such as image size, alt-text frequencies, and rare words.
The Wikipedia-based Image Text Dataset (WIT) is composed of million unique images and million texts. Different from all other listed datasets, it features multilingual texts across languages. The images and texts are collected from the Wikipedia content pages. It provides texts from multiple sources, such as reference, attribution and alt-texts, and texts in different languages for the same image.
WenLan has million image-text pairs. The web-collected pairs have gone through an elaborate cleaning process. For each data source, they also use topic models to extract topic words, and analyze the topic distribution to help with selecting desired contents.
LAION-400M has million image-text pairs, and is recently released to public. Instead of applying human designed heuristics in data cleaning, this dataset relies on the CLIP model to filter image-text pairs, where the cosine similarity scores between image and text embeddings are calculated and filtered by . As CLIP itself is also trained on noisy image-text pairs, it is yet unknown about the quality of this dataset compared to others cleaned by heuristic-based pipelines.
We summarize the characteristics of existing image-text datasets from three perspectives:
Accessibility: The datasets CC3M , CC12M , WIT and LAION-400M are released to public with the image URL and associated meta files. Other datasets are proprietary.
Scale: Except LAION-400M , the other released datasets have tens of millions of images, which are not enough for the scaling purpose.
Collection pipeline: LAION-400M is mainly filtered by the CLIP model, while all the other datasets are filtered by a series of rules including expert designed heuristics and/or complicated models. As to the data sources, WIT is collected from Wikipedia. The other datasets are crawled from Internet.
In a nutshell, we construct ALT200M to study the scaling behavior, as no such large-scale datasets are available (LAION-400M is concurrent with ours). Following prior works, we crawl images and the alt attributes from Internet, and apply minimal rule-based filters to retain as many images as possible. Our experiments show that the web-collected data can help to substantially improve the performance of captioning models.
Appendix C Detailed Hyperparameters
We include a table of important hyperparameters in Table 7 for all model sizes as defined in Table 2. All the models are pre-trained for epochs on subsets of ALT200M, and finetuned for epochs on COCO with batch size . During pre-training, the learning rate is warmed up for the first steps to the peak value, then linearly decaying to . During finetuning, the learning rate linearly decays from the initial value to without warm up. We use AdamW optimizer with weight decay . The cross-entropy loss is calculated with label smoothing .
The evaluation results on COCO “Karpathy” test split and nocaps validation set are plotted in Figure 1, 2, and 8.
Appendix D More Qualitative Examples
In this section, we provide more qualitative examples of our web-crawled dataset ALT200M and the captioning results of our LEMON model.
Figure 9 shows some examples of the image-text pairs in ALT200M. While some of the alt attributes are descriptive sentences that can serve as good training targets, e.g., Figure 9 (7), (8), (9), it is noted that some texts are not semantically well formed, e.g., Figure 9 (10). Some texts are very short phrases containing only 2 to 4 words, e.g., Figure 9 (1) - (6). We also observe that some texts do not precisely describe what content is shown in the image, but mention external knowledge or information. For example, Figure 9 (12) shows a woman pointing at the camera, but the text is “I got you”. And the text of Figure 9 (11) are likely to be extracted from news. These might put challenges for the model to learn from noisy text supervision. However, there are indeed a large variety of (fine-grained) visual objects present in the images and texts, such as burdock leaf, karate, mahjong, and great blue heron. Compared to human-annotated datasets, these web-collected data provide much richer training resources, especially for the long-tailed concepts.
After training on the ALT200M dataset, our LEMON model has achieved impressive results, even in zero-shot manner. We present more examples of generated captions in Figure 10, in addition to Figure 5 in the main paper. Compared to the baseline model trained on COCO only, the LEMON model can recognize much more fine-grained objects, as highlighted in red.