PaLI-X: On Scaling up a Multilingual Vision and Language Model
Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shakeri, Mostafa Dehghani, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, AJ Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Peter Steiner, Yang Li, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut
Introduction
The success of scaling language models makes it appealing to similarly scale Vision-Language (V&L) models, and investigate the improvements, capabilities, and emergent properties of such models. Inspired by the work in , we present PaLI-X, a multilingual vision and language model with reusable scaled-up components, consisting of a pretrained large-capacity visual encoder (using as the starting point) and a pretrained language-only encoder-decoder (using as the starting point), further trained at-scale on a vision-and-language data mixture using a combination of self-supervision and full-supervision signals.
One clear pattern that emerges from the combination of results from PaLI and the work we present in this paper is that scaling both V&L components together brings increases in performance across a wide range of tasks. We show this by comparing against the same benchmarks used for PaLI (Fig. 1, Left), and also against new benchmarks for which the new capabilities of PaLI-X are evaluated (e.g., ChartQA, AI2D, DocVQA, InfographicVQA, as well as video understanding tasks). We observe that scaling leads to large improvements over the results of the PaLI model, and also over specialized large-scale models that are trained specifically to solve certain tasks, often with the help of (often much larger) text-only LLMs . In particular, we find that increasing both the effective capacity of the vision component (which does more unilaterally) and of the language component (which also does unilaterally) is beneficial; the new PaLI-X model provides more balanced parameter allocation than any other prior work (roughly 40%-60% split of the total capacity).
Aside from confirming the impact of scale, the original contribution of PaLI-X consists in leveraging the mixture-of-objectives proposed in for vision-and-language modeling, and showing that it results in a model that improves both state-of-the-art results and the Pareto frontier for fine-tuning and few-shot configurations (Fig. 1, Right).
We also observe emergent properties based on PaLI-X’s results compared to previous models with similar architecture but smaller sizes. For instance, we report drastically improved performance on the counting ability (See Table 1 and Appendix B), both for the plain variety (count all instances of a class) and the complex variety (count instances based on a natural language description), that are not attributable to training designPlain counting is usually achievable via good object detection, while complex counting requires a fine-grained understanding of the alignment between language-based specifications and visually-based occurrences.. Additionally, we present qualitative insights into the model’s performance (Appendix A), with an emphasis on multilingual transfer learning such as the ability to detect objects using non-English labels (Fig. 2), and the ability to switch between the language of text present in the image (e.g., English) and the language of the generated image caption (e.g., Romanian).
Our technical contributions include the following:
We scale a Vision-Language model to achieve outstanding performance on a wide variety of benchmarks. We observe that scaling both the Vision & Language components is advantageous and report that performance remains unsaturated at this scale.
We show that training such a model with a mixture of objectives that combines prefix-completion and masked-token completion improves the Pareto frontier for fine-tuning vs few-shot performance at this scale.
We show that a high-capacity vision encoder (ViT-22B) can be effectively co-trained for image classification and OCR label classificationWe use OCR tokens produced by the GCP Vision API over the training images as targets. to achieve significant improvements on V&L tasks for which the understanding of text-within-image is crucial.
Overall, PaLI-X improves SoTA results via fine-tuning on 15+ benchmarks, and we show that it is the first of its kind to simultaneously adapt via multitask fine-tuning to a diverse set of benchmarks without significant performance degradation.
Related Work
Similar to large language models such as GPT4 and PaLM , the benefit of scale has also been observed in recent vision and vision-language models. Flamingo used a frozen language component and demonstrated the benefit of scaling up this part up to 70B parameters on the few-shot multimodal capabilities, while the vision encoder is fixed with 435M parameters. GIT , on the other hand, explored scaling of the vision component up to 4.8B parameter, with a 300M parameter language decoder. PaLI explored jointly scaling the vision and language component, to 4B and 17B, respectively, and showed that scaling both components benefits a wide range of vision-language tasks. All these models took advantage of vision and language unimodal pretrained models as backbones to start multimodal training. Recently, on the vision model side, a vision transformer with 22B parameter has been introduced . In this work, we make use of a ViT-22B model specifically tuned for OCR capability to explore scaling Vision-Language models to even larger parameter regime.
As first shown in , large language models are sometimes able to solve new unseen tasks at inference as long as a few examples –or shots– are provided as inputs. This is usually referred to as in-context learning . Follow-up work proposed improved ways to split and prompt the shots, such as Chain of Thought or Least-to-Most prompting . So far, the vast majority of this work has been done in the context of language inputs . In this work, we explore multimodal in-context learning with pairs of images and captions. Our work is aligned in spirit to Flamingo that uses interleaved image text pairs in the same web page and in-context tuning during pre-training. We first group the image-text pairs by url and split each group to a “shots” set and a “target” set. Then we use the few examples in the “shots” set as input features to predict the examples in the target set.
Besides solving vision-language tasks in multiple domains, recent VLMs also attempted solving these tasks at once instead of fine-tuning on each individual benchmark. Unified-IO performed multitask fine-tuning and reported solid results across 16 benchmarks. Spotlight reported that inside the UI domain, multitask fine-tuning can achieve a performance close to task-specific fine-tuning. In this work, we show that PaLI-X can be simultaneously fine-tuned with a diverse set of benchmarks in multiple domains without performance degradation.
Model
The PaLI-X model architecture follows the encoder-decoder architecture: image(s) are processed by a ViT encoder, with the resulting visual embeddings fed to an encoder-decoder backbone, along with embeddings from additional text input (e.g., question / prefix / prompt). More details are provided in Appendix A.
Visual component Our visual backbone is scaled to 22B parameters, as introduced by , the largest dense ViT model to date. To equip the model with a variety of complex vision-language tasks, we specifically focus on its OCR capabilities. To that end, we incorporate an OCR-based pretraining as follows: images from the WebLI dataset are annotated with OCR-text detected by GCP Vision API; the encoder is then further pre-trained with a mixture of the original JFT-based classification task and a new OCR-based classification task (whether or not a given token occurred in the image according to OCR results). See Appendix A for additional details on the visual component. PaLI-X is designed to take images as inputs (for few-shot and video understanding), with tasks involving a single image as the case. For , each image is independently processed by the ViT module, and the patch-level embeddings coming out of ViT are flattened and concatenated to form the visual input (See Appendix A). Note that similar to the single-image case, there is no pooling over the spatial dimension before visual embeddings are aggregated over the temporal dimension. That is, for an -frame input with -patches per frame, the resulting visual input has tokens.
Few-shot formulation In the few-shot setting, for a given target example the model receives a number of “labeled” examples (in the form of additional image, text pairs) that we refer to as shots/exemplars. The hypothesis is that information contained in these exemplars provides the model with useful context to generate predictions for the target example. Formally, the input with shots is a sequence , where and are texts and images for the shots, and and are the text (prompt) and image for the target example. PaLI-X processes this input as follows: all images, including the target one, are first independently processed by the visual encoder, and the resulting patch-level embeddings are flattened and concatenated to form the visual input sequence. After going through a projection layer, they are concatenated with the text embeddings to form the multimodal input sequence used by the encoder. We implement additional optimizations including distributing the exemplars between the encoder and the decoder, and an attention re-weighting mechanism (see Appendix B).
2 Pretraining Data and Mixture
The main pretraining data for our model is based on WebLI , consisting of roughly one billion images with alt-texts from the web and OCR annotations (using the GCP Vision API), covering over 100 languages. In addition to WebLI image, text pairs, we introduce here Episodic WebLI data, where each episode corresponds to a set of such pairs. We aim to have each episode contain loosely related images (i.e., they are clustered according to their URL field), so as to encourage attention among examples in an “episode”. We find this new dataset (with 75M episodes and around 400M images in total) important for developing the few-shot capabilities of the model.
3 Training Stages
Our model is trained in two stages. In stage 1, the visual encoder (after mixed-objective training) is kept frozen, while the rest of the parameters are trained on a total of 2.2B examples at the base resolution 224224 (native to ViT-22B), using the entire mixture. In stage 2, it continues training using only the OCR-related objectives (pix2struct and split-ocr) plus the object detection objective; this is done in several substages, during which image resolution is gradually increased to 448448, 672672 and finally 756756.
Experiments
Our results demonstrate that the larger capacity in PaLI-X scales well in both its vision and language components, and it is particularly beneficial for more challenging scene-text and document understanding tasks. Our model outperforms the SOTA on diverse vision-language tasks, with significant margins in some cases.
Experimental setup We fine-tune PaLI-X with frozen ViT-22B; the learning rate follows a linear decay from initial value 1e-4 for all fine-tuning experiments. See Appendix B for more details.
First, we present benchmarks results for the condition where external OCR systems are not used (Table 1, see Appendix B for an extended table.). The trend is that PaLI-X matches or improves SoTA results on these benchmarks, with a particularly significant improvement on the TallyQA benchmark over MoVie (specialized counting model), at +11.1 for simple counting questions (e.g., “how many giraffes”) and +18.8 for complex counting questions (e.g., “how many giraffes are drinking water”); there are significant improvements over PaLI as well, indicating that scale plays an important role in the ability of such models to perform counting tasks. We additionally note the state-of-the-art result on VQAv2 at 86.1 accuracy, achieved with an open-vocabulary generative approach, and the performance on OKVQA at 66.1 accuracy, matching the much-larger PaLM-E model performance.
Next, we examine text-heavy V&L benchmarks, for which upstream OCR systems can be used to improve performance. As shown in Table 2, PaLI-X improves SoTA for all Captioning and VQA benchmarks across the board, either without or with additional OCR input (using GCP Vision API). For instance, a significant jump of +42.9 points is observed on AI2DAs with all the other benchmarks, our training examples are carefully deduped to exclude images occurring in these benchmarks, including AI2D. Such results, therefore, are not attributable to train-test data leakage., a multiple-choice benchmark where choices are provided along with each question. Being able to have the text choices as input benefits PaLI-X compared with the previous SoTA Pix2Struct which has to render the text on the image, but this does not explain all the improvements. In a question-only configuration (no answer choice present), PaLI-X achieves 46.3 on AI2D, more than 4 points higher than Pix2Struct’s result.
In general, having access to OCR texts extracted by an external OCR pipeline boosts performance. Still, for several benchmarks (e.g., AI2D, ChartQA, OCRVQA and Widget-Cap), PaLI-X’s end-to-end performance when using its intrinsic OCR capability is close to that leveraging additional OCR input. A common feature for these benchmarks is that they have well-oriented text – diagrams, charts, book covers or user interfaces, with reasonably large font size at 756756 resolution. For tasks involving scene text in natural images (TextCaps, TextVQA, STVQA) or very high density of small texts (DocVQA, InfoVQA), results still highlight clear benefits when utilizing an external OCR model.
1.2 Multitask Fine-tuning
We simultaneously fine-tune and evaluate the pretrained checkpoints on multiple benchmarks belonging to the same category. We deduplicated every training set over the test sets of every task in the mixture to prevent the leakage of any test-set examples into the mixed training set. This is useful as it leads to a single fine-tuned model that performs all the tasks, rather than having to fine-tune each task separately. We performed such multitask fine-tuning on all Image Captioning benchmarks and most VQA benchmarks, respectively.
Table 3 shows the multitask fine-tuning result for captioning tasks. The performance on COCO is slightly decreased in the multitask setting, which is likely a result of this task needing longer training to converge. For Screen2Words, having the smallest train and dev/test sets could be responsible for the performance fluctuation. Notably, VizWiz-Cap and Widget-Cap shows improved performance from multitask fine-tuning. Overall, the average performance decreases by 1.4 points (0.2 excluding Screen2Words) with multitask fine-tuning, while offering the clear advantage of having a single checkpoint to perform all these tasks. Appendix B shows similar results for VQA tasks. We consider this outcome a positive result that establishes the on-par performance between multitask fine-tuning and single-task fine-tuning for diverse benchmarks, in contrast with previous work which argued a gap between single-task and multitask fine-tuning , or demonstrated little gap over benchmarks from the same domain .
1.3 Few-shot Evaluation
We fine-tuned the PaLI-X model on a mixture of few-shot tasks. The few-shot mixture contains Episodic mixtures, (Non-Episodic) Webli and (Non-Episodic) CC3M data. Note that all of these datasets were already used in previous stages of training, but with lower mixture proportions. During pre-training, we only use up to 4 shots, with both encoder and decoder shots (see Appendix B). For fine-tuning, we use up to 8 encoder shots and do not use decoder shots.
2 Video Captioning and Question Answering
We fine-tune and evaluate the PaLI-X model on 4 video captioning (MSR-VTT , VATEX , ActivityNet Captions , Spoken Moments in Time ) and 3 video question answering benchmarks (NExT-QA , MSR-VTT-QA , ActivityNet-QA ). A brief description of each benchmark and clarifications on their usage are provided in Appendix C.
Experimental setup We fine-tune our model (with base resolution 224224) for each task separately, use the validation split for early stopping, and report performance on the test split. We use a learning rate of for all tasks, and do not adapt any hyperparameters for specific tasks. Frames are sampled using a fixed temporal stride for each dataset (determined based on the video length distribution in that dataset such that the product of the number of frames and stride is larger than the total number of frames for half of the videos), and we experimented with including up to 8 or 16 frames per video. We did not include pooling over the spatial dimension; embeddings for 1616 patches per frame are provided as visual input to the multimodal encoder.
Results We report CIDEr score for the video captioning tasks. Video QA tasks are treated as open-ended generation tasks; we report full-string accuracy (for MSR-VTT-QA and ActivityNet-QA) and WUPS metrics (NExT-QA) in . As shown in Table 5, the 16-frames version has an edge over the 8-frame version, sometimes with a significant margin (e.g., close to a 6 point increase in CIDEr score for ActivityNet-Captions). More importantly, while PaLI-X pretraining was dominated by image-text tasks, we were able to achieve new SOTA performance for 5 out of 7 tasksAs noted in Table 5, current SOTA on NExT-QA for the open-ended QA task was achieved by Flamingo 32-shot, which had outperformed prior fine-tuning SOTA. To the best of our knowledge, PaLI-X performance on this task does outperform existing published fine-tuning performances, with the caveat that we do not have information on what Flamingo fine-tuning would have achieved on this task., and performed very close to prior SOTA on MSR-VTT-QA (47.1 vs 47.4).
3 Image classification
To test image classification capabilities we fine-tuned PaLI-X and models from on ImageNet and evaluated the resulting model on ImageNet-REAL and out-of-distribution datasets: ImageNet-R , ImageNet-A , ImageNet-Sketch , ImageNet-v2 . We used the model from the first training stage (at resolution 224) and the one from the last training stage (at resolution 756). We used the same training hyperparameters for all of runs (selected without any hyperparameter tuning; mode details in Appendix D).
The results can be seen in Table 6. We compare the results to generative model with open vocab – GIT2 (using 384 image resolution), which is the current SOTA for full fine-tuning on ImageNet. PaLI-X achieves SOTA results for generative models on Imagenet, and other datasets. We also performed zero-shot evaluation for PaLI-X and the results can be found in Appendix D.
4 Object Detection
Object detection can be easily formulated in our model as shown in pix2seq , The dataset mix used for pre-training is presented in Sec. 3; detection data was included up to and including the stage using resolution 672, after which a separate detection-specific model was fine-tuned on detection data. Before detection-specific tuning, LVIS & COCO labels were removed from all detection training datasets, allowing zero-shot evaluation on LVIS.
Bounding box mean AP on LVIS is shown in Table 7, including zero-shot performance; the detection-tuned model reaches an AP of 31 in general, and 31.4 on rare classes, and about 12 for both in zero-shot. Performance on rare classes was on par with performance on common classes, a difficult feat traditionally accomplished by complicated sampling schedules and augmentations. In our set up, it is directly enabled by PaLI-X’s diverse training mix. This could likely be further improved with investment in fine-tuning e.g. using noise-augmentation methods from pix2seq , or a further stage of high-resolution, LVIS only training. Qualitatively, we observe emergence of many interesting phenomena enabled by co-training with non-detection tasks; for example, multilingual detection, OCR bounding boxes and longer descriptions, none of which are included in detection training, are often handled well by PaLI-X. Additional results and information can be found in Appendix E.3.
Model Fairness, Biases, and Other Potential Issues
Large models, if left unchecked, have the potential to inflict harm on society – such as amplifying biases , causing disparities , or encoding narrow cultural perspectives . Hence, evaluating PaLI-X for such potential issues is important. We focus our RAI evaluation on three parts: (1) harmful associations, such as toxicity and profanity, (2) demographic parity in the model’s output, such as encoding societal stereotypes/biases, and (3) performance disparity across subgroups. This breakdown follows earlier works in the literature, such as .
Toxicity / profanity. We estimate the level of toxicity and profanity in the generated captions, including when disaggregated across subgroups. We use the FairFace dataset that comprises of images of people with ground-truth attributes: gender presentation, age and ethnicity. We generate captions and use the Perspective API (threshold ) to measure toxicity and profanity. Table 8 summarizes the results; we observe a low level of toxicity/profanity across all slices. Tables 9 and 10 provide a detailed breakdown of toxicity/profanity results for all subgroups in FairFace dataset. In Tables 11 and 12, we report similar results in the MIAP dataset, disaggregated by perceived gender and age.
Bias / Demographic Parity. We estimate the level of demographic parity (DP) in PaLI-X with respect to gender and occupation. To estimate the level of demographic parity (DP) in the model’s output, we feed an image into PaLI-X with the chosen occupation title as a prefix and record the average log-perplexity score of the captions generated by the model. To ensure that any observed parity would likely reflect unintended biases in the model itself as opposed to the evaluation dataset, we use CelebA that contains celebrity images with gender presentation annotation. Our assumption is that many occupations reflecting societal stereotypes, such as secretaries and plumbers, are quite rare in the CelebA dataset so disparities in output may reflect what is encoded in the model itself. The list of occupations is compiled based on and the US job statistics report in .
Figure 3 (top) summarizes the overall results. First, PaLI-X tends to assign a higher log-perplexity score to women than men across most occupations; i.e. men are predicted to be more likely to hold such occupations. Second, PaLI-X assigns a higher likelihood for a woman to be (‘secretary’ & ‘actor’) and a higher likelihood for a man to be (‘guard’ & ‘plumber’) at the 95% confidence level. Figure 3 (bottom) displays the corresponding correlations between perceived gender presentation and occupations within the WebLI dataset, where we use the Pearson correlation coefficient by treating each label as a binary random variable and noting that for binary random variables, zero correlation implies full independence. All absolute correlation coefficients in the data are with 99% of them being .
Performance Disparity. We present here an evaluation of how well PaLI-X performs across different subgroups using the MIAP dataset. For images containing exactly a single individual, we query PaLI-X with the question: “Is there a person in this image?” and evaluate the accuracy of its response. Note that there are no false positives in this evaluation. Table 13 summarizes the results. We observe that PaLI-X maintains a high accuracy across all subgroups.
Limitations. The analysis carried out in this section is necessarily limited, since fairness is a societal concept that cannot be reduced to statistical metrics. We expect RAI evaluations to evolve over time as new issues are detected and reported in the literature and additional datasets become available. Statistical analysis is only a single step and does not substitute for studying the broad and delayed impact of deployed models.
In addition, we rely in some parts on automated tools for inferring attributes, which are not perfectly accurate and can lead to a broad categorization of people that misidentifies real identities. We do not support the creation or application of classifiers for sensitive attributes, such as gender or ethnicity, based on visual indicators and encourage readers to delve into the comprehensive work outlining their potential risks, such as , for further insight. Also, while we use perceived gender presentation in our analysis that is provided by the data (i.e. in CelebA and FairFace), we acknowledge that people may express their gendered identities in numerous other ways.
In our evaluation, toxicity is predicted based on the generated captions only. However, without knowing the context of the image, this can introduce false positives.
Conclusions
In this work we draw more insights from further scaling vision and language models. We show that the scaling and the improved training recipe results in a model that substantially outperforms previous state-of-the-art models, leads to emergent behaviors and identifies further margins for improvements. In particular, we report that the model achieves significant improvements at document, chart, and infographic understanding, captioning, visual question answering, counting, and performs well on few-shot (in-context) captioning, video captioning and question-answering, and object detection.
Acknowledgements
We would like to thank Sarah Laszlo, Kathy Meier-Hellstern, Caroline Pantofaru, Susanna Ricco, Candice Schumann, Ken Burke, Simon Wang, Rachel Hornung, Yichang Chen, Utsav Prabhu, Abhijit Ogale, Kristina Toutanova, Weicheng Kuo, Jihyung Kil, Xiangning Chen, Liang Chen, Rich Lee, Elizabeth Adkison, James Cockerille, Eric Ni, Erica Moreira, Victor Gomes, Jeremiah Harmsen, Claire Cui, Slav Petrov, Tania Bedrax-Weiss, Joelle Barral, Tom Duerig, Paul Natsev, Fernando Pereira, Jeff Dean, and Zoubin Ghahramani for helpful discussions, feedback, and support.
Appendix A Additional Model Details and Examples
A.2 Tuning ViT-22B for better OCR capabilities
The vision encoder’s ability to understand text is crucial to several downstream tasks and general usability. JFT-based pre-training is insufficient to cover this, and so we tuned ViT-22B on WebLI-OCR data. In order to stay true to the original discriminative classification-based objective used for ViT-22B, we turn OCR into a bag-of-words prediction task. OCR texts are tokenized using the mT5 tokenizer across all languages, and the model is trained to predict whether or not a given token occurs in an image. This is treated as multilabel classification, with an expanded classification head.
In the ablation study shown in Table 22, we confirm that this this extra tuning step indeed has a significant improvement on Scene-Text understanding capabilities, demonstrated by the performance on ST-VQA and TextVQA. Meanwhile, the performance on regular VQA tasks such as those in the VQAv2 benchmark also improves.
A.3 Illustrative PaLI-X Examples
Table 14 shows representative examples of PaLI-X, illustrating improved abilities related to counting (both of the simple and complex variety), in context text-reading capabilities, and spatial awareness.
Appendix B Additional results: Image Captioning and VQA
Table 15 summarizes the Image Captioning and VQA benchmarks. For benchmarks modeled only end-to-end without OCR pipeline input (Table 1 and Table 16), fine-tuning is performed with resolution 672672. For Scene-Text and Document Understanding tasks presented in Table 2, fine-tuning is performed with resolution 756756.
B.2 Extended Tables of Image Benchmarks
An extended table of results on some Image Benchmarks is shown as Table 16.
B.3 Multi-lingual Captioning
The Crossmodal-3600 (XM3600) benchmark contains a geo-diverse set of 3600 images with human-annotated reference captions in 36 languages . Table 17 presents multilingual results for both PaLI (current SoTA on XM-3600) and PaLI-X, both finetuned with 224224 resolution. Overall, PaLI-X improves on the SoTA performance across 5 of the 7 languages we report here (and for 14 of the total 35 languages considered); notably, the performance on English is 4 CIDEr points lower compared to PaLI. The 35-language average CIDEr score is in the same ballpark between PaLI and PaLI-X, with a slight +0.5 advantage for PaLI.
B.4 TallyQA and the emergence of complex counting capability
We present in Table 18 the performance of similar models across a wide range of capacity – from 700M parameters to 55B parameters for PaLI-X. The graphs in Fig. 5 illustrate how simple counting appears to follow a more linear progression as parameter-size increases, while complex counting appears to show emergence somewhere before the datapoint provided by the performance of PaLI 17B. This corresponds to our intution that complex counting is a true multimodal task that requires additional capabilities from a model, in terms of the alignment that is required between the visual information and the prompt specification.
B.5 Details on Few-shot Modeling
Figure 6 illustrates the network flow of a few shot model. The text and prompt part of each shot is embedded and concatenated as text features for the PaLI-X model. Each shot’s images and the target image are independently encoded by the ViT component, and the ViT features are concatenated along the sequence axis as visual features. Conditioned on that sequence, the PaLI-X decoder autoregressively makes the predictions for the target image.
While images for all few-shot examples and target example are given as input to the model, text information can be provided in different ways. During inference time, all text information related to the few-shot examples is given to the encoder; in the case of a Multi-answer VQA task, for example, this includes both the prompts that contain the questions, and the expected answers. Prompt for the target example is also given to the encoder, and the decoder is tasked with generating an answer for the target example. During training, however, we increase the training efficiency by making the model predict answers for both the target example and selected shots (the decoder shots). That is, we partition the shots in two sets: encoder shots () and decoder shots (), such that . We use up to 4 shots in total during pre-training (i.e. ), and sample uniformly at random from 1 to . Text input for encoder shots contain both prompts and answers. The decoder shots, however, act as if they were target examples: their text input to the encoder contains only the prompt, and the decoder needs to predict answers for the decoder shots in addition to the target example.
Increasing the number of shots turned out to be challenging, potentially due to cross-attention to target example input tokens getting diluted by the large number of shots. To address this, we introduce an attention re-weighting mechanism. As shown in Figure 7, we explicitly boost the weights for cross attention between decoder tokens and encoded tokens from the target example (that is, the target image and the target text prompt).
Specifically, if there are shots in total, when decoding each token we multiply the cross attention weights by for the target image and text tokens from the encoder outputs. We observe this attention re-weighting technique is especially helpful when we provide the model with many shots (e.g. 32 shots). introduces a technique along similar lines to manipulate attention weights when gathering them from different threads of encoded shots at inference time.
B.5.2 Additional Few-shot Results
Figure 8 shows 3 examples on few-shot captioning and VQA tasks for qualitative analysis. The first row shows captions for the images using the images’ original language, demonstrating the cross multilingual transfer of the few-shot capability. The second row captions the images with a country’s popular food, showing that the few-shot approach can access the model’s world knowledge. The last row shows a VQA with an explanation-like scenario where we ask if the technologies in the images are “new”. Generally speaking, the shown personal computer was produced more than 40 years ago and could be regarded as old technology considering the fast pace of the current high-tech development. However, the 3 input shots provide the detailed calibration for the concept of “new” and the few-shot model successfully take the context and output “new” with plausible explanation to the very old PC.
B.5.3 Few-shot ablation results
In this section, we present and discuss some ablation results for few-shot we explored in order to inform our final design choices on PaLI-X. Unless otherwise specified, we use a 700M-parameter model with the same encoder-decoder architecture, consisting of a ViT-B/16 vision encoder and a mT5-base encoder-decoder language model.
To mitigate the computational burden that arises with many shots, we can pool (for example, average) the per-image tokens before concatenating all input tokens. This pooled image tokens model achieved a CIDEr score of 56.3 for 4-shots COCO captioning, which is substantially lower than the full model’s CIDEr score of 61.7. This highlights the importance of keeping all the tokens coming out of the ViT encoder, despite the computational overhead.
We explore per-example image-text attention, as proposed and applied in . Under this approach, the image query tokens for each example can only attend to its corresponding text tokens, while the text query tokens can attend to all tokens. By using this per-example attention model, we achieved a CIDEr score of 59.6, which is 2.1 points lower than the full attention model’s CIDEr score of 61.7 for 4-shots COCO captioning.
We report the few-shot results on COCO captioning from early-stopped PaLI-2 3B models; in this case, we did not apply normalized attention in training. We provide the test results with and without attention re-weighting during inference for a different number of encoder shots. Attention re-weighting achieves increasing CIDEr scores of 82.1, 84.3 and 84.5 with 4, 8 and 16 shots respectively. On the other hand, the model achieves 83.4, 76.5 and 66.3 without attention re-weighting. The decreasing performance may suggest that the model fails to locate the target image and text prompt among the large number of shots, whereas the attention re-weighting helps the model to focus on the target features. Accordingly, we decided to include attention re-weighting during finetuning for PaLI-X.
We explore the use of both encoder and decoder shots during pre-training. We pretrain the PaLI-2 700M model on PaLI-2 mixtures with varying number of encoder shots (between 1 and 4). The remaining shots (up to exactly 4) are used as decoder shots. Using only encoder shots leads to a 64.0 CIDEr score for 4 shots in COCO captioning. The best mix of encoder and decoder shots achieves a CIDEr score of 65.2. This suggests splitting shots leads to a more challenging pre-train task that helps the model learn more efficiently.
B.6 Finetuning hyperparameters
The hyperparamter choices for downstream finetuning experiments are summarized in Table 20. As mentioned in the Main Text, for all of the downstream finetuning experiments, we used a reduced set of hyperparameters, without heavy per-task optimization.
B.7 Multi-task finetuning
We deduplicated every training set mixture over the test sets of every task in order to prevent leakage of any test-set examples into the training set. The mixture is formed by putting the training examples of each subtask together, with heuristic adjustments for a better balance. Following the resolutions for the single-task finetuning, the multi-task captioning and VQA finetuning are done with 672 and 756 image resolutions, respectively. The multitask finetuning covers just about 5M examples, which is 20k steps with a batch size of 256. For scene-text and document understanding tasks, the multi-task finetuning uses the end-to-end setting without OCR pipeline input.
The following aspects made multitask finetuning particularly challenging: (i) all tasks used the same prompt without task-specific indicators; the model is thus required to adapt to the style of multiple benchmarks simultaneously. 2) We do not perform per-task validation set optimization. All subtasks are evaluated using the same checkpoint, but tasks converge to their optimal value at a different pace.
B.8 Ablation studies
We first show in Table 22 the advantage brought by the OCR co-training stage of ViT-22B. We pair the vanilla ViT-22B and the ViT-22B with additional OCR co-training with a small language model mT5-base and pretrain these models on 40M of WebLI-OCR data with the splitOCR objective, before finetuning on ST-VQA. Co-training on image and OCR classification has a significant advantage on ST-VQA and TextVQA. In the meantime, the performance on VQAv2, which is not very scene-text heavy, is improved as well. Moreover, we found that making the top left patch white, which helped the co-training of image classification and ocr classification on ViT-22B, is not required for the subsequent training of PaLI-X.
For ablation of the PaLI-X training procedure, we used a 5B model with UL2-3B and ViT-G with 2B parameters, which is roughly a 10:1 down-scale of the PaLI-X 55B model.
For stage 1 training, we show in Table 23 that adding image token generation does not harm the performance on the main image+language understanding tasks.
Appendix C Additional results: Video Captioning and QA
Below we give a brief description of each video data set we used for evaluation. Note that we freshly collected the data when performing the experiments, which led to different effective numbers of videos in different splits in some cases, see Table 24.
These descriptions refer to the original dataset size, but we train on (sometimes significantly) fewer videos — the exact numbers are given in Table 24. This is because not all videos in the datasets were available online at the time of writing (e.g., due to user deletion).
MSR-VTT : This dataset consists of 10K open domain video clips for video captioning, with 20 captions each. The duration of each video clip is between 10 and 30 seconds. We follow the standard splits proposed by and report results on the test set.
VATEX : VATEX includes captions for 41K videos sampled from the Kinetics-600 dataset, with 10 English captions each. We report results on the English public test set.
ActivityNet Captions : This dataset consists of 100K temporally localized sentences for 20k videos. We follow the standard split containing 50/25/25% of the dataset for training, validation and testing, and use ground truth temporal proposals at evaluation following . Note that following other works , we use the val_1 split for validation and val_2 split for testing.
Spoken Moments in Time (SMIT) : This dataset consists of long captions obtained via audio recordings for 500k short video clips. While this dataset has been traditionally only used for text to video retrieval, we find that it is a strong benchmark for captioning as it is the largest manually annotated set of videos with text captions.
ActivityNet-QA : The dataset contains 58,000 question-answer pairs for videos in the ActivityNet dataset . We report accuracy (using exact string match) on the test split. Note that we do open-ended generation for all VideoQA datasets.
MSR-VTT-QA : This dataset was created using a semi-automatic pipeline on top of the MSR-VTT dataset. We report accuracy (using exact string match) on the test split.
NExT-QA : We focus on the Open-Ended QA task, which consists of 52,044 question-answer pairs for a total of 5,440 videos (sampled from the VidOr dataset). Exactly following Next-QA and Flamingo , we report the Wu-Palmer Similarity (WUPS) on the test set.
Appendix D Additional results: Image Classification
The setup used for the experiments here uses the PaLI-X model to generate directly the (English) class name using the captioning prompt. The output is considered correct if it matches exactly the class name (apart from ImageNet-REAL, where we check if the class corresponding to the output is in the set of correct labels).
We use the same scoring technique as in PaLI to evaluate PaLI-X in zero-shot setting (without training on any Imagenet data). We use the PaLI-X model obtained after the first stage of training (using the base 224 image resolution).
The results are presented in Table 25. We compare the results to PaLI - previous zero-shot generative SOTA, and Flamingo - another generative model of similar architecture with comparable 1-shot and 5-shot results. Overall, we report that the results between PaLI and PaLI-X for 0-shot are similar.
To test image classification capabilities, we finetune PaLI-X on ImageNet and evaluate the resulting model on ImageNet-REAL and out-of-distribution datasets: ImageNet-R , ImageNet-A , ImageNet-Sketch , ImageNet-v2 .
We use the model from the first training stage (at resolution 224) and the one from the last training stage (at resolution 756). We use the same training hyperparameters for all of runs (selected without any hyperparameter tuning).
The results can be seen in Table 26. We compare the results to generative model with open vocab – GiT2 (using 384 image resolution), which is the current SOTA for full-finetuning on ImageNet. PaLI-X achieves close to SOTA results for generative models on Imagenet, and other datasets.
Appendix E Object Detection
Object detection is framed similarly to Pix2seq , with two key differences: the use of a natural language vocabulary, and class-conditioning. Prompt classes are fed to PaLI-X’s text encoder, in the format detect class1 and class2 and class3. The model is trained to only output bounding boxes corresponding to classes in this prompt. We represent bounding boxes as coordinates in the same style as pix2seq ; that is, 4 integers ymin xmin ymax xmax ranging from 0 to 999. Figure 9 shows an example input.
During training, a prompt for each example. We construct prompts from three pieces of information:
Negatives: These are the known instance negatives i.e. bounding boxes for objects definitely not present. For exhaustively labelled datasets like COCO, this is simply classes not labelled as positives. For non-exhaustively labelled datasets like LVIS, these are the classes not labelled as positives, which were presented to raters. During training sample , and use up to , where is the number of positives after sampling .
Global negatives: These are negatives which are not explicitly labelled as negatives. They are taken from a wider label space combining multiple detection datasets. For a given example, valid global negatives consist of classes from the wider label space not explicitly labelled as positives or negatives. During training, we sample and append global negatives, where is the number of positives after sampling .
By default, the combined label spaces of Visual Genome, Objects365 and OpenImagesV4 was used as the global label space, with the exception of detection finetuning, where LVIS and COCO label spaces were also added.
E.2 Preprocessing
During pre-training, data is preprocessed to remove all LVIS-rare labels, following the protocol of OwlViT . This is not done for detection finetuning. Images are randomly flipped horizontally, and randomly resized to between 0.3 and 2.0 their original sized, followed by selecting a random square crop of the current training resolution. If the image is resized to be smaller than the current resolution, it is left as is. Images are finally padded to a square.
E.3 Licenses and attribution for images used in Main Text Figure 2
Watermelon: Credit: Sarah Pflug https://burst.shopify.com/photos/cutting-watermelon.
Bowls: https://www.flickr.com/photos/ariesandrea/502826051/ CC-BY-NC-ND 2.0
Business cat Credit: Sarah Pflug, https://burst.shopify.com/photos/business-cat-in-office
Wall Credit: Matthew Henry https://burst.shopify.com/photos/man-walking-in-front-of-this-is-paradise-wall?c=urban-life