A Closer Look at the Robustness of Vision-and-Language Pre-trained Models

Linjie Li, Zhe Gan, Jingjing Liu

Introduction

Large-scale multimodal pre-training has taken innovative strides in the realm of vision-and-language (V+L) research . Pre-trained models such as ViLBERT , LXMERT and UNITER have demonstrated great generalizability over diverse V+L tasks , such as Visual Question Answering (VQA) , Visual Commonsense Reasoning (VCR) , and Referring Expression Comprehension . However, these benchmarks usually possess similar data distribution between training and test sets, with little-to-none linguistic variation in textual queries, and use only clean natural images without any visual content manipulation. Therefore, although effective for general model evaluation, these standard benchmarks lack the ability to explicitly evaluate model robustnessWe do not focus on adversarial robustness in this work, as no existing adversarial benchmark is available. Instead, we investigate existing robustness benchmarks, which are designed with challenging settings and validated by human judges. See Section 2 for a more detailed discussion..

To conduct a full dissection on model robustness, we launch a thorough investigation with systematic evaluations of V+L models over 4 generic types of robustness: (ii) robustness against linguistic variation; (iiii) robustness against logical reasoning; (iiiiii) robustness against visual content manipulation; and (iviv) robustness against answer distribution shift between training and test splits. Given the abundance of diverse datasets and splits on the popular VQA task, we take VQA as the focal point of our investigation, and compile an assemblage of 9 diverse VQA datasets that cover each type of model robustness: (ii) VQA-Rephrasings for linguistic variation; (iiii) VQA-LOL (Compose and Supplement) , VQA-Introspect and GQA for logical reasoning; (iiiiii) IV-VQA and CV-VQA for visual content manipulation; and (iviv) VQA-CP v2 and GQA-OOD for answer distribution shift. To the best of our knowledge, this is the first systematic evaluation of pre-trained V+L models through the lens of robustness. Interestingly, analysis on two archetypal V+L models reveals that by standard finetuning, pre-trained models already exhibit better robustness than many task-specific state-of-the-art methods. However, the achieved robustness is still limited, and far from human performance.

Recently, adversarial training (AT) has shown success on standard V+L tasks . Inspired by this, we investigate whether AT can also serve as an effective conduit to improve performance on robustness benchmarks aforementioned. Our evaluation of Villa (AT-enhanced pre-trained model) shows that by injecting adversarial perturbation to multimodal embeddings, PGD-based (Projected Gradient Descent) AT can help the model adapt to linguistic variation and visual content manipulation, yielding better model robustness; but with only limited effect (sometimes even hurting model performance) on datasets that exhibit salient data distribution gap between training and test sets (e.g., VQA-CP v2, GQA-OOD).

To achieve better robustness across all aspects, we propose Mango (Multimodal Adversarial Noise GeneratOr), a generic and efficient approach that introduces adversarial noise to multimodal embedding space for robustness enhancement. As shown in Figure 1(a), instead of relying on PGD to generate adversarial perturbation, Mango learns an adversarial noise generator in the form of a trained neural network to fool the model. Following , perturbation is added to the embedding space for all modalities, as our goal is the end results of AT, rather than crafting actual adversarial examples. Compared with PGD-based training which is time-consuming, Mango is lightweight and efficient, without the repetitive iterations of gradient calculation required in PGD-based approach.

To enable diverse adversarial embeddings, we further propose to randomly mask image regions and randomly insert [MASK] tokens when adding adversarial noise to image and word embeddings. Empirical results show that Mango significantly improves model robustness across all tasks considered, compared to PGD-based methods.

Our main contributions are summarized as follows. (ii) We conduct the first known study to systematically examine model robustness of prevailing V+L pre-trained models. (iiii) We propose Mango, a generic and efficient adversarial training approach to enhance V+L model robustness. (iiiiii) As summarized in Figure 3, we achieve new state of the art on 7 out of 9 robustness benchmarks, lifting model performance by +11.74 on VQA-Rephrasings, +12.50 on VQA-LOL Compose, +9.96 on VQA-LOL Supplement, +12.55 on VQA-Introspect, +0.84 on IV-VQA, +42.92 on CV-VQA, and +3.70 on GQA-OOD.

Robust VQA

We start with definition of the terminology we use throughout the paper. We follow VQA literature to unify different forms of challenging bias and out-of-distribution generalization as robustness, which is different from its definition in adversarial machine learning. Robustness does not always mean “adversarial robustness” in literature, e.g., it can also refer to model robustness towards common image corruptions . In the language of adversarial machine learning, our definition of robustness here can be unserstood as the “generalization” performance on the challenging robust VQA benchmarks.

Existing Benchmarks There has been a few independent studies on V+L robustness, mostly focusing on variations of the popular VQA task. VQA-CP , drawn from VQA v2 dataset , is the first benchmark proposed to evaluate (and reduce) question-oriented language bias in VQA models. Considerable effort has been invested on VQA-CP along 3 dimensions: (ii) compensating for question-answer distribution patterns through a regularizer based on an auxiliary model ; (iiii) taking advantage of additional supervision from human-generated attention maps ; and (iiiiii) synthesizing counterfactual examples to augment training set . Recent work shows that simple methods such as generating answers at random can already surpass state of the art on some question types. The recent GQA-OOD , another robustness-focused task, is designed based on a fine-grained reorganization of the original GQA dataset .

Besides answer distribution shift, other types of VQA model robustness are also studied: VQA-Rephrasings exposes the brittleness of VQA models to linguistic variations in questions, and proposes cyclic consistency to improve robustness; tackles antonym consistency; studies robustness against automated semantic image manipulations, and tests for prediction consistency to questions on clean images and corresponding manipulated images.

Further studies investigate robustness against logical reasoning. For instance, provides a dataset containing perception-related sub-questions per question for a new reasoning split of VQA dataset. constructs VQA-LOL with questions containing logical compositions and linguistic transformations to examine model ability in logical reasoning. GQA also falls into this category, as its large-scale rule-based questions support analysis on different reasoning skills of VQA models.

Despite the continuous effort in enhancing robustness of VQA models, these works mostly focus on either task-specific models or a single type of robustness. To provide the first comprehensive study on V+L robustness, we compile a full list of existing datasets, and group them into four robustness types: Lingual, Visual, Reason, and Answer (Table 1). Covering various respects of a ‘stress test’ for V+L models, from linguistic to visual variations, from reasoning complexity to answer distribution, this compilation serves as a unified yardstick for evaluating V+L model robustness and a guidance for future study on robust model designAlthough the benchmarks considered here are VQA tasks, these robustness types are generic, and can be extended to other V+L tasks as well. For example, Lingual and Visual are naturally applicable to any task with text and image inputs. Reason and Answer benchmarks can be constructed via logical combination of text inputs or reshuffling train/val/test splits.. As a start, we introduce a generic and effective approach that can lift model performance over all types of V+L robustness indiscriminately, which will be discussed in next section.

Mango Framework

In Sec. 3.1, we briefly review V+L pre-training. Sec. 3.2 introduces a simple baseline that injects Gaussian noise. The proposed Mango approach is detailed in Sec. 3.3.

Take VQA task as an example. Given an image-question pair (v,w)({\bm{v}},{\bm{w}}) in dataset D{\mathcal{D}}, the goal is to predict an answer that best matches ground-truth answer y{\bm{y}}. When finetuning on VQA task, a multi-layer perceptron (MLP) layer is added on top of zcls{\bm{z}}_{cls}, and Binary Cross Entropy (BCE) loss is used to supervise model training. The finetuning process can be formulated as an empirical risk minimization problem:

2 Gaussian Noise Augmentation

Randomized smoothing advocates the addition of random perturbations to model inputs, which can often yield better model performance. Recent study also shows that perturbing clean images with Gaussian noise is effective in improving model robustness against image corruptions for image classification. Inspired by this, we use Gaussian noise augmentation as a simple baseline to investigate model robustness under V+L setting. Instead of adding noise to raw image pixels as in , we add perturbations directly to the embeddings:

where σ\sigma is the standard deviation of Gaussian noise. Similarly, we add Gaussian noise to the word embeddings:

3 Adversarial Noise Generator

Adding Gaussian noise to clean image-text pairs can augment training examples to a certain level. However, as the training continues, the model can gradually adapt to the perturbations which are sampled from the same Gaussian noise distribution. To produce harder perturbations that can fool the backbone network such as UNITER or LXMERT , we propose to actively learn an adversarial noise generator. Specifically, we aim to discover an adversarial noise distribution, from which the sampled noises, when added to the multimodal embeddings, can maximally confuse the backbone network. Note that our goal is not to model the explicit density form of such a distribution, as we only care about the noise samples drawn from the distribution. To achieve this, the adversarial noise generator takes in Gaussian noise samples as input, and produces adversarial noise samples through a learned neural network.

where Rkl(p,q)=KL(p∣∣q)+KL(q∣∣p){\mathcal{R}}_{kl}(p,q)=\text{KL}(p||q)+\text{KL}(q||p), p,qp,q denote two probability distributions. The first term in Rat(θ,ϕv){\mathcal{R}}_{at}({\bm{\theta}},{\bm{\phi}}_{v}) promotes label-preserving adversarial perturbations; while the second term advocates more fine-grained label preservation, meaning that the probability distribution across all answers is used as soft label, instead of using the ground truth answer index as hard label. Similarly, we learn an adversarial noise generator (with parameters gϕwg_{{\bm{\phi}}_{w}}) that corresponds to the text modality, jointly trained with gϕvg_{{\bm{\phi}}_{v}}.The corresponding equations are omitted for simplicity.

During training, we alternate between an outer loop of the backbone network update and an inner loop of generator update. We constrain the noise samples δv{\bm{\delta}}_{v} and δw{\bm{\delta}}_{w} to be within the sphere ∣∣δv∣∣2=∣∣δw∣∣2=ϵ||{\bm{\delta}}_{v}||_{2}=||{\bm{\delta}}_{w}||_{2}=\epsilon, by scaling the generator output with a scalar. ϵ\epsilon is set as 1 in all our experiments, as perturbations with a smaller norm is more likely to preserve original semantics. For better efficiency, we also accumulate the gradients of adversarial noise generator, and only update the generator’s parameters every TT times (set to 20 or 40 in experiments) of backbone update.

The proposed adversarial noise generator is lightweight, consisting of only a few linear layers. Such a light model can easily trap in local minimum when competing with a deep backbone network. Thus, at regular intervals, we replace the learned noise generator with a new one trained from scratch. Each time, the new generator is trained against the latest learned parameters of the backbone.

Random Masking Adversarial noise generator, although produces more challenging and more diverse noise perturbations, does not alter the intrinsic statistics of training examples, such as the distribution of question lengths and image regions. In practice, we observe significant mismatch in these statistics between training and test splits of robustness benchmarks. For example, the average length of questions in VQA-LOL test split is 2-3 times longer than that in VQA v2 training split. The region distribution of images in IV-VQA and CV-VQA is very different from VQA v2 training split, due to visual content manipulation. To compensate for such statistic mismatch, we propose to randomly mask image regions (by zeroing out corresponding feature vectors) as well as randomly insert [MASK] tokens when adding adversarial noise to image and word embeddings. Empirically, this simple technique is effective in further boosting model robustness.

Comparison with PGD-based AT Although Mango is similar to Villa in terms of learning adversarial perturbations, they are different in the sense that Mango learns an adversarial noise generator to generate adversarial perturbations, instead of relying on PGD as in Villa. This makes Mango more efficient, as computing gradients of a generic lightweight noise generator is less time-consuming. Empirically, Mango also achieves better performance. The comparison on model performance and training time difference is provided in Sec. 4.2. A detailed literature review on AT is provided in Appendix.

Comparison with ANT In ANT , a similar noise generator is proposed to make neural networks robust against diverse image corruptions. However, there are two key distinctions. First, we focus on transformer models for V+L tasks, whereas focuses on convolutional networks for image classification. Second, we propose to generate adversarial noise over the embeddings of images and words, while adds adversarial noise directly on image pixels.

Experiments

In experiments, we first use Uniter as the backbone on diverse V+L tasks, then compare Mango with Uniter and Villa baselines over all 9 robustness datasets (Sec. 2), plus a VQA-v2 dataset. We focus our study on these 10 benchmarks as there is no existing robustness dataset on other tasks except for VQA.

Note that almost all 10 benchmarks provide their own training split. We follow the original papers to test model robustness under the most challenging setting (shown in Table 1), which is to evaluate models trained on the VQA training split for VQA-Rephrasings, VQA-LOL, VQA-Introspect, IV-VQA and CV-VQA. A more detailed description of all the benchmarks are provided in Appendix.

For thorough evaluation, we compare model performance on the following set of competing methods:

SOTA (task-specific): Cycle Consistency+BAN for VQA-Rephrasings, LOL for VQA-LOL Compose and Supplement, Pythia for VQA-Introspect, NSM for GQA, SAAA for CV-VQA and IV-VQA, MUTANT for VQA-CP v2, MMN for GQA-OOD, Villa for VQA v2;

UniterB{}_{\text{B}} and UniterL{}_{\text{L}}: standard finetuning of Uniter model with base and large size, respectively;

VillaB{}_{\text{B}} and VillaL{}_{\text{L}}: adversarial pre-trained and finetuned Uniter with base and large size;

MangoB{}_{\text{B}} and MangoL{}_{\text{L}}: applying adversarial noise generator on pre-trained Uniter, base and large size;

MangoVB{}_{\text{VB}} and MangoVL{}_{\text{VL}}: applying adversarial noise generator on adversarial pre-trained Uniter model (provided in the Villa paper ) with base and large size.

2 Experimental Results

Table 2 presents the results of Uniter, Villa and Mango on all robustness benchmarks. Meta-Ave (average of scores across all benchmarks) is used as the global metric.For IV-VQA and CV-VQA, we take the negative of the number of flips for calculating Meta-Ave. L2-5 in Table 2 show the performance of all the models with base size (12 layers). UniterB{}_{\text{B}} (L2) establishes a strong baseline on different types of robustness benchmarks, with a Meta-Ave of 40.98. VillaB{}_{\text{B}} (L4) further improves over this strong baseline by +1.39 Meta-Ave (42.37) via PGD-based adversarial training.

MangoB{}_{\text{B}} achieves across-the-board performance lift on all robustness benchmarks over UniterB{}_{\text{B}}, harnessing an absolute gain of +1.82 on Meta-Ave. Results on VQA-v2 show that MangoB{}_{\text{B}} also boosts performance on standard VQA benchmark. As VillaB{}_{\text{B}} performs adversarial training on both pre-training and finetuning stages, we apply our method to their adversarial pre-trained model for fair comparison. MangoVB{}_{\text{VB}} (L5) outperforms VillaB{}_{\text{B}} on 7 out of 9 robustness benchmarks, with an absolute gain +0.71 on Meta-Ave. Lastly, we compare the training speed of MangoVB{}_{\text{VB}} and VillaB{}_{\text{B}} under the same experimental setting.Experiments are conducted with the same batch size, gradient accumulation steps and number of GPUs. Our experiments show that MangoVB{}_{\text{VB}} is 25% faster than VillaB{}_{\text{B}} (1.44 vs. 1.92 second/step). This indicates that Mango is a more efficient adversarial training approach than Villa, thanks to its use of global noise generator instead of iterative PGD steps as used in Villa.

Scaling Up to Large Model Size (24 Layers) Compared to base models (L2&L4), large models (L6&L8) have more advantage on Meta-Ave (Uniter: 43.37(L) vs. 40.98(B); Villa: 44.33(L) vs. 42.37(B)), which is consistent with the observations on standard V+L benchmarks in . When applying adversarial noise to large backbone models (L7&L9), Mango further pushes the margins of performance gain across all benchmarks: an absolute gain of +1.90 over UniterL{}_{\text{L}} and +0.98 over VillaL{}_{\text{L}} on Meta-Ave.

End-to-end Comparison with SOTA Mango achieves new state of the art on 7 out of 9 benchmarks, in most cases surpassing existing methods by a significant margin. Specifically, Mango pushes state-of-the-art performance by +11.74 on VQA-Rephrasings, +12.50 on VQA-LOL Compose, +9.96 on VQA-LOL Supplement, +12.55 on VQA-Introspect, +0.84 on IV-VQA, +42.92 on CV-VQA, and +3.70 on GQA-OOD. Best results are achieved by MangoL{}_{\text{L}} or MangoVL{}_{\text{VL}}.

On VQA-CP v2 and GQA, although Mango outperforms Uniter and Villa, there is still a gap when compared to SOTA models. SOTA methods on these two benchmarks exploit additional task-specific information. Specifically, MUTANT for VQA-CP v2 is trained with excessive additional image-question pairs designed to promote positive bias; while NSM for GQA takes advantage of additional scene graph annotations, which are only provided in GQA. As the goal of our proposed method is to bring universal performance lift on all robustness benchmarks, we do not exploit these additional task-specific information introduced by MUTANT and GQA.

3 A Closer Look into Robustness

Robustness against Linguistic Variation As shown in Table 2 (‘Lingual’ column), the joint embedding learned by UniterB{}_{\text{B}} has shown its advantage of defending model robustness against linguistic variation. We contribute the performance lift from UniterB{}_{\text{B}} to excessive variations of textual inputs seen during pre-training. Comparing AT-enhanced methods, MangoVB{}_{\text{VB}} improves over VillaB{}_{\text{B}}, even though VillaB{}_{\text{B}} has already shown significant improvement over UniterB{}_{\text{B}}. We attribute the improvement from Mango to not only the adversarial data augmentation during training, but also the random masking introduced from the text modality. More analyses on each component of Mango over VQA-Rephrasings can be found in Table 5.

Robustness against Logical Reasoning We compare model performance on 4 benchmarks under the ‘Reason’ column in Table 2. Different from VQA-LOL Compose, VQA-LOL Supplement dataset consists of questions generated by heuristic rules. Semantically-close questions with different answers are included to make the task more challenging. The close-to-random performance on VQA-LOL Supplement dataset indicates that UniterB{}_{\text{B}} severely suffers from these challenging semantically-close questions.

VillaB{}_{\text{B}} brings performance lift on all 4 reasoning benchmarks. Not surprisingly, VillaB{}_{\text{B}} exhibits more robustness than UniterB{}_{\text{B}} on semantically-close questions in VQA-LOL Supplement. Our hypothesis is that the adversarial embeddings learned during VillaB{}_{\text{B}} training can mimic the effect of adding semantically-close questions as training data, and the generated adversarial perturbations are also constrained to be small to preserve the semantic meaning of the clean text embeddings.

MangoVB{}_{\text{VB}} outperforms VillaB{}_{\text{B}} on all reasoning benchmarks. Similar to VQA-Rephrasings, MangoVB{}_{\text{VB}} has more advantages over VQA-LOL Compose and VQA-LOL Supplement, whose average question length is much longer than VQA v2. By randomly inserting [MASK] tokens, MangoB{}_{\text{B}} effectively augments training data with questions of similar lengths to the test split.

Robustness against Visual Content Manipulation UniterB{}_{\text{B}} performs on par to SOTA model on IV-VQA, and significantly improves over SOTA on CV-VQA (Table 2 ‘Visual’ column). This is due to that during pre-training, UniterB{}_{\text{B}} has already be trained on diverse images, and the pre-training task of masked region modeling can also prevent UniterB{}_{\text{B}} from overfitting to visual biases. VillaB{}_{\text{B}} improves model robustness against visual content manipulation, and MangoVB{}_{\text{VB}} performs on par with VillaB{}_{\text{B}}. Our hypothesis is that by injecting adversarial perturbations at pre-training stage, the model is exposed to even more diverse images, hence easier to recover from visual biases.

Robustness against Answer Distribution Shift On out-of-distribution (OOD) benchmarks, UniterB{}_{\text{B}} performs poorly on VQA-CP v2, while improving over SOTA model on GQA-OOD (Table 2 ‘Answer’ column). As mentioned in Sec. 4.2, MUTANT is a very task-specific method, which augments VQA-CP v2 training with excessive rule-based image-question pairs to counter the training split bias. Hence, it is difficult to generalize to other robustness cases. Additional manual effort is required to generalize to other rule-based datasets such as VQA-LOL, GQA, IV-VQA and CV-VQA. Interestingly, VillaB{}_{\text{B}} improves over UniterB{}_{\text{B}} on GQA-OOD, but not on VQA-CP v2, which may be due to the fact that the generated “local” perturbations in Villa cannot cast a strong enough regularization effect on this challenging dataset. We also observe that MangoVB{}_{\text{VB}} significantly outperforms VillaB{}_{\text{B}} on both benchmarks, indicating that the generated “global” perturbations in Mango are more generalizable to challenging OOD datasets.

Evaluation on Consistency In addition to accuracy, many benchmarks consider consistency as an additional measure for evaluating model robustness. Here, we take VQA-Rephrasings and VQA-Introspect as examples to demonstrate that Mango can also help boost consistency in model predictions. Results are summarized in Table 3.

On VQA-Rephrasings, we investigate consistency in model predictions on different variants of semantically equivalent questions. Consistency is measured by a Consensus Score CS(k)CS(k). Consensus Score is the ratio of the number of subsets where all the answers are correct and the total number of subsets of size kk. For every group QQ with nn rephrasings, all subsets of size kk are sampled. The answer to a question is considered correct if it has a non-zero VQA accuracy. Mango achieves universal performance lift across all consistency measures, compared to each baseline model. The best results are achieved by MangoL{}_{\text{L}}, surpassing SOTA by +9.43, +12.27, +13.62, +14.40 on CS(k),k=1,2,3,4CS(k),k=1,2,3,4, respectively.

On VQA-Introspect, we examine consistency between the main reasoning questions and perceptual sub-questions, measured by 5 metrics. Similarly, Mango brings universal consistency improvements across all baseline models. The best performance is achieved by MangoVL{}_{\text{VL}}, surpassing SOTA by +12.55, +5.54, +2.27, +5.16, +10.10 on M✓\checkmarkS✓\checkmark, M✓\checkmarkS×\times, M×\timesS✓\checkmark, M×\timesS×\times, and S✓∣M✓\checkmark|\text{M}\checkmark, respectively.

4 Ablation Study

Noise Generation and Random Masking We select one dataset from each robustness type as a representative benchmark for ablation studies: VQA-CP v2, VQA-Rephrasings, VQA-LOL (Compose and Supplement), and IV-VQA. Results are summarized in Table 5. First, we compare with the baseline (Sec. 3.2) that simply adds Gaussian noise to either image or text modality.In our experiments, we set standard deviation to 0.5, and only perturb 50% of training data via Gaussian noise within each minibatch. Different from observations in , comparing L2/L5 with L1 indicates that adding simple Gaussian noise to multimodal embeddings is not always helpful. Especially, adding Gaussian noise on text modality brings unstable performance.

Second, we experiment with adding adversarial noise alone, without random masking. Results on L3/L6 show that universal performance improvements over Gaussian noise (L2/L5). Intuitively, adversarial noise is harder than Gaussian noise, as the adversarial noise generator learns to fool the backbone network. Model training with such augmented hard examples helps to boost model robustness.

Third, we show that by using random masking (L4/L7), which encourages more diverse adversarial embeddings, Mango is better than using adversarial noise alone (L2/L5). Randomly inserting [MASK] tokens (L7) also shifts the distribution of question lengths that the model is exposed to during training. Hence, we observe more gains on benchmarks with severe mismatches in question length between training and test sets. For example, in VQA-LOL, the questions in the test set are significantly longer than questions in the training set on average.

Lastly, we observe that adding adversarial noise on one modality is already gaining significant improvement (L4/L7). Empirically, adding adversarial noise on both modalities (L8) only performs slightly better or on par with Mango on text or image modality alone. More ablation results on model architecture are included in Appendix.

Results on LXMERT To demonstrate the versatility of Mango, we also apply Mango to a two-stream backbone, LXMERT, for generalizability test. We compare LXMERT baseline with its enhanced version with Mango (“ours” in Table 5). Evaluation is conducted over VQA-Rephrasings, VQA-LOL Compose, VQA-LOL Supplement, GQA and GQA-OOD datasets, covering 3 types of robustness. IV-VQA, CV-VQA and VQA-CP v2 are excluded in this study as the performance on these benchmarks is based on examples in VQA v2 val split, which is used to supervise LXMERT pre-training. We also report results on standard VQA v2 benchmark.

Note that LXMERT is pre-trained with VQA v2 and GQA data. Therefore, it achieves superior performance on VQA-Rephrasings, which includes questions from VQA dataset. On benchmarks whose data are unseen during pre-training, LXMERT exhibits similar robustness to Uniter. LXMERT suffers severely on semantically-close questions in VQA-LOL Supplement and logical reasoning questions in VQA-LOL Compose. A possible reason is the over-exposure to VQA questions during both pre-training and finetuning. Despite the limitations mentioned above, when applying Mango to LXMERT, we still observe universal performance lift across all benchmarks considered.

Results on other V+L tasks As aforementioned in Sec. 2, we take robust VQA as test bed (9 datasets with 10 different model settings) due to the lack of available datasets to evaluate model robustness for other V+L tasks. Mango is task-agnostic, thereby can also be applied to other standard V+L tasks. We compare MangoB{}_{\text{B}} against UniterB{}_{\text{B}} on popular V+L tasks, including NLVR2 , RefCOCO , RefCOCOg and Visual Entailment (VE) in Table 6. Mango surpasses baseline results across all 4 tasks.

Qualitative Analysis Figure 2 visualizes prediction examples from Uniter, Villa and Mango on 4 benchmarks (each for one robustness type). These visualizations illustrate Mango’s consistently accurate performance when facing challenges of: (a) uninformative leading phrase added to the question; (b) removal of irrelevant object in the image; (c) over-length logical combination of questions; and (d) imbalanced answer distribution (‘white’ appears 3 times as many as ‘blue’ in training set). More visualization examples are included in Appendix.

Conclusion

We provide the first known systematic study on the robustness of pre-trained V+L models. By examining existing models on a wide range of robustness benchmarks, we obtain a better understanding of how V+L pre-training handles various types of robust tests. We further propose Mango, a simple yet efficient adversarial training method to enhance model robustness, which advances the state of the art on 7 out of 9 robustness benchmarks by a large margin. We hope this set of results can be used as baseline for future research. A natural follow-up of this work is to investigate the adversarial robustness of pre-trained V+L models.

References

Appendix A Detailed Related Work

Early approaches to vision-and-language pre-training adopt a two-stream architecture. Later on, single-stream architecture gains popularity . Multi-task learning , adversarial training , and contrastive learning have proved useful for improving model performance. Probing analysis shows that these pre-trained models can learn essential knowledge about visual co-reference and visual relations. Recent work extends this pre-training strategy to diverse tasks such as image captioning , visual dialog , visual-language navigation , and video-text pre-training . Instead of using the conventional bottom-up-attention features , Pixel-BERT proposes end-to-end learning from image pixels to textual tokens. External knowledge, such as image tags and scene graphs , as well as weakly-supervised pre-training , are also investigated for further enhancement.

Distinct from these efforts on improving performance over standard benchmarks,Examples of standard benchmarks include VQA , VCR , NLVR2 , Image-Text Retrieval , and Referring Expressions . we focus on a different direction, evaluating and enhancing the robustness of pre-trained models. This helps us better understand how well multimodal pre-training truly advances this field, and guides us to design more robust models.

Adversarial Training As one of the most effective strategies of defending against adversarial attacks , adversarial training (AT) has been widely studied for enhancing adversarial robustness of neural networks , using adversarial examples as effective data augmentation. Recent studies show that, by injecting adversarial perturbations into feature space, AT can further improve model generalization on language understanding , visual question answering , and graph neural networks .

In our work, we investigate the use of an adversarial noise generator for robustness enhancement, inspired by , which proposes a similar noise generator to make neural networks robust against diverse image corruptions. However, there are two key distinctions. First, we focus on transformer models designed for V+L tasks, whereas focuses on convolutional networks for image classification. Second, we propose to generate adversarial noise over the embeddings of images and words, while adds adversarial noise directly on image pixels.

Appendix B More Results

We include more detailed results on IV-VQA , CV-VQA , GQA-OOD and VQA v2 , and also ablation experiments on model architecture.

On IV-VQA and CV-VQA, we decouple the inconsistency in model predictions on edited images (measured by #flips) into 3 categories: (ii) p2n: answer predicted on the edited image was wrong, but the prediction on the corresponding real image was correct; (iiii) n2p: model makes a correct prediction on the edited image, while predicting a wrong answer on real image; (iiiiii) n2n: different answers were predicted on edited and real images and both are wrong. These metrics may expose that there is brittleness even when the model makes correct predictions, indicating that models often exploit spurious correlations while making predictions. We follow to report accuracy on VQA v2 val split to serve as reference for IV-VQA, and performance on counting questions in VQA v2 val split for CV-VQA.

Similar conclusions are drawn from the results presented in Table 7. First, Mango brings consistent performance improvements across all metrics on both benchmarks, compared to Uniter. Second, Mango significantly improves over SOTA. We also observe significant improvements from Mango over Villa on CV-VQA. These results suggest that for challenging questions such as counting problems in CV-VQA, Mango is more robust than Villa.

On GQA-OOD, except for the accuracy over all GQA-OOD samples (‘All’ in Table 8), three additional metrics are considered: (ii) the accuracy on OOD samples, which are the samples of the tail of the answer class distribution (‘Tail’); (iiii) the accuracy on the head of distribution (‘Head’); and (iiiiii) Δ(head, tail)=(head - tail)/tail\Delta(\text{head, tail})=(\text{head - tail})/\text{tail} to illustrate how much the error prediction is imbalanced between frequent and rare answers (‘Δ\Delta’). More details on the statistics of head and tail examples can be found in . Mango achieves universal performance lift across all accuracy measures, compared to each baseline model. However, better accuracy does not indicate better-balanced predictions between tail and head splits. We observe that there are more performance improvements on head split than tail split. When compared to SOTA, MangoB{}_{\text{B}} surpasses MMN (SOTA with the best All) across all metrics. BAN is the SOTA method with the best Δ\Delta; however, it suffers on all accuracy measures.

On VQA v2, we use MangoB{}_{\text{B}} and UniterB{}_{\text{B}} as examples to show that our method can provide universal performance lift for each question type. This is also consistent with our observations on various robust vqa benchmarks, as they focus on different question types by design. For examples, IV-VQA speficially desgined for counting questions, VQA-LOL only includes yes/no questions.

Additional Ablations We conduct additional ablation studies to validate several model design choices, including KL-divergence Loss, retraining noise generator every TT steps (retrain NG), the architecture of NG (multiple linear layers with nonlinear activation) and the effectiveness of masking on Villa. Results are reported in Tab. 10.

A few key observations are summarized here: (ii) KL divergence loss contributes to performance improvements in Mango. (iiii) Without resetting generator parameters and retraining generator periodically results in worse performance. As explained in L409-415 of the main text, the lightweight generator may be trapped in a local optima. (iiiiii) Replacing our noise generator with a single linear layer also hurts the performance. Note that applying linear layers to a Gaussian noise only changes its mean and variance, still results in a Gaussian noise. (iviv) VillaB{}_{\text{B}} + Masking renders weaker performance than MangoVB{}_{\text{VB}}. This observation is consistent with comparison of VillaB{}_{\text{B}} in Tab. 2 and “AN” in Tab. 4, which can be considered as “MangoB{}_{\text{B}} - Masking”.

Appendix C Implementation Details

Our models are implemented based on PyTorch.https://pytorch.org/ To speed up training, we use Nvidia Apexhttps://github.com/NVIDIA/apex for mixed precision training. Gradient accumulation is applied to reduce multi-GPU communication overheads. All experiments are run on Nvidia V100 GPUs (32GB VRAM; NVLink connection). We use AadmW with β1=0.9\beta_{1}{=}0.9, β2=0.98\beta_{2}{=}0.98 and an L2 weight decay of 0.01 to optimize model training. Throughout the training, the learning rate is scheduled to warmup over the first 10% training steps followed by linear decay to 0. The peak learning rate is set to be 8e-5 and 5e-5 for base and large models, respectively. Additional hyper-parameters used to train our adversarial noise generators are listed in Table 11. Empirically, we found that model training is sensitive to adversarial noise retrain steps, pmaskimgp_{\text{mask}}^{\text{img}} and pmasktxtp_{\text{mask}}^{\text{txt}}.

Appendix D Downstream Benchmarks

In addition to dataset statistics summarized in Table 1 in main text, we provide an overview of each robustness benchmark as follows.

VQA-Rephrasings is based on VQA v2 . It contains 3 human-provided rephrasings for 40K questions on 40K images from VQA v2 val split. In addition to accuracy, consistency in model predictions to different semantically-equivalent questions is also used to measure the robustness of VQA models against linguistic variations. We follow to evaluate models trained with VQA v2 train split.

VQA-LOL is introduced to examine the logical reasoning ability of a VQA model through questions containing logical compositions and linguistic transformations (negation, disjunction, conjunction, and antonyms). It consists of two datasets: VQA-LOL Compose (logical combinations of multiple closed binary questions about the same image in VQA v2) and VQA-LOL Supplement (logical combinations of additional questions based on external object and caption annotations about the images from COCO ). Both datasets share the same train/val images as VQA v2. In total, 757K/42.5K/291K and 1.61M/91.8K/669K image-question pairs are generated for train/val/test splits of VQA-LOL Compose and VQA-LOL Supplement, respectively. In our experiments, we follow to evaluate models trained with VQA v2 train split on test split of both datasets.

VQA-Introspect is created to investigate the consistency in model predictions of a VQA model between reasoning questions and their associated low-level perception questions. It first introduces a new Reasoning split of the VQA v2 dataset and collects 238K new perception questions. These questions correspond to the set of perceptual tasks needed to effectively answer complex reasoning questions in the Reasoning split. In total, VQA-Introspect contains 167K sub-questions for 56K reasoning questions in VQA v2 train, and 72K sub-questions for 22K reasoning questions in VQA v2 val. In our experiments, we follow to evaluate models trained with VQA v1 train split on VQA-Introspect val split.

GQA contains 22M automatically generated questions based on ground-truth image scene graphs. The questions are constructed via a set of heuristic rules, which are designed to evaluate a VQA model in terms of different types of reasoning skills (e.g., spatial understanding and multi-step inference). We follow to use the balanced version of GQA, which has been designed to reduce biases in answer distribution. In the balanced version, 1.7M questions are split into 70%/10%/10% for training, validation and test sets, respectively. In our experiments, models are trained on GQA train split and we report performance on test-dev split.

IV-VQA & CV-VQA are two synthetic datasets, created by removing objects in the real VQA images. In IV-VQA, irrelevant objects are erased and model predictions before and after image manipulations are expected to be invariant. In CV-VQA, which focuses on counting questions, one relevant object is removed from the given image and model predictions on the quantity of such object are expected to be subtracted by 1. Objects of choice are based on heuristic rules and removed via inpainter-GAN . In total, 376K and 13K image-question pairs are generated for IV-VQA and CV-VQA, respectively. The detailed splits can be found in Section 3 of main text. In our experiments, we follow to evaluate models trained with VQA v2 train split on IV-VQA/CV-VQA val split.

VQA-CP v2 is an out-of-distribution (OOD) reorganization of VQA v2. It was created to examine the robustness of a VQA model in a setting where language priors cannot be relied upon for a correct prediction. The questions in VQA v2 are first assigned to one of 65 question types according to their prefix (first few words). For every question type, the prior distribution of answers is shuffled to be different in train and test splits of VQA-CP v2. Our models are trained on VQA-CP v2 train split and evaluated on test split, following .

GQA-OOD is also an OOD benchmark, created by re-organization of the GQA dataset. By utilizing fine-grained question generation templates in GQA, GQA-OOD divides questions into 37K local groups, and shifts answer distribution by selecting a subset of answer classes for each question group, according to their frequencies. Unlike VQA-CP v2, GQA-OOD features distribution shifts for both validation and test, allowing to validate models under OOD conditions. In our experiments, we follow to evaluate models trained with GQA train split on GQA-OOD test-dev split.

Appendix E More Visualizations

We provide additional visualization of model predictions in Figure 3. Mango consistently provides accurate predictions for each robustness type.