Learning from Lexical Perturbations for Consistent Visual Question Answering

Spencer Whitehead, Hui Wu, Yi Ren Fung, Heng Ji, Rogerio Feris, Kate Saenko

Introduction

Though great progress has been made towards Visual Question Answering (VQA), the approaches lack robustness and are sensitive to input variations . In particular, prior work has shown that when presented with a rephrased version of a question, VQA models often produce inconsistent answers. We conjecture that this is likely because most VQA approaches ignore interconnections between related questions and handle each example independently during training, even though learning such relationships is paramount to generalization, robustness, and compositionality.

While a few attempts have been made to improve VQA robustness via augmentation and regularization techniques , they encourage consistency among related questions at the answer prediction level, without considering the stronger form of consistency between the intermediate computation steps. Furthermore, proposed augmentations are either costly (human-generated ) or suffer from quality control issues .

In this work, we propose a novel robust VQA approach that first augments the question and then enforces consistency not only of the answer, but also of the intermediate representations computed by the model. The intuition is that (like humans) VQA models should follow the same reasoning steps to solve two differently-phrased questions that have the same meaning. For example, to answer both questions “Are the buildings tall?” and “Are the buildings short?” a model should follow the same process, i.e. detect buildings in the image, then classify their height.

How to enforce such consistency of the intermediate reasoning steps appears to be a hard problem in general. We propose that an effective way to improve this stronger notion of reasoning consistency is by maintaining the associated computation steps between questions that differ by controlled variations. We leverage the family of interpretable, compositional VQA models called Neural Module Networks (NMNs) which explicitly represent sub-tasks like object detection and spatial reasoning as modules within the network and predict sequences of weights over these modules (akin to programs) to solve each question. However, unlike existing NMN work, we regularize the model to learn not only how to compose sub-tasks, but also follow the same sequences of sub-tasks to answer two variations of the same question, illustrated in \figreffig:intuition.

In addition, we also propose a novel data augmentation algorithm based on auxiliary linguistic knowledge freely available in text-only corpora. Examples like in \figreffig:compare_examples, seem to suggest that existing VQA models ignore the relatedness between questions with mild linguistic modifications. Questions like the antonymous pair in \figreffig:compare_examples are related in that they query the same property of the same object, but differ in that they ask the opposite of each other. Existing work ignores such modifications and focuses either on human-written paraphrased VQA questions or auto-generated paraphrases using back-translation . Human paraphrases tend to add filler phrases or even change the question meaning (see \secrefsec:dataset) and are very costly to annotate, while back-translations are also hard to control and have quality problems. Motivated by this, we propose to augment data with changes at a low level, such as simple lexical substitutions. We create variations by substituting parts of the questions using rules extracted from large-scale language resources which keeps the meaning and the realistic distribution of the original questions while avoiding the cost and/or semantic incoherence commonly found with prior work.

Finally, we contribute VQA Perturbed Pairings (VQA P2), a dataset of perturbed questions derived from VQA v2 that can be used to measure robustness to the specific linguistic phenomena not previously evaluated in VQA literature, covering the usage of different synonyms (Synonymous), different phrasings (Paraphrastic), and opposite attributes (Antonymous). VQA P2 is comprised of 36.2k VQA questions, where each question has a corresponding question in VQA v2 that it differs from by controlled perturbations. VQA P2 can be easily expanded without the need for expensive human annotations , while maintaining control over the perturbations used.

To summarize, our main contributions are:

A novel VQA consistency regularization method that augments questions and enforces similar answers and reasoning steps for the original and augmented questions.

A new data augmentation method and VQA robustness benchmark, VQA P2, to diagnose the robustness of VQA models to controlled linguistic perturbations. To the best of our knowledge, this is the first approach to use large-scale linguistic resources automatically mined from human-written text to enhance VQA.

Experiments applying our approach to regularize the training of modular, compositional, and interpretable VQA models, demonstrating that it leads to significant improvement in model consistency.

Related Work

VQA. Prodigious progress has been made towards VQA. Models that learn strong feature backbones with co-attention/bilinear attention have achieved top scores on the widely adopted VQA v2 benchmark . Recently, vision and language transformer models , especially when pre-trained on large-scale vision and language datasets, have led to state-of-the-art performance on VQA accuracy. Despite this success, these models are generally not transparent and hard to diagnose. In contrast, compositional and modular approaches, like NMNs, offer a number of distinct advantages, such as being more interpretable , requiring less training data , and being able to handle complex, compositional questions . Although, there is often a trade-off between accuracy and interpretability.

VQA Benchmarks. VQA v2 has become the de facto training and evaluation data resource. Although improvements on VQA v2 test accuracy has marked encouraging progress on VQA, previous work has found that models are prone to learning superficial correlations by taking advantage of dataset bias , and the standard VQA accuracy metric does not help to quantify such bias. A number of efforts have since been proposed to serve as additional diagnostic and evaluation tool for VQA. CLEVR and GQA create compositional questions which are hard to answer if the model relies on language bias. However, the questions in these datasets are synthetic and do not match the characteristics of natural questions about real images. VQA-CP creates training and testing data splits which have different answer distributions. TDIUC categorizes and measures performance by 12 different question types. While these benchmarks can measure a model’s generalization ability to new question types or new answer distributions, they do not serve to examine how much an answer might change if we change a specific input question. More related to our work is VQA-Rephrasings dataset , which has rephrased questions associated with a subset of original VQA questions. However, as we have discussed, the open-endedness nature of human-written paraphrases makes it hard to identify the source of model inconsistenties. Unlike our dataset (VQA P2), it does not control for the specific type and degrees of question variations, making it hard to codify how questions are related to one another and learn from such relationships.

Improving VQA Robustness. Robustness of deep learning models has been intensely studied in recent years, both in the context of image classification and in a number of natural language understanding tasks . Within VQA, robustness has been studied under the perspective of grounding model predictions to visual regions that are interpretable and consistent with human annotated attention maps . While visual grounding provides a better basis for learning representations of concept words in questions, many words and expressions in VQA questions do not have a direct correspondence with visual regions (i.e., not “groundable”). More related to our work, examines model consistency under input question variations. However, these models that consider consistency among related questions only do so at the answer prediction level . Additionally, another approach applies embedding constraints between questions to improve performance . Our framework not only employs answer prediction loss, but also dissects the model and enforces similarities on the sub-tasks involved, which aligns with compositional and transferable modeling of VQA.

VQA P2 Benchmark

Our goal is to create an objective benchmark to measure the progress of robustness in VQA models, specifically the consistency of VQA predictions under different linguistic perturbations of the input questions. So the question is: what is the best method to create realistic, expressive, and semantically coherent, in an efficient fashion?

Visual question generation methods can generate a variety of questions, but the output can be incoherent for a specific image or may be limited to certain types of questions . More recently, back-translation has emerged as a method of creating paraphrased questions for VQA . However, in general, controlled generation of text is a difficult open challenge . Alternatively, template-based methods can offer more control and semantic coherence, but require extensive annotations (e.g., scene graphs ) and are limited in expressiveness by the templates used. Human-written paraphrases is the other option to generate robustness benchmark . Other than the obvious drawback of being more costly, human-written question tend to be generic, not good for diagnosing the particular source that causes answer prediction inconsistencies, as well as being often overly verbose (as shown in \secrefsec:othermethod).

Different from existing approaches, we build our dataset from the original VQA v2 data, and create variations by substituting parts of the questions using rules extracted from large-scale language resources. Doing so keeps the expressiveness and realistic distribution of the original questions and avoids the semantic incoherence commonly found with trained question generators. Utilizing large-scale linguistic resources offers an efficient means for creating questions that contain linguistic variations exhibited by humans .

Substitution Extraction. We extract lexical substitution rules from the Paraphrase Database 2.0 (PPDB) , a lexical database containing over 100 million paraphrases automatically mined from human-written text, as well as WordNet , two large-scale linguistic resources, and apply these rules to existing VQA v2 questions. We create three types of perturbations: 1) synonymous perturbations that substitute a single word with its synonym; 2) paraphrastic perturbations that substitute multi-word phrases for single/multi-word phrases with the same meaning; and 3) antonymous perturbations where adjectives or verbs are substituted with their antonyms, which explores the ability to understand opposite states of an attribute or action. We extract synonymous and paraphrastic rules from the lexical and syntactic subsets of PPDB respectively, and accept the rule if the target is “equivalent” or “entailed” by the source, according to PPDB’s lexical constraints. We gather antonymous rules from WordNet.

Rule Refinement. We automatically determine which rules can be applied by matching the source phrases and grammar requirements to the questions, discarding rules that are not applicable to any questions. We ensure that our rules are not simply adding unknown words by removing rules whose source or target contain words that don’t appear in the VQA vocabulary. We filter the synonymous and paraphrastic rules with a minimum confidence threshold to prevent low quality/frequency substitutions. Examples of these rules are shown in \tabreftab:dataset_stats, where each rule is a mapping from a source to target phrase.

Applying Substitution Rules. We apply the filtered rules to VQA v2 questions and obtain our final benchmark. Since consistency metric requires the groundtruth annotation for the each question, and VQA v2 testing sets do not offer publicly available answer annotations, we use VQA v2 validation set as the basis for our benchmark. This practice is common among methods that require per-question answer annotations .

During the rule application process, for each question, multiple substitutions for a specific type of perturbation can be made, which increases variations. For antonymous rules, we perform word sense disambiguation and limit their application to yes/no questions of the form “is/are the” and “is/are this/these” where the WordNet synsets of the source word in both the question and rule match. Antonymous rules are limited to these kinds of questions to help ensure that we only apply the antonym substitutions to attributes/states that are directly queried (e.g., “Is the window open?”), and not simply mentioned in the question (e.g., “What’s near the open window?”). For answers, synonymous and paraphrastic questions use the same answers as their VQA v2 counterparts as they share the same meaning and antonymous questions take the opposite answers. This results in a set of 36,17536,175 question-answer pairs that are perturbed counterparts of VQA v2 validation questions.

2 Comparison to Other Paraphrasing Methods

Two main alternative methods for generating paraphrases of questions are human annotations and generative approaches. \figreffig:compare shows a few examples from human-written paraphrased VQA questions and auto-generated paraphrases using back-translation . When primed with the original question, human annotators tend to give paraphrases that are longer than the original question, with a large percentage of these simply adding filler phrases that do not change the original sentence by much. For example, in the first human-written example in \figreffig:compare, the original question is essentially still preserved within the paraphrased version. Additionally, human paraphrases can introduce multiple sources of variations, such as introducing commonsense related items, sentence structural changes as well as lexical alterations. Consequently, using human-written paraphrases as a diagnostic benchmark lacks a level of precision needed to diagnose and understand model performance. On the other hand, automated methods like back-translation suffer from many quality control issues. \figreffig:compare shows how this method frequently generates mismatched phrasal replacement or semantic drift, where the generated questions no longer hold the same meaning.

In contrast, our benchmark has built-in quality control as the replacement rules extracted from linguistic resources correlate well with human judgements and are generally reliable. We target at linguistic perturbations to factor out other sources of variations when diagnosing models. Our dataset offers control over the kinds of perturbations used and allows for more precise evaluation of model consistency, meaning we can evaluate a model’s capacity for addressing specific types of perturbations offering more diagnostic insight. Additionally, due to the fast, automated nature of this process, our dataset is easily extensible to more linguistic variations and larger sets of perturbations.

Approach

Given an image-question pair, (I,Q)(I,Q), a VQA model maps the pair to a distribution over an answer set, f(I,Q)→af(I,Q)\rightarrow\mathbf{a}. Existing approaches most often treat this as a classification task, minimizing the prediction loss between the predicted answer and the ground truth answer. This standard approach, however, does not take into account possible relations among questions. The goal of our approach is to train a model to be aware of question relationships, and thereby learn to be more consistent when answering.

Modeling Question Relatedness. Typically, given an input question, a predictable set of reasoning steps are expected to answer the question. For example, in \figreffig:intuition, the sub-networks which can answer “Which country’s flag is displayed behind the TV?” should be able to decompose this task into components such as “Find(TV)”, “Transform(Behind)”, and “Describe(which country’s flag)” and learn to transfer all sub-networks to the question: “Which country’s flag appears behind the TV?” Essentially, there should be a set of elementary operations where each one solves a less complex task than the original question and related questions can share these operations, following a similar order of execution. Based on this intuition, we propose our Question-Relatedness Regularized Reasoning (Q3R) framework.

Our framework is comprised of three components: 1) a method to create linguistic variations of input questions; 2) a compositional backbone model, guided by question-based module selection; and 3) a mechanism to enforce similarities of related questions at the module level.

Our framework is agnostic to the specific designs of the components of the backbone network and can work within the general controller-module framework. We therefore select two state-of-the-art compositional models to instantiate our backbone, specifically, StackNMN , and XNM . To verify that our training scheme is not tied to either one of the network, we also present a new model by adopting modules from both networks, called HybridNet. More architectural details are in Appendix D.

Regularization Method. We propose to regularize the training of the backbone compositional model to improve its consistency and robustness against linguistic variations at the module level. Controlling and regularizing question relatedness at the module level offers several benefits. First, it provides finer control of the model’s active sub-networks than only using supervision at the output layer. Second, the related question pairs share intermediate activation similarity but do not necessarily need to match one another at a lower level (e.g., attention maps within modules), making the model less sensitive to surface-level sentence variation.

where λ\lambda is a hyperparameter to scale the loss term and dd is a distance metric between two distributions. We find that KL divergence or common vector norm losses, such as L1-norm, work well as dd in the proposed loss term.

Experiments

Data and Metrics. We use VQA v2 (Appendix C) for training and VQA P2 for our main evaluations. Since our goal is to benchmark and advance existing models’ consistency, it requires groundtruth answer labels for each test question. Since answers for the VQA v2 test sets are not publicly available, we use questions from the validation set for testing and never use it during training, as is common for evaluating on annotated subsets of VQA v2 . Our metrics are standard VQA accuracy as well as the consensus score (CS) between pairs of related questions.A non-zero CS for a pair of questions requires a model to answer both questions correctly .

Models. Existing VQA literature is vast and the goal of our work is not to exhaustively study VQA models. We focus on the consistency and robustness of models in response to linguistically perturbed data. Therefore, we pick representative models for our evaluations. BAN is a bilinear model that achieves the single-model top performance on VQA v2 without external data. BAN uses bilinear co-attention with residual connections to model the interactions between image regions and words. Transformer is a top performing architecture that acts as a multi-modal encoder. For fair comparison, we do not use large-scale pre-training. As described in \secrefsec:nmn_arch, StackNMN and XNM are examples of expert-free, end-to-end trainable NMNs. We experiment with both models as well as a hybrid, called HybridNet (see \secrefsec:nmn_arch). Experiments with BAN and Transformer examine the consistency of non-NMN state-of-the-art architectures. The NMN experiments explore the performance of different module implementations when employing our framework.

Implementation. For fair comparison, all models use the same visual features from , input textual features from , and random seed of 0. Whenever possible, we use publicly available implementations. For Q3R, λ=1.0\lambda=1.0 and r=0.2r=0.2 for all experiments and we use a L1 loss for our distance metric. Please see Appendix A for details.

We benchmark existing models on their robustness against controlled lexical perturbations, as shown in \tabreftab:accuracies. Interestingly, we see that the different classes of models have less trouble with paraphrastic changes than they do with synonymous, with the average drop in performance being −0.9-0.9 compared to −2.4-2.4. This is likely due to the fact that paraphrastic changes tend to effect transitional phrases (e.g., “be considered”), which models may ignore , whereas synonymous changes effect these as well as concept mentions (e.g., “car”) that are needed to answer the question. We see that models struggle the most with antonymous changes, dropping at least −10.1-10.1. Despite having lower VQA v2 accuracy, the NMN architectures perform better on the perturbed antonymous questions compared to BAN and Transformer. The results suggest that the Transformer architecture trained from scratch is one of the less robust architectures across the different types of perturbations. Overall, we see that existing models struggle with these controlled variations across the board, where the largest difficulties appear on the logical consistency measured with antonymous perturbations and concept mention consistency measured with synonymous perturbations. To our knowledge, this is the first study of VQA robustness analysis on different types of linguistic variations,

2 Improving Model Robustness using Q3R

We evaluate the effect of adding Q3R to different NMN architectures. \tabreftab:regularization_p2_acc shows that, on VQA P2, adding our framework (+Q3R) results in significant improvements over the base models for both accuracy and CS. We also measure validation accuracy on VQA v2 in \tabreftab:regularization_p2_acc, which shows modest yet consistent gains, although our focus is on consistency, not overall VQA v2 accuracy. This shows that models regularized by Q3R generally see improvements in consistency and overall performance, including on VQA v2.

Analysis by Perturbation Types. A benefit of having information on the specific types of linguistic variations is that we can profile a model’s performance by each type to assist understanding and diagnosis. \tabreftab:perturbation_types shows that models are generally more confident with the antonymous perturbations, likely because “yes/no” questions have higher answer prediction scores in general. We note that performance gains are significant on single-word changes and less obvious on multi-word changes.

3 Further Analysis

Additional Evaluation. As we discussed in \secrefsec:othermethod, human-written rephrasings typically contain various sources of change: such as those that involve common sense knowledge, or structural level change of sentences. Nonetheless, we are interested in observing the outcome of the additional study on VQA-Rephrasings. When trained with our Q3R framework, XNM receives a +0.7+0.7 accuracy improvement on paraphrased questions as well as a +0.4+0.4 improvement on CS score. This moderate gain may be explained by the fact that linguistic perturbations are present in human-written questions, so the model’s robustness on this dataset benefits from our training framework.

4 Discussion

Comparison to Expert Layout Supervision. Many NMNs adopt expert layouts to guide the search for optimal module paths . While this approach is suitable on synthetic datasets with simple scenes and spatial reasoning , it has limited success on realistic images . Our loss can be viewed as providing weak supervision to module layouts, avoiding the need for ground-truth module layout annotation, which is costly and not clearly defined for natural questions about real images.

Beyond NMNs. Our results seem to suggest the advantage of representing and incorporating inductive bias at the modular level, rather than just using answer-level supervision. While we have demonstrated our method using Neural Module Networks, we note that it could be generalized to improve the robustness of any other interpretable VQA architecture that involves the computation of sub-tasks, such as those based on executable symbolic programs .

Conclusion

We show that a promising direction to improve the robustness and consistency of VQA models is by modeling and learning from lexical perturbations. We propose a novel approach based on modular networks, which creates two questions related by linguistic perturbation and regularizes the visual reasoning process between them to be consistent during training. We introduce a new benchmark, VQA P2, that features categorized, controllable linguistic variations that allows us to investigate and diagnose sources of inconsistencies in model predictions, will be made publicly available. Empirical results show that existing models have difficulties with different types of linguistic variations and that our approach is effective towards improving robustness and generalization ability.

References

Appendix A Implementation Details

As noted in \secrefsec:experiments, all models use the same visual features from https://github.com/peteanderson80/bottom-up-attention and same pre-trained GloVe embeddings from Common Crawl 840B: https://nlp.stanford.edu/projects/glove/. XNM, StackNMN, and HybridNet use the implementation provided by to ensure consistency amongst different NMN models.https://github.com/shijx12/XNM-Net We implement the modules and controller of StackNMN to match the paper description and official implementation. All NMN models are trained with the Adam optimizer and have the same learning rate of 0.0008, training initialization seed of 0, and batch size of 256. Following their implementations, we use hidden dimension sizes of 512 for StackNMN and 1024 for XNM, while we use 1024 for HybridNet to match XNM. We use the recommended number of reasoning steps, T=3T=3, for XNM and use the same for StackNMN. HybridNet uses T=4T=4 to test longer reasoning sequences as well as for visualization purposes. With our Q3R framework, λ=1.0\lambda=1.0 and r=0.2r=0.2 for all experiments. For BAN, we use the original source code with the 8-glimpse model, provided by the authors and adopt their training settings.https://github.com/jnhwkim/ban-vqa We do not use the counting module nor additional training data from Visual Genome . For the Transformer model, we use LXMERT as the architecture and utilize the publicly available code with the recommended settings from the authors.https://github.com/airsplay/lxmert For fair comparison, we do not use large-scale pre-training.

Appendix B Effect of Hyper-parameters

The main hyperparameters involved with using Q3R are the weight parameter, λ\lambda, and the distance function, D(⋅)D(\cdot). Changing λ\lambda from 1 to 0.2 results in a value change of <0.1<0.1 on VQA v2 accuracy and <0.1<0.1 on CS score on VQA P2, for StackNMN. Also, the difference of applying KL-divergence based loss function and L1-norm based loss function results is also <0.1<0.1, for both VQA v2 accuracy and CS score on VQA P2.

Appendix C Additional Dataset Information

To measure the model’s robustness against question variations, it requires the availability of answer annotations during evaluation. Since the testing data of VQA v2 does not have public groundtruth information, we use the validation split of VQA v2 as the testing set for all models. For training, we use the training data of VQA v2, which contains 443,757443,757 questions for 82,78382,783 images; for testing, we use the validation split of VQA v2 (214,354214,354 questions for 40,50440,504 images) and VQA P2 (36,17536,175 question-answer pairs)

In \secrefsubsec:rephrasings, we evaluate our framework on VQA-Rephrasings . VQA-Rephrasings is a human-written paraphrased dataset that spans the 40,50440,504 validation images of VQA v2, where each image has a corresponding question group that contains 11 original question from the VQA v2 validation set and 33 rephrasings of that questions, for a total of 162,016162,016 questions.

Appendix D Model Details

Our proposed training framework is applicable to the any controller-module framework . Next, we provide details of the components of our backbone architecture, which is comprised of input encoders, a controller, and a set of functional modules.

D.2 Models

The backbone architectures can be instantiated differently to realize a variety of models with distinct module functionalities, reasoning steps, feature backbones, etc. In this work, we adopt three designs of the controller-module models to test the effects of our Q3R training framework.

StackNMN . For StackNMN, the original method uses grid features from a CNN as visual features. For fair comparison with other models, we use object features for StackNMN, which is a stronger feature backbone for VQA. The modules in this architecture largely use elementwise multiplications to fuse visual and linguistic features, compute different attention maps, or obtain answer vectors. Additionally, in Find and Transform, 1D convolutions are also used to compute weighted visual and multi-modal features. The specific module designs for StackNMN are shown in \tabreftab:modules.

HybridNet. To further investigate whether our framework can be effective regardless of the module implementations, we present a hybrid of StackNMN and XNM. Specifically, HybridNet utilizes the Transform module of StackNMN, while maintaining the rest of the design from XNM.

Appendix E Output Examples

Here are some example outputs from HybridNet with and without Q3R. In \figreffig:output_bear, the model trained without Q3R maintains its answers despite the perturbation, whereas the model trained with our framework predicts the appropriate answer and also matches visual attentions between them. Then, in \figreffig:output_scenetime, we see an example where the questions share the same meaning and the same modules are selected for both questions and both models, but the model without our framework yields inconsistent answers.