RUBi: Reducing Unimodal Biases in Visual Question Answering

Remi Cadene, Corentin Dancette, Hedi Ben-younes, Matthieu Cord, Devi Parikh

Introduction

The recent Deep Learning success in computer vision and natural language understanding allowed researchers to tackle multimodal tasks that combine visual and textual modalities . Among these tasks, Visual Question Answering (VQA) attracts increasing attention. The goal of the VQA task is to answer a question about an image. It requires a high-level understanding of the visual scene and the question, but also to ground the textual concepts in the image and to use both modalities adequately. Solving the VQA task could have tremendous impacts on real-world applications such as aiding visually impaired users in understanding their physical and online surroundings, searching through large quantities of visual data via natural language interfaces, or even communicating with robots using more efficient and intuitive interfaces.

Several large real image VQA datasets have recently emerged . Each one of them targets specific abilities that a VQA model would need to be used in real-world settings such as fine-grained recognition, object detection, counting, activity recognition, commonsense reasoning, etc. Current end-to-end VQA models achieve impressive results on most of these benchmarks and are even able to surpass the human accuracy on a specific benchmark accounting for compositional reasoning . However, it has been shown that they tend to exploit statistical regularities between answer occurrences and certain patterns in the question . While they are designed to merge information from both modalities, in practice they often answer without considering the image modality. When most of the bananas are yellow, a model does not need to learn the correct behavior to reach a high accuracy for questions asking about the color of bananas. Instead of looking at the image, detecting a banana and assessing its color, it is much easier to learn from the statistical shortcut linking the words what, color and bananas with the most occurring answer yellow.

One way to quantify the amount of statistical shortcuts from each modality is to train unimodal models. For instance, a question-only model trained on the widely used VQA v2 dataset predicts the correct answer approximately 44% of the time over the test set. VQA models are not discouraged to exploit these statistical shortcuts from the question modality, because their training set often follows the same distribution as their testing set. However, when evaluated on a test set that displays different statistical regularities, they usually suffer from a significant drop in accuracy . Unfortunately, these statistical regularities are hard to avoid when collecting real datasets. As illustrated in Figure 1, there is a crucial need to develop new strategies to reduce the amount of biases coming from the question modality in order to learn better behaviors.

We propose RUBi, a training strategy to reduce the amount of biases learned by VQA models. Our strategy reduces the importance of the most biased examples, i.e. examples that can be correctly classified without looking at the image modality. It implicitly forces the VQA model to use the two input modalities instead of relying on statistical regularities between the question and the answer. We take advantage of the fact that question-only models are by design biased towards the question modality. We add a question-only branch on top of a base VQA model during training only. This branch influences the VQA model, dynamically adjusting the loss to compensate for biases. As a result, the gradients backpropagated through the VQA model are reduced for the most biased examples and increased for the less biased. At the end of the training, we simply remove the question-only branch.

We run extensive experiments on VQA-CP v2 and demonstrate the ability of RUBi to surpass current state-of-the-art results from a significant margin. This dataset has been specifically designed to assess the capacity of VQA models to be robust to biases by the question modality. We show that our RUBi learning framework provides gains when applied on several VQA architectures such as Stacked Attention Networks and Top-Down Bottom-Up Attention . We also show that RUBi is competitive on the standard VQA v2 dataset when compared to approaches that reduce unimodal biases.

Related work

Real-world datasets display some form of inherent biases due to their collection process . As a result, machine learning models tend to reflect these biases because they capture often undesirable correlations between the inputs and the ground truth annotations . Procedures exist to identify certain kinds of biases and to reduce them. For instance, some methods are focused on gender biases , some others on the human reporting biases , and also on the shift in distribution between lab-curated data and real-world data . In the language and vision context, some works evaluate unimodal baselines or leverage language priors . In the following, we discuss about related works that assess and reduce unimodal biases learned by VQA models.

Despite being designed to merge the two input modalities, it has been found that VQA models often rely on superficial correlations between inputs from one modality and the answers without considering the other modality . An interesting way to quantify the amount of unimodal biases that can potentially be learned by a VQA model consists in training models using only one of the two modalities . The question-only model is a particularly strong baseline because of the large amount of statistical regularities that can be leveraged from the question modality. With the RUBi learning strategy, we take advantage of this baseline model to prevent VQA models from learning question biases.

Unfortunately, biased models that exploit statistical shortcuts from one modality usually reach impressive accuracy on most of the current benchmarks. VQA-CP v2 and VQA-CP v1 were recently introduced as diagnostic datasets containing different answer distributions for each question-type between train and test splits. Consequentially, models biased towards the question modality fail on these benchmarks. We use the more challenging VQA-CP v2 dataset extensively in order to show the ability of our approach to reduce the learning of biases coming from the question modality.

Balancing datasets to avoid unimodal biases

Once the unimodal biases have been identified, one method to overcome these biases is to create more balanced datasets. For instance, the synthetic datasets for VQA minimize question-conditional biases via rejection sampling within families of related questions to avoid simple shortcuts to the correct answer.

Doing rejection sampling in real VQA datasets is usually not possible due to the cost of annotations. Another solution is to collect complementary examples to increase the difficulty of the task. For instance, VQA v2 has been introduced to weaken language priors in the VQA v1 dataset by identifying complementary images. For a given VQA v1 question, VQA v2 also contains a similar image with a different answer to the same question. However, even with this additional balancing, statistical biases from the question remain and can be leveraged . That is why we propose an approach to reduce unimodal biases during training. It is designed to learn unbiased models from biased datasets. Our learning strategy dynamically modifies the loss values to reduce biases from the question. By doing so, we reduce the importance of certain examples, similarly to the rejection sampling approach, while increasing the importance of complementary examples which are already in the training set.

Architectures and learning strategies to reduce unimodal biases

In parallel of these previous works on balancing datasets, an important effort has been carried out to design VQA models to overcome biases from datasets. proposed a hand-designed architecture called Grounded VQA model (GVQA). It breaks the task of VQA down into a first step of locating and recognizing the visual regions needed to answer the question, and a second step of identifying the space of plausible answers based on a question-only branch. This approach requires training multiple sub-models separately. In contrast, our learning strategy is end-to-end. Their complex design is not straightforward to apply on different architectures while our approach is model-agnostic. While we rely on a question-only branch, we remove it at the end of the training.

The work most related to ours in terms of approach is . The authors propose a learning strategy to overcome language priors in VQA models. They first introduce an adversary question-only branch. It takes as input the question encoding from the VQA model and produces a question-only loss. They use a gradient negation of this loss to discourage the question encoder to capture unwanted biases that could be exploited by the VQA model. They also propose a loss based on the difference of entropies between the VQA model and the question-only branch output distributions. These two losses are only backpropagated to the question encoder. In contrast, our learning strategy targets the full VQA model parameters to reduce the impact of unwanted biases more effectively. Instead of relying on these two additional losses, we use the question-only branch to dynamically adapt the value of the classification loss in order to reduce the learning of biases in the VQA model. A visual comparison between and RUBi can be found in Figure 5 in the supplementary materials.

Reducing Unimodal Biases Approach

Each one of them can be defined to instantiate most of the state of the art models, such as to cite a few.

The classical learning strategy of VQA models, depicted in Figure 2, consists in minimizing the standard cross-entropy criterion over a dataset of size nn.

VQA models are inclined to learn unimodal biases from the datasets . This can be shown by evaluating models on datasets that have different distributions of answers for the test set, such as VQA-CP v2. In other words, they rely on statistical regularities from one modality to provide accurate predictions without having to consider the other modality. As an extreme example, strongly biased models towards the question modality always output yellow to the question what color is the banana. They do not learn to use the image information because there are too few examples in the dataset where the banana is not yellow. Once trained, their inability to use the two modalities adequately makes them inoperable on data coming from different distributions such as real-world data. Our contribution consists in modifying this cost function to avoid the learning of these biases.

1 RUBi learning strategy

During training, the branch acts as a proxy preventing any VQA model of the form presented in Equation (1) from learning biases. At the end of the training, we simply remove the branch and use the predictions from the base VQA model.

Preventing biases by masking predictions

Before passing the predictions of our base VQA model to the loss function defined in Equation (2), we merge them with a mask of length ∣A∣|\mathcal{A}| containing a scalar value between 0 and 1 for each answer. This mask is obtained by passing the output of the neural network nnq\mathit{nn}_{q} through a sigmoid function σ\sigma. The goal of this mask is to dynamically alter the loss by modifying the predictions of the VQA model. To obtain the new predictions, we simply compute an element-wise product ⊙\odot between the mask and the original predictions as defined in the following equation.

Our method modifies the predictions in this specific way to prevent the VQA model to learn biases from the question. To better understand the impact of our approach on the learning, we examine two scenarios. First, we reduce the importance of the most biased examples, i.e. examples that can be correctly classified without using the image modality. To do so, the question-only branch outputs a mask to increase the score of the correct answer while decreasing the scores of the others. As a result, the loss is much lower for these biased examples. In other words, the gradients backpropagated through the VQA model are smaller, thereby reducing the importance of these examples in the learning. As illustrated in the first row of Figure 3, given the question what color is the banana, the mask takes a high value of 0.8 for the answer yellow which is the most likely answer for this question in the training set. On the other hand, the value for the other answers green and white are smaller. We see that the mask influences the VQA model to produce new predictions where the score associated with the answer yellow increases from 0.8 to 0.94. Compared to the classical learning approach, the loss is smaller with RUBi and decreases from 0.22 to 0.06. Secondly, we increase the importance of examples that cannot be answered without using both modalities. For these examples, the question-only branch outputs a mask that increases the score of the wrong answer. As a result, the loss is much higher and the VQA model is encouraged to learn from these examples. We illustrate this behavior in the second row of Figure 3 for the same question about the color of the banana. When the image contains a green banana, RUBi increases the loss from 0.69 to 1.20.

Joint learning procedure

We jointly optimize the parameters of the base VQA model and its question-only branch using the gradients computed from two losses. The main loss LQM\mathcal{L}_{QM} refers to the cross-entropy loss associated with the predictions of fQM(vi,qi)f_{QM}(v_{i},q_{i}) from Equation 4. We backpropagate this loss to optimize all the parameters θQM\theta_{QM} which contributed to this loss. θQM\theta_{QM} is the union of the parameters of the base VQA model, the encoders, and the neural network nnq\mathit{nn}_{q} of the question-only branch. In our setup, we share the parameters of the question encoder eqe_{q} between the VQA model and the question-only branch. The question-only loss LQO\mathcal{L}_{QO} is a cross-entropy loss associated with the predictions of fQ(qi)f_{Q}(q_{i}) from Equation 3. We use this loss to only optimize θQO\theta_{QO}, union of the parameters of cqc_{q} and nnq\mathit{nn}_{q}. By doing so, we further improve the question-only branch ability to capture biases. Note that we do not backpropagate this loss to the question encoder eqe_{q} preventing it from directly learning question biases. We obtain our final loss LRUBi\mathcal{L}_{\text{RUBi}} by summing the two losses together in the following equation:

2 Baseline architecture

Experiments

We train and evaluate our models on VQA-CP v2 . This dataset was developed to evaluate the models robustness to question biases. We follow the same training and evaluation protocol as , who also propose a learning strategy to reduce biases. For each model, we report the standard VQA evaluation metric . We also evaluate our models on the standard VQA v2 . Further implementation details are included in the supplementary materials, as well as results on VQA-CP v1 and grounding experiments on VQA-HAT .

1 Results

In Table 1, we compare our approach consisting of our baseline architecture trained with RUBi on VQA-CP v2 against the state of the art. To be fair, we only report approaches that use the strong visual features from . We compute the average accuracy over 5 experiments with different random seeds. Our RUBi approach reaches an average overall accuracy of 47.11% with a low standard deviation of ±\pm0.51. This accuracy corresponds to a gain of +5.94 percentage points over the current state-of-the-art UpDn + Q-Adv + DoE. It also corresponds to a gain of +15.88 over GVQA , which is a specific architecture designed for VQA-CP. RUBi reaches a +8.65 improvement over our baseline model trained with the classical cross-entropy. In comparison, the second best approach UpDn + Q-Adv + DoE only achieves a +1.43 gain in overall accuracy over their baseline UpDn. In addition, our approach does not significantly reduce the accuracy over our baseline for the answer type Other, while the second best approach reduces it by 10.57 point.

Additional baselines

We compare our results to two sampling-based training methods. In the Balanced Sampling method, we sample the questions such that the answer distribution is uniform. In the Question-Type Balanced Sampling method, we sample the questions such that for every question type, the answer distribution is uniform, but the question type distribution remains the same overall Both methods are tested with our baseline architecture. We can see that the Question-Type Balanced Sampling improves the result from 38.46 in accuracy to 42.11. This gain is already +0.94 higher than the previous state of the art method , but remains significantly lower than our proposed method.

Architecture agnostic

RUBi can be used on existing VQA models without changing the underlying architecture. In Table 3, we experimentally demonstrate the generality and effectiveness of our learning scheme by showing results on two additional architectures, Stacked Attention Networks (SAN) and Bottom-Up and Top-Down Attention (UpDn) . First, we show that applying RUBi on these architectures leads to important gains over the baselines trained with their original learning strategy. We report a gain of +11.73 accuracy point for SAN and +4.5 for UpDn. This lower gap in accuracy may show that UpDn is less driven by biases than SAN. This is consistent with results from . Secondly, we show that these architectures trained with RUBi obtain better accuracy than with the state-of-the-art strategy from . We report a gain of +3.4 with SAN + RUBi over SAN + Q-Adv + DoE, and +3.06 with UpDn + RUBi over UpDn + Q-Adv + DoE. Full results splitted by question type are available in the supplementary materials.

Impact on VQA v2

We report the impact of our method on the standard VQA v2 dataset in Table 3. VQA v2 train, val and test sets follow the same distribution, contrarily to VQA-CP v2 train and test sets. In this context, we usually observe a drop in accuracy using approaches focused on reducing biases. This is due to the fact that exploiting unwanted correlations from the VQA v2 train set is not discouraged and often leads to a higher accuracy on the test set. Nevertheless, our RUBi approach leads to a comparable drop to what can be seen in the state-of-the-art. We report a drop of 1.94 percentage points with respect to our baseline, while report a drop of 3.78 between GVQA and their SAN baseline. report drops of 0.05, 0.73 and 2.95 for their three learning strategies with the UpDn architecture which uses the same visual features as RUBi. As shown in this section, RUBi improves the accuracy on VQA-CP v2 from a large margin, while maintaining competitive performance on the standard VQA v2 dataset compared to similar approaches.

Validation of the masking strategy

We compare different fusion techniques to combine the output of nnq\mathit{nn}_{q} with the output from the VQA model. We report a drop of 7.09 accuracy point on VQA-CP v2 by replacing the sigmoid with a ReLU on our best scoring model. Using an element-wise sum instead of an element-wise product leads to a further performance drop. These results confirm the effectiveness of our proposed masking method which relies on a sigmoid and an element-wise sum.

Validation of the question-only loss

In Table 4, we validate the ability of the question-only loss LQO\mathcal{L}_{QO} to reduce the question biases. The absence of LQO\mathcal{L}_{QO} implies that the question-only classifier cqc_{q} is never used, and nnq\mathit{nn}_{q} only receives gradients from the main loss LQM\mathcal{L}_{QM}. Using LQO\mathcal{L}_{QO} leads to consistent gains on all three architectures. We report a gain of +0.89 for our Baseline architecture, +0.22 for SAN, +4.76 for UpDn.

2 Qualitative analysis

To better understand the impact of our RUBi approach, we compare in Figure 4 the answer distribution on VQA-CP v2 for some specific question patterns. We also display interesting behaviors on some examples using attention maps extracted as in . In the first row, we show the ability of RUBi to reduce biases for the is this person skiing question pattern. Most examples in the train set have the answer yes, while in the test set, they have the answer no. Nevertheless, RUBi outputs 80% of no, while the baseline almost always outputs yes. Interestingly, the best scoring region from the attention map of both models is localized on the shoes. To get the answer right, RUBi seems to reason about the absence of skis in this region. It seems that our baseline gets it wrong by not seeing that the skis are not locked under the ski boots. This unwanted behavior could be due to the question biases. In the second row, similar behaviors occur for the what color are the bananas question pattern. 80% of the answers from the train set are yellow, while most of them are green in the test set. We show that the amount of green and white answers from RUBi are much closer to the ones from the test set than with our baseline. In the example, it seems that RUBi relies on the color of the banana, while our baseline misses it. In the third row, it seems that RUBi is able to ground the textual concepts such as top part of the fire hydrant and color on the right visual region, while the baseline relies on the correlations between the fire hydrant, the yellow color of its core and the answer yellow. Similarly on the fourth row, RUBi grounds color, star, fire hydrant on the right region, while our baseline relies on correlations between color, fire hydrant, the yellow color of the top part region and the answer yellow. Interestingly, there is no similar question that involves the color of a star on a fire hydrant in the training set. It shows the capacity of RUBi to generalize to unseen examples by composing and grounding existing visual and textual concepts from other kinds of question patterns.

Conclusion

We propose RUBi to reduce unimodal biases learned by Visual Question Answering (VQA) models. RUBi is a simple learning strategy designed to be model agnostic. It is based on a question-only branch that captures unwanted statistical regularities from the question modality. This branch influences the base VQA model to prevent the learning of unimodal biases from the question. We demonstrate a significant gain of +5.94 percentage point in accuracy over the state-of-the-art result on VQA-CP v2, a dataset specifically designed to account for question biases. We also show that RUBi is effective with different kinds of common VQA models. In future works, we would like to extend our approach on other multimodal tasks.

Acknowledgments

We would like to thank the reviewers for valuable and constructive comments and suggestions. We additionally would like to thank Abhishek Das and Aishwarya Agrawal for their help.

The effort from Sorbonne University was partly supported within the Labex SMART supported by French state funds managed by the ANR within the Investissements d’Avenir programme under reference ANR-11-LABX-65, and partly funded by grant DeepVision (ANR-15-CE23-0029-02, STPGP-479356-15), a joint French/Canadian call by ANR & NSERC.

References

Supplementary materials

In Table 5, we report results on the VQA-CP v1 dataset . Our RUBi approach consistently leads to significant gains over the classical learning strategy with a gain of +9.8 overall accuracy point with our baseline architecture, +19.2 with SAN and +7.66 with UpDn. Additionally, RUBi leads to a gain of +2.65 over the adversarial regularization method (AdvReg) from with SAN. A visual comparison between RUBi and can be found in Figure 5. Finally, all three architectures trained with RUBi reach a higher accuracy than GVQA which has been hand-designed to overcome biases.

Detailed results on VQA-CP v2

In Table 6, we report the full results of our experiments for SAN and UpDn architectures on the VQA-CP v2 dataset.

Quantitative study of the grounding ability on VQA-HAT

We conduct additional studies to evaluate the grounding ability of models trained with RUBi. We follow the experimental protocol of VQA-HAT . We train our models on VQA v1 train set and evaluate them using rank-correlation on the VQA-HAT val set, which is a subset of the VQA v1 val set. This metric compares attention maps computed from a model against human annotations indicating which regions humans found relevant for answering the question. In Table 7, we report a gain of +0.012 with our baseline architecture trained with RUBi, a gain of +0.019 with SAN and a loss of -0.003 with UpDn architecture. In future works, we would like to go beyond these early results in order to further evaluate the impact on grounding induced by RUBi.

Qualitative study of the grounding ability on VQA-HAT

We display in Figure 6 and Figure 7 some manually selected VQA triplets associated to the human attention maps provided by VQA-HAT and the attention maps computed from our baseline architecture when trained with and without RUBi. In Figure 6, we observe that the attention maps with RUBi are closer to the human attention maps than without RUBi. On the contrary, we observe in Figure 7 some failure to improve grounding ability.

2 Implementation details

We use the pretrained Faster R-CNN by to extract object features. We use the setup that extracts 36 regions for each image. We do not fine-tune the image extractor.

Question encoder

We use the same preprocessing as in . We apply a lower case transformation and remove the punctuation. We only consider the most frequent 3000 answers for both VQA v2 and VQA CP v2. We then use a pretrained Skip-thought encoder with a two-glimpses self attention mechanism. The final embedding is of size 4800. We fine-tune the question encoder during training.

Baseline architecture

Our baseline architecture is a simplified version of the MuRel architecture . First, it computes a bilinear fusion between the question vector and the visual features for each region. The bilinear fusion module is a BLOCK composed of 15 chunks, each of rank 15. The dimension of the projection space is 1000, and the output dimension is 2048. The output of the bilinear fusion is aggregated using a max pooling over nvn_{v} regions. The resulting vector is then fed into a MLP classifier composed of three layers of size (2048, 2048, 3000), with ReLU activations. It outputs the predictions over the space of the 3000 answers.

Question-only branch

The RUBi question-only branch feeds the question into a first MLP composed of three layers, of size (2048, 2048, 3000), with ReLU activations. First, this output vector goes through a sigmoid to compute the mask that will alter the predictions of the VQA model. Secondly, this same output vector goes through a single linear layer of size 3000. We use these question-only predictions to compute the question-only loss.

Optimization process

We train all our models with the Adam optimizer. We train our baseline architecture with the learning rate scheduler of . We use a learning rate of 1.5×10−41.5\times 10^{-4} and a batch size of 256. During the first 7 epochs, we linearly increase the learning rate to 6×10−46\times 10^{-4}. After epoch 14, we apply a learning rate decay strategy which multiplies the learning rate by 0.25 every two epochs. We train our models until convergence as we do not have a validation set for VQA-CP v2. For the UpDn and SAN architectures, we follow the optimization procedure described in .

Software and hardware

We use pytorch 1.1.0 to implement our algorithms in order to benefit from the GPU acceleration. We use four NVidia Titan Xp GPU in this study. We use a single GPU for each experiments. We use a dedicated SSD to load the visual features using multiple threads. A single experiment from Table 1 with the baseline architecture trained with or without RUBi takes less than five hours to run.