MUREL: Multimodal Relational Reasoning for Visual Question Answering

Remi Cadene, Hedi Ben-younes, Matthieu Cord, Nicolas Thome

Introduction

Since the success of Convolutional Neural Networks (ConvNets) at the ILSVRC 2012 challenge , Deep Learning has become the baseline approach for any computer vision problem. Beyond their outstanding performances for perception tasks, e.g. classification or detection, deep ConvNets have also been successfully used for new artificial intelligence tasks like Visual Question Answering (VQA) . VQA requires a high level understanding of images and questions, and is often considered to be a good proxy for visual reasoning. However, it is not straightforward to use ConvNets in a context where a high level of reasoning is required. The question of leveraging the perception power of deep CNNs for reasoning tasks is crucial if we want to go further in visual scene understanding .

It is also not trivial to define nor evaluate a model’s capacity to reason about the visual modality in VQA. To fill this need, synthetic datasets have been released, e.g. CLEVR , which specific structure controls the exact reasoning primitives required to give the answer . However, methods that tackle the VQA problem on real data struggle to integrate this explicit reasoning procedure. Instead, state-of-the-art methods often rely on the much simpler attentional framework . Despite its effectiveness, this mechanism restricts visual reasoning to a soft selection of regions that are relevant to answer the question. This arguably limits the modeling power of such models to bridge the gap between the perceptual strengths of ConvNets and the high-level reasoning demand for VQA.

In this paper, we propose MuRel, a multimodal relational network that goes one step further towards reasoning about questions and images. Our first contribution is to introduce the MuRel cell, an atomic reasoning primitive enabling to represent rich interactions between question and image regions. It is based on a vectorial representation that explicitly models relations between regions. Our second contribution is to embed this MuRel cell into an iterative reasoning process, which progressively refines the internal network representation to answer the question. The rationale of MuRel is illustrated in Figure 1: for the question ”what is she eating”, our model focuses on two main regions (the head and the donut) with important visual cues and semantic relations between them to provide the correct answer (”donut”). The visual reasoning of our MuRel system is formed by this multi-step relational module that discards useless information to focus on the relevant regions.

In the experiments, we show additional results for explaining the behaviour of MuRel. We also provide various ablative studies to validate the relevance of the MuRel cell and the iterative reasoning process, and show that MuRel is highly competitive or even outperforms state-of-the-art results on three of the most common VQA datasets: the VQA 2.0 dataset , VQA-CP v2 and TDIUC .

Related work and contributions

Recently, the deep learning community started to tackle complex visual reasoning problems such as relationship detection , object recognition , abstract reasoning or visual causality , while more theoretical work attempt to formalize relational reasoning .

But the most popular image reasoning task is certainly Visual Question Answering (VQA), which has been a hot research topic for the last five years . Since the seminal work of , different sub-problems have been identified for the resolution of VQA. In particular, explicit reasoning techniques have been developed relying on the synthetic CLEVR dataset . Meanwhile, real-data VQA systems are the test bed for more practical approaches based on high quality visual representations or multimodal fusion schemes.

The research efforts towards VQA models that are able to reason about a visual scene is mainly conducted using the CLEVR dataset . This artificial dataset provides questions that require spatial and relational reasoning on simple images coming from a visual world with low variability. An important line of work attempts to solve this task through explicit reasoning. In such methods , a neural network reads the question and generates a program, corresponding to a graph of elementary neural operations that process the image. However, there are two major downsides to these techniques. First, their performance strongly depends on whether or not program annotations are used to learn the program generator; and second, they can be matched or surpassed by simpler models that implicitly learn to reason without requiring program annotation. In particular, FiLM modulates the visual feature map with an affine transformation whose parameters depend on the question. In more recent work, the MAC network draws inspiration from the Model-View-Controller paradigm to design the trainable MAC cell on which the network iterates. Finally, in , they reason over all the possible pairs of objects in the picture, thus introducing relationship modeling in visual question answering.

VQA on real data

An important part of the research in VQA is focused on designing functions that can represent high-level correlations between two vector spaces. Among these multimodal fusion algorithms, the most effective ones use second order (or higher ) interactions, made tractable through sketching methods , or with more success using the tensor decomposition framework .

This line of work is often considered orthogonal to visual reasoning contributions. In a setup involving real data, complex methods such as explicit or relational reasoning are much more challenging to implement than with artificial images and questions. This is certainly why the most widely used reasoning framework involves soft attention mechanisms . Given a question, these models assign an importance score to each region, and use them to weight-sum pool the visual representations. Multiple attention maps (also called glimpses) can even be computed in parallel or sequentially . More complex attention strategies have been explored, such as the Structured Attention , where a locally-connected graphical structure is considered to infer the region saliency scores. also leverages a graphical structure between regions to address weaknesses of the soft-attention mechanism, improving the VQA model’s ability to count. In , the image representation is computed using pairwise semantic attention and spatial graph convolutions. The soft attention framework is questioned in , where regions are hardly selected based on the norm of their feature. Finally, recent work of simultaneously attends over regions and word tokens through a bilinear attention network.

Importantly, the type of visual features used to feed the VQA system has an large impact on performance. While early work have been using fixed-grid representation given by a fully-convolutional network (such as ResNet-152 ), performance can be improved using predictions from an object detector . Recently, a crucial component in the VQA Challenge 2018 winning entry was the mix of multiple types of visual features .

MuRel contributions

In this work, we move away from the classical attention framework widely used in real-data VQA systems. Instead, we use a vectorial representation, more expressive than scalar attention maps, to model the semantic interaction between each region’s visual content and the question. In addition, we include a notion of spatial and semantic context in the representations by representing pairs of image regions through interactions between their visual embeddings and spatial coordinates. Differently than the approach followed in where a locally connected graph structure is built, we use the relations between all possible pairs of regions.

Our MuRel network embodies an iterative process with inspiration from works driven by the synthetic reasoning CLEVR dataset, e.g., MAC or FiLM , which we adapt to the real data VQA purpose. In particular, we improve the interactions between image regions and questions by using richer bilinear fusion models and by explicitly incorporating relations between regions.

MuRel approach

Our VQA approach is depicted in Figure 3. Given an image v∈Iv\in\mathcal{I} and a question q∈Qq\in\mathcal{Q} about this image, we want to predict an answer a^∈A\hat{a}\in\mathcal{A} that matches the ground truth answer a⋆a^{\star}. As very common in VQA, the prediction a^\hat{a} is given by classification scores:

In Section 3.1, we present the MuRel cell, a neural module that learns to perform elementary reasoning operations by blending question information into the set of spatially-grounded visual representations. Next, in Section 3.2, we leverage the power of this cell using the MuRel network, a VQA architecture that iterates through a MuRel cell to reason about the scene with respect to a question.

We want to include question information within each visual representation si\bm{s}_{i}. Multiple multimodal fusion strategies have been recently proposed to model the relevant interactions between two modalities. One of the most efficient technique is the one proposed by , based on the Tucker decomposition of third-order tensors. This bilinear fusion model learns to focus on the relevant correlations between input dimensions. It models rich and fine-grained multimodal interactions, while keeping a relatively low number of parameters. Each input vector si\bm{s}_{i} is fused with the question embedding q\bm{q} using the same bilinear fusion:

where Θ\Theta are the trainable parameters of the fusion module. Each dimension mm of mi\bm{m}_{i} can be written as a bilinear function in the form ∑s,qws,q,msisqq\sum_{s,q}w^{s,q,m}\bm{s}_{i}^{s}\bm{q}^{q}. Thanks to the Tucker decomposition, the tensor {ws,q,m}\{w_{s,q,m}\} is factorized into the list of parameters Θ\Theta. We set the number of dimensions in mi\bm{m}_{i} to dvd_{v} to facilitate the use of residual connections throughout our architecture.

In classical attention models, the fusion between image region and question features s\bm{s} and q\bm{q} only learns to encode whether a region is relevant. In the MuRel cell, the local multimodal information is represented within a richer vectorial form mi\bm{m}_{i} which can encode more complex correlations between both modalities. This allows to store more specific information about what precise characteristic of a particular region is important in a given textual context.

Pairwise interactions

To answer certain types of question, it can be necessary to reason over multiple object that interact together. More generally, we want each representation to be aware of the spatial and semantic context around it. Given that our features are structured as a bag of localized vectors , modeling the visual context of each region is not straightforward. Similarly to the recent work of , we opt for a pairwise relationship modeling where each region receives a message based on its relations to its neighbours. In their work, a region’s neighbours correspond to the KK most similar regions, whereas in the MuRel cell the neighbourhood is composed of every region in the image. Besides, instead of using scalar pairwise attention and graph convolutions with Gaussian kernels as they do, we merge spatial and semantic representations to build relationship vectors. In particular, we compute a context vector eˇi\check{\bm{e}}_{i} for every region. It consists in an aggregation of all the pairwise links ri,j\bm{r}_{i,j} coming into ii. We define it as eˇi=max⁡jri,j\check{\bm{e}}_{i}=\max_{j}\bm{r}_{i,j}, where ri,j\bm{r}_{i,j} is a vector containing information about the content of both regions, but also about their relative spatial positioning. We use the max⁡\max operator in the aggregation function to reduce the noise that can be induced by average or sum poolings, which oblige all the regions to interact with each other. To encode the relationship vector, we use the following formulation:

Through the B⁡(.,.;Θb)\operatorname*{B}(.,.;\Theta_{b}) operator, the cell is free to learn spatial concepts such as on top of, left, right, etc. In parallel, B⁡(.,.;Θs)\operatorname*{B}(.,.;\Theta_{s}) encodes correlations between multimodal vectors (si,sj)(\bm{s}_{i},\bm{s}_{j}), corresponding to semantic visual concepts conditionned on the question representation. By summing up both spatial and semantic fusions, the network can learn high-level relational concepts such as wear, hold, etc.

The context representation eˇi\check{\bm{e}}_{i} that contains an aggregation of the messages ri,j\bm{r}_{i,j} provided by its neighbours updates the multimodal vector mi\bm{m}_{i} in an additive manner:

This formulation of the pairwise modelling is actually closer to the Graph Networks , where the notion of relational inductive biases is formalized.

Finally, the MuREL cell’s output is computed as a residual function of its input, to avoid the vanishing gradient problem. Each visual feature si\bm{s}_{i} is updated as: s^i=si+xi\hat{\bm{s}}_{i}=\bm{s}_{i}+\bm{x}_{i}.

The chain of operations that updates the set of localized region embeddings {si}i∈[1,N]\{\bm{s}_{i}\}_{i\in[1,N]} using the multimodal fusion with q\bm{q} and the pairwise modeling operator is noted:

2 MuRel network

Mimicking a simple form of progressive reasoning, our model leverages the power of bilinear fusions to iteratively merge visual information into context-aware visual embeddings. As we can see in Figure 3, a MuRel cell iteratively updates the region state vectors {si}\{\bm{s}_{i}\}, each time refining the representations with contextual and question information. More specifically, for each step t=1..Tt=1..T where TT is the total number of steps, a MuRel cell processes and updates the state vectors following Equation (6):

The state vectors are initialized with the features outputted by the object detector; for each region ii, si0=vi\bm{s}_{i}^{0}=\bm{v}_{i}.

The MuRel network represents each region regarding the question, but also using its own visual context. This representation is done iteratively, through multiple steps of a MuRel cell. The residual nature of this module makes it possible to align multiple cells without being subject to gradient vanishing. Moreover, the weights of our model are shared across the cells, which enables compact parametrization and good generalization.

The scene representation s\bm{s} is merged with the question embedding q\bm{q} to compute a score for every possible answer y^=B⁡(s,q;Θy)\hat{\bm{y}}=\operatorname*{B}\left(\bm{s},\bm{q};\Theta_{y}\right). Finally, a^\hat{a} is the answer with maximum score in y^\hat{\bm{y}}.

Our model can also be leveraged to define visualization schemes finer than mere attention maps. Especially, we can highlight important relations between image regions for answering a specific question. At the end of the MuRel network, the visual features {siT}\{\bm{s}_{i}^{T}\} are aggregated using a max⁡\max operation, yielding a dv−d_{v}-dimensional vector s\bm{s}. Thus, we can compute a contribution map by measuring to what extent each region contributes to the final vector. To do so, we compute the point-wise c=arg⁡ ⁣max⁡i{siT}∈[1,N]dv\bm{c}=\arg\!\max_{i}\{\bm{s}_{i}^{T}\}\in[1,N]^{d_{v}}, and measure the occurrence frequency of each region in this vector c\bm{c}. This provides a value for each region that estimates its contribution to the final vector. Interestingly, this process can be done after each cell, and not exclusively at the last one. Intuitively, it measures what the contribution map would have been if the iterative process had stopped at this point. As we can see in Figures 1,3,5, these relevance scores match human intuition and can be used to explain the model’s decision, even if the network has not been trained with any selection mechanism.

Similarly, we are able to visualize the pairwise relationships involved in the prediction of the MuRel cell. The first step is to find i⋆i^{\star}, which is the region that is the most impacted by the pairwise modeling. It is the region such that ∥eˇixi∥2\|\frac{\check{\bm{e}}_{i}}{\bm{x}_{i}}\|_{2} is maximal (cf. Equation (4)). This bounding box is shown in green in all our visualizations. We then measure the contribution of every other region to i⋆i^{\star} using the occurrence frequencies in arg⁡ ⁣max⁡jri,j\arg\!\max_{j}\bm{r}_{i,j}. We show in red the regions whose contribution to i⋆i^{\star} is above a certain threshold (0.2 in our visualizations). If there is no such region, the green box is not shown.

Connection to previous work

We can draw a comparison between our MuRel network and the FiLM network proposed in . Beyond the fact that their model is built for the synthetic CLEVR dataset and ours processes real data, some connections can be found between both models. In their work, the image passes through multiple residual cells, whereas we only have one cell through which we iterate. In FiLM, the multimodal interaction is modeled with a feature-wise affine modulation, while we use a bilinear fusion strategy which seems better suited to real world data. Finally, both MuRel and FiLM leverage the spatial structure of the image representation to model the relations between regions. In FiLM, the image is represented with a fully-convolutional network which outputs a feature map disposed in a fixed spatial grid. With this structure on image features, the relations between regions are modeled with a 3×33\times 3 convolution inside each residual block. Thus, the representation of each region depends on its neighbours in the locally-connected graph induced by the fixed grid structure. In our MuRel network, the image is represented as a set of localized features. This makes the relational modeling non trivial. As we want to model relations between regions that are potentially far apart, we consider that the set of regions forms a complete graph, where each region is connected to all the others.

Experiments

Datasets: We validate the benefits of the MuRel cell and the MuRel network on three recent datasets. VQA 2.0 is the most used dataset. It comes with a training set, a validation set and an online testing set. We provide a fine grained analysis on the validation set, while we compare MuRel to the state-of-the-art models on the testing set. Then, we use VQA Changing Priors v2 to demonstrate the generalization capacity of MuRel. VQA-CP v2 uses the same data as in VQA 2.0, but proposes different distribution of answers per question between training and validation splits. Finally, we use the TDIUC dataset to construct a more detailed analysis of our model’s performance on 12 well-defined types of question. TDIUC is currently the biggest dataset for visual question answering.

Hyper-parameters: We use standard features extraction, preprocessings and loss function . We use the recent Bottom-up features provided by to represent our image as a set of 36 localized regions. For the question embedding, we use the pretrained Skip-thought encoder from . Inspired by recent works, we use Adam as optimizer with a learning scheduler . More details about the experimental setup are given in appendix.

2 Model validation

We compare MuRel against models trained on the same Bottom-up features which are required to reach the best performances.

In Table 1, we compare MuRel against a strong attentional model based on bilinear fusions , which encompasses a multi-glimpses attentional process . The goal of this experiments is to compare our approach with strong baselines for real VQA in controlled conditions. In addition to using the same bottom-up features, which are crucial for fair comparisons, we also dimension the attention-based baseline to have an equivalent amount of learned parameters than MuRel (∼\sim60 millions including those from the GRU encoder). Also, we train it following the same experimental setup to insure competitiveness. MuRel reaches a higher accuracy on the three datasets. We report a significant gain of +1.70 on VQA 2.0 and +1.50 on VQA CP v2. Not only these results validate the ability of MuRel to better model interactions between the question and the image, but also to generalize when the distribution of the answers per question are completely different between the training and validation set as in VQA CP v2. A gain of +1.24 on TDIUC demonstrates the richer modeling capacity of MuRel in a fine-grained context of 12 well delimited question types.

Ablation study

In Table 2, we compare three ablated instances of MuRel to its complete form. First, we validate the benefits of the pairwise module. Adding it to a vanilla MuRel without iterative process leads to higher accuracy on every datasets. In fact, between line 1 and 2, we report a gain of +0.44 on VQA 2.0, +0.24 on VQA CP v2 and +0.36 on TDIUC. Secondly, we validate the interest of the iterative process. Between line 1 et 3, we report a gain of +0.59 on VQA 2.0, +0.49 on VQA CP v2 and +0.42 on TDIUC. Notably, this modification does not add any parameters, because we iterate over a single MuRel cell. Unsharing the weights by using a different MuRel cell for each step gives similar results. Finally, the pairwise module and the iterative process are added to create the complete MuRel network. This instance (in line 4) reaches the highest accuracy on the three datasets. Interestingly, the gains provided by the combination of the two methods are sometimes larger than those of each one separately. For instance, we report a gain of +1.01 on VQA 2.0 between line 1 and 4. This attests to the complementary of the two modules.

Number of reasoning steps

In Figure 4, we perform an analysis of the iterative process. We train four different MuRel networks on the VQA 2.0 train split, each with a different number of iterations over the MuRel cell. Performance is reported on val split. Networks with two and three steps respectively provides a gain of +0.30 and +0.57 in overall accuracy on VQA 2.0 over the network with a single step. An interesting aspect of the iterative process of MuRel is that the four networks have exactly the same amount of parameters, but the accuracy significantly varies with respect to the number of steps. While the accuracy for the answer type involving numbers keeps increasing, we report a decrease in overall accuracy at four reasoning steps. Counting is a challenging task: not only does the model need to detect every occurrence of the desired object, but also the representation computed after the final aggregation must keep the information of the number of detected instances. The complexity of this question may require deeper relational modeling, and thus benefit from a higher number of iterations over the MuRel cell.

3 State of the art comparison

In Table 3, we compare MuRel to the most recent contributions on the VQA 2.0 dataset. For fairness considerations, all the scores correspond to models trained on the VQA 2.0 train+val split, using the Bottom-up visual features . Interestingly, our model surpasses both MUTAN and MLB , which correspond to some of the latest development in visual attention and bilinear models. This tends to indicate that VQA models can benefit from retaining local information in mulitmodal vectors instead of scalar coefficients. Moreover, our model greatly improves over the recent method proposed in where the regions are structured using pairwise attention scores, which are leveraged through spatial graph convolutions. This shows the interest of our spatial-semantic pairwise modeling between all possible pairs of regions. Finally, even though we did not extensively tune the hyperparameters of our model, our overall score on the test-dev split is highly competitive with state-of-the-art methods. In particular, we are comparable to Pythia who won the VQA Challenge 2018. Please note that they improve their overall scores up to 70.01% when they include multiple types of visual features and more training data. Also, we did not report the score of 69.52% obtained by BAN as they train their model on extra data from the Visual Genome dataset .

TDIUC

One of the core aspect of VQA models lies in their ability to address different tasks. The TDIUC dataset enables a detailed analysis of the strengths and limitations of a model by evaluating its performance on different types of question. We show in Table 4 a detailed comparison of recent models to our MuRel. We obtain state-of-the-art results on the Overall Accuracy and the arithmetic mean of per-type accuracies (A-MPT), and surpass by a significant margin the second best model proposed by . Interestingly, we improve over this model even though it uses a combination of Bottom-up and fixed-grid features, as well as a supervision on the question types (hence its 100% result on the Absurd task). MuRel notably surpasses all previous methods on the Positional reasoning (+5.9 over MCB), Counting (+8.53 over QTA) questions. These improvements are likely due to the pairwise structure induced within the MuRel cell, which makes the answer prediction depend on the spatial and semantic relations between regions. The effectiveness of our per-region context modelling is also demonstrated by our the improvement on Scene recognition questions. For these questions, representing the image as a collection of independent objects shows lower performance than replacing each of them in its spatial and semantic context. Interestingly, our results on the harmonic mean of per-type accuracies (H-MPT) are lower than state-of-the-art. For MuRel, this harmonic metric is significantly harmed by our low score of 21.43% on the Utility and Affordances task. As these questions concern the possible usages of objects present in the scene (such as Can you eat the yellow object?), and are not directly related to the visual understanding of the scene.

VQA-CP v2

This dataset has been proposed to evaluate and reduce the question-oriented bias in VQA models. In particular, the distributions of answers with respect to question types differ from train to val splits. In Table 5, we report the scores of two recent baselines , on which we improve significantly. In particular, we demonstrate an important gain over GVQA , whose architecture is designed to focus on Yes/No questions. However, since both methods do not use the Bottom-up features, the fairness of the comparison can be questioned. So we also train an attention model similar to using these Bottom-up region representation. We observe that MuRel provides a substantial gain over this strong attention baseline. Given the distribution mismatch between train and val splits, models that only focus on linguistic biases to answer the question are systematically penalized on their val scores. This property of VQA-CP v2 implies that the pairwise iterative structure of MuRel is less prone to question-based overfitting than classical attention architectures.

4 Qualitative results

In Figure 5 we illustrate the behaviour of a MuRel network with three shared cells. Iterations through the MuRel cell tend to gradually discard regions, keeping only the most relevant ones. As explained in Section 3.2, the regions that are most involved in the pairwise modeling process are shown in green and red. Both region contributions and pairwise links match human intuition. In the first row, the most relevant relations according to our model are between the player’s hand, containing the WII controller, and the screen, which explains the prediction bowling. In the third row, the model answers kite using the relation between the man’s hand and the kite he is holding. Finally, in the last row, our model is able to address a third question on the same image than in Figure 1 and 3. Here, the relation between the head of the woman and her hat is used to provide the right answer. As VQA models are often subject to linguistic bias , this type of visualization shows that the MuRel network actually relies on the visual information to answer questions.

Conclusion

In this paper, we introduced MuRel, a multimodal relational network for Visual Question Answering task. Our system is based on rich representations of visual image regions that are progressively merged with the question representation. We also included region relations with pairwise combinations in our fusion, and the whole system can be leveraged to define visualization schemes helping to interpret the decision process of MuRel.

We validated our approach on three challenging datasets: VQA 2.0, VQA-CP v2 and TDIUC. We exhibited various ablation studies, clearly demonstrating the gain of our vectorial representation to model the attention, the use of pairwise combination, and the multi-step iterations in the whole process. Our final MuRel network is very competitive and outperforms state-of-the-art results on two of the most widely used datasets.

References