Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering

Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, Dacheng Tao

I Introduction

Thanks to recent advances of deep neural networks (DNN) in computer vision and natural language processing, computers are expected to be able to automatically understand the semantics of images and natural languages in the near future. Such advances also continue to redefine and drive research in image-text retrieval , image captioning , and visual question answering .

Compared with image-text retrieval and image captioning which just require the underlying algorithms to search or generate a free-form text description for a given image, visual question answering (VQA) is a more challenging task that requires fine-grained understanding of the semantics of both the images and the questions as well as supports complex reasoning to predict the best-matching answer accurately. In some aspects, the VQA task can be treated as a generalization of image captioning and image-text retrieval. Thus building effective VQA algorithms which can achieve performance that is close to that of human beings, is an important step towards the general artificial intelligence.

To support the VQA task, we need to address the following three issues effectively (see the example in Fig. 1): (1) extracting discriminative features for image and question representations; (2) combining the visual features from the image and the textual features from the question to generate the fused image-question features; (3) using the fused image-question features to learn a multi-class classifier for predicting the best-matching answer correctly. Deep neural networks (DNNs) are very effective and flexible, most of the existing VQA approaches tackle these three issues in one single DNNs model and train the model in an end-to-end fashion through back-propagation.

For feature-based image representation, directly using the global features extracted from the whole image may introduce noisy information (i.e., irrelevant features) that are irrelevant to the given question, e.g., the given question may strongly relate to only a small part of the image (i.e., image attention region) rather than the whole image. Therefore, it is intuitive to introduce visual attention mechanism into the VQA task to adaptively learn the most relevant image regions for a given question. Modeling visual attention may significantly improve performance . On the other hand, the questions interpreted in natural languages may also contain colloquialisms that can be treated as noise, thus it is very important to model the question attention simultaneously. Unfortunately, most existing approaches only model the image attention without considering the question attention. Motivated by these observations, we design a deep network architecture for the VQA task by using a co-attention learning module to jointly learn the attentions for both the image and the question, which may allow us to extract more discriminative features for image and question representations.

For multi-modal feature fusion, most existing approaches simply use linear models (e.g., concatenation or element-wise addition) to integrate the visual feature from the image with the textual feature from the question even their distributions may vary dramatically . Such linear models may not be able to generate expressive image-question features that are able to fully capture the complex correlations between multi-modal features. In contrast to linear pooling, bilinear pooling has recently been used to integrate different CNN features for fine-grained image recognition . Unfortunately, such bilinear pooling approach may output high-dimensional features for image-question representation and the underlying deep networks for feature extraction may contain huge number of model parameters, which may seriously limit its applicability for VQA. To tackle these problems effectively, Multi-modal Compact Bilinear (MCB) pooling and Multi-modal Low-rank Bilinear (MLB) pooling have been developed to reduce the computational complexity of the original bilinear pooling model and make it practicable for VQA. However, MCB needs very high-dimensional feature to guarantee good performance and MLB needs a great many training iterations to converge to a satisfactory solution. To tackle these problems, we propose a Multi-modal Factorized Bilinear pooling approach (MFB) which enjoys the dual benefits of compact output features of MLB and robust expressive capacity of MCB. Moreover, we extend the bilinear MFB model to a generalized high-order setting and proposed a Multi-modal Factorized High-order pooling (MFH) method to achieve more effective fusion of multi-modal features by exploiting their complex correlations sufficiently. By introducing more complex high-order interactions between multi-modal features, our MFH method can achieve more discriminative image-question representation and further result in significant improvement on the VQA performance.

For answer prediction, some datasets like VQA provide multiple answers for each image-question pair and such diverse answers are typically annotated by different users. As the answers are represented in natural languages, for a given question, different users may provide diverse answers or expressions which have same or similar meaning, thus such diverse answers may have strong correlations and they are not independent at all. For example, both a little dog and a puppy could be the correct answers for the same question. Motivated by these observations, it is important to design an appropriate mechanism to model the complex correlations between multiple diverse answers for the same question. In MCB , an answer sampling strategy was proposed to randomly pick an answer from a set of candidates during the training course. In this way, the complex correlations between multiple diverse answers could be eventually learned by the model with sufficient training iterations. In this paper, we formulate the problem of answer prediction as a label distribution learning problem. The answers for an image-question pair in the training dataset are converted to a probability distribution over all possible answers. We use the Kullback-Leibler divergence (KLD) as the loss function to achieve more accurate characterization of the consistency between the probability distribution of the predicted answers and the probability distribution of the ground truth answers given by the annotators. Compared with the answer sampling method in MCB , using the KLD loss can achieve faster convergence rate and obtain slightly better accuracy on answer prediction.

In summary, we have made the following contributions in this study:

A co-attention learning architecture is designed to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features (i.e., noisy information) effectively and obtain more discriminative features for image and question representations.

A Multi-modal Factorized Bilinear Pooling (MFB) approach is developed to achieve more effective fusion of the visual features from the image and the textual features from the question. By supporting more effective exploitation of the complex correlations between multi-modal features, our MFB approach can significantly outperform the existing bilinear pooling approaches.

A generalized Multi-modal Factorized High-order pooling (MFH) approach is developed by cascading multiple MFB blocks. Compared with MFB, MFH captures more complex correlations of multi-modal feature to achieve more discriminative image-question representation and further result in significant improvement on the VQA performance.

The KL divergence (KLD) is used as the loss function to achieve more accurate characterization of the consistency between the predicted answers and the annotated answers, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction.

Extensive experiments over multiple VQA datasets are conducted to explain the reason why our approaches are effective. Our experimental results demonstrate that: (a) our proposed approaches can achieve the state-of-the-art performance on the real-world VQA datasets; and (b) the normalization techniques are extremely important in bilinear pooling models.

The rest of the paper is organized as follows: In section II, we review the related work of VQA approaches, especially the ones introducing the bilinear pooling. In section III, we revisit the bilinear model and its factorized extension. Then, we propose the bilinear MFB model and reveal the fact that MFB is a generalization form of MLB. Based on MFB, we further propose its generalized high-order extension MFH. In section IV, we propose the co-attention learning network architecture for VQA based on MFB or MFH. In section V, we analyze the importance of modeling answer correlation in VQA and propose a solution with KLD loss. In section VI, we introduce our extensive experimental results for algorithm evaluation and multiple real-word VQA datasets are used to evaluate our proposed approaches. Finally, we conclude this paper in section VII.

II Related Work

In this section, we briefly review the most relevant research on VQA, especially those studies that use multi-modal bilinear models.

Malinowski et al. made an early attempt at solving the VQA task . Since then, solving the VQA task has received increasing attention from the communities of computer vision and natural language processing. Most existing VQA approaches can be classified into the following three categories: (a) the coarse joint-embedding models ; (b) the fine-grained joint-embedding models with attention ; (c) the external knowledge based models .

The coarse joint-embedding models are the most straightforward solution for VQA. Image and question are first represented as global features and then integrated to predict the answer. Zhou et al. proposed a baseline approach for the VQA task by using the concatenation of the image CNN features and the question BoW (bag-of-words) features, and a linear classifier is learned to predict the answer . Wang et al. perform a detailed analysis on the modeling of questions using CNN to obtain better question representations for VQA . Some approaches introduce more complex deep models, e.g., LSTM networks or residual networks , to tackle the VQA task in an end-to-end fashion.

One limitation of joint-embedding models is that their global features may contain noisy information (i.e., irrelevant features), and such noisy global features may not be able to answer the fine-grained problems correctly (e.g., “what color are the cat’s eyes?”) . Therefore, recent VQA approaches introduce the visual attention mechanism into the VQA task by adaptively learning the local fine-grained image features for a given question. Chen et al. proposed a question-guided attention map that projects the question embeddings to the visual space and formulates a configurable convolutional kernel to search the image attention region . Yang et al. proposed a stacked attention network to learn the attention iteratively . Some approaches introduce off-the-shelf object detectors or object proposals as the candidates of the attention regions and then use the question to identify the relevant ones. Fukui et al. proposed multi-modal compact bilinear pooling to integrate the visual features from the image spatial grids with the textual features from the questions to predict the attention . As the VQA task need to fully understand the semantic of the question in natural language, it is necessary to learn the textual attention for question simultaneously. Inspired by the works from the NLP community , some approaches perform attention learning on both the images and the questions. Lu et al. proposed a co-attention learning framework to alternately learn the image attention and the question attention . Nam et al. proposed a multi-stage co-attention learning model to refine the attentions based on memory of previous attentions .

Despite the joint embedding models can deliver impressive VQA performance, they are not good enough for answering the questions that require complex reasoning or knowledge of common senses. Therefore, introducing external knowledge is beneficial for VQA. However, existing approaches have either only been applied to specific datasets , or have been ineffective on benchmark datasets . Thus they still have rooms for further exploration and development.

II-B Multi-modal Bilinear Models for VQA

Multi-modal feature fusion plays a critical and fundamental role in VQA. After the image and the question representations are obtained, concatenation or element-wise summations are most frequently used for multi-modal feature fusion. Since the distributions of two feature sets in different modalities (i.e.,the visual features from images and the textual features from questions) may vary significantly, the representation capacity of the simply-fused features may be insufficient, limiting the final prediction performance.

Fukui et al. first introduced the bilinear model to solve the problem of multi-modal feature fusion in VQA . In contrast to the aforementioned approaches, they proposed the Multi-modal Compact Bilinear pooling (MCB), which uses the outer product of two feature vectors in different modalities to produce a very high-dimensional feature for quadratic expansion . To reduce the computational cost, they used a sampling-based approximation approach that exploits the property that the projection of two vectors can be represented as their convolution. The MCB model outperformed the simple fusion approaches and demonstrated superior performance on the VQA dataset . Nevertheless, MCB usually needs high-dimensional features (e.g., 16,000-D) to guarantee robust performance, which may seriously limit its applicability for VQA due to limitations in GPU memory.

III Generalized Multi-modal Factorized High-order Pooling

In this section, we first revisit the multi-modal bilinear models and then introduce the Multi-modal Factorized Bilinear pooling (MFB) model. We give detailed explanation on the implementation of our MFB model and further analyze its relationship with the existing MLB approach . By treating our MFB model as the basic building block, we extend the idea of bilinear pooling to a generalized high-order pooling and we further propose a Multi-modal High-order pooling (MFH) model by simply cascading multiple MFB blocks to capture more complex high-order interactions between multi-modal features.

Inspired by the matrix factorization tricks for uni-modal data , the projection matrix WiW_{i} in Eq.(2) can be factorized as two low-rank matrices:

Relationship to MLB. Eq.(4) shows that the MLB in Eq.(1) is a special case of the proposed MFB with k=1k=1, which corresponds to the rank-1 factorization. Figuratively speaking, MFB can be decomposed into two stages (see in Fig. 2(b)): first, the features from different modalities are expanded to a high-dimensional space and then integrated with element-wise multiplication. After that, sum pooling followed by the normalization layers are performed to squeeze the high-dimensional feature into the compact output feature, while MLB directly projects the features to the low-dimensional output space and performs element-wise multiplication. Therefore, with the same dimensionality for the output features, we can conjecture that MLB may suffer from insufficient representation.

III-B From Bilinear Pooling to Generalized High-order Pooling

From the previous work like , we have witnessed that the bilinear pooling models have superior representation capacity than the traditional linear pooling models. This inspires us that exploiting the complex interactions among the feature dimensions is beneficial for capturing the common semantics of multi-modal features . Therefore, a natural idea is to extend the second-order bilinear pooling to the generalized high-order pooling to further enhance the representation capacity of fused features. In this section, we introduce a generalized Multi-modal Factorized High-order pooling (MFH) model by cascading multiple MFB blocks.

As shown in Fig. 2(b), the MFB module can be separated into the expand stage and the squeeze stage as follows.

To make pp MFB blocks cascadable, we slightly modify the original MFBexp stage in Eq.(5) as follows:

The overall flowchart of the MFH approach is illustrated in Fig. 3. With the increase of pp, the model size and the dimensionality of the output feature for MFH grow linearly. In order to control the model complexity and the training time that we can afford, we use p<4p<4 in our experiments. It is worth noting that the propoed MFB model in section III-A is a special case of our MFHp model with p=1p=1.

IV Network Architectures for VQA

The goal of the VQA task is to answer a question about an image. The inputs to the model contain an image and a corresponding question about the image. Our model extracts the representations for both the image and the question, integrates multi-modal features by using the MFB or MFH module in Fig. 2(b), treats each individual answer as one class and performs multi-class classification to predict the correct answer. In this section, two network architectures are introduced. The first one is the baseline with one MFB or MFH module, which is used to perform ablation analysis with different hyper-parameters for comparison with other baseline approaches. The second one introduces co-attention learning to achieve more effective characterization of the fine-grained correlations between multi-modal features, which may result in a model with better representation capability.

The multi-modal features (that are extracted from the image and the question) are fed to the MFB or MFH module to generate the fused image-question feature zz. Finally, zz is fed to an NN-way classifier to predict the best-matching answer. Therefore, all the weights except the ones for the ResNet (due to the limitation of GPU memory) are optimized jointly in an end-to-end manner. The whole network architecture is illustrated in Fig. 4.

IV-B The Co-Attention Model

For a given image, different questions could result into an entire different set of answers. Therefore, an image attention model, which can predict the relevance between each spatial grid of the image with the question, is beneficial for predicting the best-matching answer accurately. From the results reported in MCB , one can see that incorporating such image attention mechanism allows the model to effectively learn which image region is important for the question, clearly contributing to better performance than the models without using attention. However, their attention model only focuses on learning the image attention while completely ignoring the question attention. Since the questions are interpreted in natural languages, the contribution of each word is definitely different. Therefore, we develop a co-attention learning approach named MFB+CoAtt or MFH+CoAtt (see Fig. 5) to jointly learn the attentions for both the question and the image.

Specifically, 14×\times14 (196) spatial grids of the image (res5c feature maps in ResNet) are used to represent the input image and TT output features from the LSTM networks are used to represent each word in the input question. After that, the TT question features are fed into a question attention module and output an attentive question representation. This attentive question representation is fed into an image attention module (with 196 image features), and MFB or MFH is used to generate a fused image-question representation. Such fused image-question representation is further used to learn a multi-class classifier for answer prediction. In our excrements, we find that using MFH rather than MFB in the image attention module does not improve the prediction accuracy significantly while inducing much higher computational cost. Therefore, in most of our experiments (unless in the final model ensemble experiment), the MFH module is only used in the feature fusion stage for integrating the attentive features extracted from the image and the question.

Both the image attention module and question attention module consist of sequential 1 ×\times 1 convolutional layers and ReLU layers followed by the softmax normalization layers to predict the attention weight for each input feature. The attentive feature are obtained by the weighted sum of the input features. To further improve the representation capacity of the attentive feature, multiple attention maps are generated to enhance the learned attention map, and these attention maps are concatenated to output the attentive image features.

It is worth noting that the question attention in our network architecture is learned in a self-attentive manner by using the question feature itself. This is different from the image attention module which is learned by using both the image features and question features. The reason is that we assume that the question attention (i.e., the key words of the question) can be inferred without seeing the image, as humans do.

V Answer Correlation Modeling

In most existing VQA approaches, the answering stage is formulated as a multi-class classification problem and each answer refers to an individual class. In practice, this assumption may not hold for the VQA task because the answers with the same or similar meaning can be expressed diversely by different annotators. For example, both the answers ‘a little dog’ and ‘a puppy’ could be correct for a given image-question pair. Therefore, it is crucial to model the answer correlations in the VQA task so that the learned model could be more robust.

Note that KL-divergence loss contains an additional constant term compared to the multi-label cross-entropy loss. They are equivalent during optimization.

VI Experiments

We have conducted several experiments to evaluate the performance of our MFB models for the VQA task by using the VQA datasets to verify our approach. We first perform ablation analysis on the MFB and MFH baseline models to verify the superior performance of the proposed approaches over existing state-of-the-art methods such as MCB and MLB . We then provide detailed analysis of the reasons why our models outperform their counterparts. Finally, we choose the optimal hyper-parameters for the MFB or MFH module and train the models with co-attention for fair comparison with the state-of-the-art approaches on the real-world VQA datasets. The corresponding source codes and pre-trained models are released onlinehttps://github.com/yuzcccc/vqa-mfb.

We have evaluated the performances of our proposed approaches over multiple VQA datasets. In addition, we have compared our proposed approaches with the state-of-the-art algorithms.

The VQA-1.0 dataset consists of approximately 200,000 images from the MS-COCO dataset , with 3 questions per image and 10 answers per question. The data set is split into three: train (80k images and 240k question-answer pairs), val (40k images and 120k question-answer pairs), and test (80k images and 240k question-answer pairs). Additionally, there is a 25%\% test subset named test-dev. Two tasks are provided to evaluate performance: Open-Ended (OE) and Multiple-Choices (MC). We use the tools provided by Antol et al. to evaluate the accuracy on the two tasks. Specifically, the accuracy of a predicted answer aa is calculated as follows:

where Count(a)\textrm{Count}(a) is the count of the answer aa voted by different annotators.

VI-A2 VQA-2.0

The VQA-2.0 dataset is the updated version of the VQA dataset. Compared with the VQA dataset, it contains more training samples (440k question-answer pairs for training and 214k pairs for validation), and is more balanced to weaken the potential that an overfitted model may achieve good results. Specifically, for every question there are two images in the dataset that result in two different answers to the question. At this point only the train and validation sets are available. Therefore, we report the results of the Open-Ended task on validation set with the model trained on train set. The evaluation criterion on this dataset is same as the one used in the VQA-1.0 dataset.

VI-B Experimental Setup

For the VQA and VQA 2.0 datasets, we use the Adam solver with β1=0.9\beta_{1}=0.9, β2=0.99\beta_{2}=0.99. The base learning rate is set to 0.0007 and decays every 40k iterations using an exponential rate of 0.5 for MFB and 0.25 for MFH. All the models are trained up to 100k iterations. Dropouts are used after each LSTM layer (dropout ratio p=0.3p=0.3) and MFB and MFH modules (p=0.1p=0.1). The number of answers N=3000N=3000. For all experiments (except for the ones shown in Table II, which use the train and val sets together as the training set like the comparative approaches), we train on the train set, validate on the val set, and report the results on the test-dev and test-standard setsthe submission attempts for the test set are strictly limited. Therefore, we report most of our results on the test-dev set and the best results on the test-standard set. The batch size is set to 200 for the models without the attention mechanism, and set to 64 for the models with attention (due to GPU memory limitation).

All experiments are implemented with the Caffe toolbox and performed on the workstations with NVIDIA TitanX GPUs.

VI-C Ablation Study on the VQA-1.0 Dataset

We design the following ablation experiments to verify the efficacy of our MFB and MFH modules, as well as the advantage of the KLD loss in modeling answer correlations.

First, MFB significantly outperforms all the baseline multi-modal fusion models. MFB is at least 2 points higher than the compared baseline models of similar sizes: the MFB(kk=5,oo=200) model outperforms the EltwiseProd model by 2.1 points, and the MFB(kk=5,oo=1000) model outperforms the EltwiseProd+FC+ReLU model by 2.2 points. These results demonstrates the advantage of second-order bilinear pooling models over the first-order pooling models on learning discriminative multi-modal feature representations.

Second, MFB outperforms other multi-modal bilinear pooling approaches. With 5/6 parameters, MFB(kk=5,oo=1000) achieves an improvement of about 1.0 points compared with MCB and MLB. Moreover, with only 1/3 parameters and 2/3 GPU memory usage, MFB(kk=5,oo=200) obtains similar results to MCB. These characteristics allows us to train our model on a memory limited GPU with larger batch-size. In Fig. 7, we show the courses of validation, from which it can be seen that MFB significantly outperforms the two other methods in terms of accuracy on the validation set. Furthermore, it can be seen from the accuracy curve of MCB that its performance gradually falls after 25,000 iterations, indicating that it suffer from overfitting with the high-dimensional output features. In comparison, the performance of MFB is relatively robust.

Third, when koko is fixed to a constant, e.g., 5000, the number of factors kk affects the performance. Increasing kk from 1 to 5, produces a 0.5 points performance gain. When k=10k=10, the performance has approached saturation. This phenomenon can be explained by the fact that a large kk corresponds to using a large window to sum pool the features, which can be treated as a compressed representation and may lose some information. When kk is fixed, increasing oo does not produce any further improvement. This suggests that high-dimensional output features may be easier to overfit. Similar results can be seen in MCB . In summary, k=5k=5 and o=1000o=1000 may be a suitable combination for our MFB model on the VQA dataset, so we use these settings in our follow-up experiments.

VI-C2 Answer Correlation Modeling Strategies

In Fig. 8, the validation accuracies of MFB and MFB+CoAtt models w.r.t. different answer sampling strategies are demonstrated respectively. Max Prob means using the most frequent answer of the sample as the unique label and formulate the optimization for VQA as the traditional multi-class problem with single label. This strategy refer to the baseline approach that does not consider answer correlation. Answer Sampling is the strategy used in MCB , which random sample an answer from the candidate answer set at each time. KLD is the strategy proposed in section V of this paper.

From the results, we have the following observations. First, modeling answer correlation bring remarkable improvement on the VQA-1.0 dataset. The Answer Sampling and KLD strategies which model the answer correlation, significantly outperform the Max Prob strategy. Second, compared with the Answer Sampling strategy, the proposed KLD strategy has the merits of faster convergence rate and slightly better accuracy, especially on the complex MFB+CoAtt model.

VI-D Results on the VQA-1.0 Dataset

Table II compares our approaches with the current state-of-the-art. The table is split into four parts over the rows: the first summarizes the methods without introducing the attention mechanism; the second includes the methods with attention; the third illustrates the results of approaches with external pre-trained word embedding models, e.g., GloVe or Skip-thought Vectors (StV) ; and the last includes the models trained with the external large-scale Visual Genome dataset additionally. To best utilize model capacity, the training data set is augmented so that both the train and val sets are used as the training set. Also, to better understand the question semantics, pre-trained GloVe word vectors are concatenated with the learned word embedding. The MFB model corresponds to the MFB baseline model. The MFB+Att model indicates the model that replaces the MCB with our MFB in the MCB+Att model . The MFB+CoAtt model represents the network shown in Fig. 5. The MFB+CoAtt+GloVe model additionally concatenates the learned word embedding with the pre-trained GloVe vectors. The MFB+CoAtt+GloVe+VG model further introduce the data from the Visual Genome dataset into the training set.

From Table II, we have the following observations.

First, the model with MFB outperforms other comparative approaches significantly. The MFB baseline outperforms all other existing approaches without the attention mechanism for both the OE and MC tasks, and even surpasses some approaches with attention. When attention is introduced, MFB+Att consistently outperforms current next-best model MCB+Att, highlighting the efficacy and robustness of the proposed MFB.

Second, the co-attention model further improve the performance over the attention model with only considering the image attention. By additionally introducing the self-attention module for questions, MFB+CoAtt delivers an improvement of 0.5 points on the OE task compared to the MFB+Att model in terms of overall accuracy. Moreover, for each question type (i.e., Y/N, Num or Others), the improvement of MFB+CoAtt over MFB+Att is significant, indicating the effect of the self-attention module in our co-attention learning framework.

Third, by replacing MFB with MFH, the performance of all of our models further enjoy an improvement of about 0.7∼\sim1.1 points steadily. The performance of a single MFH+CoAtt+GloVe model has even surpassed the best published results with an ensemble of 7 MLB or MFB models shown in Table III on the test-standard set.

Finally, with external pre-trained GloVe model and the Visual Genome dataset, the performance of our models are further improved. The MFH+CoAtt+GloVe+VG model significantly outperforms the best reported results with a single model on both the OE and MC task.

In Table III, we compare our model with the state-of-the-art results with model ensemble. Similar with , we train 7 individual MFB (or MFH)+CoAtt+GloVe models and average the prediction scores of them. 4 of the 7 models additionally introduce the Visual Genome dataset into the training set. All the reported results are fetched from the leaderboard of the VQA-1.0 datasetthe Standard tab in http://www.visualqa.org/roe.html. For fair comparison, only the published results are demonstrated. From the results, the ensemble of MFB models outperforms the next best result by 1.5 points on the OE task and by 2.2 points on the MC task respectively. Furthermore, the result of the ensemble of MFH models obtain a further improvement of 0.8 points and achieve the new state-of-the-art. Finally, compared with the results obtained by human, there is still a lot of room for improvement to approach the human-level.

To demonstrate the effects of co-attention learning, we visualize the learned question and image attentions of some image-question pairs from the val set in Fig. 9. The examples are randomly picked from different question types. It can seen that the learned question and image attentions are usually closely focus on the key words and the most relevant image regions. From the incorrect examples, we can also draw conclusions about the weakness of our approach, which are perhaps common to all VQA approaches: 1) some key words in the question are neglected by the question attention module, which seriously affects the learned image attention and final predictions (e.g., the word catcher in the first example and the word bottom in the third example); 2) even the intention of the question is well understood, some visual contents are still unrecognized (e.g., the flags in the second example) or misclassified (the meat in the fourth example), leading to the wrong answer for the counting problem. These observations are useful to guide further improvement for the VQA task in the future.

VI-E Results on the VQA-2.0 Dataset

Table IV demonstrates our results on the VQA-2.0 dataset (a.k.a, VQA challenge 2017). We compare our models with the results of baseline models (including the MCB model which is the champion of VQA Challenge 2016) and the results of the top-ranked teams on the leaderboard. We use the same training strategies aforementioned for this dataset.

From the results, our single MFB and MFH models (with CoAtt+GloVe but without the Visual Genome data argumentation) significantly surpass all the baseline approaches. If we neglect the tiny difference between the results on test-dev and test-standard sets, MFB and MFH is about 2.7 points and 3.5 points higher than the MCB model respectively. Finally, with an ensemble 9 models, we report the accuracy of 68.02%\% on the test-dev set and 68.16%\% on the test-challenge set respectively http://visualqa.org/roe_2017.html, which ranks the second place (tied with another team) in VQA Challenge 2017. The details of the 9 models are illustrated in Table V.

In the solution of the champion team, they introduce the region-based visual features extracted from the Faster R-CNN model which is pre-trained on the large-scale Visual Genome dataset . Using these visual features instead of the convolutional features from the ResNet model brings surprisingly good performance even with a simple VQA model. By using their visual features as the backbone for our models with MFH, we are in the first place on the real-time leaderboard of the VQA-2.0 dataset up to now (15 March, 2018). We report the overall accuracy 70.92%\% on the test-standard set of VQA-2.0 with 8 models while they report the accuracy 70.34%\% with up to 30 models .

VII Conclusions

In this paper, a network architecture with co-attention learning is designed to model both the image attention and the question attention simultaneously, so that we can reduce the irrelevant features effectively and extract more discriminative features for image and question representations. A Multi-modal Factorized Bilinear pooling (MFB) approach is developed to achieve more effective fusion of the visual features from the images and the textual features from the questions, and a generalized high-order model called MFH is developed to capture more complex interactions between multi-modal features. Compared with the existing bilinear pooling methods, our proposed MFB and MFH approaches can achieve significant improvement on the VQA performance because they can achieve more effective exploitation of the complex correlations between multi-modal features. By using the KL divergence as the loss function, our proposed answer prediction approach can achieve faster convergence rate and obtain better performance as compared with the state-of-the-art strategies. Our experimental results have demonstrated that our approaches have achieved the state-of-the-art or comparable performance on two large-scale real-world VQA datasets.

References