Multimodal Residual Learning for Visual QA

Jin-Hwa Kim, Sang-Woo Lee, Dong-Hyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, Byoung-Tak Zhang

Introduction

Visual question-answering tasks provide a testbed to cultivate the synergistic proposals which handle multidisciplinary problems of vision, language and integrated reasoning. So, the visual question-answering tasks let the studies in artificial intelligence go beyond narrow tasks. Furthermore, it may help to solve the real world problems which need the integrated reasoning of vision and language.

Deep residual learning not only advances the studies in object recognition problems, but also gives a general framework for deep neural networks. The existing non-linear layers of neural networks serve to fit another mapping of F(x)\mathcal{F}(\mathbf{x}), which is the residual of identity mapping x\mathbf{x}. So, with the shortcut connection of identity mapping x\mathbf{x}, the whole module of layers fit F(x)+x\mathcal{F}(\mathbf{x})+\mathbf{x} for the desired underlying mapping H(x)\mathcal{H}(\mathbf{x}). In other words, the only residual mapping F(x)\mathcal{F}(\mathbf{x}), defined by H(x)−x\mathcal{H}(\mathbf{x})-\mathbf{x}, is learned with non-linear layers. In this way, very deep neural networks effectively learn representations in an efficient manner.

Many attentional models utilize the residual learning to deal with various tasks, including textual reasoning and visual question-answering . They use an attentional mechanism to handle two different information sources, a query and the context of the query (e.g. contextual sentences or an image). The query is added to the output of the attentional module, that makes the attentional module learn the residual of query mapping as in deep residual learning.

In this paper, we propose Multimodal Residual Networks (MRN) to learn multimodality of visual question-answering tasks exploiting the excellence of deep residual learning . MRN inherently uses shortcuts and residual mappings for multimodality. We explore various models upon the choice of the shortcuts for each modality, and the joint residual mappings based on element-wise multiplication, which effectively learn the multimodal representations not using explicit attention parameters. Figure 2 shows inference flow of the proposed MRN.

Additionally, we propose a novel method to visualize the attention effects of each joint residual mapping. The visualization method uses back-propagation algorithm for the difference between the visual input and the output of the joint residual mapping. The difference is back-propagated up to an input image. Since we use the pretrained visual features, the pretrained CNN is augmented for visualization. Based on this, we argue that MRN is an implicit attention model without explicit attention parameters.

Our contribution is three-fold: 1) extending the deep residual learning for visual question-answering tasks. This method utilizes multimodal inputs, and allows a deeper network structure, 2) achieving the state-of-the-art results on the Visual QA dataset for both Open-Ended and Multiple-Choice tasks, and finally, 3) introducing a novel method to visualize spatial attention effect of joint residual mappings from the collapsed visual feature using back-propagation.

Related Works

Deep residual learning allows neural networks to have a deeper structure of over-100 layers. The very deep neural networks are usually hard to be optimized even though the well-known activation functions and regularization techniques are applied . This method consistently shows state-of-the-art results across multiple visual tasks including image classification, object detection, localization and segmentation.

This idea assumes that a block of deep neural networks forming a non-linear mapping F(x)\mathcal{F}(\mathbf{x}) may paradoxically fail to fit into an identity mapping. To resolve this, the deep residual learning adds x\mathbf{x} to F(x)\mathcal{F}(\mathbf{x}) as a shortcut connection. With this idea, the non-linear mapping F(x)\mathcal{F}(\mathbf{x}) can focus on the residual of the shortcut mapping x\mathbf{x}. Therefore, a learning block is defined as:

where x\mathbf{x} and y\mathbf{y} are the input and output of the learning block, respectively.

2 Stacked Attention Networks

Stacked Attention Networks (SAN) explicitly learns the weights of visual feature vectors to select a small portion of visual information for a given question vector. Furthermore, this model stacks the attention networks for multi-step reasoning narrowing down the selection of visual information. For example, if the attention networks are asked to find a pink handbag in a scene, they try to find pink objects first, and then, narrow down to the pink handbag.

For the attention networks, the weights are learned by a question vector and the corresponding visual feature vectors. These weights are used for the linear combination of multiple visual feature vectors indexing spatial information. Through this, SAN successfully selects a portion of visual information. Finally, an addition of the combined visual feature vector and the previous question vector is transferred as a new input question vector to next learning block.

Here, ql\mathbf{q}^{l} is a question vector for ll-th learning block and V\mathbf{V} is a visual feature matrix, whose columns indicate the specific spatial indexes. F(q,V)\mathcal{F}(\mathbf{q},\mathbf{V}) is the attention networks of SAN.

Multimodal Residual Networks

Deep residual learning emphasizes the importance of identity (or linear) shortcuts to have the non-linear mappings efficiently learn only residuals . In multimodal learning, this idea may not be readily applied. Since the modalities may have correlations, we need to carefully define joint residual functions as the non-linear mappings. Moreover, the shortcuts are undetermined due to its multimodality. Therefore, the characteristics of a given task ought to be considered to determine the model structure.

We infer a residual learning in the attention networks of SAN. Since Equation 18 in shows a question vector transferred directly through successive layers of the attention networks. In the case of SAN, the shortcut mapping is for the question vector, and the non-linear mapping is the attention networks.

In the attention networks, Yang et al. assume that an appropriate choice of weights on visual feature vectors for a given question vector sufficiently captures the joint representation for answering. However, question information weakly contributes to the joint representation only through coefficients p\mathbf{p}, which may cause a bottleneck to learn the joint representation.

The coefficients p\mathbf{p} are the output of a nonlinear function of a question vector q\mathbf{q} and a visual feature matrix V\mathbf{V} (see Equation 15-16 in Yang et al. ). The Vi\mathbf{V}_{i} is a visual feature vector of spatial index ii in 14×1414\times 14 grids.

Lu et al. propose an element-wise multiplication of a question vector and a visual feature vector after appropriate embeddings for a joint model. This makes a strong baseline outperforming some of the recent works . We firstly take this approach as a candidate for the joint residual function, since it is simple yet successful for visual question-answering. In this context, we take the global visual feature approach for the element-wise multiplication, instead of the multiple (spatial) visual features approach for the explicit attention mechanism of SAN. (We present a visualization technique exploiting the element-wise multiplication in Section 5.2.)

Based on these observations, we follow the shortcut mapping and the stacking architecture of SAN ; however, the element-wise multiplication is used for the joint residual function F\mathcal{F}. These updates effectively learn the joint representation of given vision and language information addressing the bottleneck issue of the attention networks of SAN.

2 Multimodal Residual Networks

MRN consists of multiple learning blocks, which are stacked for deep residual learning. Denoting an optimal mapping by H(q,v)\mathcal{H}(\mathbf{q},\mathbf{v}), we approximate it using

The first (linear) approximation term is Wq′(1)qW^{(1)}_{\mathbf{q}^{\prime}}\mathbf{q} and the first joint residual function is given by F(1)(q,v)\mathcal{F}^{(1)}(\mathbf{q},\mathbf{v}). The linear mapping Wq′W_{\mathbf{q}^{\prime}} is used for matching a feature dimension. We define the joint residual function as

where σ\sigma is tanh⁡\tanh, and ⊙\odot is element-wise multiplication. The question vector and the visual feature vector directly contribute to the joint representation. We justify this choice in Sections 4 and 5.

For a deeper residual learning, we replace q\mathbf{q} with H1(q,v)H_{1}(\mathbf{q},\mathbf{v}) in the next layer. In more general terms, Equations 4 and 5 can be rewritten as

where LL is the number of learning blocks, H0=qH_{0}=\mathbf{q}, Wq′=Πl=1LWq′(l)W_{\mathbf{q}^{\prime}}=\Pi_{l=1}^{L}W_{\mathbf{q}^{\prime}}^{(l)}, and WF(l)=Πm=l+1LWq′(m)W_{\mathcal{F}^{(l)}}=\Pi_{m=l+1}^{L}W_{\mathbf{q}^{\prime}}^{(m)}. The cascading in Equation 6 can intuitively be represented as shown in Figure 2. Notice that the shortcuts for a visual part are identity mappings to transfer the input visual feature vector to each layer (dashed line). At the end of each block, we denote HlH_{l} as the output of the ll-th learning block, and ⊕\oplus is element-wise addition.

Experiments

We choose the Visual QA (VQA) dataset for the evaluation of our models. Other datasets may not be ideal, since they have limited number of examples to train and test , or have synthesized questions from the image captions .

The questions and answers of the VQA dataset are collected via Amazon Mechanical Turk from human subjects, who satisfy the experimental requirement. The dataset includes 614,163 questions and 7,984,119 answers, since ten answers are gathered for each question from unique human subjects. Therefore, Antol et al. proposed a new accuracy metric as follows:

The questions are answered in two ways: Open-Ended and Multiple-Choice. Unlike Open-Ended, Multiple-Choice allows additional information of eighteen candidate answers for each question. There are three types of answers: yes/no (Y/N), numbers (Num.) and others (Other). Table 3 shows that Other type has the most benefit from Multiple-Choice.

The images come from the MS-COCO dataset, 123,287 of them for training and validation, and 81,434 for test. The images are carefully collected to contain multiple objects and natural situations, which is also valid for visual question-answering tasks.

2 Implementation

Torch framework and rnn package are used to build our models. For efficient computation of variable-length questions, TrimZero is used to trim out zero vectors . TrimZero eliminates zero computations at every time-step in mini-batch learning. Its efficiency is affected by a batch size, RNN model size, and the number of zeros in inputs. We found out that TrimZero was suitable for VQA tasks. Approximately, 37.5% of training time is reduced in our experiments using this technique.

We follow the same preprocessing procedure of DeeperLSTM+NormalizedCNN (Deep Q+I) by default. The number of answers is 1k, 2k, or 3k using the most frequent answers, which covers 86.52%, 90.45% and 92.42% of questions, respectively. The questions are tokenized using Python Natural Language Toolkit (nltk) . Subsequently, the vocabulary sizes are 14,770, 15,031 and 15,169, respectively.

By default, we follow Deep Q+I. The common embedding size of the joint representation is 1,200. The learnable parameters are initialized using a uniform distribution from −0.08-0.08 to 0.080.08 except for the pretrained models. The batch size is 200, and the number of iterations is fixed to 250k. The RMSProp is used for optimization, and dropouts are used for regularization. The hyperparameters are fixed using test-dev results. We compare our method to state-of-the-arts using test-standard results.

3 Exploring Alternative Models

Figure 3 shows alternative models we explored, based on the observations in Section 3. We carefully select alternative models (a)-(c) for the importance of embeddings in multimodal learning , (d) for the effectiveness of identity mapping as reported by , and (e) for the confirmation of using question-only shortcuts in the multiple blocks as in . For comparison, all models have three-block layers (selected after a pilot test), using VGG-19 features and 1k answers, then, the number of learning blocks is explored to confirm the pilot test. The effect of the pretrained visual feature models and the number of answers are also explored. All validation is performed on the test-dev split.

Results

The VQA Challenge, which released the VQA dataset, provides evaluation servers for test-dev and test-standard test splits. For the test-dev, the evaluation server permits unlimited submissions for validation, while the test-standard permits limited submissions for the competition. We report accuracies in percentage.

The test-dev results of the alternative models for the Open-Ended task are shown in Table 2. (a) shows a significant improvement over SAN. However, (b) is marginally better than (a). As compared to (b), (c) deteriorates the performance. An extra embedding for a question vector may easily cause overfitting leading to the overall degradation. And, the identity shortcuts in (d) cause the degradation problem, too. Extra parameters of the linear mappings may effectively support to do the task.

(e) shows a reasonable performance, however, the extra shortcut is not essential. The empirical results seem to support this idea. Since the question-only model (50.39%) achieves a competitive result to the joint model (57.75%), while the image-only model gets a poor accuracy (28.13%) (see Table 2 in ). Eventually, we chose model (b) as the best performance and relative simplicity.

The effects of other various options, Skip-Thought Vectors for parameter initialization, Bayesian Dropout for regularization, image captioning model for postprocessing, and the usage of shortcut connections, are explored in Appendix A.1.

To confirm the effectiveness of the number of learning blocks selected via a pilot test (L=3L=3), we explore this on the chosen model (b), again. As the depth increases, the overall accuracies are 58.85% (L=1L=1), 59.44% (L=2L=2), 60.53% (L=3L=3) and 60.42% (L=4L=4).

The ResNet-152 visual features are significantly better than VGG-19 features for Other type in Table 2, even if the dimension of the ResNet features (2,048) is a half of VGG features’ (4,096). The ResNet visual features are also used in the previous work ; however, our model achieves a remarkably better performance with a large margin (see Table 3).

The number of target answers slightly affects the overall accuracies with the trade-off among answer types. So, the decision on the number of target answers is difficult to be made. We chose Res, 2k in Table 2 based on the overall accuracy (for Multiple-Choice task, see Appendix A.1).

Our chosen model significantly outperforms other state-of-the-art methods for both Open-Ended and Multiple-Choice tasks in Table 3. However, the performance of Number and Other types are still not satisfactory compared to Human performance, though the advances in the recent works were mainly for Other-type answers. This fact motivates to study on a counting mechanism in future work. The model comparison is performed on the test-standard results.

2 Qualitative Analysis

In Equation 5, the left term σ(Wqq)\sigma(W_{\mathbf{q}}\mathbf{q}) can be seen as a masking (attention) vector to select a part of visual information. We assume that the difference between the right term V:=σ(W2σ(W1v))\mathcal{V}:=\sigma(W_{2}\sigma(W_{1}\mathbf{v})) and the masked vector F(q,v)\mathcal{F}(\mathbf{q},\mathbf{v}) indicates an attention effect caused by the masking vector. Then, the attention effect Latt=12∥V−F∥2\mathcal{L_{\textit{att}}}=\frac{1}{2}\lVert\mathcal{V}-\mathcal{F}\rVert^{2} is visualized on the image by calculating the gradient of Latt\mathcal{L_{\textit{att}}} with respect to a given image I\mathcal{I}, while treating F\mathcal{F} as a constant.

This technique can be applied to each learning block in a similar way.

Since we use the preprocessed visual features, the pretrained CNN is augmented only for this visualization. Note that model (b) in Table 2 is used for this visualization, and the pretrained VGG-19 is used for preprocessing and augmentation. The model is trained using the training set of the VQA dataset, and visualized using the validation set. Examples are shown in Figure 4 (more examples in Appendix A.2-4).

Unlike the other works that use explicit attention parameters, MRN does not use any explicit attentional mechanism. However, we observe the interpretability of element-wise multiplication as an information masking, which yields a novel method for visualizing the attention effect from this operation. Since MRN does not depend on a few attention parameters (e.g. 14×1414\times 14), our visualization method shows a higher resolution than others . Based on this, we argue that MRN is an implicit attention model without explicit attention mechanism.

Conclusions

The idea of deep residual learning is applied to visual question-answering tasks. Based on the two observations of the previous works, various alternative models are suggested and validated to propose the three-block layered MRN. Our model achieves the state-of-the-art results on the VQA dataset for both Open-Ended and Multiple-Choice tasks. Moreover, we have introduced a novel method to visualize the spatial attention from the collapsed visual features using back-propagation.

We believe our visualization method brings implicit attention mechanism to research of attentional models. Using back-propagation of attention effect, extensive research in object detection, segmentation and tracking are worth further investigations.

The authors would like to thank Patrick Emaase for helpful comments and editing. This work was supported by Naver Corp. and partly by the Korea government (IITP-R0126-16-1072-SW.StarLab, KEIT-10044009-HRI.MESSI, KEIT-10060086-RISF, ADD-UD130070ID-BMRR).

References

A Appendix

A.2 More Examples

A.3 Comparative Analysis

A.4 Failure Examples