A Focused Dynamic Attention Model for Visual Question Answering
Ilija Ilievski, Shuicheng Yan, Jiashi Feng
Introduction
Visual question answering (VQA) is an active research direction that lies in the intersection of computer vision, natural language processing, and machine learning. Even though with a very short history, it already has received great research attention from multiple communities. Generally, the VQA investigates a generalization of traditional QA problems where visual input (e.g., an image) is necessary to be considered. More concretely, VQA is about how to provide a correct answer to a human posed question concerning contents of one presented image or video.
VQA is a quite challenging task and undoubtedly important for developing modern AI systems. The VQA problem can be regarded as a Visual Turing Test , and besides contributing to the advancement of the involved research areas, it has other important applications, such as blind person assistance and image retrieval. Coming up with solutions to this task requires natural language processing techniques for understanding the questions and generating the answers, as well as computer vision techniques for understanding contents of the concerned image. With help of these two core techniques, the computer can perform reasoning about the perceived contents and posed questions.
Recently, VQA is advanced significantly by the development of machine learning methods (in particular the deep learning ones) that can learn proper representations of questions and images, align and fuse them in a joint question-image space and provide a direct mapping from this joint representation to a correct answer.
For example, consider the following image-question pair: an image of an apple tree with a basket of apples next to it, and a question “How many apples are in the basket?”. Answering this question requires VQA methods to first understand the semantics of the question, then locate the objects (apples) in the image, understand the relation between the image objects (which apples are in the basket), and finally count them and generate an answer with the correct number of apples.
The first feasible solution to VQA problems was provided by Malinowski and Fritz in , where they used a semantic language parser and a Bayesian reasoning model, to understand the meaning of questions and to generate the proper answers. Malinowski and Fritz also constructed the first VQA benchmark dataset, named as DAQUAR, which contains 1,449 images and 12,468 questions generated by humans or automatically by following a template and extracting facts from a database . Shortly after, Ren et al. released the TORONTO–QA dataset, which contains a large number of images (123,287) and questions (117,684), but the questions are automatically generated and thus can be answered without complex reasoning. Nevertheless, the release of the TORONTO–QA dataset was important since it provided enough data for deep learning models to be trained and evaluated on the VQA problem . More recently, Antol et al. published the currently largest VQA dataset. It consists of three human posed questions and ten answers given by different human subjects, for each one of the 204,721 images found in the Microsoft COCO dataset . Answering the 614,163 questions requires complex reasoning, common sense, and real-world knowledge, making the VQA dataset suitable for a true Visual Turing Test. The VQA authors split the evaluation on their dataset on two tasks: an open-ended task, where the method should generate a natural language answer, and a multiple-choice task, where for each question the method should chose one of the 18 different answers.
The current top performing methods employ deep neural network model that predominantly uses the convolutional neural network (CNN) architecture to extract image features and a Long Short-Term Memory (LSTM) network to extract the representations for questions. The CNN and LSTM representation vectors are then usually fused by concatenation or element-wise multiplication . Other approaches additionally incorporate some kind of attention mechanism over the image features .
Properly modeling the image contents is one of the critical factors for solving VQA problems well. A common practice with existing VQA methods on modeling image contents is to extract global features for the overall image. However, only using global feature is arguably insufficient to capture all the necessary visual information and provide full understanding of image contents such as multiple objects, spatial configuration of the objects and informative background. This issue can be relieved to some extent by extracting features from object proposals – the image regions that possibly contain objects of interest. However, using features from all image regions may provide too much noise or overwhelming information irrelevant to the question and thus hurt the overall VQA performance.
In this work, we propose a question driven attention model that is able to automatically identify and focus on image regions relevant for the current question. We name our proposed model Focused Dynamic Attention (FDA) for Visual Question Answering. With the FDA model, computers can select and recognize the image regions in a well-aligned sequence with the key words containing in a given question. Recall the above VQA example. To answer the question of “How many apples are in the basket?”, FDA would first localize the regions corresponding to the key words “apples” and “basket” (with the help of a generic object detector) and extract description features from these regions of interest. Then VQA compliments the features from selected image regions with a global image feature providing contextual information for the overall image, and reconstruct a visual representation by encoding them with a Long Short-Term Memory (LSTM) unit.
We evaluate and compare the performance of our proposed FDA model on two types of VQA tasks, i.e., the open-ended task and the multiple-choice task, on the VQA dataset – the largest VQA benchmark dataset. Extensive experiments demonstrate that FDA brings substantial performance improvement upon well-established baselines.
The main contributions of this work can be summarized as follows:
We introduce a focused dynamic attention mechanism that learns to use the question word order to shift the focus from one image object, to another.
We describe a model that fuses local and global context visual features with textual features.
We perform an extensive evaluation, comparing to all existing methods, and achieve state-of-the-art accuracy on the open-ended, and on the multiple-choice VQA tasks.
The rest of the paper is organized as follows. In Section 2 we review the current VQA models, and compare them to our model. We formulate the problem and explain our motivation in Section 3. We describe our model in Section 4 and in Section 5 we evaluate and compare it with the current state-of-the-art models. We conclude our work in Section 6.
Related Work
VQA has received great research attention recently and a couple of methods have been developed to solve this problem. The most similar model to ours is the Stacked Attention Networks (SAN) proposed by Yang et al. . Both models use attention mechanism that combines the words and image regions. However, use convolutional neural network to put attention over the image regions, based on the question word unigrams, bigrams, and trigrams. Further, their attention mechanism is not using object bounding boxes, which makes the attention less focused.
Another model that uses attention mechanism in solving VQA problems is the ABC-CNN model described in . ABC-CNN uses the question embedding to configure convolutional kernels that will define an attention weighted map over the image features. The advantage of our FDA model over ABC-CNN is two fold. First, FDA employs an LSTM network to encode the image region features in a order that corresponds to the question word order. Second, FDA does not put handcrafted weights on the image features (a practice showed to hurt the learning process in our experiments). Instead, FDA extracts CNN features directly from the cropped image regions of interest. In this sense, FDA is more efficient than ABC-CNN in visual contents modeling.
Yet another attention model for visual question answering is proposed in . The work, is closely related to the work by , in that it also applies a weighted map over the image and the question word features. However, similar to our work, they use object proposals from to select image regions instead of the whole image. Different from that work, our proposed FDA model also employs the information embedded in the order of the question words and focuses on the corresponding object bounding boxes. In contrast, the model proposed in straightforwardly concatenate all the image region features with the question word features and feed them all at once to a two layer network.
Jiang et al. propose another model that combines the CNN image features and an LSTM network for encoding the multimodal representation, with the addition of a Compositional Memory units which fuse the image and word feature vectors .
Ma et al. in take an interesting approach and use three convolutional neural networks to represent not only the image, but also the question, and their common representation in a multimodal space. The multimodal representation is then fed to a SoftMax layer to produce the answer.
Another interesting approach worth mentioning is the work by Andreas et al. . They use a semantic grammar parser to parse the question and propose neural network layouts accordingly. They train a model to learn to compose a network from one of the proposed network layouts using several types of neural modules, each specifically designed to address the different sub-tasks of the VQA problem (e.g. counting, locating an object, etc.).
Method Overview
In this section, we briefly describe the motivation and give formal problem formulation.
The visual question answering problem can be represented as predicting the best answer given an image and a question . Common practice is to use the 1,000 most common answers in the training set and thus simplify the VQA task to a classification problem. The following equation represents the problem mathematically:
where is the set of all possible answers and are the model weights.
2 Motivation
The baseline methods from show only modest increase in accuracy when including the image features (4.98% for open-ended questions, and 2.42% for multiple-choice question). We believe that the image contains a lot more information and should increase the accuracy much more. Thus, we focus on improving the image features and design a visual attention mechanism, which learns to focus on the question related image regions.
The proposed attention mechanism is loosely inspired on the human visual attention mechanism. Humans shift the focus from one image region to another, before understanding how the regions relate to each other and grasping the meaning of the whole image. Similarly, we feed our model image regions relevant for the question at hand, before showing the whole image.
Focused Dynamic Attention for VQA
The FDA model is composed of question and image understanding components, attention mechanism, and a multimodal representation fusion network (Figure 1). In this section we describe them individually.
Following a common practice, our FDA model uses an LSTM network to encode the question in a vector representation . The LSTM network learns to keep in its state the feature vectors of the important question words, and thus provides the question understanding component with a word attention mechanism.
2 Image Understanding
Following prior work , we use a pre-trained convolutional neural network (CNN) to extract image feature vectors. Specifically, we use the Deep Residual Networks model used in ILSVRC and COCO 2015 competitions, which won the 1st places in: ImageNet classification, ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation . We extract the weights of the layer immediately before the final SoftMax layer and regard them as visual features. We extract such features for the whole image (global visual features) and for the specific image regions (local visual features). However, contrary to the existing approaches, we employ an LSTM network to combine the local and global visual features into a joint representation.
3 Focused Dynamic Attention Mechanism
We introduce a focused dynamic attention mechanism that learns to focus on image regions related to the question words.
The attention mechanism works as follows. For each image objectDuring training we use the ground truth object bounding boxes and labels. At test time we use the precomputed bounding boxes from and classify them with to obtain the object labels. it uses word2vec word embeddings to measure the similarity between the question words and the object label. Next, it selects objects with similarity score greater than 0.5 and extracts the feature vectors of the objects bounding boxes with a pre-trained ResNet model . Following the question word order, it feeds the LSTM network with the corresponding object feature vectors. Finally, it feeds the LSTM network with the feature vector of the whole image and it uses the resulting LSTM state as a visual representation. Thus, the attention mechanism enables the model to combine the local and global visual features into a single representation, necessary for answering complex visual questions.
Figure 1 illustrates the focused dynamic attention mechanism with an example.
4 Multimodal Representation Fusion
We regard the final state of the two LSTM networks as a question and image representation. We start fusing them into single representation by applying Tanh on the question representation and ReLUDefined as . on the image representation Applying different activation functions gave slightly worse overall results. We proceed by doing an element-wise multiplication of the two vector representations and the resulting vector is fed to a fully-connected neural network. Finally a SoftMax layer classify the multimodal representation into one of the possibleWe follow and use the 1000 most common answers answers.
Evaluation
In this section we detail the model implementation and compare our model against the current state-of-the-art methods.
For all experiments we use the Visual Question Answering (VQA) dataset , which is the largest and most complex image dataset for the visual question answering task. The dataset contains three human posed questions and ten answers given by different human subjects, for each one of the 204,721 images found in the Microsoft COCO dataset . Figure 2 shows two representative examples found in the dataset. The evaluation is done on following two test splits test-dev and test-std and on following two tasks:
An open-ended task, where the method should generate a natural language answer;
A multiple-choice task, where for each question the method should chose one of the 18 different answers.
We evaluate the performance of all the methods in the experiments using the public evaluation server for fair evaluation.
2 Baseline Model
We compare our model against the baseline models provided by the VQA dataset authors , which currently achieve the best performance on the test-standard split for the multiple-choice task. The model, first described in , is a standard implementation of an LSTMCNN VQA model. It uses an LSTM to encode the question and CNN features to encode the image. To answer a question it multiplies the last LSTM state with the image CNN features and feeds the result into a SoftMax layer for classification into one of the 1,000 most common answers. The implementation in uses a deeper two layer LSTM network for encoding the question, and normalized image CNN features, which showed crucial for achieving the state-of-the-art.
3 Model Implementation and Training Details
We transform the question words into a vector form by multiplying one-hot vector representation with a word embedding matrix. The vocabulary size is 12,602 and the word embeddings are 300 dimensional. We feed a pre-trained ResNet network and use the 2,048 dimensional weight vector of the layer before the last fully-connected layer.
The word and image vectors are feed into two separate LSTM networks. The LSTM networks are standard implementation of one layer LSTM network , with a 512 dimensional state vector. The final state of the question LSTM is passed through Tanh, while the final state of the image LSTM is passed through ReLUDefined as .. We do element-wise multiplication on the resulting vectors, to obtain a multimodal representation vector, which is then fed to a fully-connected neural network.
4 Model Evaluation and Comparison
We compare our model with the baselines provided by the VQA authors . The results for the open-ended task are listed in Table 1 and the results for the multiple-choice task are given in Table 2. In the tables, the “Question” and “Image” baselines are only using the question words and the image, respectively. The “Q+I” is a baseline that combines the two, but do not use an LSTM network. “LSTM Q+I” and “D-LSTM” are LSTM models, with one and two layers accordingly. Comparing the performance of baselines we can observe the accuracy increase with the addition of information from each modality.
From Table 1, one can observe that our proposed FDA model achieves the best performance on this benchmark dataset. It outperforms the state-of-the-art (SAN) with a margin of around . The SAN model also employs attention to focus on specific regions. However, their attention model (without access to the automatically generated bounding boxes) is focusing on more spread regions which may include cluttered and noisy background. In contrast, FDA only focuses on the selected regions and extracts more clean information for answering the questions. This is the main reason that FDA can outperform SAN although these two methods are both based on attention models.
The advantage of employing focused dynamic attention in FDA is more significant when solving the multiple-choice VQA problems. From Table 2, one can observe that our proposed FDA model achieves the best ever performance on the VQA dataset. In particular, it improves the performance of the state-of-the-art (D-LSTM) by a margin of which is quite significant for this challenging task. The D-LSTM method employs a deeper network to enhance the discriminative capacity of the visual features. However, they do not identify the informative regions for answering the questions. In contrast, FDA incorporates the automatic region localization by employing a question-driven attention model. This is helpful for filtering out irrelevant noise, and establishing the correspondence between regions and candidate answers. Thus FDA gives substantial performance improvement.
5 Qualitative Results
We qualitatively evaluate our model on a set of examples where complex reasoning and focusing on the relevant local visual features are needed for answering the question correctly.
Figure 3 shows particularly difficult examples (the predominant image color is not the correct answer) of “What color” type of questions. But, by focusing on the question related image regions, the FDA model is still able to produce the correct answer.
In Figure 4 we show examples where the model focuses on different regions from the same image, depending on the words in the question. Focusing on the right image region is crucial when answering unusual questions for an image (Row 1), questions about small image objects (Row 2), or when the most dominant image object partly occludes the question related region and can lead to a wrong answer (Row 3).
Representative examples of questions that require image object identification are shown in Figure 5. We can observe that the focused attention enables the model to answer complex questions (Row 1, left) and counting questions (Row 1, right). The question guided image object identification greatly simplifies the answering of questions like the ones shown in Row 2 and Row 3.
Conclusion
In this work, we proposed a novel Focused Dynamic Attention (FDA) model to solve the challenging VQA problems. FDA is built upon a generic object-centric attention model for extracting question related visual features from an image as well as a stack of multiple LSTM layers for feature fusion. By only focusing on the identified regions specific for proposed questions, FDA was shown to be able to filter out overwhelming irrelevant informations from cluttered background or other regions, and thus substantially improved the quality of visual representations in the sense of answering proposed questions. By fusing cleaned regional representation, global context and question representation via LSTM layers, FDA provided significant performance improvement over baselines on the VQA benchmark datasets, for both the open-ended and multiple-choices VQA tasks. Excellent performance of FDA clearly demonstrates its stronger ability of modeling visual contents and also verifies paying more attention to visual part in VQA tasks could essentially improve the overall performance. In the future, we are going to further explore along this research line and investigate different attention methods for visual information selection as well as better reasoning model for interpreting the relation between visual contents and questions.