Image Captioning and Visual Question Answering Based on Attributes and External Knowledge

Qi Wu, Chunhua Shen, Anton van den Hengel, Peng Wang, Anthony Dick

Introduction

Vision-to-Language problems present a particular challenge in Computer Vision because they require translation between two different forms of information. In this sense the problem is similar to that of machine translation between languages. In machine language translation there have been a series of results showing that good performance can be achieved without developing a higher-level model of the state of the world. In , for instance, a source sentence is transformed into a fixed-length vector representation by an ‘encoder’ RNN, which in turn is used as the initial hidden state of a ‘decoder’ RNN that generates the target sentence.

Despite the supposed equivalence between an image and a thousand words, the manner in which information is represented in each data form could hardly be more different. Human language is designed specifically so as to communicate information between humans, whereas even the most carefully composed image is the culmination of a complex set of physical processes over which humans have little control. Given the differences between these two forms of information, it seems surprising that methods inspired by machine language translation have been so successful. These RNN-based methods which translate directly from image features to text, without developing a high-level model of the state of the world, represent the current state of the art for key Vision-to-Language (V2L) problems, such as image captioning and visual question answering.

This approach is reflected in many recent successful works on image captioning, such as . Current state-of-the-art captioning methods use a CNN as an image ‘encoder’ to produce a fixed-length vector representation , which is then fed into the ‘decoder’ RNN to generate a caption.

Visual Question Answering (VQA) is a more recent challenge than image captioning. It is distinct from many problems in Computer Vision because the question to be answered is not determined until run time . In this V2L problem an image and a free-form, open-ended question about the image are presented to the method which is required to produce a suitable answer . As in image captioning, the current state of the art in VQA relies on passing CNN features to an RNN language model. However, visual question answering is a significantly more complex problem than image captioning, not least because it requires accessing information not present in the image. This may be common sense, or specific knowledge about the image subject. For example, given an image, such as Figure 1, showing ‘a group of people enjoying a sunny day at the beach with umbrellas’, if one asks a question ‘why do they have umbrellas?’, to answer this question, the machine must not only detect the scene ‘beach’, but must know that ‘umbrellas are often used as points of shade on a sunny beach’. Recently, Antol et al. also have suggested that VQA is a more “AI-complete” task since it requires multimodal knowledge beyond a single sub-domain.

The contributions of this paper are two-fold. First, we propose a fully trainable attribute-based neural network founded upon the CNN+RNN architecture, that can be applied to multiple V2L problems. We do this by inserting an explicit representation of attributes of the scene which are meaningful to humans. Each semantic attribute corresponds to a word mined from the training image descriptions, and represents higher-level knowledge about the content of the image. A CNN-based classifier is trained for each attribute, and the set of attribute likelihoods for an image form a high-level representation of image content. An RNN is then trained to generate captions, or question answers, on the basis of the likelihoods. Our attribute-based model yields significantly better performance than current state-of-the-art approaches in the task of image captioning.

Based on the proposed attribute-based V2L model, our second contribution is to introduce a method of incorporating knowledge external to the image, including common sense, into the VQA process. In this work, we fuse the automatically generated description of an image with information extracted from an external knowledge base (KB) to provide an answer to a general question about the image (See Figure 5). The image description takes the form of a set of captions, and the external knowledge is text-based information mined from a Knowledge Base. Specifically, for each of the top-kk attributes detected in the image we generate a query which may be applied to a Resource Description Framework (RDF) KB, such as DBpedia. RDF is the standard format for large KBs, of which there are many. The queries are specified using Semantic Protocol And RDF Query Language (SPARQL). We encode the paragraphs extracted from the KB using Doc2Vec , which maps paragraphs into a fixed-length feature representation. The encoded attributes, captions, and KB information are then input to an LSTM which is trained so as to maximise the likelihood of the ground truth answers in a training set. We further propose a question-guided knowledge selection scheme to improve the quality of the extracted KB information. The knowledge that is not related to the question is filtered out. The approach that we propose here combines the generality of information that using a KB allows with the generality of questions that the LSTM allows. In addition, it achieves an accuracy of 70.98% on the Toronto COCO-QA , while the latest state of the art is 61.60%. On the VQA evaluation server (which does not publish ground truth answers for its test set), we also produce the state-of-the-art result, which is 59.50%.

A preliminary version of this work was published at CVPR 2016 . The new material in this paper comprises further experiments on two additional VQA datasets. More ablation models of the original model are implemented and studied. More importantly, a new model (A+C+Selected-K-LSTM) is introduced for the visual question answering task, leading to a new state-of-the-art result.

Related work

Using attribute-based models as a high-level representation has shown potential in many computer vision tasks such as object recognition, image annotation and image retrieval. Farhadi et al. were among the first to propose to use a set of visual semantic attributes to identify familiar objects, and to describe unfamiliar objects. Vogel and Schiele used visual attributes describing scenes to characterize image regions and combined these local semantics into a global image description. Su et al. defined six groups of attributes to build intermediate-level features for image classification. Li et al. introduced the concept of an ‘object bank’ which enables objects to be used as attributes for scene representation.

2 Image Captioning

The problem of annotating images with natural language at the scene level has long been studied in both computer vision and natural language processing. Hodosh et al. proposed to frame sentence-based image annotation as the task of ranking a given pool of captions. Similarly, posed the task as a retrieval problem, but based on co-embedding of images and text in the same space. Recently, Socher et al. used neural networks to co-embed image and sentences together and Karpathy et al. co-embedded image crops and sub-sentences.

Attributes have been used in many image captioning methods to fill the gaps in predetermined caption templates. Farhadi et al. , for instance, used detections to infer a triplet of scene elements which is converted to text using a template. Li et al. composed image descriptions given computer vision based inputs such as detected objects, modifiers and locations using web-scale nn-grams. Zhu et al. converted image parsing results into a semantic representation in the form of Web Ontology Language, which is converted to human readable text. A more sophisticated CRF-based method use of attribute detections beyond triplets was proposed by Kulkarni et al . The advantage of template-based methods is that the resulting captions are more likely to be grammatically correct. The drawback is that they still rely on hard-coded visual concepts and suffer the implied limits on the variety of the output. Fang et al. won the 2015 COCO Captioning Challenge with an approach that is similar to ours in as much as it applies a visual concept (i.e., attribute) detection process before generating sentences. They first learned 10001000 independent detectors for visual words based on a multi-instance learning framework and then used a maximum entropy language model conditioned on the set of visually detected words directly to generate captions.

In contrast to the aforementioned two-stage methods, the recent dominant trend in V2L is to use an architecture which connects a CNN to an RNN to learn the mapping from images to sentences directly. Mao et al., for instance, proposed a multimodal RNN (m-RNN) to estimate the probability distribution of the next word given previous words and the deep CNN feature of an image at each time step. Similarly, Kiros et al. constructed a joint multimodal embedding space using a powerful deep CNN model and an LSTM that encodes text. Karpathy and Li also proposed a multimodal RNN generative model, but in contrast to , their RNN is conditioned on the image information only at the first time step. Vinyals et al. combined deep CNNs for image classification with an LSTM for sequence modeling, to create a single network that generates descriptions of images. Chen et al. learn a bi-directional mapping between images and their sentence-based descriptions, which allows to reconstruct visual features given an image description. Xu et al. proposed a model based on visual attention. Jia et al. applied additional retrieved sentences to guide the LSTM in generating captions.

Interestingly, this end-to-end CNN-RNN approach ignores the image-to-word mapping which was an essential step in many of the previous image captioning systems detailed above . The CNN-RNN approach has the advantage that it is able to generate a wider variety of captions, can be trained end-to-end, and outperforms the previous approach on the benchmarks. It is not clear, however, what the impact of bypassing the intermediate high-level representation is, and particularly to what extent the RNN language model might be compensating. Donahue et al. described an experiment, for example, using tags and CRF models as a mid-layer representation for video to generate descriptions, but it was designed to prove that LSTM outperforms an SMT-based approach . It remains unclear whether the mid-layer representation or the LSTM leads to the success. Our paper provides several well-designed experiments to answer this question.

We thus here show not only a method for introducing a high-level representation into the CNN-RNN framework, and that doing so improves performance, but we also investigate the value of high-level information more broadly in V2L tasks. This is of critical importance at this time because V2L has a long way to go, particularly in the generality of the images and text it is applicable to.

3 Visual Question Answering

Malinowski et al. may be the first to study the VQA problem. They proposed a method that combines semantic parsing and image segmentation with a Bayesian approach to sampling from nearest neighbors in the training set. Tu et al. built a query answering system based on a joint parse graph from text and videos. Geman et al. proposed an automatic ‘query generator’ that is trained on annotated images and produces a sequence of binary questions from any given test image. Each of these approaches places significant limitations on the form of question that can be answered.

Most recently, inspired by the significant progress achieved using deep neural network models in both computer vision and natural language processing, an architecture which combines a CNN and RNN to learn the mapping from images to sentences has become the dominant trend. Both Gao et al. and Malinowski et al. used RNNs to encode the question and output the answer. Whereas Gao et al. used two networks, a separate encoder and decoder, Malinowski et al. used a single network for both encoding and decoding. Ren et al. focused on questions with a single-word answer and formulated the task as a classification problem using an LSTM. Antol et al. proposed a large-scale open-ended VQA dataset based on COCO, which is called VQA. Inspired by Xu et al. who encode visual attention in the Image Captioning, propose to use the spatial attention to help answering visual questions. formulate the VQA as a classification problem and restrict the answer only can be drawn from a fixed answer space.

Our framework also exploits both CNN and RNNs, but in contrast to preceding approaches which use only image features extracted from a CNN in answering a question, we employ multiple sources, including image content, generated image captions and mined external knowledge, to feed to an RNN to answer questions. Large-scale Knowledge Bases (KBs), such as Freebase and DBpedia , have been used successfully in several natural language Question Answering (QA) systems . However, VQA systems exploiting KBs are still relatively rare.

Zhu et al. used a hand-crafted KB primarily containing image-related information such as category labels, attribute labels and affordance labels, but also some quantities relating to their specific question format such as GPS coordinates and similar. Instead of building a problem-specific KB, we use a pre-built large-scale KB (DBpedia ) from which we extract information using a standard RDF query language. DBpedia has been created by extracting structured information from Wikipedia, and is thus significantly larger and more general than a hand-crafted KB. Rather than having a user pose their question in a formal query language, our VQA system is able to encode questions written in natural language automatically. This is achieved without manually specified formalization, but rather depends on processing a suitable training set. The result is a model which is very general in the forms of question that it will accept. The quality of the information in the KB is one of the primary issues in this approach to VQA. The problem is that KBs constructed by analysing Wikipedia and similar are patchy and inconsistent at best, and hand-curated KBs are inevitably very topic specific. Using visually-sourced information is a promising approach to solve this problem , but has a way to go before it might be usefully applied within our approach. After inspecting the database shows that the ‘comment’ field is the most generally informative about an attribute, as it contains a general text description of it. We therefore find this is still a feasible solution.

Image Captioning using Attributes

Our image captioning model is summarized in Figure 2. The model includes an image analysis part and a caption generation part. In the image analysis part, we first use supervised learning to predict a set of attributes, based on words commonly found in image captions We solve this as a multi-label classification problem and train a corresponding deep CNN by minimizing an element-wise logistic loss function. Secondly, a fixed length vector Vatt(I){V_{att}}(I) is created for each image II, whose length is the size of the attribute set. Each dimension of the vector contains the prediction probability for a particular attribute. In the captioning generation part, we apply an LSTM-based sentence generator. In the baseline model, as in we use a pre-trained CNN to extract image features \mboxCNN(I)\mbox{CNN}(I) which are fed into the LSTM directly. For the sake of completeness a fine-tuned version of this approach is also implemented.

Our first task is to describe the image content in terms of a set of attributes. An attribute vocabulary is first constructed. Unlike , that use a vocabulary from separate hand-labeled training data, our semantic attributes are extracted from training captions and can be any part of speech, including object names (nouns), motions (verbs) or properties (adjectives). The direct use of captions guarantees that the most salient attributes for an image set are extracted. We use the cc (c=256c=256) most common words in the training captions to determine the attribute vocabulary Vatt\mathcal{V}_{att}. Similar to , the top 1515 most frequent closed-class words such as ‘a’,‘on’,‘of’ are removed since they are in nearly every caption. In contrast to , our vocabulary is not tense or plurality sensitive, for instance, ‘ride’ and ‘riding’ are classified as the same semantic attribute, similarly ‘bag’ and ‘bags’. This significantly decreases the size of our attribute vocabulary. The full list of attributes can be found in the supplementary material. Our attributes represent a set of high-level semantic constructs, the totality of which the LSTM then attempts to represent in sentence form. Generating a sentence from a vector of attribute likelihoods exploits a much larger set of candidate words which are learned separately, allowing for greater flexibility in the generated text.

Given this attribute vocabulary, we can associate each image with a set of attributes according to its captions. We then wish to predict the attributes given a test image. Because we do not have ground truth bounding boxes for attributes, we cannot train a detector for each using the standard approach. Fang et al. solved a similar problem using a Multiple Instance Learning framework to detect visual words from images. Motivated by the relatively small number of times that each word appears in a caption, we instead treat this as a multi-label classification problem. To address the concern that some attributes may only apply to image sub-regions, we follow Wei et al. in designing a region-based multi-label classification framework that takes an arbitrary number of sub-region proposals as input, then a shared CNN is associated with each proposal, and the CNN output results from different proposals are aggregated with max pooling to produce the final prediction.

Figure 3 summarizes the attribute prediction network. The model is a VggNet structure followed by a max-pooling operation on the regions with a multi-label loss. The CNN model is first initialized from the VggNet pre-trained on ImageNet. The shared CNN is then fine-tuned on the target multi-label dataset (our image-attribute training data). In this step, the input is the global image and the output of the last fully-connected layer is fed into a cc-way softmax over the cc class labels. The cc here represents the attributes vocabulary size. In contrast to who employs the squared loss, we find that element-wise logistic loss function performs better. Suppose that there are NN training examples and yi=[yi1,yi2,...,yic]\bm{y_{i}}=[y_{i1},y_{i2},...,y_{ic}] is the label vector of the ithi^{th} image, where yij=1y_{ij}=1 if the image is annotated with attribute jj, and yij=0y_{ij}=0 otherwise. If the predictive probability vector is pi=[pi1,pi2,...,pic]\bm{p_{i}}=[p_{i1},p_{i2},...,p_{ic}], the cost function to be minimized is

During the fine-tuning process, the parameters of the last fully connected layer (i.e. the attribute prediction layer) are initialized with a Xavier initialization . The learning rates of ‘fc6’ and ‘fc7’ of the VggNet are initialized as 0.001 and the last fully connected layer is initialized as 0.01. All the other layers are fixed during training. We executed 40 epochs in total and decreased the learning rate to one tenth of the current rate for each layer after 10 epochs. The momentum is set to 0.9. The dropout rate is set to 0.5.

To predict attributes based on regions, we first extract hundreds of proposal windows from an image. However, considering the computational inefficiency of deep CNNs, the number of proposals processed needs to be small. Similar to , we first apply the normalized cut algorithm to group the proposal bounding boxes into mm clusters based on the IoU scores matrix. The top kk hypotheses in terms of the predictive scores reported by the proposal generation algorithm are kept and fed into the shared CNN. We also include the whole image in the hypothesis group. As a result, there are mk+1mk+1 hypotheses for each image. We set m=10,k=5m=10,k=5 in all experiments. We use Multiscale Combinatorial Grouping (MCG) for the proposal generation. Finally, a cross hypothesis max-pooling is applied to integrate the outputs into a single prediction vector Vatt(I){V_{att}}(I).

Since we formulate the attribute prediction as a multi-label problem, our attributes prediction network can be replaced by any other multi-label classification framework and it also can be benefit from the development of the multi-label classification researches. For example, to address the computational inefficiency of using a large numbers of proposed regions, we can apply an ‘R-CNN’ architecture so that we do not need to compute the convolutional feature map multiple times. The Regional Proposal Network can predict region proposal and attributes together so that we do not need the external region proposal tools. We even can consider the attributes dependencies by using the recently proposed CNN-RNN model . However, we leave them as the further work.

2 Caption Generation Model

Similar to , we propose to train a caption generation model by maximizing the probability of the correct description given the image. However, rather than using image features directly as in typically the case, we use the semantic attribute prediction value Vatt(I){V_{att}}(I) from the previous section as the input. Suppose that {S1,...,SL}\{S_{1},...,S_{L}\} is a sequence of words. The log-likelihood of the words given their context words and the corresponding image can be written as:

where p(St∣S1:t−1,Vatt(I))p(S_{t}|S_{1:t-1},{V_{att}}(I)) is the probability of generating the word StS_{t} given attribute vector Vatt(I){V_{att}}(I) and previous words S1:t−1S_{1:t-1}. We employ the LSTM , a particular form of RNN, to model this.

The LSTM is a memory cell encoding knowledge at every time step for what inputs have been observed up to this step. We follow the model used in . Letting σ\sigma be the sigmoid nonlinearity, the LSTM updates for time step tt given inputs xtx_{t}, ht−1h_{t-1}, ct−1c_{t-1} are:

Here, it,ft,ct,oti_{t},f_{t},c_{t},o_{t} are the input, forget, memory, output state of the LSTM. The various WW matrices are trained parameters and ⊙\odot represents the product with a gate value. hth_{t} is the hidden state at time step tt and is fed to a Softmax, which will produce a probability distribution pt+1p_{t+1} over all words and indicate the word at time step t+1t+1.

Training details: The LSTM model for image captioning is trained in an unrolled form. More formally, the LSTM takes the attributes vector Vatt(I){V_{att}}(I) and a sequence of words S=(S0,...,SL,SL+1)S=(S_{0},...,S_{L},S_{L+1}), where S0S_{0} is a special start word and SL+1S_{L+1} is a special END token. Each word has been represented as a one-hot vector StS_{t} of dimension equal to the size of words dictionary. The words dictionaries are built based on words that occur at least 5 times in the training set, which lead to 2538, 7414, and 8791 words on Flickr8k, Flickr30k and MS COCO datasets separately. Note it is different from the semantic attributes vocabulary Vatt\mathcal{V}_{att}. The training procedure is as following: At time step t=−1t=-1, we set x−1=WeaVatt(I)x_{-1}=W_{ea}{V_{att}}(I), hinitial=0⃗h_{initial}=\vec{0} and cinitial=0⃗c_{initial}=\vec{0}, where WeaW_{ea} is the learnable attributes embedding weights. This gives us an initial LSTM hidden state h−1h_{-1} which can be used in the next time step. From t=0t=0 to t=Lt=L, we set xt=WesStx_{t}=W_{es}S_{t} and the hidden state ht−1h_{t-1} is given by the previous step, where WesW_{es} is the learnable word embedding weights. The probability distribution pt+1p_{t+1} over all words is then computed by the LSTM feed-forward process. Finally, on the last step when SL+1S_{L+1} represents the last word, the target label is set to the END token.

Our training objective is to learn parameters WeaW_{ea}, WesW_{es} and all parameters in LSTM by minimizing the following cost function:

where NN is the number of training examples and L(i)L^{(i)} is the length of the sentence for the ii-th training example. pt(St(i))p_{t}(S_{t}^{(i)}) corresponds to the activation of the Softmax layer in the LSTM model for the ii-th input and θ{\bm{\theta}} represents model parameters, λθ⋅∣∣θ∣∣22\lambda_{\bm{\theta}}\cdot||{\bm{\theta}}||_{2}^{2} is a regularization term. We use SGD with mini-batches of 100 image-sentence pairs. The attributes embedding size, word embedding size and hidden state size are all set to 256 in all the experiments. The learning rate is set to 0.001 and clip gradients is 5. The dropout rate is set to 0.5.

To infer the sentence given an input image, we use Beam Search, i.e., we iteratively consider the set of bb best sentences up to time tt as candidates to generate sentences at time t+1t+1, and only keep the best bb results. We set the bb to 5. Figure 4 shows some examples of the predicted attributes and generated captions. More results can be found in the supplementary material.

A VQA Model with External Knowledge

The key differentiator of our VQA model is that it is able to usefully combine image information with that extracted from a Knowledge Base, within the LSTM framework. The novelty lies in the fact that this is achieved by representing both of these disparate forms of information as text before combining them. Figure 5 summarises how this is achieved: given an image, an attribute-based representation Vatt(I){V_{att}}(I) (in Section 3.1) is first generated and it will used as one of input sources of our VQA-LSTM model. The second input source are those captions generated in section 3.2. Rather than inputing the generated words directly, the hidden state vector of the caption-LSTM after it has generated the last word in each caption is used to represent its content. Average-pooling is applied over the 5 hidden-state vectors, to obtain a vector representation Vcap(I){V_{cap}}(I) for the image II. The third input source is the textual knowledge which is mined from a large-scale knowledge base, the DBpedia. More details are shown in the following section.

The external data source that we use here is DBpeida as a source of general background information, although any such KB could equally be applied. DBpeida is a structured database of information extracted from Wikipedia. The whole DBpedia dataset describes 4.584.58 million entities, of which 4.224.22 million are classified in a consistent ontology. The data can be accessed using an SQL-like query language for RDF called SPARQL. Given an image and its predicted attributes, we use the top-fiveWe only use top-5 attributes to query the KB because, based on observation of training data, an image typically contains 5-8 attributes. We also tested with top-10, but no improvements were observed. most strongly predicted attributes to generate DBpedia queries. There are a range of problems with DBpedia and similar, however, including the sparsity of the information, and the inconsistency of its representation. Inspecting the database shows that the ‘comment’ field is the most generally informative about an attribute, as it contains a general text description of it. We therefore retrieve the comment text for each query term. The KB+SPARQL combination is very general, however, and could be applied problem specific KBs, or a database of common sense information, and can even perform basic inference over RDF. Figure 6 shows an example of the query language and returned text.

Since the text returned by the SPARQL query is typically much longer than the captions generated in the section 3.2, we turn to Doc2Vec to extract the semantic meaningsWe investigated to use an LSTM to encode the mined paragraphs, but we observed little performance improvement, despite the additional training overhead.. Doc2Vec, also known as Paragraph Vector, is an unsupervised algorithm that learns fixed-length feature representations from variable-length pieces of texts, such as sentences, paragraphs, and documents. Le et al. proved that it can capture the semantics of paragraphs. A Doc2Vec model is trained to predict words in the document given the context words. We collect 100,000 documents from DBpedia to train a model with vector size 500. To obtain the knowledge vector Vknow(I){V_{know}}(I) for image II, we combine the 5 returned paragraphs in to a single large paragraph, before semantic features using our pre-trained Doc2Vec model.

2 Question-guided Knowledge Selection

We incrementally implemented a question-guided knowledge selection scheme to rule out the noise information, since we observed that some mined knowledge are not necessary for answering the given question. For example, if the question is asking about the ‘dog’ in the image, it does not make sense to input a piece of ‘bird’ knowledge into the model, although the image does have a ‘bird’ inside.

Given a question QQ and mined nn knowledge paragraphs using above KB+SPARQL combination, we first use our pre-trained Doc2Vec model to extract the semantic feature V(Q)V(Q) of the question and the feature V(Ki)V(K_{i}) for each single knowledge paragraph, where i∈ni\in n. Then, we find the kk closest knowledge paragraphs to the question based on the cosine similarity between the V(Q)V(Q) and V(Ki)V(K_{i}). Finally, we combine the kk selected knowledge paragraphs in to a single one and use the Doc2Vec model to extract its semantic feature. In our experiments, we set n=10,k=5n=10,k=5.

3 An Answer Generation Model with Multiple Inputs

We propose to train a VQA model by maximizing the probability of the correct answer given the image and question. We want our VQA model to be able to generate multiple word answers, so we formulate the answering process as a word sequence generation procedure. Let Q={q1,...,qn}Q=\{q_{1},...,q_{n}\} represent the sequence of words in a question, and A={a1,...,al}A=\{a_{1},...,a_{l}\} the answer sequence, where nn and ll are the length of question and answer, respectively. The log-likelihood of the generated answer can be written as:

where p(at∣a1:t−1,I,Q)p(a_{t}|a_{1:t-1},I,Q) is the probability of generating ata_{t} given image information II, question QQ and previous words a1:t−1a_{1:t-1}. We employ an encoder LSTM to take the semantic information from image II and the question QQ, while using a decoder LSTM to generate the answer. Weights are shared between the encoder and decoder LSTM.

In the training phase, the question QQ and answer AA are concatenated as {q1,...,qn,a1,...,al,al+1}\{q_{1},...,q_{n},a_{1},...,a_{l},a_{l+1}\}, where al+1a_{l+1} is a special END token. Each word is represented as a one-hot vector of dimension equal to the size of the word dictionary. The training procedure is as follows: at time step t=0t=0, we set the LSTM input:

where WeaW_{ea}, WecW_{ec}, WekW_{ek} are learnable embedding weights for the vector representation of attributes, captions and external knowledge, respectively. Given the randomly initialized hidden state, the encoder LSTM feeds forward to produce hidden state h0h_{0} which encodes all of the input information. From t=1t=1 to t=nt=n, we set xt=Wesqtx_{t}=W_{es}q_{t} and the hidden state ht−1h_{t-1} is given by the previous step, where WesW_{es} is the learnable word embedding weights. The decoder LSTM runs from time step n+1n+1 to l+1l+1. Specifically, at time step t=n+1t=n+1, the LSTM layer takes the input xn+1=Wesa1x_{n+1}=W_{es}a_{1} and the hidden state hnh_{n} corresponding to the last word of the question, where a1a_{1} is the start word of the answer. The hidden state hnh_{n} thus encodes all available information about the image and the question. The probability distribution pt+1p_{t+1} over all answer words in the vocabulary is then computed by the LSTM feed-forward process. Finally, for the final step, when al+1a_{l+1} represents the last word of the answer, the target label is set to the END token.

Our training objective is to learn parameters WeaW_{ea}, WecW_{ec}, WekW_{ek}, WesW_{es} and all the parameters in the LSTM by minimizing the following cost function:

where NN is the number of training examples, and n(i)n^{(i)} and l(i)l^{(i)} are the length of question and answer respectively for the ii-th training example. Let pt(at(i))p_{t}(a_{t}^{(i)}) correspond to the activation of the Softmax layer in the LSTM model for the ii-th input and θ{\bm{\theta}} represent the model parameters. Note that λθ⋅∣∣θ∣∣22\lambda_{\bm{\theta}}\cdot||{\bm{\theta}}||_{2}^{2} is a regularization term, where λθ=0.5×10−8\lambda_{\theta}=0.5\times 10^{-8}. We use Stochastic gradient Descent (SGD) with mini-batches of 100 image-QA pairs. The attributes, internal textual representation, external knowledge embedding size, word embedding size and hidden state size are all 256 in all experiments. The learning rate is set to 0.001 and clip gradients is 5. The dropout rate is set to 0.5.

Experiments

We report image captioning results on the popular Flickr8k , Flickr30k and Microsoft COCO dataset . These datasets contain 8,000, 31,000 and 123,287 images respectively, and each image is annotated with 5 sentences. In our reported results, we use pre-defined splits for Flickr8k.Because most of previous works in image captioning are not evaluated on the official split for Flickr30k and MS COCO, for fair comparison, we report results with the widely used publicly available splits in the work of .We further tested on the actually MS COCO test set consisting of 40775 images (human captions for this split are not available publicly), and evaluated them on the COCO evaluation server.

1.2 Evaluation

Metrics: We report results with the frequently used BLEU metric and sentence perplexity (PPL\mathcal{PPL}). For MS COCO dataset, we additionally evaluate our model based on the metrics of METEOR and CIDEr .

Baselines: To verify the effectiveness of our high-level attributes representation, we provide a baseline method. The baseline framework is same as the one proposed in section 3.2, except that the attributes vector Vatt(I){V_{att}}(I) is replaced by the last hidden layer of CNN directly. For the VggNet+LSTM, we use the second fully connected layer (fc7) as the image features, which has 4096 dimensions. In VggNet-PCA+LSTM, PCA is applied to decrease the feature dimension from 4096 to 1000. VggNet+ft+LSTM applies a VggNet that has been fine-tuned on the target dataset, based on the task of image-attributes classification.

Our Approaches: We evaluate several variants of our approach: Att-GT+LSTM models use ground-truth attributes as the input while Att-RegionCNN+LSTM uses the attributes vector Vatt(I){V_{att}}(I) predicted by the region based attributes prediction network in section 3.1. We also evaluate an approach Att-SVM+LSTM with linear SVM predicted attributes vector. We use the second fully connected layer of the fine-tuned VggNet to feed the SVM. To verify the effectiveness of the region based attributes prediction in the captioning task, the Att-GlobalCNN+LSTM is implemented by using the global image for attributes prediction.

Results: Table I and II report image captioning results on Flickr8k, Flickr30k and Microsoft COCO dataset. It is not surprising that Att-GT+LSTM model performs best, since ground truth attributes labels are used. We report these results here just to show the advances of adding an intermediate image-to-word mapping stage. Ideally, if we could train a perfectly accurate attribute predictor, we could obtain an outstanding improvement compared to both baseline and state-of-the-art methods. Indeed, apart from using ground truth attributes, our Att-RegionCNN+LSTM models generate the best results on all the three datasets over all evaluation metrics. Especially comparing with baselines, which do not contain an attributes prediction layer, our final models bring significant improvements, nearly 15% for B-1 and 30% for CIDEr on average. VggNet+ft+LSTM models perform better than other baselines because of the fine-tuning on the target dataset. However, they do not perform as well as our attributes-based models. Att-SVM+LSTM and Att-GlobalCNN+LSTM under-perform Att-RegionCNN+LSTM, indicating that region-based attributes prediction provides useful detail beyond whole image classification. Our final model also outperforms the current state-of-the-art listed in the tables. We also evaluated an approach (not shown in table) that combines CNN features and attributes vector together as the input of the LSTM, but we found this approach is not as good as using attributes vector only in the same setting. In any case, above experiments show that an intermediate image-to-words stage (i.e. attributes prediction layer) bring us significant improvements.

We further generated captions for the images in the COCO test set containing 40,775 images and evaluated them on the COCO evaluation server. These results are shown in Table III. We achieve 0.73 on B-1, and surpass human performances on 13 of the 14 metrics reported. Other state-of-the-art methods are also shown for comparison.

Human Evaluation: We additionally perform a human evaluation on our proposed model, to evaluate the caption generation ability. We randomly sample 1000 results from the COCO validation dataset, generated by our proposed model Att-RegionCNN+LSTM and the baseline model VggNet+LSTM. Following the human evaluation protocol of the MS COCO Captioning Challenge 2015, two evaluation metrics are applied. M1 is the percentage of captions that are evaluated as better or equal to human caption and M2 is the percentage of captions that pass the Turing Test. Table IV summarizes the human evaluation results. We can see our model outperforms the baseline model on both metrics. We did not evaluate on the test split because the human ground truth is not publicly available.

Table V summarizes some properties of recurrent layers employed in some recent RNN-based methods. We achieve state-of-the-art using a relatively low dimensional visual input feature and recurrent layer. Lower dimension of visual input and RNN normally means less parameters in the RNN training stage, as well as lower computation cost.

2 Evaluation on Visual Question Answering

We evaluate our model on four recent publicly available visual question answering datasets. DAQURA-ALL is proposed in . There are 7,795 training questions and 5,673 test questions. DAQURA-REDUCED is a reduced version of DAQURA-ALL. There are 3,876 training questions and only 297 test questions. This dataset is constrained to 37 object categories and uses only 25 test images. Two large-scale VQA data are constructed both based on MS COCO images. The Toronto COCO-QA Dataset contains 78,736 training and 38,948 testing examples, which are generated from 117,684 images. All of the question-answer pairs in this dataset are automatically converted from human-sourced image descriptions. Another benchmarked dataset is VQA , which is a much larger dataset and contains 614,163 questions and 6,141,630 answers based on 204,721 MS COCO images. We randomly choose 5000 images from the validation set as our val set, with the remainder testing. The human ground truth answers for the actual VQA test split are not available publicly and only can be evaluated via the VQA evaluation server. Hence, we also apply our final model on a test split and report the overall accuracy. Table VI displays some dataset statistics.

Metrics: Following , the accuracy value (the proportion of correctly answered test questions), and the Wu-Palmer similarity (WUPS) are used to measure performance. The WUPS calculates the similarity between two words based on the similarity between their common subsequence in the taxonomy tree. If the similarity between two words is greater than a threshold then the candidate answer is considered to be right. We report on thresholds 0.9 and 0.0, following .

Evaluations: To illustrate the effectiveness of our model, we provide two baseline models and several state-of-the-art results in table VII and VIII. The Baseline method is implemented simply by connecting a CNN to an LSTM. The CNN is a pre-trained (on ImageNet) VggNet model from which we extract the coefficients of the last fully connected layer. We also implement a baseline model VggNet+ft-LSTM, which applies a vggNet that has been fine-tuned on the COCO dataset, based on the task of image-attributes classification. We also present results from a series of cut down versions of our approach for comparison. Att-LSTM uses only the semantic level attribute representation Vatt{V_{att}} as the LSTM input. To evaluate the contribution of the internal textual representation and external knowledge for the question answering, we feed the image caption representation Vcap{V_{cap}} and knowledge representation Vknow{V_{know}} with the Vatt{V_{att}} separately, producing two models, Att+Cap-LSTM and Att+Know-LSTM. We also tested the Cap+Know-LSTM, for the experiment completeness. Att+Cap+Know-LSTM combines all the available information. Our final model is the A+C+Selected-K-LSTM, which uses the selected knowledge information (see section 4.2) as the input. GUESS simply selects the modal answer from the training set for each of 4 question types. VIS+BOW performs multinomial logistic regression based on image features and a BOW vector obtained by summing all the word vectors of the question. VIS+LSTM has one LSTM to encode the image and question, while 2-VIS+BLSTM has two image feature inputs, at the start and the end. Malinowskiet al. propose a neural-based approach and Ma et al. encodes both images and questions with a CNN. Yang et al. use a stacked attention networks to infer the answer progressively.

All of our proposed models outperform the Baseline method. And our final model A+C+Selected-K-LSTM achieves the best state-of-the-art on the DAQURA-Reduced set. Att+Cap+Know-LSTM performs not as good as A+C+Selected-K-LSTM, which shows the effectiveness of our question-based knowledge selection scheme.

2.2 Results on Toronto COCO-QA

Table IX reports the results on Toronto COCO-QA. All of our proposed models outperform the Baseline and all of the comparator state-of-the-art methods. Our final model A+C+Selected-K-LSTM achieves the best results. It surpasses the baseline by nearly 20% and outperforms the previous state-of-the-art methods around 10%. Att+Cap-LSTM clearly improves the results over the Att-LSTM model. This proves that internal textual representation plays a significant role in the VQA task. The Att+Know-LSTM model does not perform as well as Att+Cap-LSTM , which suggests that the information extracted from captions is more valuable than that extracted from the KB. Cap+Know-LSTM also performs better than Att+Know-LSTM. This is not surprising because the Toronto COCO-QA questions were generated automatically from the MS COCO captions, and thus the fact that they can be answered by training on the captions is to be expected. This generation process also leads to questions which require little external information to answer. The comparison on the Toronto COCO-QA thus provides an important benchmark against related methods, but does not really test the ability of our method to incorporate extra information. It is thus interesting that the additional external information provides any benefit at all.

Table X shows the per-category accuracy for different models. Surprisingly, the counting ability (see question type ‘Number’) increases when both captions and external knowledge are included. This may be because some ‘counting’ questions are not framed in terms of the labels used in the MS COCO captions. Ren et al.also observed similar cases. In they mentioned that “there was some observable counting ability in very clean images with a single object type but the ability was fairly weak when different object types are present”. We also find there is a slight increase for the ‘color’ questions when the KB is used. Indeed, some questions like ‘What is the color of the stop sign?’ can be answered directly from the KB, without the visual cue.

2.3 Results on VQA

Antol et al. provide the VQA dataset which is intended to support “free-form and open-ended Visual Question Answering”. They also provide a metric for measuring performance: min⁡{# humans that said answer3,1}\min\{\frac{\text{\# humans that said answer}}{3},1\} thus 100%100\% means that at least 3 of the 10 humans who answered the question gave the same answer.

Evaluation: There are several splits for VQA dataset, such as the validation set, test-develop and test-standard set. We first tested several aspects of our models on the validation set (we randomly choose 5000 images from the validation set as our val set, with the remainder testing).

Inspecting Table XI, results the on VQA validation set, we see that the attribute-based Att-LSTM is a significant improvement over our VggNet+LSTM baseline. We also evaluate another baseline, the VggNet+ft+LSTM, which uses the penultimate layer of the attributes prediction CNN (in Section 3.1) as the input to the LSTM. Its overall accuracy on the VQA is 50.01, which is still lower than our proposed models (detailed results of different question types are not shown in Table XI due to the limited space.) Adding either image captions or external knowledge further improves the result. Our final model A+C+S-K-LSTM produces the best results, outperforming the baseline VggNet-LSTM by 11%.

Figure 7 relates the performance of the various models on five categories of questions. The ‘object’ category is the average of the accuracy of question types starting with ‘what kind/type/sport/animal/brand…’, while the ‘number’ and ‘color’ category corresponds to the question type ‘how many’ and ‘what color’. The performance comparison across categories is of particular interest here because answering different classes of questions requires different amounts of external knowledge. The ’Where’ questions, for instance, require knowledge of potential locations, and ’Why’ questions typically require general knowledge about people’s motivation. ’Number’ and ’Color’ questions, in contrast, can be answered directly. The results show that for ’Why’ questions, adding the KB improves performance by more than 50%50\% (Att-LSTM achieves 7.77%7.77\% while Att+Know-LSTM achieves 11.88%11.88\%), and that the combined A+C+K-LSTM achieves 13.53%13.53\%. We further improve it to 13.76%13.76\% by using the question-guided knowledge selected model A+C+S-K-LSTM. Compared with the Att-LSTM model, the performance gain of the Cap+Know-LSTM model mainly come from the ‘why’ and ‘where’ started questions. This means that the external knowledge we employed in the model provide useful information to answer such questions. The figure 1 shows an real example produced by our model. More questions that require common-sense knowledge to answer can be found in the supplementary materials.

We have also tested on the VQA test-dev and test-standard consisting of 60,864 and 244,302 questions (for which ground truth answers are not published) using our final A+C+S-K-LSTM model, and evaluated them on the VQA evaluation server. Table XII shows the server reported results. The results on the Test-dev can be found in the supplementary material.

Antol et al. provide several results for this dataset. In each case they encode the image with the final hidden layer from VggNet, and questions and captions are encoded using a BOW representation. A softmax neural network classifier with 2 hidden layers and 1000 hidden units (dropout 0.5) in each layer with tanh non-linearity is then trained, the output space of which is the 1000 most frequent answers in the training set. They also provide an LSTM model followed by a softmax layer to generate the answer. Two version of this approach are used, one which is given only the question and the image, and one which is given only the question (see for details). Our final model outperforms all the listed approaches according to the overall accuracy. Figure 8 provides some indicative results. More results can be found in the supplementary material.

Conclusions

In this paper, we first examined the importance of introducing an intermediate attribute prediction layer into the predominant CNN-LSTM framework, which was neglected by almost all previous work. We implemented an attribute-based model which can be applied to the task of image captioning. We have shown that an explicit representation of image content improves V2L performance, in all cases. Indeed, at the time of submitting this paper, our image captioning model outperforms the state-of-the-art on several captioning datasets.

Secondly, in this paper we have shown that it is possible to extend the state-of-the-art RNN-based VQA approach so as to incorporate the large volumes of information required to answer general, open-ended, questions about images. The knowledge bases which are currently available do not contain much of the information which would be beneficial to this process, but nonetheless can still be used to significantly improve performance on questions requiring external knowledge (such as ’Why’ questions). The approach that we propose is very general, however, and will be applicable to more informative knowledge bases should they become available. We further implement a knowledge selection scheme which reflects both of the content of the question and the image, in order to extract more specifically related information. Currently our system is the state-of-the-art on three VQA datasets and produces the best results on the VQA evaluation server.

Further work includes generating knowledge-base queries which reflect the content of the question and the image, in order to extract more specifically related information. The Knowledge Base itself also can be improved. For instance, Open-IE provides more general common-sense knowledge such as ‘cats eat fish’. Such knowledge will help answer high-level questions.

Acknowledgements

This research was in part supported by the Data to Decisions Cooperative Research Centre. Correspondence should be addressed to C. Shen.

References