Straight to the Facts: Learning Knowledge Base Retrieval for Factual Visual Question Answering

Medhini Narasimhan, Alexander G. Schwing

Introduction

When answering questions given a context, such as an image, we seamlessly combine the observed content with general knowledge. For autonomous agents and virtual assistants which naturally participate in our day to day endeavors, where answering of questions based on context and general knowledge is most natural, algorithms which leverage both observed content and general knowledge are extremely useful.

To address this challenge, in recent years, a significant amount of research has been devoted to question answering in general and Visual Question Answering (VQA) in particular. Specifically, the classical VQA tasks require an algorithm to answer a given question based on the additionally provided context, given in the form of an image. For instance, significant progress in VQA was achieved by introducing a variety of VQA datasets with strong baselines . The images in these datasets cover a broad range of categories and the questions are designed to test perceptual abilities such as counting, inferring spatial relationships, and identifying visual cues. Some challenging questions require logical reasoning and memorization capabilities. However, the majority of the questions can be answered by solely examining the visual content of the image. Hence, numerous approaches to solve these problems focus on extracting visual cues using deep networks.

We note that many of the aforementioned methods focus on the visual aspect of the question answering task, i.e., the answer is predicted by combining representations of the question and the image. This clearly contrasts the described human-like approach, which combines observations with general knowledge. To address this discrepancy, in very recent meticulous work, Wang et al. introduced a ‘fact-based’ VQA task (FVQA), an accompanying dataset, and a knowledge base of facts extracted from three different sources, namely WebChild , DBPedia , and ConceptNet . Different from the classical VQA datasets, Wang et al. argued that such a dataset can be used to develop algorithms which answer more complex questions that require a combination of observation and general knowledge. In addition to the dataset, Wang et al. also developed a model which leverages the information present in the supporting facts to answer questions about an image.

To this end, Wang et al. design an approach which extracts keywords from the question and retrieves facts that contain those keywords from the knowledge base. Clearly, synonyms and homographs pose challenges which are hard to recover from.

To address this issue, we develop a learning based retrieval method. More specifically, our approach learns a parametric mapping of facts and question-image pairs to an embedding space. To answer a question, we use the fact that is most aligned with the provided question-image pair. As illustrated in Fig. 1, our approach is able to accurately answer both more visual questions as well as more fact based questions. For instance, given the image illustrated on the left hand side along with the question, “Which object in the image can be used to eat with?”, we are able to predict the correct answer, “fork.” Similarly, the proposed approach is able to predict the correct answer for the other two examples. Quantitatively we demonstrate the efficacy of the proposed approach on the recently introduced FVQA dataset, outperforming state-of-the-art by more than 5%5\% on the top-1 accuracy metric.

Related Work

We develop a framework for visual question answering that benefits from a rich knowledge base. In the following, we first review classical visual question answering tasks before discussing visual question answering methods that take advantage of knowledge bases.

Visual Question Answering. In recent years, a significant amount of research has been devoted to developing techniques which can answer a question about a provided context such as an image. Of late, visual question answering has also been used to assess reasoning capabilities of state-of-the-art predictors. Using a variety of datasets , models based on multi-modal representation and attention , deep network architectures , and dynamic memory nets have been developed. Despite these efforts, assessing the reasoning capabilities of present day deep network-based approaches and differentiating them from mere memorization of training set statistics remains a hard task. Most of the methods developed for visual question answering focus exclusively on answering questions related to observed content. To this end, these methods use image features extracted from networks such as the VGG-16 trained on large image datasets such as ImageNet . However, it is unlikely that all the information which is required to answer a question is encoded in the features extracted from the image, or even the image itself. For example, consider an image containing a dog, and a question about this image, such as “Is the animal in the image capable of jumping in the air?”. In such a case, we would want our method to combine common sense and general knowledge about the world, such as the ability of a healthy dog to jump, along with features and observations from the image, such as the presence of the dog. This motivates us to develop methods that can use knowledge bases encoding general knowledge.

Knowledge-based Visual Question Answering. There has been interest in the natural language processing community in answering questions based on knowledge bases (KBs) using either semantic parsing or information retrieval methods. However, knowledge based visual question answering is still relatively unexplored, even though this is appealing from a practical standpoint as this decouples the reasoning by the neural network from the storage of knowledge in the KB. Notable examples in this direction are work by Zhu et al. , Wu et al. , Wang et al. , Krishnamurthy and Kollar , and Narasimhan et al. .

The works most related to our approach include Ask Me Anything (AMA) by Wu et al. , Ahab by Wang et al. , and FVQA by Wang et al. . AMA describes the content of an image in terms of a set of attributes predicted about the image, and multiple captions generated about the image. The predicted attributes are used to query an external knowledge base, DBpedia , and the retrieved paragraphs are summarized to form a knowledge vector. The predicted attribute vector, the captions, and the database-based knowledge vector are passed as inputs to an LSTM that learns to predict the answer to the input question as a sequence of words. A drawback of this work is that it does not perform any explicit reasoning and ignores the possible structure in the KB. Ahab and FVQA, on the other hand, attempt to perform explicit reasoning. Ahab converts an input question into a database query, and processes the returned knowledge to form the final answer. Similarly, FVQA learns a mapping from questions to database queries through classifying questions into categories and extracting parts from the question deemed to be important. While both of these methods rely on fixed query templates, this very structure offers some insight into what information the method deems necessary to answer a question about a given image. Both these methods use databases with a particular structure: those that contain facts about visual concepts represented as tuples, for example, (Cat, CapableOf, Climbing), and (Dog, IsA, Pet). We develop our method on the dataset released as part of the FVQA work, referred to as the FVQA dataset , which is a subset of three structured databases – DBpedia , ConceptNet , and WebChild . The method presented in FVQA produces a query as an output of an LSTM which is fed the question as an input. Facts in the knowledge base are filtered on the basis of visual concepts such as objects, scenes, and actions extracted from the input image. The predicted query is then applied on the filtered database, resulting in a set of retrieved facts. A matching score is then computed between the retrieved facts and the question to determine the most relevant fact. The most correct fact forms the basis of the answer for the question.

In contrast to Ahab and FVQA, we propose to directly learn an embedding of facts and question-image pairs into a space that permits to assess their compatibility. This has two important advantages over prior work: 1) by avoiding the generation of an explicit query, we eliminate errors due to synonyms, homographs, and incorrect prediction of visual concept type and answer type; and 2) our technique is easy to extend to any knowledge base, even one with a different structure or size. We also do not require any ad-hoc filtering of knowledge, and can instead learn to transform extracted visual concepts into a vector close to a relevant fact in the learned embedding space. Our method also naturally produces a ranking of facts deemed to be useful for the given question and image.

Learning Knowledge Base Retrieval

In the following, we first provide an overview of the proposed approach for knowledge based visual question answering before discussing our embedding space and learning formulation.

Overview. Our developed approach is outlined in Fig. 2. The task at hand is to predict an answer yy for a question QQ given an image xx by using an external knowledge base KB, which consists of a set of facts fif_{i}, i.e., KB={f1,…,f∣KB∣}\text{KB}=\left\{f_{1},\ldots,f_{|\text{KB}|}\right\}. Each fact fif_{i} in the knowledge base is represented as a Resource Description Framework (RDF) triplet of the form fi=(ai,ri,bi)f_{i}=(a_{i},r_{i},b_{i}), where aia_{i} is a visual concept in the image, bib_{i} is an attribute or phrase associated with the visual entity aia_{i}, and ri∈Rr_{i}\in{\cal R} is a relation between the two entities. The dataset contains ∣R∣=13|{\cal R}|=13 relations r∈R={r\in{\cal R}=\{Category, Comparative, HasA, IsA, HasProperty, CapableOf, Desires, RelatedTo, AtLocation, PartOf, ReceivesAction, UsedFor, CreatedBy}\}. Example triples of the knowledge base in our dataset are (Umbrella, UsedFor, Shade), (Beach, HasProperty, Sandy), (Elephant, Comparative-LargerThan, Ant).

To answer a question QQ correctly given an image xx, we need to retrieve the right supporting fact and choose the correct entity, i.e., either aa or bb. Importantly, entity aa is always derived from the image and entity bb is derived from the fact base. Consequently we refer to this choice as the answer source s∈{Image,KnowledgeBase}s\in\left\{\text{Image},\text{KnowledgeBase}\right\}. Using this formulation, we can extract the answer yy from a predicted fact f^=(a^,r^,b^)\hat{f}=(\hat{a},\hat{r},\hat{b}) and a predicted answer source s^\hat{s} using

It remains to answer, how to predict a fact f^\hat{f} and how to infer the answer source s^\hat{s}. The latter is a binary prediction task and we describe our approach below. For the former, we note that the knowledge base contains a large number of facts. We therefore consider it infeasible to search through all the facts fif_{i} ∀i∈{1,…,∣KB∣}\forall i\in\{1,\ldots,|\text{KB}|\} using an expensive evaluation based on a deep net. We therefore split this task into two parts: (1) Given a question, we train a network to predict the relation r^\hat{r}, that the question focuses on. (2) Using the predicted relation, r^\hat{r}, we reduce the fact space to those containing only the predicted relation.

Subsequently, to answer the question QQ given image xx, we only assess the suitability of the facts which contain the predicted relation r^\hat{r}. To assess the suitability, we design a score function S(gF(fi),gNN(x,Q))S(g^{\text{F}}(f_{i}),g^{\text{NN}}(x,Q)) which measures the compatibility of a fact representation gF(fi)g^{\text{F}}(f_{i}) and an image-question representation gNN(x,Q)g^{\text{NN}}(x,Q). Intuitively, the higher the score, the more suitable the fact fif_{i} for answering question QQ given image xx.

Formally, we hence obtain the predicted fact f^\hat{f} via

where we search for the fact f^\hat{f} maximizing the score SS among all facts fif_{i} which contain relation r^\hat{r}, i.e., among all fif_{i} with i∈{j:rel⁡(fj)=r^}i\in\{j:\operatorname{rel}(f_{j})=\hat{r}\}. Hereby we use the operator rel⁡(fi)\operatorname{rel}(f_{i}) to indicate the relation of the fact triplet fif_{i}. Given the predicted fact using Eq. (2) we obtain the answer yy from Eq. (1) after predicting the answer source s^\hat{s}.

This approach is outlined in Fig. 2. Pictorially, we illustrate the construction of an image-question embedding gNN(x,Q)g^{\text{NN}}(x,Q), via LSTM and CNN net representations that are combined via an MLP. We also illustrate the fact embedding gF(fi)g^{\text{F}}(f_{i}). Both of them are combined using the score function S(⋅,⋅)S(\cdot,\cdot), to predict a fact f^\hat{f} from which we extract the answer as described in Eq. (1).

In the following, we first provide details about the score function SS, before discussing prediction of the relation r^\hat{r} and prediction of the answer source s^\hat{s}.

Scoring the facts. Fig. 2 illustrates our approach to score the facts in the knowledge base, i.e., to compute S(gF(fi),gNN(x,Q))S(g^{\text{F}}(f_{i}),g^{\text{NN}}(x,Q)). We obtain the score in three steps: (1) computing of a fact representation gF(fi)g^{\text{F}}(f_{i}); (2) computing of an image-question representation gNN(x,Q)g^{\text{NN}}(x,Q); (3) combination of the fact and image-question representation to obtain the final score SS. We discuss each of those steps in the following.

(1) Computing a fact representation. To obtain the fact representation gF(fi)g^{\text{F}}(f_{i}), we concatenate two vectors, the averaged GloVe-100 representation of the words of entity aia_{i} and the averaged GloVe-100 representation of the words of entity bib_{i}. Note that this fact representation is non-parametric, i.e., there are no trainable parameters.

(2) Computing an image-question representation. We compute the image-question representation gNN(x,Q)g^{\text{NN}}(x,Q), by combining a visual representation gwV(x)g_{w}^{V}(x), obtained from a standard deep net, e.g., ResNet or VGG, with a visual concept representation gwC(x)g_{w}^{C}(x), and a sentence representation gwQ(Q)g_{w}^{Q}(Q), of the question QQ, obtained using a trainable recurrent net. For notational convenience we concatenate all trainable parameters into one vector ww. Making the dependence on the parameters explicit, we obtain the image-question representation via gwNN(x,Q)=gwNN(gwV(x),gwQ(Q),gwC(x)).g^{\text{NN}}_{w}(x,Q)=g^{\text{NN}}_{w}(g_{w}^{V}(x),g_{w}^{Q}(Q),g_{w}^{C}(x)).

More specifically, for the question embedding gwQ(Q)g^{Q}_{w}(Q), we use an LSTM model . For the image embedding gwV(x)g^{V}_{w}(x), we extract image features using ResNet-152 pre-trained on the ImageNet dataset . In addition, we also extract a visual concept representation gwC(x)g_{w}^{C}(x), which is a multi-hot vector of size 1176 indicating the visual concepts which are grounded in the image. The visual concepts detected in the images are objects, scenes, and actions. For objects, we use the detections from two Faster-RCNN models that are trained on the Microsoft COCO 80-object and the ImageNet 200-object datasets. In total, there are 234 distinct object classes, from which we use that subset of labels that coincides with the FVQA dataset. The scene information (such as pasture, beach, bedroom) is extracted by the VGG-16 model trained on the MIT Places 365-class dataset . Again, we use a subset of Places to construct the 1176-dimensional multi-hot vector gwC(x)g_{w}^{C}(x). For detecting actions, we use the CNN model proposed in which is trained on the HICO and MPII datasets. The HICO dataset contains labels for 600 human-object interaction activities while the MPII dataset contains labels for 393 actions. We use a subset of actions, namely those which coincide with the ones in the FVQA dataset.

All the three vectors gwV(x),gwQ(Q),gwC(x)g_{w}^{V}(x),g_{w}^{Q}(Q),g_{w}^{C}(x) are concatenated and passed to the multi-layer perceptron gwNN(⋅,⋅,⋅)g^{\text{NN}}_{w}(\cdot,\cdot,\cdot).

(3) Combination of fact and image-question representation. For each fact representation gF(fi)g^{\text{F}}(f_{i}), we compute a score

where gwNN(x,Q)g_{w}^{\text{NN}}(x,Q) is the image question representation. Hence, the score SS is the cosine similarity between the two normalized representations and represents the fit of fact fif_{i} to the image-question pair (x,Q)(x,Q).

Predicting the relation. To predict the relation r^∈R=hw1r(Q)\hat{r}\in{\cal R}=h_{w_{1}}^{r}(Q), from the obtained question QQ, we use an LSTM net. More specifically, we first embed and then encode the words of the question QQ, one at a time, and linearly transform the final hidden representation of the LSTM to predict r^\hat{r}, from ∣R∣|{\cal R}| possibilities using a standard multinomial classification. For the results presented in this work, we trained the relation prediction parameters w1w_{1} independently of the score function. We leave a joint formulation to future work.

Predicting the answer source. Prediction of the answer source s^=hw2s(Q)\hat{s}=h_{w_{2}}^{s}(Q) from a given question QQ is similar to relation prediction. Again, we use an LSTM net to embed and encode the words of the question QQ before linearly transforming the final hidden representation to predict s^∈{Image,KnowledgeBase}\hat{s}\in\{\text{Image},\text{KnowledgeBase}\}. Analogous to relation prediction, we train this LSTM net’s parameters w2w_{2} separately and leave a joint formulation to future work.

Learning. As mentioned before, we train the parameters ww (score function), w1w_{1} (relation prediction), and w2w_{2} (answer source prediction) separately. To train w1w_{1}, we use a dataset D1={(Q,r)}{\cal D}_{1}=\{(Q,r)\} containing pairs of question and the corresponding relation which was used to obtain the answer. To learn w2w_{2}, we use a dataset D2={(Q,s)}{\cal D}_{2}=\{(Q,s)\}, containing pairs of question and the corresponding answer source. For both classifiers we use stochastic gradient descent on the classical cross-entropy and binary cross-entropy loss respectively. Note that both the datasets are readily available from .

To train the parameters of the score function we adopt a successive approach operating in time steps t={1,…,T}t=\{1,\ldots,T\}. In each time step, we gradually increase the difficulty of the dataset D(t){\cal D}^{(t)} by mining hard negatives. More specifically, for every question QQ, and image xx, D(0){\cal D}^{(0)} contains the ‘groundtruth’ fact f∗f^{\ast} as well as 99 randomly sampled ‘non-groundtruth’ facts. After having trained the score function on this dataset we use it to predict facts for image-question pairs and create a new dataset D(1){\cal D}^{(1)} which now contains, along with the groundtruth fact, another 99 non-groundtruth facts that the score function assigned a high score to.

Given a dataset D(t){\cal D}^{(t)}, we train the parameters ww of the representations involved in the score function Sw(gF(fi),gwNN(x,Q))S_{w}(g^{\text{F}}(f_{i}),g^{\text{NN}}_{w}(x,Q)), and its image, question, and concept embeddings by encouraging that the score of the groundtruth fact f∗f^{\ast} is larger than the score of any other fact. More formally, we aim for parameters ww which ensure the classical margin, i.e., an SVM-like loss for deep nets:

where L(f∗,f)L(f^{\ast},f) is the task loss (aka margin) comparing the groundtruth fact f∗f^{\ast} to other facts ff. In our case L≡1L\equiv 1. Since we may not find parameters ww which ensure feasibility ∀(f,x,Q)∈D(t)\forall(f,x,Q)\in{\cal D}^{(t)}, we introduce slack variables ξ(f,x,Q)≥0\xi_{(f,x,Q)}\geq 0 to obtain after reformulation:

Instead of enforcing the constraint ∀(f,x,Q)\forall(f,x,Q) in the dataset D(t){\cal D}^{(t)}, it is equivalent to require

Using this constraint, we find the parameters ww by solving

For applicability of the standard sub-gradient descent techniques, we reformulate the program given in Eq. (6) to read as

which can be optimized using standard deep net packages. The proposed approach for learning the parameters ww is summarized in Alg. 1. In the following we now assess the suitability of the proposed approach.

Evaluation

In the following, we assess the proposed approach. We first provide details about the proposed dataset before presenting quantitative results for prediction of relations from questions, prediction of answer-source from questions, and prediction of the answer and the supporting fact. We also discuss mining of hard negatives. Finally, we show qualitative results.

Dataset and Knowledge Base. We use the publicly available FVQA dataset and its knowledge base to evaluate our model. This dataset consists of 2,190 images, 5,286 questions, and 4,126 unique facts corresponding to the questions. The knowledge base, consisting of 193,449 facts, were constructed by extracting the top visual concepts for all the images in the dataset and querying for those concepts in the three knowledge bases, WebChild , ConceptNet , and DBPedia . The dataset consists of 5 train-test folds, and all the scores we report are averaged across all splits.

Predicting Answer Source from Questions. We assess the accuracy of predicting the answer source ss given a question QQ. To predict the source of the answer, we use an LSTM architecture as discussed in detail in Sec. 3. Note that for predicting the answer source, the size of the LSTM embedding and word embeddings was set to 64 each. Table 2 summarizes the accuracy of the prediction results of our model. We observe the prediction accuracy of the proposed approach to be close to perfect.

Predicting the Correct Answer. Our score function based model to retrieve the supporting fact is described in detail in Sec. 3. For the image embedding, we pass the 2048 dimensional feature vector returned by ResNet through a fully-connected layer and reduce it to a 64 dimensional vector. For the question embedding, we use an LSTM with a hidden layer of size 128. The two are then concatenated into a vector of size 192 and passed through a two layer perceptron with 256 and 128 nodes respectively. Note that the baseline doesn’t use image features apart from the detected visual concepts.

The multi-hot visual concept embedding is passed through a fully-connected layer to form a 128 dimensional vector. This is then concatenated with the output of the perceptron and passed through another layer with 200 output nodes. We found a late fusion of the visual concepts to results in a better model as the facts explicitly contain these terms.

Fact embeddings are constructed using GloVe-100 vectors each, for entities aa and bb. If aa or bb contain multiple words, an average of all the embeddings is computed. We use cosine distance between the MLP and the fact embeddings to score the facts. The highest scoring fact is chosen as the answer. Ties are broken randomly.

Based on the answer source prediction which is computed using the aforementioned LSTM model, we choose either entity aa or bb of the fact to be the answer. See Eq. (1) for the formal description. Accuracy is computed based on exact match between the chosen entity and the groundtruth answer.

To assess the importance of particular features we investigate 5 variants of our model with varying features: two oracle approaches ‘gtgt Question + Image + Visual Concepts’ and ‘gtgt Question + Visual Concepts’ which make use of groundtruth relation type and answer type data. More specifically, ‘gtgt Question + Image + Visual Concepts’ and ‘gtgt Question + Visual Concepts’ use the groundtruth relations and answer sources respectively. We have three approaches using a variety of features as follows: ‘Question + Image + Visual Concepts,’ ‘Question + Visual Concepts,’ and ‘Question + Image.’ We drop either the Image embeddings from ResNet or the Visual Concept embeddings to obtain two other models, ‘Question + Visual Concepts’ and ‘Question + Image.’

Table 5 shows the accuracy of our model in predicting an answer and compares our results to other FVQA baselines. We observe the proposed approach to outperform the state-of-the-art ensemble technique by more than 3%3\% and the strongest baseline without ensemble by over 5%5\% on the top-1 accuracy metric. Moreover we note the importance of visual concepts to accurately predict the answer. By including groundtruth information we assess the maximally possible top-1 and top-3 accuracy. We observe the difference to be around 8%8\%, suggesting that there is some room for improvement.

Question to Supporting Fact. To provide a complete assessment of the proposed approach we illustrate in Table 5 the top-1 and top-3 accuracy scores in retrieving the supporting facts of our model compared to other FVQA baselines. We observe the proposed approach to improve significantly both the top-1 and top-3 accuracy by more than 20%20\%. We think this is a significant improvement towards efficiently including knowledge bases into visual question answering.

Mining Hard Negatives. We trained our model over three iterations of hard negative mining, i.e., T=2T=2. In iteration 1 (t=0t=0), all the 193,449 facts were used to sample the 99 negative facts during train. At every 10th epoch of training, negative facts which received high scores were saved. In the next iteration, the trained model along with the negative facts is loaded and we ensure that the 99 negative facts are now sampled from the hard negatives. Table 5 shows the Top-1 and Top-3 accuracy for predicting the supporting facts over each of the three iterations. We observe significant improvements due to the proposed hard negative mining strategy. While naïve training of the proposed approach yields only 20.17%20.17\% top-1 accuracy, two iterations improve the performance to 64.5%64.5\%.

Synonyms and Homographs. Here we show the improvements of our model compared to the baseline with respect to synonyms and homographs. To this end, we run additional tests using Wordnet to determine the number of question-fact pairs which contain synonyms. The test data contains 1105 such pairs out of which our model predicts 91.6% (1012) correctly, whereas the FVQA model predicts 78.0% (862) correctly. In addition, we manually generated 100 synonymous questions by replacing words in the questions with synonyms (e.g., “What in the bowl can you eat?” is rephrased to “What in the bowl is edible?”). Tests on these 100 new samples find that our model predicts 89 of these correctly, whereas the key-word matching FVQA technique gets 61 of these right. With regards to homographs, the test set has 998 questions which contain words that have multiple meanings across facts. Our model predicts correct answers for 79.4% (792), whereas the FVQA model gets 66.3% (662) correct.

Qualitative Results. Fig. 3 shows the Visual Concepts (VCs) detected for a few samples along with the top 3 facts retrieved by our model. Providing these predicted VCs as input to our fact-scoring MLP helps improve supporting fact retrieval as well as answer accuracy by a large margin of over 30%30\% as seen in Tables 5 and 5. As can be seen in Fig. 3, there is a close alignment between relevant facts and predicted VCs, as VCs provide a high-level overview of the salient content in the images.

In Fig. 4, we show success and failure cases of our method. There are 3 steps to producing the correct answer using our method: (1) correctly predicting the relation, (2) retrieving supporting facts containing the predicted relation, and relevant to the image, and (3) choosing the answer from the predicted answer source (Image/Knowledge Base). The top two rows of images show cases where all the 3 steps were correctly executed by our proposed method. Note that our method works for a variety of relations, objects, answer sources, and varying difficulty. It is correctly able to identify the object of interest, even when it is not the most prominent object in the image. For example, in the middle image of the first row, the frisbee is smaller than the dog in the image. However, we were correctly able to retrieve the supporting fact about the frisbee using information from the question, such as ‘capable of’ and ‘flying.’

A mistake in any of the 3 steps can cause our method to produce an incorrect answer. The bottom row of images in Fig. 4 displays prototypical failure modes. In the leftmost image, we miss cues from the question such as ‘round,’ and instead retrieve a fact about the person. In the middle image, our method makes a mistake at the final step and uses information from the wrong answer source. This is a very rare source of errors overall, as we are over 97%97\% accurate in predicting the answer source, as shown in Table 2. In the rightmost image, our method makes a mistake at the first step of predicting the relation, making the remaining steps incorrect. Our relation prediction is around 75%75\%, and 92%92\% accurate by the top-1 and top-3 metrics, as shown in Table 2, and has some scope for improvement. For qualitative results regarding synonyms and homographs we refer the interested reader to the supplementary material.

Conclusion

In this work, we addressed knowledge-based visual question answering and developed a method that learns to embed facts as well as question-image pairs into a space that admits efficient search for answers to a given question. In contrast to existing retrieval based techniques, our approach learns to embed questions and facts for retrieval. We have demonstrated the efficacy of the proposed method on the recently introduced and challenging FVQA dataset, producing state-of-the-art results. In the future, we hope to address extensions of our work to larger structured knowledge bases, as well as unstructured knowledge sources, such as online text corpora.

Acknowledgments: This material is based upon work supported in part by the National Science Foundation under Grant No. 1718221, Samsung, and 3M. We thank NVIDIA for providing the GPUs used for this research. We also thank Arun Mallya and Aditya Deshpande for their help.

References