Latent Variable Models for Visual Question Answering
Zixu Wang, Yishu Miao, Lucia Specia
Introduction
As a classic multi-modal machine learning problem, Visual Question Answering (VQA) systems are tasked with providing a correct textual answer given an image and a textual question. Current VQA models are trained to learn the relationship between areas in an image and the question, and to choose the correct answer from a vocabulary of answer candidates, i.e., they are modelled as a classification problem. The majorities of popular VQA models are created in a deterministic manner and explore solely information from the given image-question pair. There are other approaches attempting to incorporate extra information, such as image captions and mutated inputs . However, it in turn restricts the practical applications as the extra information is required to be explicitly available during testing.
In this paper, we propose an approach to explore additional information as latent variables in VQA: we employ latent variables for VQA to exploit extra information (i.e. image captions and answer categories) to complement limited textual information from image and question pairs. We assume a realistic setting where this information – esp. captions – may only be available during the training phase. To that end, we introduce a continuous latent variable as the caption representation to capture the essential information from this modality. Moreover, the answer category is modelled as a discrete latent variable, which acts as an inductive bias to benefit the learning of answer prediction, and can be integrated out during testing. The motivation is that the generative framework is able to incorporate many other types of information as continuous or discrete latent variables, and as such it effectively leverages additional resources to constrain the original image-question distribution while omitting them in testing. This grants the models with stronger generalisation ability compared to its deterministic counterparts, which generally require off-the-shelf pipelines to model the information from external modalities.
Intuitively, image captions describe diverse aspects of an image and include attributes and relations of objects in a more informative way. In our work, a continuous latent variable is employed for capturing the caption distributions and constraining the generative distribution conditioned on image and question pairs. In this way, the joint multimodal representations from images and question can benefit from the caption modality during training, and it requires no explicit caption inputs in testing. Similarly, there exists a strong connection between a question and answer pair when the question provides informative signals on its type or the category of possible answers. For example, “How many”, “Where is” and “what is” normally connect to numbers, locations, and objects respectively. We propose a discrete latent variable is employed for modelling answer categories and providing better inductive bias from the question and answer pairs.
A novel generative VQA framework combining the modularity of latent variables with the flexibility to introduce extra information as continuous and/or discrete latent variables.
A method to incorporate additional information which does not rely on building multiple deterministic pipelines, aiming at learning the underlying compositional, relational, and hierarchical structures of multiple modalities. The models benefit from the extra information during training without providing explicit inputs in testing.
The improvements over deterministic baseline models (e.g. UpDn and VL-BERT ) in experiments with the VQA v2.0 dataset demonstrate the effectiveness of our proposed latent variable models. Our qualitative analysis also indicates that using extra resources (i.e. captions and answer categories) as latent variables captures complementary information during training and benefits the VQA performance in testing.
Model
We first present an overview of our general model structure, followed by the encoders for different modalities, and the proposed corresponding latent variables.
In a VQA task, images and questions are normally used to learn a joint multimodal distribution for answer predictions. We postulate that the joint representation can be improved by other multimodal information. Hence, we introduce captions and answer categories to our VQA model as continuous and discrete latent variables respectively to encourage a better learning in the joint distribution of image and question pairs during training. A notable advantage of the latent variable models is that they do not explicitly require captions or answer categories during testing, and therefore can be easily extended to condition on any other useful information.
Firstly we introduce the notations used in the general VQA model. , , are used to denote the input image, question, and answer instances respectively. The image feature , question representation , and answer representation are extracted from the image encoder, question encoder, and answer encoder. The VQA task is constructed as a classification problem to output the most likely answer from a fixed set of answers based on the content of the image and question :
In our latent variable model, we introduce image captions to the training phase. Similarly, we extract the caption features by a caption encoder. However, instead of directly feeding in the caption features into to the model, we employ a continuous latent distribution to be the caption representations. Here is modelled as variational distribution. We then build a generative distribution to infer the caption information by conditioning on image and question pairs, which is optimised during training via neural variational inference. We originally experimented using as the vairational distribution. However, this distribution is quite close (i.e. small KL divergence) to the generative distribution , which weakens the learning signal from KL divergence.
In addition, we introduce a discrete latent variable for modelling answer category inferred via , which is also conditioned on image and question pairs.
Hence, the training of the latent variable model is carried out by the samples . During testing, the answer is predicted from the image and question pair :
where the discrete latent variable is directly integrated out, and the is the Monte-Carlo sample from .
2 Continuous Latent Variable: Caption
As captions are modelled by a continuous latent variable, we only have explicit captions during training. Here we present the generative distribution that is conditioned on images and questions during testing, and the variational distribution that is conditioned on explicit captions during training. Therefore, the caption encoder is only used in the training phase.
Generative Distribution - . We use a latent distribution to model the joint multimodal distributions of images and questions. Compared to its deterministic counterpart using concatenated multimodal features, we parameterise the stochastic distribution with .
Variational Distribution - . We first apply a RNN model to embed the caption inputs and a latent variable to model the caption semantics and distributions, where .
3 Discrete Latent Variable: Answer Category
Assuming that each image and question pair can be projected to an answer category to help find a correct answer, we are able to encourage the model to distinguishing candidates across answer categories instead of only the spurious relationships between questions and answers via simple linguistic features. Therefore, in order to leverage this useful inductive bias, we propose a discrete latent variable to model the answer category given an image and question pair . In particular, for each answer category , we have a conditional independent distribution over the answers in the certain answer category.
We trained an answer category classifier using joint image-question pairs as the input, given the true labels as shown at left bottom in Figure 1. We then use the category distribution to modify the answer distribution through element-wise production to get more precise answer distribution.
Datasets & Setup
We use the VQA v2.0 dataset for our proposed latent variable model. The answers are balanced in order to minimise the effectiveness of dataset priors. We report the results on validation set and test-standard set through the official evaluation server. The source of image captions in our work is the MSCOCO dataset .
We use answer categories from the annotations of . The answers in the VQA v2.0 dataset are annotated with a set of 15 categories for the top 500 answers that makes up the 82% Although the category definitions cannot cover all types of answers, and false prediction during testing might be observed, the latent variable can still maintain the robustness in predicting correct answers by summing over all the probabilities of predicted categories. of the VQA v2.0 dataset; and the other answers are treated as an additional category.
Experiments
In this section, we first describe the experimental results of our latent variable model, compared with both a UpDn (Bottom-up Top-down) baseline model and a state-of-the-art pre-trained visual-linguistic model (VL-BERT); then we conduct qualitative analysis to validate the effectiveness of proposed components.
We compare the results of our latent variable model with the baseline model (UpDn), a state-of-the-art visual-linguistic pre-training model (VL-BERT), and three other related VQA models; where uses generated captions to assist answer predictions and explore the interactions between visual and linguistic inputs.
As demonstrated in Table 1, our latent variable model outperforms when acting as an extension. In particular, our latent variable model outperforms UpDn by 0.69% accuracy on test-dev set and by 0.62% accuracy on test-standard set. In addition, our model improves the performance by 0.24% accuracy than its VL-BERT counterpart on test-dev set and by 0.15% accuracy on test-standard set. These results indicate the effectiveness of including captions and answer categories as latent variables, to promote the distribution of image and caption pairs to be closer to the captions’ space, and to learn a better distinction among different kinds of answers, or different answers within the same answer category.
The result of (68.37%) is a very strong baseline, which follows a traditional deterministic approach. However, their model is trained to generate captions that can be used at test time, while in our case only image and question pairs are required for answer prediction. and both achieve comparable performance (70.34% and 71.27% on test-standard set, respectively) to VL-BERT (72.22%) without pre-training, by dynamically modulating the intra-modality information and exploring the latent interaction between modalities. Our latent variable model has the overall best result when combined with the strong pre-training VL-BERT, which indicates both the effectiveness of the visual-linguistic pre-training framework, and the incorporation of continuous (captions) and discrete (answer categories) latent variables.
Compared to the results on the standard baseline (UpDn), the improvements achieved by our proposed model on the VL-BERT framework is smaller. This is because VL-BERT has been pre-trained on massive image captioning data, where the learning of visual features have largely benefited from the modality of captions already. Nevertheless, based upon the strong baseline model, our proposed model can still improve performance slightly, which further indicates the effectiveness of the latent variable framework.
The state-of-the-art performance on VQA v2.0 among pre-training frameworks is achieved by LXMERT, Oscar and Uniter . They have been extensively pre-trained using massive datasets on languages and vision tasks (including VQA) in a multi-task learning fashion. Our work is not directly comparable, and is not aimed at improving and beating the state-of-the-art performance. Instead, it is focused on exploring the potential of latent variable models to represent additional useful information in multimodal learning and to contribute to pre-trained vision-language frameworks. We draw attention to the advantage of using generative framework on the VQA task. In this case, we can employ more information during training (which is omitted in testing) to regularise the original multimodal distribution. This can be demonstrated by the improvements on VL-BERT brought by the latent variables.
2 Qualitative Analysis
We perform an ablation study to qualitatively analyse the effect of the components introduced in our work brought by the continuous (image caption) and discrete (answer category) latent variables, as shown in Table 2.
The introduction of captions as a continuous latent variable improves the classification performance, with an additional modality as input to benefit the learning of multimodal representations. According to the breakdown numbers in 2, the improvements brought by the latent variables of captions and answer categories are 0.70 and 0.36 respectively for All questions altogether. The combined strategy reaches 0.94 which indicates that the benefits from the two latent variables are complementary. Note that neither the captions nor the answer categories is available during testing; we only make use of these modalities during training.
To further investigate potential benefit of the captions, we design an experiment that feed in ground truth captions via variational distribution for caption representations instead of inferring them from question and answer pairs (i.e. use to replace ). We test this out in the validation dataset and obtain 64.24 (‘Ours w/ caption’) compared to 64.09 (‘Ours w/o caption’). It shows that having explicit captions as input gives slightly better performance. However, the captions in these experiments are ground truth, which means that if we were to use instead automatically generated captions from an image captioning pipeline, the numbers might drop due to the possible captioning errors. Primarily, our proposed model (‘Ours w/o caption’) achieves the performance on par with with the model with ground truth captions, which demonstrates the effectiveness of the strategy that incorporates extra modality by latent variables.
2.2 Effect of Answer Category
As it can be observed from Table 2, after introducing answer category as an additional discrete latent variable, our proposed model can also be improved over the UpDn baseline, where the largest improvement can be observed for the “Yes/No” type. For the question types “Num” and “Other”, the results of UpDn+category are lower than the baseline. This may be due to the multiple answer candidates under the two categories. For example, although the category classier can accurately predict the answer category (e.g., “count”, “color”, etc.), it can be still difficult to distinguish among the answers - {“9”, “20”, …, “many” for “count”; “black”, “brown”, …, “black and white” for “color”}. We highlight that the contribution of answer categories as a discrete latent variable is to introduce an inductive bias which helps predict the correct answer categories and answers given a specific image and question.
In order to further elaborate the effectiveness of answer categories, we extract examples where our model predicted the correct answers while the UpDn baseline failed to do so, as shown in Figure 2. For the two cases in the top row, both models predict answers under the same and correct answer categories, hence the answer space is similar; however, our latent variable model can effectively distinguish and learn the difference among the answers which fall within the same category. The bottom row of Figure 2 shows two cases where the two models predict answers in different answer categories, and therefore they are also very different in meaning. Our model not only outputs the highest probability for the correct answer category, but also makes the correct final prediction.
Conclusions
In this paper, we propose to tackle VQA under the framework of latent variable models, employing captions and answer categories as the continuous and the discrete latent variables respectively to constrain the original image-question distribution while omitting the extra information during the test phase. Our experimental results and qualitative analysis show the effectiveness of the latent variables in boosting answering performance at test time when only image and question pairs are available. This framework could be easily generalised to incorporate other types of information or modalities to enhance VQA and other tasks.