Learning by Asking Questions

Ishan Misra, Ross Girshick, Rob Fergus, Martial Hebert, Abhinav Gupta, Laurens van der Maaten

Introduction

Machine learning models have led to remarkable progress in visual recognition. However, while the training data that is fed into these models is crucially important, it is typically treated as predetermined, static information. Our current models are passive in nature: they rely on training data curated by humans and have no control over this supervision. This is in stark contrast to the way we humans learn — by interacting with our environment to gain information. The interactive nature of human learning makes it sample efficient (there is less redundancy during training) and also yields a learning curriculum (we ask for more complex knowledge as we learn).

In this paper, we argue that next-generation recognition systems need to have agency — the ability to decide what information they need and how to get it. We explore this in the context of visual question answering (VQA; ). Instead of training on a fixed, large-scale dataset, we propose an alternative interactive VQA setup called learning-by-asking (LBA): at training time, the learner receives only images and decides what questions to ask. Questions asked by the learner are answered by an oracle (human supervision). At test-time, LBA is evaluated exactly like VQA using well understood metrics.

The interactive nature of LBA requires the learner to construct meta-knowledge about what it knows and to select the supervision it needs. If successful, this facilitates more sample efficient learning than using a fixed dataset, because the learner will not ask redundant questions.

We explore the proposed LBA paradigm in the context of the CLEVR dataset , which is an artificial universe in which the number of unique objects, attributes, and relations are limited. We opt for this synthetic setting because there is little prior work on asking questions about images: CLEVR allows us to perform a controlled study of the algorithms needed for asking questions. We hope to transfer the insights obtained from our study to a real-world setting.

Building an interactive learner that can ask questions is a challenging task. First, the learner needs to have a “language” model to form questions. Second, it needs to understand the input image to ensure the question is relevant and coherent. Finally (and most importantly), in order to be sample efficient, the learner should be able to evaluate its own knowledge (self-evaluate) and ask questions which will help it to learn new information about the world. The only supervision the learner receives from the interaction is the answer to the questions it poses.

We present and study a model for LBA that combines ideas from visually grounded language generation , curriculum learning , and VQA. Specifically, we develop an epsilon-greedy learner that asks questions and uses the corresponding answers to train a standard VQA model. The learner focuses on mastering concepts that it can rapidly improve upon, before moving to new types of questions. We demonstrate that our LBA model not only asks meaningful questions, but also matches the performance of human-curated data. Our model is also sample efficient and by interactively asking questions it reduces the number of training samples needed to obtain the baseline question-answering accuracy by 40%.

Related Work

Visual question answering (VQA) is a surrogate task designed to assess a system’s ability to thoroughly understand images. It has gained popularity in recent years due to the release of several benchmark datasets . Motivated by the well-studied difficulty of analyzing results on real-world VQA datasets , Johnson et al. recently proposed a more controlled, synthetic VQA dataset that we adopt in this work.

Current VQA approaches follow a traditional supervised learning paradigm. A large number of image-question-answer triples are collected and a subset of this data is randomly selected for training. Learning-by-asking (LBA) uses an alternative and more challenging setting: training images are drawn from a distribution, but the learner decides what question it needs to ask to learn the most. The learner receives only answer level supervision from these interactions. It must learn to formulate questions as well as model its own knowledge to remove redundancy in question-asking. LBA also has the potential to generalize to open-world scenarios.

There is also significant progress on building models for VQA using LSTMs with convolutional networks , stacked attention networks , module networks , relational networks , and others . LBA is independent of the backbone VQA model and can be used with any existing architecture.

Visual question generation (VQG) was recently proposed as an alternative to image captioning . Our work is related to VQG in the sense that we require the learner to generate questions about images, however, our objective in doing so is different. Whereas VQG focuses on asking questions that are relevant to the image content, LBA requires the learner to ask questions that are both relevant and informative to the learner when answered. A positive side effect is that LBA circumvents the difficulty of evaluating the quality of generated questions (which also hampers image captioning ), because the question-answering accuracy of our final model directly correlates with the quality of the questions asked. Such evaluation has also been used in recent works in the language community .

Active learning (AL) involves a collection of unlabeled examples and a learner that selects which samples will be labeled by an oracle . Common selection criteria include entropy , boosting the margin for classifiers and expected informativeness . Our setting is different from traditional AL settings in multiple ways. First, unlike AL where an agent selects the image to be labeled, in LBA the agent selects an image and generates a question. Second, instead of asking for a single image level label, our setting allows for richer questions about objects, relationships etc. for a single image. While did use simple predefined template questions for AL, templates offer limited expressiveness and a rigid query structure. In our approach, questions are generated by a learned language model. Expressive language models, like those used in our work, are likely necessary for generalizing to real-world settings. However, they also introduce a new challenge: there are many ways to generate invalid questions, which the learner must learn to discard (see Figure 2).

Exploratory learning centers on settings in which an agent explores the environment to acquire supervision ; it has been studied in the context of, among others, computer games and navigation , multi-user games , inverse kinematics , and motion planning for humanoids . Exploratory learning problems are generally framed with reinforcement learning in which the agent receives (delayed) rewards, which are used to learn a policy that maximizes the expected rewards. A key difference in the LBA setting is that it does not have sparse delayed rewards. Contextual multi-armed bandits are another class of reinforcement learning algorithms that more closely resemble the LBA setting. However, unlike bandits, online performance is irrelevant in our setting: our aim is not to minimize regret, but to minimize the error of the final VQA model produced by the learner.

Learning by Asking

The challenge of the LBA setting implies that, at training time, the learner must decide which question to ask about an image and the only supervision the oracle provides are the answers. As the number of oracle requests is constrained by a budget BB, the learner must ask questions that maximize (in expectation) the learning signal from each image-question pair sent to the oracle.

Approach

We propose a LBA agent built from three modules: (1) a question proposal module that generates a set of question proposals for an input image; (2) a question answering module (or VQA model) that predicts answers from (I,q)(\mathbf{I},q) pairs; and (3) a question selection module that looks at both the answering module’s state and the proposal module’s questions to pick a single question to ask the oracle. After receiving the oracle’s answer, the agent creates a tuple (I,q,a)(\mathbf{I},q,a) that is used as the online learning signal for all three modules. Each of the modules is described in a separate subsection below; the interactions between them are illustrated in Figure 3.

For the CLEVR universe, the oracle is a program interpreter that uses the ground-truth scene information to produce answers. As this oracle only understands questions in the form of programs (as opposed to natural language), our question proposal and answering modules both represent questions as programs. However, unlike , we do not exploit prior knowledge of the CLEVR programming language in any of the modules; instead, it is treated as a simple means that is required to communicate with the oracle. See supplementary material for examples of programs and details on the oracle.

When the LBA model asks an invalid question, the oracle returns a special answer indicating (1) that the question was invalid and (2) whether or not all the objects that appear in the question are present in the image.

The question proposal module aims to generate a diverse set of questions (programs) that are relevant to a given image. We found that training a single model to meet both these requirements resulted in limited diversity of questions. Thus, we employ two subcomponents: (1) a question generation model gg that produces questions qg∼g(q∣I)q_{g}\sim g(q|\mathbf{I}); and (2) a question relevance model r(I,qg)r(\mathbf{I},q_{g}) that predicts whether a generated question qgq_{g} is relevant to an image I\mathbf{I}. Figure 2 shows examples of irrelevant questions that need to be filtered by rr. The question generation and relevance models are used repeatedly to produce a set of question proposals, Qp⊆Q\mathcal{Q}_{p}\subseteq\mathcal{Q}.

Our question relevance model, r(I,q)r(\mathbf{I},q), takes the questions from the generator gg as input and filters out irrelevant questions to construct a set of question proposals, Qp\mathcal{Q}_{p}. The special answer provided by the oracle whenever an invalid question is asked (as described above) serves as the online learning signal for the relevance model. Specifically, the model is trained to predict (1) whether or not a image-question pair is valid and (2) whether or not all objects that are mentioned in the question are all present in the image. Questions for which both predictions are positive (i.e., that are deemed by the relevance model to be valid and to contain only objects that appear in the image) are put in the question proposal set, Qp\mathcal{Q}_{p}. We sample from the generator until we have 50 question proposals per image that are predicted to be valid by r(I,q)r(\mathbf{I},q).

2 Question Answering Module (VQA Model)

Our question answering module is a standard VQA model, v(a∣I,q)v(a|\mathbf{I},q), that learns to predict the answer aa given an image-question pair (I,q)(\mathbf{I},q). The answering module is trained online using the supervision signal from the oracle.

A key requirement for selecting good questions to ask the oracle is the VQA model’s capability to self-evaluate its current state. We capture the state of the VQA model at LBA round tt by keeping track of the model’s question-answering accuracy st(a)s_{t}(a) per answer aa on the training data obtained so far. The state captures information on what the answering module already knows; it is used by the question selection module.

3 Question Selection Module

The question selection module defines a policy, π(Qp;I,s1,…,t)\pi(\mathcal{Q}_{p};\mathbf{I},s_{1,\dots,t}), that selects the most informative question to ask the oracle from the set of question proposals Qp\mathcal{Q}_{p}. To select an informative question, the question selection module uses the current state of the answering module (how well it is learning various concepts) and the difficulty of each of the question proposals. These quantities are obtained from the state st(a)s_{t}(a) and the beliefs of the current VQA model, v(a∣I,q)v(a|\mathbf{I},q) for an image-question pair, respectively.

The state st(a)s_{t}(a) contains information about the current knowledge of the answering module. The difference in the state values at the current round, tt, and a past round, t−Δt-\Delta, measures how fast the answering module is improving for each answer. Inspired by curriculum learning , we use this difference to select questions on which the answering module can improve the fastest. Specifically, we compute the expected accuracy improvement under the answer distribution for each question qp∈Qpq_{p}\in\mathcal{Q}_{p}:

We use the expected accuracy improvement as an informativeness value that the learner uses to pick a question that helps it improve rapidly (thereby enforcing a curriculum). In particular, our selection policy, π(Qp;I,s1,…,t)\pi(\mathcal{Q}_{p};\mathbf{I},s_{1,\dots,t}), uses the informativeness scores to select the question to ask the oracle using an epsilon-greedy policy . The greedy part of the selection policy is implemented via argmax⁡qp∈Qph(qp;I,s1,…,t)\operatorname*{argmax}_{q_{p}\in\mathcal{Q}_{p}}h(q_{p};\mathbf{I},s_{1,\dots,t}), and we set ϵ ⁣= ⁣0.1\epsilon\!=\!0.1 to encourage exploration. Empirically, we find that our policy automatically discovers an easy-to-hard curriculum (see Figures 6 and 8). In all experiments, we set Δ ⁣= ⁣20\Delta\!=\!20; whenever t ⁣< ⁣Δt\!<\!\Delta, we set st−Δ(a) ⁣= ⁣0s_{t-\Delta}(a)\!=\!0.

4 Training Phases

5 Implementation Details

The LSTM in gg has 512 hidden units. After a linear projection, the image features are fed as its first hidden state. We input a discrete variable representing the question type as the first token into the LSTM before starting generation. Following , we use a prefix-tree program representation for the questions.

We implement the relevance model, rr, and the VQA model, vv, using the stacked attention network architecture using the implementation of . The only modification we make is to concatenate the spatial coordinates to the image features before computing attention as in . We do not share weights between rr and vv.

Our models use image features from a ResNet-101 pre-trained on ImageNet , in particular, from the conv4_23 layer of that network. We use ADAM with a fixed learning rate of 5e ⁣− ⁣45e\!-\!4 to optimize all models. Additional implementation details are presented in the supplementary material.

Experiments

CNN+LSTM encodes the image using a CNN, the question using an LSTM, and predicts answers using an MLP.

CNN+LSTM+SA extends CNN+LSTM with the stacked attention (SA) model described in Section 4.2. This is the same as our default answering module vv.

FiLM uses question features from a GRU to modulate the image features in each CNN layer.

In Figure 4, we compare the quality of the LBA-generated questions to CLEVR train by measuring the question-answering accuracy of VQA models trained on both datasets. The figure shows (top) CLEVR val accuracy and (bottom) CLEVR-Humans accuracy. From these plots, we draw four observations.

(1) Using the bootstrap set alone (leftmost point) yields poor accuracy and LBA provides a significant learning signal.

(2) The quality of the LBA-generated training data is at least as good as that of the CLEVR train. This is an impressive result given that CLEVR train has the dual advantage of matching the distribution of CLEVR val and being human curated for training VQA models. Despite these advantages, LBA matches and sometimes surpasses its performance. More importantly, LBA shows better generalization on CLEVR-Humans which has a different answer distribution (see Figure 9).

(3) LBA data is sometimes more sample efficient than CLEVR train: for instance, on both CLEVR val and CLEVR-Humans. The CNN+LSTM+SA model only requires 60% of (I,q,a)(\mathbf{I},q,a) LBA tuples to achieve the accuracy of the same model trained on all of CLEVR train.

(4) Finally, we also observe that our LBA agents have low variance at each sampled point during training. The shaded error bars show one standard deviation computed from 5 independent runs using different random seeds. This is an important property for drawing meaningful conclusions from interactive training environments (c.f., ).

Qualitative results. Figure 6 shows five samples from the LBA-generated data at various iterations tt. They provide insight into the curriculum discovered by our LBA agent. Initially, the model asks simple questions about colors (row 1) and shapes (row 2). It also makes basic mistakes (rightmost column of rows 1 and 2). As the answering module vv improves, the selection policy π\pi asks more complex questions about spatial relationships and counts (rows 3 and 4).

2 Analysis: Question Proposal Module

Analyzing the generator gg. We evaluate the diversity of the generated questions by looking at the distribution of corresponding answers. In Figure 5.1 (top) we use the final LBA model to generate 10 questions for each image in the training set. We plot the histogram of the answers to these questions for generators with and without “question type” conditioning. The histogram shows that conditioning the generator gg on question type leads to better coverage of the answer space. We also note that about 4% of the generated questions have invalid programming language syntax.

We observe in the top two rows of Table 1 that the increased question diversity translates into improved question-answering accuracy. Diversity is also controlled by the sampling temperature, τ\tau, used in gg. Rows 3-5 show that a lower temperature, which gives less diverse question proposals, negatively impacts final accuracy.

Analyzing the relevance model rr. Figure 5.1 (bottom) displays the percentage of invalid questions sent to the oracle at different time steps during online LBA training. The invalid question rate decreases during training from 25% to 5%, even though question complexity appears to be increasing (Figure 6). This result indicates that the relevance model rr improves significantly during training.

We can also decouple the effect of the relevance model rr from the rest of our setup by replacing it with a “perfect” relevance model (the oracle) that flawlessly filters all invalid questions. Table 1 (row 6) shows that the accuracy and sample efficiency differences between the “perfect” relevance model and our relevance model are small, which suggests our model performs well.

3 Analysis: Question Answering Module

4 Analysis: Question Selection Module

To investigate the role of the selection policy in LBA, we compare four alternatives: (1) random selection from the question proposals; (2) using the prediction entropy of the answering module vv for each proposal after four forward passes with dropout (like in ); (3) using the variation ratio of the prediction; and (4) our curriculum policy from Section 4.3. We run LBA training with five different random seeds and report the mean accuracy and stdev of a CNN+LSTM+SA model for each selection policy in Figure 7. In line with results from prior work , the entropy-based policies perform worse than random selection. By contrast, our curriculum policy substantially outperforms random selection of questions. Figure 8 plots the normalized informativeness score hh (Equation 1) and the training question-answering accuracy (s(a)s(a) grouped by per answer type). These plots provide insight into the behavior of the curriculum selection policy, π\pi. Specifically, we observe a delayed pattern: a peak in the the informativeness score (blue arrow) for an answer type is followed by an uptick in the accuracy (blue arrow) on that answer type. We also observe that the policy’s informativeness score suggests an easy-to-hard ordering of questions: initially (after 64k requests), the selection policy prefers asking the easier color questions, but it gradually moves on to size and shape questions and, eventually, to the difficult count questions. We emphasize that this easy-to-hard curriculum is learned automatically without any extra supervision.

5 Varying the Size of the Bootstrap Data

Discussion and Future Work

This paper introduces the learning-by-asking (LBA) paradigm and proposes a model in this setting. LBA moves away from traditional passively supervised settings where human annotators provide the training data in an interactive setting where the learner seeks out the supervision it needs. While passive supervision has driven progress in visual recognition , it does not appear well suited for general AI tasks such as visual question answering (VQA). Curating large amounts of diverse data which generalizes to a wide variety of questions is a difficult task. Our results suggest that interactive settings such as LBA may facilitate learning with higher sample efficiency. Such high sample efficiency is crucial as we move to increasingly complex visual understanding tasks.

An important property of LBA is that it does not tie the distribution of questions and answers seen at training time to the distribution at test time. This more closely resembles the real-world deployment of VQA systems where the distribution of user-posed questions to the system is unknown and difficult to characterize beforehand . The CLEVR-Humans distribution in Figure 9 is an example of this. This issue poses clear directions for future work : we need to develop VQA models that are less sensitive to distributional variations at test time; and not evaluate them under a single test distribution (as in current VQA benchmarks).

A second major direction for future work is to develop a “real-world” version of a LBA system in which (1) CLEVR images are replaced by natural images and (2) the oracle is replaced by a human annotator. Relative to our current approach, several innovations are required to achieve this goal. Most importantly, it requires the design of an effective mode of communication between the learner and the human “oracle”. In our current approach, the learner uses a simple programming language to query the oracle. A real-world LBA system needs to communicate with humans using diverse natural language. The efficiency of LBA learners may be further improved by letting the oracle return privileged information that does not just answer an image-question pair, but that also explains why this is the right or wrong answer . We leave the structural design of this privileged information to future work.

Acknowledgments: The authors would like to thank Arthur Szlam, Jason Weston, Saloni Potdar and Abhinav Shrivastava for helpful discussions and feedback on the manuscript; Soumith Chintala and Adam Paszke for their help with PyTorch.

References

Appendix A Hyperparameters for Models

Relevance model rr: We use the stacked-attention network implementation from in our experiments. We use Adam to minimize the cross-entropy loss of this model, using a fixed learning rate of 5e−45e-4. We use L2L_{2} weight decay of 4e−54e-5 as a regularizer. The model uses image features from an ImageNet pre-trained ResNet-101 conv4_23 layer (10241024 channels). Following , we concatenate spatial coordinates (two channels xx and yy) to these features to finally get 10261026 channel image features of spatial resolution 14×1414\times 14.

VQA model vv: The stacked-attention network serves as the default choice for vv. We use the same hyperparameters as in the relevance model rr, but do not share weights between vv and rr.

Appendix B Details on Oracle

Appendix C Extra Experiment: More Questions or More Images?

Appendix D Translation for CLEVR Humans

We show a few examples of CLEVR images, questions, programs and answers in Figure LABEL:fig:clevr_ex. We show examples with short programs for ease of visualization.