Neural Modular Control for Embodied Question Answering

Abhishek Das, Georgia Gkioxari, Stefan Lee, Devi Parikh, Dhruv Batra

Introduction

Abstraction is an essential tool for navigating our daily lives. When seeking a late night snack, we certainly do not spend time planning out the mechanics of walking and are thankfully also unburdened of the effort of recalling to beat our heart along the way. Instead, we conceptualize our actions as a series of higher-level semantic goals – exit bedroom; go to kitchen; open fridge; find snack; – each of which is executed through specialized coordination of our perceptual and sensorimotor skills. This ability to abstract long, complex sequences of actions into semantically meaningful subgoals is a key component of human cognition and it is natural to believe that artificial agents can benefit from applying similar mechanisms when navigating our world.

We study such hierarchical control in the context of a recently proposed task – Embodied Question Answering (EmbodiedQA) – where an embodied agent is spawned at a random location in a novel environment (e.g. a house) and asked to answer a question (‘What color is the piano in the living room?’). To do so, the agent must navigate from egocentric vision alone (without access to a map of the environment), locate the entity in question (‘piano in the living room’), and respond with the correct answer (e.g. ‘red’). From a reinforcement learning (RL) perspective, EmbodiedQA presents challenges that are known to make learning particularly difficult – partial observability, planning over long time horizons, and sparse rewards – the agent may have to navigate through multiple rooms in search for the answer, executing hundreds of primitive motion actions along the way (forward; forward; turn-right; …) and receiving a reward based only on its final answer.

To address this challenging learning problem, we develop a hierarchical Neural Modular Controller (NMC) – consisting of a master policy that determines high-level subgoals, and sub-policies that execute a series of low-level actions to achieve these subgoals. Our NMC model constructs a hierarchy that is arguably natural to this problem – navigation to rooms and objects vs. low-level motion actions. For example, NMC seeks to break down a question ‘What color is the piano in the living room?’ to the series of subgoals exit-room; find-room[living]; find-object[piano]; answer; and execute this plan with specialized neural ‘modules’ corresponding to each subgoal. Each module is trained to issue a variable length series of primitive actions to achieve its titular subgoal – e.g. the find-object[piano] module is trained to navigate the agent to the input argument piano within the current room. Disentangling semantic subgoal selection from sub-policy execution results in easier to train models due to shorter time horizons. Specifically, this hierarchical structure introduces:

Compressed Time Horizons: The master policy makes orders of magnitude fewer decisions over the course of a navigation than a ‘flat model’ that directly predicts primitive actions – allowing the answering reward in EmbodiedQA to more easily influence high-level motor control decisions.

Modular Pretraining: As each module corresponds to a specific task, they can be trained independently before being combined with the master policy. Likewise, the master policy can be trained assuming ideal modules. We do this through imitation learning sub-policies.

Interpretability: The predictions made by the master policy correspond to semantic subgoals and exposes the reasoning of the agent to inspection (‘What is the agent trying to do right now?’) in a significantly more interpretable fashion than just its primitive actions.

First, we learn and evaluate master and sub-policies for each of our subgoals, trained using behavior cloning on expert trajectories, reinforcement learning from scratch, and reinforcement learning after behavior cloning. We find that reinforcement learning after behavior cloning dramatically improves performance over each individual training regime. We then evaluate our combined hierarchical approach on the EQA benchmark in House3D environments. Our approach significantly outperforms prior work both in navigational and question answering performance – our agent is able to navigate closer to the target object and is able to answer questions correctly more often.

Related Work

Our work builds on and is related to prior work in hierarchical reinforcement and imitation learning, grounded language learning, and embodied question-answering agents in simulated environments.

Hierarchical Reinforcement and Imitation Learning. Our formulation is closely related to Le et al. , and can be seen as an instantiation of the options framework , wherein a global master policy proposes subgoals – to be achieved by local sub-policies – towards a downstream task objective . Relative to other work on automatic subgoal discovery in hierarchical reinforcement learning , we show that given knowledge of the problem structure, simple heuristics are quite effective in breaking down long-range planning into sequential subgoals. We make use of a combination of hierarchical behavior cloning and actor-critic to train our modular policy.

Neural Module Networks and Policy Sketches. At a conceptual-level, our work is analogous to recent work on neural module networks (NMNs) for visual question answering. NMNs first predict a ‘program’ from the question, consisting of a sequence of primitive reasoning steps, which are then executed on the image to obtain the answer. Unlike NMNs, where each primitive reasoning module has access to the entire image (completely observable) our setting is partially observable – each sub-policy only has access to first-person RGB – making active re-evaluation of subgoals after executing each sub-policy essential. Our work is also closely related to policy sketches , which are symbolic descriptions of subgoals provided to the agent without any grounding or sub-policy for executing them. There are two key differences w.r.t. to our work. First, an important framework difference – Andreas et al. assume access to a policy sketch at test time, i.e. for every task to be performed. In EmbodiedQA, this would correspond to the agent being provided with a high-level plan (exit-room; find-room[living]; …) for every question it is ever asked, which is an unrealistic assumption in real-world scenarios with a robot. In contrast, we assume that subgoal supervision (in the form of expert demonstrations and plans) are available on training environments but not on test, and the agent must learn to produce its own subgoals. Second, a subtle but important implementation difference – unlike , our sub-policy modules accept input arguments that are embeddings of target rooms and objects (e.g.find-room[living], find-object[piano]). This results in our sub-policy modules being shared not just across tasks (questions) as in , but also across instantiations of similar navigation sub-policies – i.e., find-object[piano] and find-object[chair] share parameters that enable data efficient learning without exhaustively learning separate policies for each.

Grounded Language Learning. Beginning with SHRDLU , there has been a rich progression of work in grounding language-based goal specifications into actions and pixels in physically-simulated environments. Recent deep reinforcement learning-based approaches to this explore it in 2D gridworlds , simple visual and textual environments, perceptually-realistic 3D home simulators , as well as real indoor scenes . Our hierarchical policy learns to ground words from the question into two levels of hierarchical semantics. The master policy grounds words into subgoals (such as find-room[kitchen]), and sub-policies ground these semantic targets (such as cutting board, bathroom) into primitive actions and raw pixels, both parameterized as neural control policies and trained end-to-end.

Embodied Question-Answering Agents. Finally, hierarchical policies for embodied question answering have previously been proposed by Das et al. in the House3D environment , and by Gordon et al. in the AI2-THOR environment . Our hierarchical policy, in comparison, is human-interpretable, i.e. the subgoal being pursued at every step of navigation is semantic, and due to the modular structure, can navigate over longer paths than prior work, spanning multiple rooms.

Neural Modular Control

We now describe our approach in detail. Recall that given a question, the goal of our agent is to predict a sequence of navigation subgoals and execute them to ultimately find the target object and respond with the correct answer. We first present our modular hierarchical policy. We then describe how we extract optimal plans from shortest path navigation trajectories for behavior cloning. And finally, we describe how the various modules are combined and trained with a combination of imitation learning (behavior cloning) and reinforcement learning.

Notation. Recall that NMC has 2 levels in the hierarchy – a master policy that generates subgoals and sub-policies for each of these subgoals. We use ii to index the sequence of subgoals and tt to index actions generated by sub-policies. Let S={s}\mathcal{S}=\{s\} denote the set of states, G={g}\mathcal{G}=\{g\} the set of variable-time subgoals with elements g=⟨gtask,gargument⟩g=\langle g_{\text{task}},g_{\text{argument}}\rangle, e.g. g=⟨g=\langleexit-room,None⟩\rangle, or g=⟨g=\langlefind-room,bedroom⟩\rangle. Let A={a}\mathcal{A}=\{a\} be the set of primitive actions (forward, turn-left, turn-right). The learning problem can then be succinctly put as learning a master policy πθ:S→G\pi_{\theta}:\mathcal{S}\rightarrow\mathcal{G} parameterized by θ\theta and sub-policies πϕg:S→A∪{stop}\pi_{\phi_{g}}:\mathcal{S}\rightarrow\mathcal{A}\cup\{\text{{stop}}\} parameterized by ϕg, ∀g∈G\phi_{g},\,\forall g\in\mathcal{G}, where the stop action terminates a sub-policy and returns control to the master policy.

While navigating an environment, control alternates between the master policy selecting subgoals and sub-policies executing these goals through a series of primitive actions. More formally, given an initial state s0s_{0} the master policy predicts a subgoal g0∼πθ(g∣s0)g_{0}\sim\pi_{\theta}(g|s_{0}), the corresponding sub-policy executes until some time T0{T}_{0} when either (1) the sub-policy terminates itself by producing the stop token aT0∼πϕg0(a∣sT0)=stopa_{T_{0}}\sim\pi_{\phi_{g_{0}}}(a|s_{T_{0}})=\text{{stop}} or (2) a maximum number of primitive actions has been reached. Either way, this returns the control back to the master policy which predicts another subgoal and repeats this process until termination. This results in a state-subgoal trajectory:

for the master policy. Notice that the terminal state of the ithi^{\text{th}} sub-policy sTis_{T_{i}} forms the state for the master policy to predict the next subgoal gi+1g_{i+1}. For the (i+1)th(i+1)^{\text{th}} subgoal gi+1g_{i+1}, the low-level trajectory of states and primitive actions is given by:

Note that by concatenating all sub-policy trajectories in order (σg0,σg1,…,σgT)(\sigma_{g_{0}},\sigma_{g_{1}},\ldots,\sigma_{g_{\mathcal{T}}}), the entire trajectory of states and primitive actions can be recovered.

Subgoals ⟨Tasks, Arguments⟩\langle\text{Tasks, Arguments}\rangle. As mentioned above, each subgoal is factorized into a task and an argument g=⟨gtask,gargument⟩g=\langle g_{\text{task}},g_{\text{argument}}\rangle. There are 4 possible tasks – exit-room, find-room, find-object, and answer. Tasks find-object and find-room accept as arguments one of the 50 objects and 12 room types in EQA v1 dataset respectively; exit-room and answer do not accept any arguments. This gives us a total of 50+12+1+1=6450+12+1+1=64 subgoals. ⟨\langleexit-room,none⟩\rangle, ⟨\langleanswer,none⟩\rangle, } 0 args ⟨\langlefind-object,couch⟩\rangle, ⟨\langlefind-object,cup⟩\rangle, …, ⟨\langlefind-object,xbox⟩\rangle, } 50 args ⟨\langlefind-room,living⟩\rangle, ⟨\langlefind-room,bedroom⟩\rangle, …, ⟨\langlefind-room,patio⟩\rangle. } 12 args

Descriptions of these tasks and their success criteria are provided in Table 1.

Sub-policies. To take advantage of the comparatively lower number of subgoal tasks, we decompose sub-policy parameters ϕg\phi_{g} into ϕgtask\phi_{g_{\text{task}}} and ϕgargument\phi_{g_{\text{argument}}}, where ϕgtask\phi_{g_{\text{task}}} are shared across the same task and ϕgargument\phi_{g_{\text{argument}}} is an argument specific embedding. Parameter sharing enables us to learn the shared task in a sample-efficient manner, rather than exhaustively learning separate sub-policies for each combination.

Perception and Question Answering. To ensure fair comparisons to prior work, we use the same perception and question answering models as used by Das et al. . The perception model is a simple convolutional neural network trained to perform auto-encoding, semantic segmentation, and depth estimation from RGB frames taken from House3D . Like , we use the bottleneck layer of this model as a fixed feature extractor. We also use the same post-navigational question-answering model as , which encodes the question with a 2-layer LSTM and performs dot-product based attention between the question encoding and the image features from the last five frames along the navigation path right before the answer module is called. This post-navigational answering module is trained using visual features along the shortest path trajectories and then frozen. By keeping these parts of the architecture identical to , our experimental comparisons can focus on the differences only due to our contributions, the Neural Modular Controller.

2 Hierarchical Behavior Cloning from Expert Trajectories

The questions in EQA v1 dataset (e.g. ‘What color is the fireplace?’) are constructed to inquire about attributes (color, location, etc.) of specific target objects (‘fireplace’). This notion of a target enables the construction of an automatically generated expert trajectory (s0∗,a0∗,…,sT∗,aT∗)(s_{0}^{*},a_{0}^{*},\ldots,s_{T}^{*},a_{T}^{*}) – the states and actions along the shortest path from the agent spawn location to the object of interest specified in the question. Notice that these shortest paths may only be used as supervision on training environments but may not be utilized during evaluation on test environments (where the agent must operate from egocentric vision alone).

Specifically, we would like to use these expert demonstrations to pre-train our proposed NMC navigator using behavior cloning. However, these trajectories (s0∗,a0∗,…,sT∗,aT∗)(s_{0}^{*},a_{0}^{*},\ldots,s_{T}^{*},a_{T}^{*}) correspond to a series of primitive actions. To provide supervision for both the master policy and sub-policies, these shortest-path trajectories must be annotated with a sequence of subgoals and segmented into their respective temporal extents, resulting in Σ∗\Sigma^{*} and (σgi∗)(\sigma_{g_{i}}^{*}).

We automate this ‘lifting’ of annotation up the hierarchy by leveraging the object and room bounding boxes provided by the House3D . Essentially, a floor plan may be viewed as an undirected graph with rooms as nodes and doorways as edges connecting a pair of adjacent rooms. An example trajectory is shown in Fig. 2(a) for the question ‘What color is the fireplace?’. The agent is spawned in a bedroom, the shortest path exits into the hall, enters the living room, and approaches the fireplace. We convert this trajectory to the subgoal sequence (exit-room, find-room[living], find-object[fireplace], answer) by recording the transitions on the shortest path from one room to another, which also naturally provides us with temporal extents of these subgoals.

We follow a couple of subtle but natural rules: (1) find-object is tagged only when the agent has reached the destination room containing the target object; and (2) exit-room is tagged only when the ‘out-degree’ of the current room in the floor-plan-graph is exactly 1 (i.e. either the current room has exactly one doorway or the current room has two doorways but the agent came in through one). Rule (2) ensures a semantic difference between exit-room and find-room – informally, exit-room means ‘get me out of here’ and find-room[name] means ‘look for room name’.

Tab. 1 summarizes these subgoals and the heuristics used to automatically extract them from navigational paths. Fig. 2(b) shows the proportions of these subgoals in expert trajectories as a function of the distance from target object. Notice that when the agent is close to the target, it is likely to be within the same room as the target and thus find-object dominates. On the other hand, when the agent is far away from the target, find-room and exit-room dominate.

We perform this lifting of shortest paths for all training set questions in EQA v1 dataset , resulting in NN expert trajectories {Σn∗}n=1N\{\Sigma_{n}^{*}\}_{n=1}^{N} for the master policy and K(>>N)K(>>N) trajectories {σgk∗}k=1K\{\sigma_{g_{k}}^{*}\}_{k=1}^{K} for sub-policies. We can then perform hierarchical behavior cloning by minimizing the sum of cross-entropy losses over all decisions in all expert trajectories. As is typical in maximum-likelihood training of directed probabilistic models (e.g. hierarchical Bayes Nets), full supervision results in decomposition into independent sub-problems. Specifically, with a slight abuse of notation, let (si∗,gi+1∗)∈Σ∗(s_{i}^{*},g_{i+1}^{*})\in\Sigma^{*} denote an iterator over all state-subgoal tuples in Σ∗\Sigma^{*}, and ∑(si∗,gi+1∗)∈Σ∗\displaystyle\sum_{(s_{i}^{*},g_{i+1}^{*})\in\Sigma^{*}} denote a sum over such tuples.

Now, the independent learning problems can be written as:

Intuitively, each sub-policy independently maximizes the conditional probability of actions observed in the expert demonstrations, and the master policy essentially trains assuming perfect sub-policies.

3 Asynchronous Advantage Actor-Critic (A3C) Training

After the independent behavior cloning stage, the policies have learned to mimic expert trajectories; however, they have not had to coordinate with each other or recover from their own navigational errors. As such, we fine-tune them with reinforcement learning – first independently and then jointly.

Reward Structure. The ultimate goal of our agent is to answer questions accurately; however, doing so requires navigating the environment sufficiently well in search of the answer. We mirror this structure in our reward RR, decomposing it into a sum of a sparse terminal reward RterminalR_{\text{terminal}} for the final outcome and a dense, shaped reward RshapedR_{\text{shaped}} determined by the agent’s progress towards its goals. For the master policy πθ\pi_{\theta}, we set RterminalR_{\text{terminal}} to be 1 if the model answers the question correctly and 0 otherwise. The shaped reward RshapedR_{\text{shaped}} at master-step ii is based on the change of navigable distance to the target object before and after executing subgoal gig_{i}. Each sub-policy πϕg\pi_{\phi_{g}} also has a terminal 0/1 reward RterminalR_{\text{terminal}} for stopping in a successful state, e.g. Exit-room ending outside the room it was called in (see Tab. 1 for all success definitions). Like the master policy, RshapedR_{\text{shaped}} at time tt is set according to the change in navigable distance to the sub-policy target (e.g. a point just inside a living room for find-room[living]) after executing the primitive action ata_{t}. Further, sub-policies are also penalized a small constant (-0.02) for colliding with obstacles.

Policy Optimization. We update the master and sub-policies to maximize expected discounted future rewards J(πθ)J(\pi_{\theta}) and J(πϕg)J(\pi_{\phi_{g}}) respectively through the Asynchronous Advantage Actor Critic policy-gradient algorithm. Specifically, for the master policy, the gradient of the expected reward is written as:

where cθ(sTi)c_{\theta}(s_{T_{i}}) is the estimated value of sTis_{T_{i}} produced by the critic for πθ\pi_{\theta}. To further reduce variance, we follow and estimate Q(sTi,gi)≈Rθ(sTi)+γcθ(sTi+1)Q(s_{T_{i}},g_{i})\approx R_{\theta}(s_{T_{i}})+\gamma c_{\theta}(s_{T_{i+1}}) such that Q(sTi,gi)−cθ(sTi)Q(s_{T_{i}},g_{i})-c_{\theta}(s_{T_{i}}) computes a generalized advantage estimator (GAE). Similarly, each sub-policy πϕg\pi_{\phi_{g}} is updated according to the gradient

Recall from Section 3.1 that these critics share parameters with their corresponding policy networks such that subgoals with a common task also share a critic. We train each policy network independently using A3C with GAE with 8 threads across 4 GPUs. After independent reinforcement fine-tuning of the sub-policies, we train the master policy further using the trained sub-policies rather than expert subgoal trajectories.

Initial states and curriculum. Rather than spawn agents at fixed distances from target, from where accomplishing the subgoal may be arbitrarily difficult, we sample locations along expert trajectories for each question or subgoal. This ensures that even early in training, policies are likely to have a mix of positive and negative reward episodes. At the beginning of training, all points along the trajectory are equally likely; however, as training progresses and success rate improves, we reduce the likelihood of sampling points nearer to the goal. This is implemented as a multiplier α\alpha on available states [s0,s1,...,sαT][s_{0},s_{1},...,s_{\alpha T}], initialized to 1.01.0 and scaled by 0.90.9 whenever success rate crosses a 40%40\% threshold.

Experiments and Results

Dataset. We benchmark performance on the EQA v1 dataset , which contains ∼9,000{\sim}9,000 questions in 774774 environments – split into 7129(648)7129(648) / 853(68)853(68) / 905(58)905(58) questions (environments) for training/validation/testing respectivelyNote that the size of the publicly available dataset on embodiedqa.org/data is larger than the one reported in the original version of the paper due to changes in labels for color questions.. These splits have no overlapping environments between them, thus strictly checking for generalization to novel environments. We follow the same splits.

Evaluating sub-policies. We begin by evaluating the performance of each sub-policy with regard to its specialized task. For clarity, we break results down by subgoal task rather than for each task-argument combination. We compare sub-policies trained with behavior cloning (BC), reinforcement learning from scratch (A3C), and reinforcement fine-tuning after behavior cloning (BC+A3C). We also compare to a random agent that uniformly samples actions including stop to put our results in context. For each, we report the success rate (as defined in Tab. 1) on the EQA v1 validation set which consists of 6868 novel environments unseen during training. We spawn sub-policies at randomly selected suitable rooms (i.e. Find-object[sofa] will only be executed in a room with a sofa) and allow them to execute for a maximum episode length of 50 steps or until they terminate.

Fig. 3 shows success rates for the different subgoal tasks over the course of training. We observe that:

Behavior cloning (BC) is more sample-efficient than A3C from scratch. Sub-policies trained using BC improve significantly faster than A3C for all tasks, and achieve higher success rates for Exit-room and Find-room. Interestingly, this performance gap is larger for tasks where a random policy does worse – implying that BC helps more as task complexity increases.

Reinforcement Fine-Tuning with A3C greatly improves over BC training alone. Initializing A3C with a policy trained via behavior cloning results in a model that significantly outperforms either approach on its own, nearly doubling the success rate of behavior cloning for some tasks. Intuitively, mimicking expert trajectories in behavior cloning provides dense feedback for agents about how to navigate the world; however, agents never have to face the consequences of erroneous actions e.g. recovering from collisions with objects – a weakness that A3C fine-tuning addresses.

Evalating master policy. Next, we evaluate how well the master policy performs during independent behavior cloning on expert trajectories i.e. assuming perfect sub-policies, as specified in Eq. 59a. Even though there is no overlap between training and validation environments, the master policy is able to generalize reasonably and gets ∼48%\sim 48\% intersection-over-union (IoU) with ground truth subgoal sequences on the validation set. Note that a sequence of sub-goals that is different from the one corresponding to the shortest path may still be successful at navigating to the target object and answering the question correctly. In that sense, IoU against ground truth subgoal sequences is a strict metric. Fig. 3(d) shows the training and validation cross-entropy loss curves for the master policy.

Evalating NMC. Finally, we put together the master and sub-policies and evaluate navigation and question answering performance on EmbodiedQA. We compare against the PACMAN model proposed in . For accurate comparison, both PACMAN and NMC use the same publicly available and frozen pretrained CNNgithub.com/facebookresearch/EmbodiedQA, and the same visual question answering model – pretrained to predict answers from last 55 observations of expert trajectories, following . Agents are evaluated by spawning 1010, 3030, or 5050 primitive actions away from target, which corresponds to distances of 1.151.15, 4.874.87, and 9.649.64 meters from target respectively, denoted by d0\mathbf{d_{0}} in Tab. 2. When allowed to run free from this spawn location, dT\mathbf{d_{T}} measures final distance to target (how far is the agent from the goal at termination), and dΔ=dT−d0\mathbf{d_{\Delta}}=\mathbf{d_{T}}-\mathbf{d_{0}} evaluates change in distance to target (how much progress does the agent make over the course of its navigation). Answering performance is measured by accuracy\mathbf{accuracy} (i.e. did the predicted answer match ground-truth). Note that report a number of additional metrics (percentage of times the agent stops, retrieval evaluation of answers, etc.). Accuracies for PACMAN are obtained by running the publicly available codebase accompanying , and numbers are different than those reported in the original version of due to changes in the dataset1.

As shown in Tab. 2, we evaluate two versions of our model – 1) NMC (BC) naively combines master and sub-policies without A3C finetuning at any level of hierarchy, and 2) NMC (BC+A3C) is our final model where each stage is trained with BC+A3C, as described in Sec. 3. As expected, NMC (BC) performs worse than NMC (BC + A3C), evident in worse navigation dT\mathbf{d_{T}}, dΔ\mathbf{d_{\Delta}} and answering accuracy\mathbf{accuracy}. PACMAN (BC) and NMC (BC) go through the same training regime, and there are no clear trends as to which is better – PACMAN (BC) has better dΔ\mathbf{d_{\Delta}} and answering accuracy\mathbf{accuracy} at T−10T_{-10} and T−50T_{-50}, but worse at T−30T_{-30}. No A3C finetuning makes it hard for sub-policies to recover from erroneous primitive actions, and for master policy to adapt to sub-policies. A3C finetuning significantly boosts performance, i.e. NMC (BC + A3C) outperforms PACMAN with higher dΔ\mathbf{d_{\Delta}} (makes more progress towards target), lower dT\mathbf{d_{T}} (terminates closer to target), and higher answering accuracy\mathbf{accuracy}. This gain primarily comes from the choice of subgoals and the master policy’s ability to explore over this space of subgoals instead of primitive actions (as in PACMAN), enabling the master policy to operate over longer time horizons, critical for sparse reward settings as in EmbodiedQA.

Conclusion

We introduced Neural Modular Controller (NMC), a hierarchical policy for EmbodiedQA consisting of a master policy that proposes a sequence of semantic subgoals from question (e.g. ‘What color is the sofa in the living room?’ →\rightarrow Find-room[living], Find-object[sofa], Answer), and specialized sub-policies for executing each of these tasks. The master and sub-policies are trained using a combination of behavior cloning and reinforcement learning, which is dramatically more sample-efficient than each individual training regime. In particular, behavior cloning provides dense feedback for how to navigate, and reinforcement learning enables policies to deal with consequences of their actions, and recover from errors. The efficacy of our proposed model is demonstrated on the EQA v1 dataset , where NMC outperforms prior work both in navigation and question answering.

This work was supported in part by NSF, AFRL, DARPA, Siemens, Google, Amazon, ONR YIPs and ONR Grants N00014-16-1-{\{2713,2793}\}. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government, or any sponsor.

References