Chasing Ghosts: Instruction Following as Bayesian State Tracking

Peter Anderson, Ayush Shrivastava, Devi Parikh, Dhruv Batra, Stefan Lee

Introduction

One long-term challenge in AI is to build agents that can navigate complex 3D environments from natural language instructions. In the Vision-and-Language Navigation (VLN) instantiation of this task , an agent is placed in a photo-realistic reconstruction of an indoor environment and given a natural language navigation instruction, similar to the example in Figure 1. The agent must interpret this instruction and execute a sequence of actions to navigate efficiently from its starting point to the corresponding goal. This task is challenging for existing models , particularly as the test environments are unseen during training and no prior exploration is permitted in the hardest setting.

To be successful, agents must learn to ground language instructions to both visual observations and actions. Since the environment is only partially-observable, this in turn requires the agent to relate instructions, visual observations and actions through memory. Current approaches to the VLN task use unstructured general purpose memory representations implemented with recurrent neural network (RNN) hidden state vectors . However, these approaches lack geometric priors and contain no mechanism for reasoning about the likelihood of alternative trajectories – a crucial skill for the task, e.g., ‘Would this look more like the goal if I was on the other side of the room?’. Due to this limitation, many previous works have resorted to performing inefficient first-person search through the environment using search algorithms such as beam search . While this greatly improves performance, it is clearly inconsistent with practical applications like robotics since the resulting agent trajectories are enormously long – in the range of hundreds or thousands of meters.

To address these limitations, it is essential to move towards reasoning about alternative trajectories in a representation of the environment – where there are no search costs associated with moving a physical robot – rather than in the environment itself. Towards this, we extend the Matterport3D simulator to provide depth outputs, enabling us to investigate the use of a semantic spatial map in the context of the VLN task for the first time. We propose an instruction-following agent incorporating three components: (1) a mapper that builds a semantic spatial map of its environment from first-person views; (2) a filter that determines the most probable trajectory(ies) and goal location(s) in the map, and (3) a policy that executes a sequence of actions to reach the predicted goal.

From a modeling perspective, our key contribution is the filter that formulates instruction following as a problem of Bayesian state tracking . We notice that a visually-grounded navigation instruction typically contains a description of expected future observations and actions on the path to the goal. For example, consider the instruction ‘walk out of the bathroom, turn left, and go on to the bottom of the stairs and wait near the coat rack’ shown in Figure 1. When following this instruction, we would expect to immediately observe a bathroom, and at the end a coat rack near a stairwell. Further, in reaching the goal we can anticipate performing certain actions, such as turning left and continuing that way. Based on this intuition, we use a sequence-to-sequence model with attention to extract sequences of latent vectors representing observations and actions from a natural language instruction.

Faced with a known starting state, a (partially-observed) semantic spatial map generated by the mapper, and a sequence of (latent) observations and actions, we now quite naturally interpret our instruction following task within the framework of Bayesian state tracking. Specifically, we formulate an end-to-end differentiable histogram filter with learnable observation and motion models, and we train it to predict the most likely trajectory taken by a human demonstrator. We emphasize that we are not tracking the state of the actual agent. In the VLN setting, the pose of the agent is known with certainty at all times. The key challenge lies in determining the location of the natural-language-specified goal state. Leveraging the machinery of Bayesian state estimation allows us to reason in a principled fashion about what a (hallucinated) human demonstrator would do when following this instruction – by explicitly modeling the demonstrator’s trajectory over multiple time steps in terms of a probability distribution over map cells. The resulting model encodes both strong geometric priors (e.g., pinhole camera projection) and strong algorithmic priors (e.g., explicit handling of uncertainty, which can be multi-modal), while enabling explainability of the learned model. For example, we can separately examine the motion model, the observation model, and their interaction during filtering.

Empirically, we show that our filter-based approach significantly outperforms a strong LingUNet baseline when tasked with predicting the goal location in VLN given a partially-observed semantic spatial map. On the full VLN task (incorporating the learned policy as well), our approach achieves a success rate on the test server of 32.7% (29.9% SPL ), a credible result for a new class of model trained exclusively with imitation learning and without data augmentation. Although our policy network is specific to the Matterport3D simulator environment, the rest of our pipeline is general and operates without knowledge of the simulator’s navigation graph (which has been heavily utilized in previous work ). We anticipate this could be an advantage for sim-to-real transfer (i.e., in real robot scenarios where a navigation graph is not provided, and could be non-trivial to generate).

Extend the existing Matterport3D simulator used for VLN to support depth image outputs.

Implement and investigate a semantic spatial memory in the context of VLN for the first time.

Propose a novel formulation of instruction following / goal prediction as Bayesian state tracking of a hypothetical human demonstrator.

Show that our approach outperforms a strong baseline for goal location prediction.

Demonstrate credible results on the full VLN task with the addition of a simple reactive policy, with less reliance on navigation constraints than prior work.

Related work

Vision-and-Language Navigation Task. The VLN task , based on the Matterport3D dataset , builds on a rich history of prior work on situated instruction-following tasks beginning with SHRDLU . Despite the task’s difficulty, a recent flurry of work has seen significant improvements in success rates and related metrics . Key developments include the use of instruction-generation (‘speaker’) models for trajectory re-ranking and data augmentation , which have been widely adopted. Other work has focused on developing modules for estimating progress towards the goal and learning when to backtrack . However, comparatively little attention has been paid to the memory architecture of the agent. LSTM memory has been used in all previous work.

Memory architectures for navigation agents. Beyond the VLN task, various categories of memory structures for deep neural navigation agents can be identified in the literature, including unstructured, addressable, metric and topological. General purpose unstructured memory representations, such as LSTM memory , have been used extensively in both 2D and 3D environments . However, LSTM memory does not offer context-dependent storage or retrieval, and so does not naturally facilitate local reasoning when navigating large or complex environments . To overcome these limitations, both addressable and topological memory representations have been proposed for navigating in mazes and for predicting free space. However, in this work we elect to use a metric semantic spatial map – which preserves the geometry of the environment – as our agent’s memory representation since reasoning about observed phenomena from alternative viewpoints is an important aspect of the VLN task. Semantic spatial maps are grid-based representations containing convolutional neural network (CNN) features which have been recently proposed in the context of visual navigation , interactive question answering , and localization . However, there has been little work on incorporating these memory representations into tasks involving natural language. The closest work to ours is Blukis et al. , however our map construction is more sophisticated as we use depth images and do not assume that all pixels lie on the ground plane. Furthermore, our major contribution is formulating instruction-following as Bayesian state tracking.

Preliminaries: Bayes filters

A Bayes filter is a framework for estimating a probability distribution over a latent state s{\boldsymbol{s}} (e.g., the pose of a robot) given a history of observations o{\boldsymbol{o}} and actions a\boldsymbol{a} (e.g., camera observations, odometry, etc.). At each time step tt the algorithm computes a posterior probability distribution bel(st)=p(st∣a1:t,o1:t){bel({\boldsymbol{s}}_{t})=p({\boldsymbol{s}}_{t}\mid\boldsymbol{a}_{1:t},{\boldsymbol{o}}_{1:t})} conditioned on the available data. This is also called the belief.

Taking as a key assumption the Markov property of states, and conditional independence between observations and actions given the state, the belief bel(st)bel({\boldsymbol{s}}_{t}) can be recursively updated from bel(st−1)bel({\boldsymbol{s}}_{t-1}) using two alternating steps to efficiently combine the available evidence. These steps may be referred to as the prediction based on action at\boldsymbol{a}_{t} and the observation update using observation ot{\boldsymbol{o}}_{t}.

Prediction. In the prediction step, the filter processes the action at\boldsymbol{a}_{t} using a motion model p(st ∣ st−1,at)p({\boldsymbol{s}}_{t}~{}\mid~{}{\boldsymbol{s}}_{t-1},\boldsymbol{a}_{t}) that defines the probability of a state st{\boldsymbol{s}}_{t} given the previous state st−1{\boldsymbol{s}}_{t-1} and an action at\boldsymbol{a}_{t}. In particular, the updated belief bel‾(st)\overline{bel}({\boldsymbol{s}}_{t}) is obtained by integrating (summing) over all prior states st−1{\boldsymbol{s}}_{t-1} from which action at\boldsymbol{a}_{t} could have lead to st{\boldsymbol{s}}_{t}, as follows:

Observation update. During the observation update, the filter incorporates information from the observation ot{\boldsymbol{o}}_{t} using an observation model p(ot∣st)p({\boldsymbol{o}}_{t}\mid{\boldsymbol{s}}_{t}) which defines the likelihood of an observation ot{\boldsymbol{o}}_{t} given a state st{\boldsymbol{s}}_{t}. The observation update is given by:

where η\eta is a normalization constant and Equation 2 is derived from Bayes rule.

Differentiable implementations. To apply Bayes filters in practice, a major challenge is to construct accurate probabilistic motion and observation models for a given choice of belief representation bel(st)bel({\boldsymbol{s}}_{t}). However, recent work has demonstrated that Bayes filter implementations – including Kalman filters , histogram filters and particle filters – can be embedded into deep neural networks. The resulting models may be seen as new recurrent architectures that encode algorithmic priors from Bayes filters (e.g., explicit representations of uncertainty, conditionally independent observation and motion models) yet are fully differentiable and end-to-end learnable.

Agent model

In this section, we describe our VLN agent that simultaneously: (1) builds a semantic spatial map from first-person views; (2) determines the most probable goal location in the current map by filtering likely trajectories taken by a human demonstrator from the start location (i.e., the ‘ghost’); and (3) executes actions to reach the predicted goal. Each of these functions is the responsibility of a separate module which we refer to as the mapper, filter, and policy, respectively. We begin with the mapper.

Inputs. As with previous work on VLN task , we provide the agent with a panoramic view of its environment at each time stepThe panoramic setting is chosen for comparison with prior work – not as a requirement of our architecture. comprised of a set of RGB images It={It,1,It,2,…,It,K}{{\cal I}_{t}=\{I_{t,1},I_{t,2},\dots,I_{t,K}\}}, where It,kI_{t,k} represents the image captured in direction kk. The agent also receives the associated depth images Dt={Dt,1,Dt,2,…,Dt,K}{{\cal D}_{t}=\{D_{t,1},D_{t,2},\dots,D_{t,K}\}} and camera poses Pt={Pt,1,Pt,2,…,Pt,K}{{\cal P}_{t}=\{P_{t,1},P_{t,2},\dots,P_{t,K}\}}. We additionally assume that the camera intrinsics and the ground plane are known. In the VLN task, these inputs are provided by the simulator, in other settings they could be provided by SLAM systems etc.

Map update. To integrate map observations Ft{\cal F}_{t} into our semantic spatial map Mt{\cal M}_{t}, we use a convolutional implementation of a Gated Recurrent Unit (GRU) . In preliminary experiments we found that using convolutions in both the input-to-state and state-to-state transitions reduced the variance in the performance of the complete agent by sharing information across neighboring map cells. However, since both the map Mt{\cal M}_{t} and the map update Ft{\cal F}_{t} are sparse, we use a sparsity-aware convolution operation that evaluates only observed pixels and normalizes the output . We also mask the GRU map update to prevent bias terms from accumulating in the unobserved regions.

2 Filter

At the beginning of each episode the agent is placed at a start location s0∗=(x0,y0,θ0){{\boldsymbol{s}}^{*}_{0}=(x_{0},y_{0},\theta_{0})}, where θ\theta represents the agent’s heading and xx and yy are coordinates in the world frame as previously described. The agent is given an instruction X{\cal X} describing the trajectory to an unknown goal coordinate sT∗=(xT,yT,⋅){{\boldsymbol{s}}^{*}_{T}=(x_{T},y_{T},\cdot)}. As an intermediate step towards actually reaching the goal, we wish to identify likely goal locations in the partially-observed semantic spatial map M{\cal M} generated by the mapper.

Our approach to this problem is based on the observation that a natural language navigation instruction typically conveys a sequence of expected future observations and actions, as previously discussed. Based on this observation, we frame the problem of determining the goal location sT∗{\boldsymbol{s}}^{*}_{T} as a tracking problem. As illustrated in Figure 2 and described further below, we implement a Bayes filter to track the pose st∗{\boldsymbol{s}}^{*}_{t} of a hypothetical human demonstrator (i.e., the ‘ghost’) from the start location to the goal. As inputs to the filter, we provided a series of latent observations ot{\boldsymbol{o}}_{t} and actions at\boldsymbol{a}_{t} extracted from the navigation instruction X{\cal X}. The output of the filter is the belief over likely goal locations bel(sT)bel({\boldsymbol{s}}_{T}).

Observations and actions. To transform the instruction X{\cal X} into a latent representation of observations o{\boldsymbol{o}} and actions a\boldsymbol{a}, we use a sequence-to-sequence model with attention . We first tokenize the instruction into a sequence of words X={x1,x2,…,xl}{{\cal X}=\{{\boldsymbol{x}}_{1},{\boldsymbol{x}}_{2},\dots,{\boldsymbol{x}}_{l}\}} which are encoded using learned word embeddings and a bi-directional LSTM to output a series of encoder hidden states {e1,e2,…,el}{\{\boldsymbol{e}_{1},\boldsymbol{e}_{2},\dots,\boldsymbol{e}_{l}\}} and a final hidden state e\boldsymbol{e} representing the output of a complete pass in each direction. We then use an LSTM decoder to generate a series of latent observation and action vectors {o1,o2,…,oT}{\{{\boldsymbol{o}}_{1},{\boldsymbol{o}}_{2},\dots,{\boldsymbol{o}}_{T}\}} and {a1,a2,…,aT}{\{\boldsymbol{a}_{1},\boldsymbol{a}_{2},\dots,\boldsymbol{a}_{T}\}} respectively. Here, ot{\boldsymbol{o}}_{t} is given ot=[e^to,ht]{\boldsymbol{o}}_{t}=[\hat{\boldsymbol{e}}^{o}_{t},{\boldsymbol{h}}_{t}], where ht{\boldsymbol{h}}_{t} is the hidden state of the decoder LSTM, and e^to\hat{\boldsymbol{e}}^{o}_{t} is the attended instruction representation computed using a standard dot-product attention mechanism . The action vectors at\boldsymbol{a}_{t} are computed analogously, using the same decoder LSTM but with a separate learned attention mechanism. The only input to the decoder LSTM is a positional encoding of the decoding time step tt. While the correct number of decoding time steps TT is unknown, in practice we always run the filter for a fixed number of time steps equal to the maximum trajectory length in the dataset (which is 6 steps in the navigation graph).

Motion model. We implement the motion model p(st∣st−1,at,M){p({\boldsymbol{s}}_{t}\mid{\boldsymbol{s}}_{t-1},\boldsymbol{a}_{t},{\cal M})} as a convolution over the belief bt−1\boldsymbol{b}_{t-1}. This ensures that agent motion is consistent across the state space while explicitly enforcing locality, i.e., the agent cannot move further than half the kernel size in a single time step. Similarly to Jonschkowski and Brock , the prediction step from Equation 1 is thus reformulated as:

where conv is a small 3-layer CNN with ReLU activations operating on the semantic spatial map M{\cal M} and the spatially-tiled action vector at\boldsymbol{a}_{t}, MM is the motion kernel size and the softmax function enforces the prior that g(at,M)g(\boldsymbol{a}_{t},{\cal M}) represents a probability mass function. Note that we include M{\cal M} in the input so that the motion model can learn that the agent is unlikely to move through obstacles.

Observation model. We require an observation model p(ot∣st,M){p({\boldsymbol{o}}_{t}\mid{\boldsymbol{s}}_{t},{\cal M})} to define the likelihood of a latent observation ot{\boldsymbol{o}}_{t} conditioned on the agent’s state st{\boldsymbol{s}}_{t} and the map M{\cal M}. A generative observation model like this would be hard to learn, since it is not clear how to generate high-dimensional latent observations and normalization needs to be done across observations, not states. Therefore, we follow prior work and learn a discriminative observation model that takes ot{\boldsymbol{o}}_{t} and M{\cal M} as inputs and directly outputs the likelihood of this observation for each state. As detailed further in Section 4.4, this observation model is trained end-to-end without direct supervision of the likelihood.

To implement our observation model we use LingUNet , a language-conditioned image-to-image network based on U-Net . Specifically, we use the LingUNet implementation from Blukis et al. with 3 cascaded convolution and deconvolution operations. The spatial dimensionality of the LingUNet output matches the input image (in this case, M{\cal M}), and number of output channels is selected to match the number of heading bins Θ\Theta. Outputs are restricted to the range $$ using a sigmoid function. The observation update from Equation 2 is re-defined as:

where η\eta is a normalization constant and ⊙\odot represents element-wise multiplication.

Goal prediction. In summary, to identify goal locations in the partially-observed spatial map M{\cal M}, we initialize the belief b0\boldsymbol{b}_{0} with the known starting state s0{\boldsymbol{s}}_{0}. We then iteratively: (1) Generate a latent observation ot{\boldsymbol{o}}_{t} and action at\boldsymbol{a}_{t}, (2) Compute the prediction step using Equation 3, and (3) Compute the observation update using Equation 5. We stop after TT filter update time steps. The resulting belief bT\boldsymbol{b}_{T} represents the posterior probability distribution over goal locations.

3 Policy

The final component of our agent is a simple reactive policy network. It operates over a global action space defined by the complete set of panoramic viewpoints observed in the current episode (including both visited viewpoints, and their immediate neighbors). Our agent thus memorizes the local structure of the observed navigation graph to enable it to return to any previously observed location in a single action. The probability distribution over actions is defined by a softmax function, where the logit associated with each viewpoint ii is given by yi=MLP([b1:T,i,vi])y_{i}=\text{MLP}([\boldsymbol{b}_{1:T,i},{\boldsymbol{v}}_{i}]), where MLP is a two-layer neural network, b1:T,i\boldsymbol{b}_{1:T,i} is a vector containing the belief at each time step 1:T1:T in a gaussian neighborhood around viewpoint ii, and vi{\boldsymbol{v}}_{i} is a vector containing the distance from the agent’s current location to viewpoint ii, and an indicator variable for whether ii has been previously visited. If the policy chooses to revisit a previously visited viewpoint, we interpret this as a stop action. Note that our policy does not have direct access to any representation of the instruction, or the semantic map M{\cal M}. Although our policy network is specific to the Matterport3D simulator environment, the rest of our pipeline is general and operates without knowledge of the simulator’s navigation graph.

4 Learning

Our entire agent model is fully differentiable, from policy actions back to image pixels via the semantic spatial map, geometric feature projection function, etc. Training data for the model consists of instruction-trajectory pairs (X,s1:T∗)({\cal X},{\boldsymbol{s}}^{*}_{1:T}). In all experiments we train the filter using supervised learning by minimizing the KL-divergence between the predicted belief b1:T\boldsymbol{b}_{1:T} and the true trajectory from the start to the goal s1:T∗{\boldsymbol{s}}^{*}_{1:T}, backpropagating gradients through the previous belief bt−1\boldsymbol{b}_{t-1} at each step. Note that the predicted belief b1:T\boldsymbol{b}_{1:T} is independent of the agent’s actual trajectory s1:T{\boldsymbol{s}}_{1:T} given the map M{\cal M}. In the goal prediction experiments (Section 5.2), the model is trained without a policy and so the agent’s trajectory s1:T{\boldsymbol{s}}_{1:T} is generated by moving towards the goal with 50% probability, or randomly otherwise. In the full VLN experiments (Section 5.3), we train the filter concurrently with the policy. The policy is trained with cross-entropy loss to maximize the likelihood of the ground-truth target action, defined as the first action in the shortest path from the agent’s current location st{\boldsymbol{s}}_{t} to the goal sT∗{\boldsymbol{s}}^{*}_{T}. In this regime, trajectories are generated by sampling an action from the policy with 50% probability, or selecting the ground-truth target action otherwise. In both sets of experiments we train all parameters end-to-end (except for the pretrained CNN). We have verified that the stand-alone performance of the filter is not unduly impacted by the addition of the policy, but we leave the investigation of more sophisticated RL training regimes to future work.

Implementation details. We provide further implementation details in the supplementary material. PyTorch code will be released to replicate all experimentshttps://github.com/batra-mlp-lab/vln-chasing-ghosts and a video overview is also availablehttps://www.youtube.com/watch?v=eoGbescCNP0.

Experiments

Simulator. We use the Matterport3D Simulator based on the Matterport3D dataset containing RGB-D images, textured 3D meshes and other annotations captured from 11K panoramic viewpoints densely sampled throughout 90 buildings. Using this dataset, the simulator implements a visually-realistic first-person environment that allows the agent to look in any direction while moving between panoramic viewpoints along edges in a navigation graph. Viewpoints are 2.25m apart on average.

Depth outputs. As the Matterport3D Simulator supports RGB output only, we extend it to support depth outputs which are necessary to accurately project CNN features into the semantic spatial map. Our simulator extension projects the undistorted depth images from the Matterport3D dataset onto cubes aligned with the provided ‘skybox’ images, such that each cube-mapped pixel represents the euclidean distance from the camera center. We then adapt the existing rendering pipeline to render depth images from these cube-maps, converting depth values from euclidean distance back to distance from the camera plane in the process. To fill missing depth values corresponding to shiny, bright, transparent, and distant surfaces, we apply a simple cross-bilateral filter based on the NYUv2 implementation . We additionally implement various other performance improvements, such as caching, which boosts the frame-rate of the simulator up to 1000 FPS, subject to GPU performance and CPU-GPU memory bandwith. We have incorporated these extensions into the original simulator codebase.https://github.com/peteanderson80/Matterport3DSimulator

R2R instruction dataset. We evaluate using the Room-to-Room (R2R) dataset for Vision-and-Language Navigation (VLN) . The dataset consists of 22K open-vocabulary, crowd-sourced navigation instructions with an average length of 29 words. Each instruction corresponds to a 5–24m trajectory in the Matterport3D dataset, traversing 5–7 viewpoint transitions. Instructions are divided into splits for training, validation and testing. The validation set is further split into two components: val-seen, where instructions and trajectories are situated in environments seen during training, and val-unseen containing instructions situated in environments that are not seen during training. All the test set instructions and trajectories are from environments that are unseen in training and validation.

2 Goal prediction results

We first evaluate the goal prediction performance of our proposed mapper and filter architecture in a policy-free setting using fixed trajectories. Trajectories are generated by an agent that moves towards the goal with 50% probability, or randomly otherwise. As an ablation, we also report results for our model excluding heading from the agent’s filter state, i.e., st=(x,y){\boldsymbol{s}}_{t}=(x,y), to quantify the value of encoding the agent’s orientation in the motion and observation models. We compare to two baselines as follows:

LingUNet baseline. As a strong neural net baseline, we compare to LingUNet – a language-conditioned variant of the U-Net image-to-image architecture – that has recently been applied to goal location prediction in the context of a simulated quadrocopter instruction-following task . We choose LingUNet because existing VLN models do not explicitly model the goal location or the map, and are thus not capable of predicting the goal location from a provided trajectory. Following Blukis et al. we train a 5-layer LingUNet module conditioned on the sentence encoding e\boldsymbol{e} and the semantic map M{\cal M} to directly predict the goal location distribution (as well as a path visitation distribution, as an auxilliary loss) in a single forward pass. As we implement our observation model using a (smaller, 3-layer) LingUNet, the LingUNet baseline resembles an ablated single-step version of our model that dispenses with the decoder generating latent observations and actions as well as the motion model. Note that we use the same mapper architecture for our filter and for LingUNet.

Hand-coded baseline. We additionally compare to hand-coded goal prediction baseline designed to exploit biases in the R2R dataset and the provided trajectories. We first calculate the mean straight-line distance from the start position to the goal across the entire training set, which is 7.6m. We then select as the predicted goal the position (x,y)(x,y) in the map at a radius of 7.6m from the start position that has the greatest observed map area in an Gaussian-weighted neighborhood of (x,y)(x,y).

Results. As illustrated in Table 1, our proposed filter architecture that explicitly models belief over trajectories that could be taken by a human demonstrator outperforms a strong LingUNet baseline at predicting the goal location (with an average success rate of 45% vs. 32% in unseen environments). This finding holds at all time steps (i.e., regardless of the sparsity of the map). We also demonstrate that removing the heading θ\theta from the agent’s state in our model degrades this success rate to 39%, demonstrating the importance of relative orientation to instruction understanding. For instance, it is unlikely for an agent following the true path to turn 180 degrees midway through (unless this is commanded by the instruction). Similarly, without knowing heading, the model can represent instructions such as ‘go past the table’ but not ‘go past with the table on your left’. Finally, the poor performance of the handcoded baseline confirms that the goal location cannot be trivially predicted from the trajectory.

3 Vision-and-Language Navigation results

Having established the efficacy of our approach for goal prediction from a partial map, we turn to the full VLN task that requires our agent to take actions to actually reach the goal.

Evaluation. In VLN, an episode is successful if the final navigation error is less than 3m. We report our agent’s average success rate at reaching the goal (SR), and SPL , a recently proposed summary measure of an agent’s navigation performance that balances navigation success against trajectory efficiency (higher is better). We also report trajectory length (TL) and navigation error (NE) in meters, as well as oracle success (OS), defined as the agent’s success rate under an oracle stopping rule.

Results. In Table 2, we present our results in the context of state-of-the-art methods; however, as noted by the RL and Aug columns in the table, these approaches include reinforcement learning and complex data augmentation and pretraining strategies. These are non-trivial extensions that are the result of a community effort and are orthogonal to our own contribution. We also use a less powerful CNN (ResNet-34 vs. ResNet-152 in prior work). For the most direct comparison, we consider the ablated models in the lower panel of Table 2 to be most appropriate. We find these results promising given this is the first work to explore such a drastically different model class (i.e., maintaining a metric map and a probability distribution over alternative trajectories in the map). Our model also exhibits less overfitting than other approaches – performing equally well on both seen (val-seen) and unseen (val-unseen) environments.

Further, our filtering approach allows us greater insight into the model. We examine a qualitative example in Figure 3. On the left, we can see the agent attends to appropriate visual and direction words when generating latent observations and actions, supporting the intuition in Figure 1. On the right, we can see the growing confidence our goal predictor places on the correct location as more of the map is explored – despite the increasing number of visible alternatives. We provide further examples (including insight into the motion and observation models) in the supplementary video.

Conclusion

We show that instruction following can be formulated as Bayesian state tracking in a model that maintains a semantic spatial map of the environment, and an explicit probability distribution over alternative possible trajectories in that map. To evaluate our approach we choose the complex problem of Vision-and-Language Navigation (VLN). This represents a significant departure from existing work in the area, and required augmenting the Matterport3D simulator with depth. Empirically, we show that our approach outperforms recent alternative approaches to goal location prediction, and achieves credible results on the full VLN task without using RL or data augmentation – while offering reduced overfitting to seen environments, unprecedented intepretability and less reliance on the simulator’s navigation constraints.

We thank Abhishek Kadian and Prithviraj Ammanabrolu for their help in the initial stages of the project. The Georgia Tech effort was supported in part by NSF, AFRL, DARPA, ONR YIPs, ARO PECASE. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government, or any sponsor.

References

SUPPLEMENTARY MATERIAL

Implementation Details

Simulator. In experiments, we set the Matterport3D simulator to generate 320×256320\times 256 pixel images with a 6060 degree vertical field of view. To capture more of the floor and nearby obstacles (and less of the roof) we set the camera elevation to 3030 degrees down from horizontal. At each panoramic viewpoint location in the simulator we capture a horizontal sweep containing 12 images at 30 degree increments, which are projected into the map in a single time step as described in Section 4.1 of the main paper.

Mapper. For our CNN implementation we use a ResNet-34 architecture that is pretrained on ImageNet . We found that fine-tuning the CNN while training our model mainly improved performance on the Val-Seen set, and so we left the CNN parameters fixed in the reported experiments. To extract the visual feature representation v{\boldsymbol{v}} we concatenate the output from the CNN’s last 2 layers to provide a 16×20×768{16\times 20\times 768} representation. The dimensionality of our map representation M{\cal M} is fixed at 128×96×96{128\times 96\times 96} and each cell represents a square region with side length 0.50.5m (the entire map is thus 48m×48m48\text{m}\times 48\text{m}). In the mapper’s convolutional GRU we use 3×33\times 3 convolutional filters and we train with spatial dropout of 0.50.5 in both the input-to-state and state-to-state transitions with fixed dropout masks for the duration of each episode.

Filter. In the instruction encoder we use a hidden state size of 256256 for both the forward and backward encoders, and a word embedding size of 300300. We use a motion kernel size MM of 7, but we upscale the motion kernel g(at,M)g(\boldsymbol{a}_{t},{\cal M}) by a scale factor of 2×2\times before applying it such that the agent can move a maximum of 3.53.5m in a single time step.

Training. In training, we use the Adam optimizer with an initial learning rate of 1e-3, weight decay of 1e-7, and a batch size of 5. In the goal prediction experiment, all models are trained for 8K iterations, after which all models have converged. In the full VLN experiment, our models are trained for 17.5K iterations, and we pick the iteration with the highest SPL performance on Val-Unseen to report and submit to the test server. Training the model takes around 1 day for goal prediction, and 2.5 days for the full VLN task, using a single Titan X GPU.

Visualizations. In the main paper and the supplementary videohttps://www.youtube.com/watch?v=eoGbescCNP0, we depict top-down floorplan visualizations of Matterport environments to provide greater insight into the model’s behavior. These visualizations are rendered from textured meshes in the Matterport3D dataset , using the provided GAPS software which was modified to render using an orthographic projection.

Visualizations

Observations and actions. In this section we provide further visualizations of the attention weights in the sequence decoders that generate latent observations and actions (refer to Section 4.2 of the main paper). Instructions are examples from the Val-Unseen set. In general, the attention models for the motion model (generating latent action vectors a\boldsymbol{a}) and the observation model (generating latent observation vectors o{\boldsymbol{o}}) specialize in different ways. The motion model focuses attention on action words, while the observation model focuses on visual words, as illustrated in Figures 4 and 5. The sequential ordering of the instructions (e.g., attention weights showing a diagonal structure from top-left to bottom-right) is also evident.