Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments

Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, Yoav Artzi

Introduction

Consider the visual challenges of following natural language instructions in a busy urban environment. Figure 1 illustrates this problem. The agent must identify objects and their properties to resolve mentions to traffic light and American flags, identify patterns in how objects are arranged to find the flow of traffic, and reason about how the relative position of objects changes as it moves to go past objects. Reasoning about vision and language has been studied extensively with various tasks, including visual question answering , visual navigation , interactive question answering , and referring expression resolution . However, existing work has largely focused on relatively simple visual input, including object-focused photographs or simulated environments . While this has enabled significant progress in visual understanding, the use of real-world visual input not only increases the challenge of the vision task, it also drastically changes the kind of language it elicits and requires fundamentally different reasoning.

In this paper, we study the problem of reasoning about vision and natural language using an interactive visual navigation environment based on Google Street View.\hrefhttps://developers.google.com/maps/documentation/streetview/introhttps://developers.google.com/maps/documentation/streetview/intro We design the task of first following instructions to reach a goal position, and then resolving a spatial description at the goal by identifying the location in the observed image of Touchdown, a hidden teddy bear. Using this environment and task, we release Touchdown, Touchdown is the \hrefhttps://en.wikipedia.org/wiki/Touchdown_(mascot)unofficial mascot of Cornell University. a dataset for navigation and spatial reasoning with real-life observations.

We design our task for diverse use of spatial reasoning, including for following instructions and resolving the spatial descriptions. Navigation requires the agent to reason about its relative position to objects and how these relations change as it moves through the environment. In contrast, understanding the description of the location of Touchdown requires the agent to reason about the spatial relations between observed objects. The two tasks also diverge in their learning challenges. While in both learning requires relying on indirect supervision to acquire spatial knowledge and language grounding, for navigation, the training data includes demonstrated actions, and for spatial description resolution, annotated target locations. The task can be addressed as a whole, or decomposed to its two portions.

The key data collection challenge is designing a scalable process to obtain natural language data that reflects the richness of the visual input while discouraging overly verbose and unnatural language. In our data collection process, workers write and follow instructions. The writers navigate in the environment and hide Touchdown. Their goal is to make sure the follower can execute the instruction to find Touchdown. The measurable goal allows us to reward effective writers, and discourages overly verbose descriptions.

We collect 9,3269{,}326 examples of the complete task, which decompose to the same number of navigation tasks and 27,57527{,}575 spatial description resolution (SDR) tasks. Each example is annotated with a navigation demonstration and the location of Touchdown. Our linguistically-driven analysis shows the data requires significantly more complex reasoning than related datasets. Nearly all examples require resolving spatial relations between observable objects and between the agent and its surroundings, and each example contains on average 5.35.3 commands and refers to 10.710.7 unique entities in its environment.

We empirically study the navigation and SDR tasks independently. For navigation, we focus on the performance of existing models trained with supervised learning. For SDR, we cast the problem of identifying Touchdown’s location as an image feature reconstruction problem using a language-conditioned variant of the UNet architecture . This approach significantly outperforms several strong baselines.

Related Work and Datasets

Jointly reasoning about vision and language has been studied extensively, most commonly focusing on static visual input for reasoning about image captions and grounded question answering . Recently, the problem has been studied in interactive simulated environments where the visual input changes as the agent acts, such as interactive question answering and instruction following . In contrast, we focus on an interactive environment with real-world observations.

The most related resources to ours are R2R and Talk the Walk . R2R uses panorama graphs of house environments for the task of navigation instruction following. It includes 9090 unique environments, each containing an average of 119119 panoramas, significantly smaller than our 29,64129{,}641 panoramas. Our larger environment requires following the instructions closely, as finding the goal using search strategies is unlikely, even given a large number of steps. We also observe that the language in our data is significantly more complex than in R2R (Section 5). Our environment setup is related to Talk the Walk, which uses panoramas in small urban environments for a navigation dialogue task. In contrast to our setup, the instructor does not observe the panoramas, but instead sees a simplified diagram of the environment with a small set of pre-selected landmarks. As a result, the instructor has less spatial information compared to Touchdown. Instead the focus is on conversational coordination.

SDR is related to the task of referring expression resolution, for example as studied in ReferItGame and Google Refexp . Referring expressions describe an observed object, mostly requiring disambiguation between the described object and other objects of the same type. In contrast, the goal of SDR is to describe a specific location rather than discriminating. This leads to more complex language, as illustrated by the comparatively longer sentences of SDR (Section 5). Kitaev and Klein proposed a similar task to SDR, where given a spatial description and a small set of locations in a fully-observed simulated 3D environment, the system must select the location described from the set. We do not use distractor locations, requiring a system to consider all areas of the image to resolve a spatial description.

Environment and Tasks

We use Google Street View to create a large navigation environment. Each position includes a 360∘ RGB panorama. The panoramas are connected in a graph-like structure with undirected edges connecting neighboring panoramas. Each edge connects to a panorama in a specific heading. For each panorama, we render perspective images for all headings that have edges. Our environment includes 29,64129{,}641 panoramas and 61,31961{,}319 edges from New York City. Figure 2 illustrates the environment.

We design two tasks: navigation and spatial description resolution (SDR). Both tasks require recognizing objects and the spatial relations between them. Navigation focuses on egocentric spatial reasoning, where instructions refer to the agent’s relationship with its environment, including the objects it observes. The SDR task displays more allocentric reasoning, where the language requires understanding the relations between the observed objects to identify the target location. While navigation requires generating a sequence of actions from a small set of possible actions, SDR requires choosing a specific pixel in the observed image. Both tasks present different learning challenges. The navigation task could benefit from reward-based learning, while the SDR task defines a supervised learning problem. The two tasks can be addressed separately, or combined by completing the SDR task at the goal position at the end of the navigation.

The agent’s goal is to follow a natural language instruction and reach a goal position. Let S\mathcal{S} be the set of all states. A state s∈Ss\in\mathcal{S} is a pair (I,α)(\textbf{I},\alpha), where I is a panorama and α\alpha is the heading angle indicating the agent heading. We only allow states where there is an edge connecting to a neighboring panorama in the heading α\alpha. Given a navigation instruction xˉn\bar{x}_{n} and a start state s1∈Ss_{1}\in\mathcal{S}, the agent performs a sequence of actions. The set of actions A\mathcal{A} is {FORWARD,LEFT,RIGHT,STOP}\{{\tt{FORWARD}},{\tt{LEFT}},{\tt{RIGHT}},{\tt{STOP}}\}. Given a state ss and an action a∈Aa\in\mathcal{A}, the state is deterministically updated using a transition function T:S×A→ST:\mathcal{S}\times\mathcal{A}\rightarrow\mathcal{S}. The FORWARD{\tt{FORWARD}} action moves the agent along the edge in its current heading. Formally, if the environment includes the edge (Ii,Ij)(\textbf{I}_{i},\textbf{I}_{j}) at heading α\alpha in Ii\textbf{I}_{i}, the transition is T((Ii,α),FORWARD)=(Ij,α′)T((\textbf{I}_{i},\alpha),{\tt{FORWARD}})=(\textbf{I}_{j},\alpha^{\prime}). The new heading α′\alpha^{\prime} is the heading of the edge in Ij\textbf{I}_{j} with the closest heading to α\alpha. The LEFT{\tt{LEFT}} (RIGHT{\tt{RIGHT}}) action changes the agent heading to the heading of the closest edge on the left (right). Formally, if the position panorama I has edges at headings α>α′>α′′\alpha>\alpha^{\prime}>\alpha^{\prime\prime}, T((I,α),LEFT)=(I,α′)T((\textbf{I},\alpha),{\tt{LEFT}})=(\textbf{I},\alpha^{\prime}) and T((I,α),RIGHT)=(I,α′′)T((\textbf{I},\alpha),{\tt{RIGHT}})=(\textbf{I},\alpha^{\prime\prime}). Given a start state s1s_{1} and a navigation instruction xˉn\bar{x}_{n}, an execution eˉ\bar{e} is a sequence of state-action pairs ⟨(s1,a1),...,(sm,am)⟩\langle(s_{1},a_{1}),...,(s_{m},a_{m})\rangle, where T(si,ai)=si+1T(s_{i},a_{i})=s_{i+1} and am=STOPa_{m}={\tt{STOP}}.

We use three evaluation metrics: task completion, shortest-path distance, and success-weighted edit distance. Task completion (TC) measures the accuracy of completing the task correctly. We consider an execution correct if the agent reaches the exact goal position or one of its neighboring nodes in the environment graph. Shortest-path distance (SPD) measures the mean distance in the graph between the agent’s final panorama and the goal. SPD ignores turning actions and the agent heading. Success weighted by edit distance (SED) is 1N∑i=1NSi(1−lev(eˉ,eˉ^)max⁡(∣eˉ∣,∣eˉ^∣))\frac{1}{N}\sum_{i=1}^{N}S_{i}(1-\frac{{\rm lev}(\bar{e},\hat{\bar{e}})}{\max(|\bar{e}|,|\hat{\bar{e}}|)}), where the summation is over NN examples, SiS_{i} is a binary task completion indicator, eˉ\bar{e} is the reference execution, eˉ^\hat{\bar{e}} is the predicted execution, lev(⋅,⋅){\rm lev}(\cdot,\cdot) is the Levenshtein edit distance, and ∣⋅∣|\cdot| is the execution length. The edit distance is normalized and inversed. We measure the distance and length over the sequence of panoramas in the execution, and ignore changes of orientation. SED is related to success weighted by path length (SPL) , but is designed for instruction following in graph-based environments, where a specific correct path exists.

2 Spatial Description Resolution (SDR)

Given an image I and a natural language description xˉs\bar{x}_{s}, the task is to identify the point in the image that is referred to by the description. We instantiate this task as finding the location of Touchdown, a teddy bear, in the environment. Touchdown is hidden and not visible in the input. The image I is a 360∘ RGB panorama, and the output is a pair of (x,y)(x,y) coordinates specifying a location in the image.

We use three evaluation metrics: accuracy, consistency, and distance error. Accuracy is computed with regard to an annotated location. We consider a prediction as correct if the coordinates are within a slack radius of the annotation. We measure accuracy for radiuses of 40, 80, and 120 pixels and use euclidean distance. Our data collection process results in multiple images for each sentence. We use this to measure consistency over unique sentences, which is measured similar to accuracy, but with a unique sentence considered correct only if all its examples are correct . We compute consistency for each slack value. We also measure the mean euclidean distance between the annotated location and the predicted location.

Data Collection

We frame the data collection process as a treasure-hunt task where a leader hides a treasure and writes directions to find it, and a follower follows the directions to find the treasure. The process is split into four crowdsourcing tasks (Figure 3). The two main tasks are writing and following. In the writing task, a leader follows a prescribed route and hides Touchdown the bear at the end, while writing instructions that describe the path and how to find Touchdown. The following task requires following the instructions from the same starting position to navigate and find Touchdown. Additional tasks are used to segment the instructions into the navigation and target location tasks, and to propagate Touchdown’s location to panoramas that neighbor the final panorama. We use a customized Street View interface for data collection. However, the final data uses a static set of panoramas that do not require the Street View interface.

We generate routes by sampling start and end positions. The sampling process results in routes that often end in the middle of a city block. This encourages richer language, for example by requiring to describe the goal position rather than simply directing to the next intersection. The route generation details are described in the Supplementary Material. For each task, the worker is placed at the starting position facing north, and asked to follow a route specified in an overhead map view to a goal position. Throughout, they write instructions describing the path. The initial heading requires the worker to re-orient to the path, and thereby familiarize with their surroundings better. It also elicits interesting re-orientation instructions that often include references to the direction of objects (e.g., flow of traffic) or their relation to the agent (e.g., the umbrellas are to the right). At the goal panorama, the worker is asked to place Touchdown in a location of their choice that is not a moving object (e.g., a car or pedestrian) and to describe the location in their instructions. The worker goal is to write instructions that a human follower can use to correctly navigate and locate the target without knowing the correct path or location of Touchdown. They are not permitted to write instructions that refer to text in the images, including street names, store names, or numbers.

Task II: Target Propagation to Panoramas

The writing task results in the location of Touchdown in a single panorama in the Street View interface. However, resolving the spatial description to the exact location is also possible from neighboring panoramas where the target location is visible. We use a crowdsourcing task to propagate the location of Touchdown to neighboring panoramas in the Street View interface, and to the identical panoramas in our static data. This allows to complete the task correctly even if not stopping at the exact location, but still reaching a semantically equivalent position. The propagation in the Street View interface is used for our validation task. The task includes multiple steps. At each step, we show the instruction text and the original Street View panorama with Touchdown placed, and ask for the location for a single panorama, either from the Street View interface or from our static images. The worker can indicate if the target is occluded. The propagation annotation allows us to create multiple examples for each SDR, where each example uses the same SDR but shows the environment from a different position.

Task III: Validation

We use a separate task to validate each instruction. The worker is asked to follow the instruction in the customized Street View interface and find Touchdown. The worker sees only the Street View interface, and has no access to the overhead map. The task requires navigation and identifying the location of Touchdown. It is completed correctly if the follower clicks within a 9090-pixel radiusThis is roughly the size of Touchdown. The number is not directly comparable to the SDR accuracy measures due to different scaling. of the ground truth target location of Touchdown. This requires the follower to be in the exact goal panorama, or in one of the neighboring panoramas we propagated the location to. The worker has five attempts to find Touchdown. Each attempt is a click. If the worker fails, we create another task for the same example to attempt again. If the second worker fails as well, the example is discarded.

Task IV: Segmentation

We annotate each token in the instruction to indicate if it describes the navigation or SDR tasks. This allows us to address the tasks separately. First, a worker highlights a consecutive prefix of tokens to indicate the navigation segment. They then highlight a suffix of tokens for the SDR task. The navigation and target location segments may overlap (Figure 4).

Workers and Qualification

We require passing a qualification task to do the writing task. The qualifier task requires correctly navigating and finding Touchdown for a predefined set of instructions. We consider workers that succeed in three out of the four tasks as qualified. The other three tasks do not require qualification. Table 1 shows how many workers participated in each task.

Payment and Incentive Structure

The base pay for instruction writing is \0.60.Fortargetpropagation,validation,andsegmentationwepaid. For target propagation, validation, and segmentation we paid\0.150.15, \0.25,and, and\0.120.12. We incentivize the instruction writers and followers with a bonus system. For each instruction that passes validation, we give the writer a bonus of \0.25andthefollowerabonusofand the follower a bonus of\0.100.10. Both sides have an interest in completing the task correctly. The size of the graph makes it difficult, and even impossible, for the follower to complete the task and get the bonus if the instructions are wrong.

Data Statistics and Analysis

Workers completed 11,01911{,}019 instruction-writing tasks, and 12,66412{,}664 validation tasks. 89.1%89.1\% examples were correctly validated, 80.1%80.1\% on the first attempt and 9.0%9.0\% on the second.Several paths were discarded due to updates in Street View data. While we allowed five attempts at finding Touchdown during validation tasks, 64%64\% of the tasks required a single attempt. The value of additional attempts decayed quickly: only 1.4%1.4\% of the tasks were only successful after five attempts. For the full task and navigation-only, Touchdown includes 9,3269{,}326 examples with 6,5266{,}526 in the training set, 1,3911{,}391 in the development set, and 1,4091{,}409 in the test set. For the SDR task, Touchdown includes 9,3269{,}326 unique descriptions and 25,57525{,}575 examples with 17,88017{,}880 for training, 3,8363{,}836 for development, and 3,8593{,}859 for testing. We use our initial paths as gold-standard demonstrations, and the placement of Touchdown by the original writer as the reference location. Table 2 shows basic data statistics. The mean instruction length is 108.0108.0 tokens. The average overlap between navigation and SDR is 11.411.4 tokens. Figure 5 shows the distribution of text lengths. Overall, Touchdown contains a larger vocabulary and longer navigation instructions than related corpora. The paths in Touchdown are longer than in R2R , on average 35.235.2 panoramas compared to 6.06.0. SDR segments have a mean length of 29.829.8 tokens, longer than in common referring expression datasets; ReferItGame expressions 4.44.4 tokens on average and Google RefExp expressions are 8.58.5.

We perform qualitative linguistic analysis of Touchdown to understand the type of reasoning required to solve the navigation and SDR tasks. We identify a set of phenomena, and randomly sample 2525 examples from the development set, annotating each with the number of times each phenomenon occurs in the text. Table 3 shows results comparing Touchdown with R2R.See the Supplementary Material for analysis of SAIL and Lani. Sentences in Touchdown refer to many more unique, observable entities (10.710.7 vs 3.73.7), and almost all examples in Touchdown include coreference to a previously-mentioned entity. More examples in Touchdown require reasoning about counts, sequences, comparisons, and spatial relationships of objects. Correct execution in Touchdown requires taking actions only when certain conditions are met, and ensuring that the agent’s observations match a described scene, while this is rarely required in R2R. Our data is rich in spatial reasoning. We distinguish two types: between multiple objects (allocentric) and between the agent and its environment (egocentric). We find that navigation segments contain more egocentric spatial relations than SDR segments, and SDR segments require more allocentric reasoning. This corresponds to the two tasks: navigation mainly requires moving the agent relative to its environment, while SDR requires resolving a point in space relative to other objects.

Spatial Reasoning with LingUNet

We cast the SDR task as a language-conditioned image reconstruction problem, where we predict a distribution of the location of Touchdown over the entire observed image.

We use the LingUNet architecture , which was originally introduced for goal prediction and planning in instruction following. LingUNet is a language-conditioned variant of the UNet architecture , an image-to-image encoder-decoder architecture widely used for image segmentation. LingUNet incorporates language into the image reconstruction phase to fuse the two modalities. We modify the architecture to predict a probability distribution over the input panorama image.

We process the description text tokens xˉs=⟨x1,x2,…,xl⟩\bar{x}_{s}=\langle x_{1},x_{2},\dots,x_{l}\rangle using a bi-directional Long Short-term Memory (LSTM) recurrent neural network to generate ll hidden states. The forward computation is hif=BiLSTM(φ(xi),hi−1f),i=1,…,l\textbf{h}_{i}^{f}={\rm BiLSTM}(\varphi(x_{i}),\textbf{h}_{i-1}^{f}),i=1,\dots,l, where φ\varphi is a learned word embedding function. We compute the backward hidden states hib\textbf{h}_{i}^{b} similarly. The text representation is an average of the concatenated hidden states x=1l∑i=1l[hif;hib]\textbf{x}=\frac{1}{l}\sum_{i=1}^{l}[h^{f}_{i};h^{b}_{i}]. We map the RGB panorama I to a feature representation F0\textbf{F}_{0} with a pre-trained ResNet18 .

LingUNet performs mm levels of convolution and deconvolution operations. We generate a sequence of feature maps Fk=\textscCNNk(Fk−1),k=1,…,m\textbf{F}_{k}={\textsc{CNN}}_{k}(\textbf{F}_{k-1}),k=1,\dots,m with learned convolutional layers \textscCNNk{\textsc{CNN}}_{k}. We slice the text representation x to mm equal-sized slices, and reshape each with a linear projection to a 1×11\times 1 filter Kk\textbf{K}_{k}. We convolve each feature map Fk\textbf{F}_{k} with Kk\textbf{K}_{k} to obtain a text-conditioned feature map Gk=\textscConv(Kk,Fk)\textbf{G}_{k}=\textsc{Conv}(\textbf{K}_{k},\textbf{F}_{k}). We use mm deconvolution operations to generate feature maps of increasing size to create H1\textbf{H}_{1}:

We compute a single value for each pixel by projecting the channel vector for each pixel using a single-layer perceptron with a ReLU non-linearity. Finally, we compute a probability distribution over the feature map using a SoftMax. The predicted location is the mode of the distribution.

2 Experimental Setup

The evaluation metrics are described in Section 3.2 and the data in Section 5.

We use supervised learning. The gold label is a Gaussian smoothed distribution. The coordinate of the maximal value of the distribution is the exact coordinate where Touchdown is placed. We minimize the KL-divergence between the Gaussian and the predicted distribution.

Systems

We evaluate three non-learning baselines: (a) Random: predict a pixel at random; (b) Center: predict the center pixel; (c) Average: predict the average pixel, computed over the training set. In addition to a two-level LingUNet (m=2m=2), we evaluate three learning baselines: Concat, ConcatConv, and Text2Conv. The first two compute a ResNet18 feature map representation of the image and then fuse it with the text representation to compute pixel probabilities. The third uses the text to compute kernels to convolve over the ResNet18 image representation. The Supplementary Material provides further details.

3 Results

Table 4 shows development and test results. The low performance of the non-learning baselines illustrates the challenge of the task. We also experiment with a UNet architecture that is similar to our LingUNet but has no access to the language. This result illustrates that visual biases exist in the data, but only enable relatively low performance. All the learning systems outperform the non-learning baselines and the UNet, with LingUNet performing best.

Figure 7 shows pixel-level predictions using LingUNet. The distribution prediction is visualized as a heatmap overlaid on the image. LingUNet often successfully solves descriptions anchored in objects that are unique in the image, such the fire hydrant at the top image. The lower example is more challenging. While the model correctly reasons that Touchdown is on a light just above the doorway, it fails to find the exact door. Instead, the probability distribution is shared between multiple similar locations, the space above three other doors in the image.

Navigation Baselines

We evaluate three non-learning baselines: (a) Stop: agent stops immediately; (b) Random: Agent samples non-stop actions uniformly until reaching the action horizon; and (c) Frequent: agent always takes the most frequent action in the training set (FORWARD). We also evaluate two recent navigation models: (a) GA: gated-attention ; and (b) RConcat: a recently introduced model for landmark-based navigation in an environment that uses Street View images . We represent the input images with ResNet18 features similar to the SDR task.

We use asynchronous training using multiple clients to generate rollouts on different partitions of the training data. We compute the gradients and updates using Hogwild! and Adam learning rates . We use supervised learning by maximizing the log-likelihood of actions in the reference demonstrations.

The details of the models, learning, and hyperparameters are provided in the Supplementary Material.

2 Results

Table 5 shows development and test results for our three valuation metrics (Section 3.1). The Stop, Frequent and Random illustrate the complexity of the task. The learned baselines perform better. We observe that RConcat outperforms GA across all three metrics. In general though, the performance illustrates the challenge of the task. Appendix F includes additional navigation experiments, including single-modality baselines.

Complete Task Performance

We use a simple pipeline combination of the best models of the SDR and navigation tasks to complete the full task. Task completion is measured as finding Touchdown. We observe an accuracy of 4.54.5% for a threshold of 8080px. In contrast, human performance is significantly higher. We estimate human performance using our annotation statistics . To avoid spam and impossible examples, we consider only examples that were successfully validated. We then measure the performance of workers that completed over 3030 tasks for these valid examples. This includes 5555 workers. Because some examples required multiple tries to validate this set includes tasks that workers failed to execute but were later validated. The mean performance across this set of workers using the set of valid tasks is 92%92\% accuracy.

Data Distribution and Licensing

We release the environment graph as panorama IDs and edges, scripts to download the RGB panoramas using the Google API, the collected data, and our code at \hrefhttp://touchdown.aitouchdown.ai. These parts of the data are released with a CC-BY 4.04.0 license. Retention of downloaded panoramas should follow Google’s policies. We also release ResNet18 image features of the RGB panoramas through a request form. The complete license is available with the data.

Conclusion

We introduce Touchdown, a dataset for natural language navigation and spatial reasoning using real-life visual observations. We define two tasks that require addressing a diverse set of reasoning and learning challenges. Our linguistically-driven analysis shows the data presents complex spatial reasoning challenges. This illustrates the benefit of using visual input that reflects the type of observations people see in their daily life, and demonstrates the effectiveness of our goal-driven data collection process.

Acknowledgements

This research was supported by a Google Faculty Award, NSF award CAREER-1750499, NSF Graduate Research Fellowship DGE-1650441, and the generosity of Eric and Wendy Schmidt by recommendation of the Schmidt Futures program. We wish to thank Jason Baldridge for his extensive help and advice, and Valts Blukis and the anonymous reviewers for their helpful comments.

References

Appendix A Route Generation

We generate each route by randomly sampling two panoramas in the environment graph, and querying the Google Direction API\hrefhttps://developers.google.com/maps/documentation/directions/starthttps://developers.google.com/maps/documentation/directions/start to obtain a route between them that follows correct road directions. Although the routes follow the direction of allowed traffic, the panoramas might still show moving against the traffic in two-way streets depending on which lane was used for the original panorama collection by Google. The route is segmented into multiple routes with length sampled uniformly between 3535 and 4545. We do not discard the suffix route segment, which may be shorter. Some final routes had gaps due to our use of the API. If the number of gaps is below three, we heuristically connect the detached parts of the route by adding intermediate panoramas, otherwise we remove the route segment. Each of the route segments is used in a separate instruction-writing task. Because panoramas and route segments are sampled randomly, the majority of route segments stop in the middle of a block, rather than at an intersection. This explicit design decision requires instruction-writers to describe exactly where in the block the follower should stop, which elicits references to a variety of object types, rather than simply referring to the location of an intersection.

Appendix B Additional Data Analysis

We perform linguistically-driven analysis to two additional navigation datasets: SAIL and LANI , both using simulated environments. Both datasets include paragraphs segmented into single instructions. We performs our analysis at the paragraph level. We use the same categories as in Section 5. Table 6 shows the analysis results. In general, in addition to the more complex visual input, Touchdown displays similar or increased linguistic diversity compared to Lani and SAIL. Lani contains a similar amount of coreference, egocentric spatial relations, and temporal conditions, and more examples than Touchdown of imperatives and directions. SAIL contains a similar number of imperatives, and more examples of counts than Touchdown. We also visualize some of the common nouns and modifiers observed in our data (Figure 8).

Appendix C SDR Pixel-level Predictions

Figures 9–14 show SDR pixel-level predictions for comparing the four models we used: LingUNet, Concat, ConcatConv, and Concat. Each figure shows the SDR description to resolve followed by the model outputs. We measure accuracy at a threshold of 8080 pixels. Red-overlaid pixels visualize the Gaussian smoothed annotated target location. Green-overlaid pixels visualize the model’s probability distribution over pixels.

Appendix D SDR Experimental Setup Details

We use learned word vectors of size 300300. For all models, we use a single-layer, bi-directional recurrent neural network (RNN) with long short-term memory (LSTM) cells to encode the description into a fixed-size vector representation. The hidden layer in the RNN has 600600 unit. We compute the text embedding by averaging the RNN hidden states.

We provide the model with the complete panorama. We embed the panorama by slicing it into eight images, and projecting each image from a equirectangular projection to a perspective projection. Each of the eight projected images is of size 800×460800\times 460. We pass each image separately through a ResNet18 pretrained on ImageNet , and extract features from the fourth to last layer before classification; each slice’s feature map is of size 128×100×58128\times 100\times 58. Finally, the features for the eight image slices are concatenated into a single tensor of size 128×100×464128\times 100\times 464.

We concatenate the text representation along the channel dimension of the image feature map at each feature pixel and apply a multi-layer perceptron (MLP) over each pixel to obtain a real-value score for every pixel in the feature map. The multilayer perceptron includes two fully-connected layers with biases and ReLu\rm ReLu non-linearities on the output of the first layer. The hidden size of each layer is 128128. A SoftMax layer is applied to generate the final probability distribution over the feature pixels.

ConcatConv

The network structure is the same as Concat, except that after concatenating the text and image features and before applying the MLP, we mix the features across the feature map by applying a single convolution operation with a kernel of size 5×55\times 5 and padding of 22. This operation does not change the size of the image and text tensor. We use a the same MLP architecture as in Concat on the outputs of the convolution, and compute a distribution over pixels with a SoftMax.

Text2Conv

Given the text representation and the featurized image, we use a kernel conditioned on the text to convolve over the image. The kernel is computed by projecting the text representation into a vector of size 409,600409{,}600 using a single learned layer without biases or non-linearities. This vector is reshaped into a kernel of size 5×5×128×1285\times 5\times 128\times 128, and used to convolve over the image features, producing a tensor of the same size as the featurized image. We use a the same MLP architecture as in Concat on the outputs of this operation, and compute a distribution over pixels with a SoftMax.

LingUNet

We apply two convolutional layers to the image features to compute F1\mathbf{F}_{1} and F2\mathbf{F}_{2}. Each uses a learned kernel of size 5×55\times 5 and padding of 22. We split the text representation into two vectors of size 300300, and use two separate learned layers to transform each vector into another vector of size 16,38416{,}384 that is reshaped to 1×1×128×1281\times 1\times 128\times 128. The result of this operation on the first half of the text representation is K1\mathbf{K}_{1}, and on the second is K2\mathbf{K}_{2}. The layers do not contain biases or non-linearities. These two kernels are applied to F1\mathbf{F}_{1} and F2\mathbf{F}_{2} to compute G1\mathbf{G}_{1} and G2\mathbf{G}_{2}. Finally, we use two deconvolution operations in sequence on G1\mathbf{G}_{1} and G2\mathbf{G}_{2} to compute H1\mathbf{H}_{1} and H2\mathbf{H_{2}} using learned kernels of size 5×55\times 5 and padding of 22.

D.2 Learning

We initialize parameters by sampling uniformly from [−0.1,0.1][-0.1,0.1]. During training, we apply dropout to the word embeddings with probability 0.50.5. We compute gradient updates using Adam , and use a global learning rate of 0.00050.0005 for LingUNet, and 0.0010.001 for all other models. We use early stopping with patience with a validation set containing 7%7\% of the training data to compute accuracy at a threshold of 8080 pixels after each epoch. We begin with a patience of 44, and when the accuracy on the validation set reaches a new maximum, patience resets to 44.

D.3 Evaluation

We compare the predicted location to the gold location by computing the location of the feature pixel corresponding to the gold location in the same scaling as the predicted probability distribution. We scale the accuracy threshold appropriately.

D.4 LingUNet Architecture Clarifications

Our LingUNet implementation for SDR task differs slightly from the original implementation . We set the stride for both convolution and devonvolution operations to be 11, whereas in the original LingUNet architecture the stride is set to 22. Experiments with the original implementation show equivalent performance.

Appendix E Navigation Experimental Setup Details

We use learned word vectors of size 3232 for all models. We map the instruction xˉn\bar{x}_{n} to a vector x\mathbf{x} using a single-layer uni-directional RNN with LSTM cells with 256256 hidden units. The instruction representation x\mathbf{x} is the hidden state of the final token in the instruction.

We generate ResNet18 features for each 360∘360^{\circ} panorama It\textbf{I}_{t}. We center the feature map according agent’s heading αt\alpha_{t}. We crop a 128×100×100128\times 100\times 100 sized feature map from the center. We pre-compute mean value along the channel dimension for every feature map and save the resulting 100×100100\times 100 features. This pre-computation allows for faster learning. We use the saved features corresponding to It\textbf{I}_{t} and the agent’s heading αt\alpha_{t} as I^t\hat{\textbf{I}}_{t}.

GA

E.2 Learning

We initialize parameters by sampling uniformly from [−0.1,0.1][-0.1,0.1]. We set the horizon to 5555 during learning, and use an horizon of 5050 during testing. We stop training using SPD performance on the development set. We use early stopping with patience, beginning with a patience value of 55 and resetting to 55 every time we observe a new minimum SPD error. The global learning rate is fixed at 0.000250.00025.

Appendix F Additional Navigation Experiments

We also experiment with raw RGB images similar to Mirowski et al. . We project and resize each 360∘360^{\circ} panorama It\textbf{I}_{t} to a 60∘60^{\circ} perspective image I^t\hat{\textbf{I}}_{t} of size 3×84×843\times 84\times 84, where the center of the panorama is the agent’s heading αt\alpha_{t}. Table 7 shows the development and test results using RGB images. We observe better performance using ResNet18 features compared to RGB images.

F.2 Single-modality Experiments

We study the importance of each of the two modalities, language and vision, for the navigation task. We separately remove the embeddings of the language (NO-TEXT) and visual observations (NO-IMAGE) for both models. Table 8 shows the results. We observe no meaningful learning in the absence of the vision modality, whereas limited performance is possible without the natural language input. These results show both modalities are necessary, and also indicate that our navigation baselines (Table 5) benefit relatively little from the text.