A Recurrent Vision-and-Language BERT for Navigation

Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, Stephen Gould

Introduction

Asking a robot to navigate in complex environments following human instructions has been a long-term goal in AI research. Recently, a great variety of vision-and-language navigation (VLN) setups have been introduced for relevant studies and a large number of works explore different methods to leverage visual and language clues to assist navigation. For example, in the popular R2R navigation task , enhancing the learning of visual-textual correspondence is essential for the agent to correctly interpret the instruction and perceive the environment.

On the other hand, recent work on vision-and-language pre-training has achieved significant improvement over a wide range of visiolinguistic problems. Instead of designing complex and monolithic models for different tasks, those methods pre-train a multi-layer Transformer on a large number of image-text pairs to learn generic cross-modal representations , known as V&L BERT (Bidirectional Encoder Representations from Transformers ). Such advances have inspired us to employ V&L BERT for VLN, replacing the complicated modules for modelling cross-modal relationships and allowing the learning of navigation to adequately benefit from the pre-trained visual-textual knowledge. Unlike recent works on VLN, which apply a pre-trained V&L BERT only for encoding language or for measuring the instruction-path compatibility , we propose to use existing V&L BERT models themselves for learning to navigate.

However, an essential difference between VLN and other vision-and-language tasks is that VLN can be considered as a partially observable Markov decision process, in which future observations are dependent on the agent’s current state and action. Meanwhile, at each navigational step, the visual observation only corresponds to partial instruction, requiring the agent to keep track of the navigation progress and correctly localise the relevant sub-instruction to gain useful information for decision making. Another difficulty of applying V&L BERT for VLN is the high demand on computational power; since the navigational episode could be very long, performing self-attention on a long visual and textual sequence at each time step will cost an excessive amount of (GPU) memory during training.

To address the aforementioned problems, we propose a recurrent vision-and-language BERT for navigation, or simply VLN ↻\circlearrowright BERT. Instead of employing large-scale datasets for pre-training which usually require thousands of GPU hours, the aim of this work is to allow the learning of VLN to adequately benefit from pre-trained V&L BERT. Based on the previously proposed V&L BERT models, we implement a recurrent function in their original architecture (Fig. 1) to model and leverage the history-dependent state representations, without explicitly defining a memory buffer or applying any external recurrent modules such as an LSTM . To reduce the memory consumption, we control the self-attention to consider the language tokens as keys and values but not queries during navigation, which is similar to the cross-modality encoder in LXMERT . Such design greatly reduces the memory usage so that the entire model can be trained on a single GPU without performance degeneration. Furthermore, as in the original V&L BERT, our proposed model has the potential of multi-task learning, it is able to address other vision and language problems along with the navigation task.

We employ two datasets to evaluate the performance of our VLN ↻\circlearrowright BERT, R2R and REVERIE . The chosen datasets are different in terms of the provided visual clues, the instructions and the goal. Our agent, initialised from a pre-trained V&L BERT and fine-tuned on the two datasets, achieves state-of-the-art results. We also initialise our model with the PREVALENT , a LXMERT-like model pre-trained for VLN. On the test split of R2R , it improves the Success Rate absolutely by 8% and achieves 57% Success weighted by Path Length (SPL). For the remote referring expression task in REVERIE , our agent obtains 23.99% navigation SPL and 13.51% Remote Grounding SPL. These results indicate the strong generalisation ability of our proposed VLN ↻\circlearrowright BERT as well as the potential of using it for merging the learning of VLN with other vision and language tasks.

Related Work

Learning navigation with visual-linguistic clues has drawn significant research interests. The recent R2R and Touchdown datasets introduce human natural language as guidance and apply photo-realistic environments for navigation. Following these work, dialog-based navigation such as CVDN , VNLA and HANNA , navigation for localising a remote object such as REVERIE , VLN in continuous environment , and multilingual navigation with spatial-temporal grounding such as RxR have been proposed for further research.

One crucial challenge in VLN is to understand the visual-textual correspondence for decision making. To achieve this, Self-Monitoring and RCM adopt cross-modal attention to highlight the relevant observations and instruction at each step. Speaker-Follower and EnvDrop learn on the augmented training data via a self-supervised manner. FAST resorts self-correction navigation, while APS samples adversarial paths for training to enhance the model’s ability to generalise. AuxRN applies several auxiliary losses to learn comprehensive representations, Qi et al. and Wang et al. also design loss functions to encourage the agent to follow the instructions to take the shortest paths. More recently, Hong et al. propose a graph network to model the intra- and inter-modal relationships among the contextual and visual clues. The great improvements achieved by these methods encourage researchers to explore simpler and more powerful visiolinguistic learning network for VLN.

Visual BERT Pre-Training

Following the success of pre-trained BERT on a wide range of natural language processing tasks , the model has been extended to process visual tokens and to pre-train on large-scale image/video-text pairs for learning generic visual-linguistic representations. Previous research introduce two-stream BERT models which encode texts and images separately, and fuse the two modalities in a later stage , as well as one-stream BERT models which directly perform inter-modal grounding . Although video BERT approaches have been proposed to learn the correspondence between texts and video frames , we are the first to integrate recurrence into BERT to learn partially-observable and temporal-dependent inputs. In terms of VLN pre-training, PRESS fine tunes a pre-trained language BERT to encode instructions , PREVALENT trains a V&L BERT on a large amount of image-text-action triplets from scratch to learn navigation-oriented textual representations , and VLN-BERT fine-tunes a ViLBERT on instruction-trajectory pairs to measure their compatibility in beam search setting. Unlike all previous work, our VLN ↻\circlearrowright BERT can augment various V&L BERT models with recurrent function, it is a navigator network by itself that can be directly trained for navigation.

V&L Multi-Task Learning

Instead of building a monolithic model for different V&L tasks, numbers of previous work explore multi-task learning with a unified model for utilising the common and the complementary knowledge to reduce the domain gap . Very recently, 12-in-1 trains a single ViLBERT on 12 different datasets across four categories of V&L tasks, including visual question answering, referring expressions, multi-modal verification and caption-based image retrieval. In Vision-and-Language Navigation, Wang et al. propose a multitask navigation model to address the R2R navigation and the Navigation from Dialog History (NDH) problems. Comparing to previous methods, our VLN ↻\circlearrowright BERT uses a single network to address the navigation task and the remote referring expression task seamlessly in REVERIE .

Proposed Model

In this section, we first define the vision-and-language navigation task, then we revisit the BERT model and present the architecture of our proposed VLN ↻\circlearrowright BERT .

The problem of VLN can be formulated as follows: Given a natural language instruction U\boldsymbol{U} which contains a sequence of words, at each time step tt, the agent observes the environment and infers an action ata_{t} that transfers the agent from state st\boldsymbol{s}_{t} to a new state st+1\boldsymbol{s}_{t+1}. The state consists of the navigational history and the current spatial position defined by a triplet ⟨Ct,θt,ϕt⟩\langle\boldsymbol{C}_{t},\theta_{t},\phi_{t}\rangle, where Ct\boldsymbol{C}_{t} is a viewpoint on the pre-defined connectivity graph of the environment , and θt\theta_{t} and ϕt\phi_{t} are the angles of heading and elevation, respectively. The agent needs to execute a sequence of actions to navigate on the connectivity graph and eventually decides to stop at the target position to complete the task.

2 Revisit BERT

Formally, the kk-th attention head at the ll-th layer performs self-attention over Xl−1\boldsymbol{X}_{l-1} as

where ReLU is the Rectified Linear Unit activation function and LayerNorm is layer normalisation .

Based on this architecture, BERT has been extended to V&L BERT , which takes the concatenation of language tokens and visual tokens as input, and pre-trains on image-text corpus to learn generic visiolinguistic representations.

3 Recurrent VLN BERT

The idea of our VLN ↻\circlearrowright BERT can be adapted to a wide range of Transformer-based networks. In this section we apply the recently proposed one-stream V&L BERT model OSCAR for demonstration. We modify the model to enable the learning of navigation and the associated referring expression (REF) task. As shown in Fig. 2, at each time step, the network takes four sets of tokens as input; the previous state token st−1\boldsymbol{s}_{t-1}, the language tokens X\boldsymbol{X}, the visual tokens V ⁣t\boldsymbol{V}_{\!t} for scene, and the visual tokens O ⁣t\boldsymbol{O}_{\!t} for objects (only in REVERIE ). Then, it performs self-attention over these cross-modal tokens to capture the textual-visual correspondence for inferring the action probabilities pta\boldsymbol{p}^{a}_{t} and the object grounding probabilities pto\boldsymbol{p}^{o}_{t} (only in REVERIE ):

At initialisation (t=0t{=}0), a sequence of words consisting of the classification token [CLS], the language tokens of the instruction U\boldsymbol{U} and the separation token [SEP] will be fed into VLN ↻\circlearrowright BERT, where [CLS] and [SEP] are pre-defined in BERT models. In the pre-training of OSCAR , the [CLS] token is applied for aggregating relevant visiolinguistic clues from the input sequence for contrastive learning. Here, we defined the embedded [CLS] token as the initial state representation s0\boldsymbol{s}_{0}, to inherit such function — initialise an agent’s state which is aware of the entire navigation task.

During navigation steps (t>0t{>}0), unlike the state token st\boldsymbol{s}_{t} or the visual tokens V ⁣t\boldsymbol{V}_{\!t} and O ⁣t\boldsymbol{O}_{\!t} which performs self-attention with respect to the entire input sequence, the language tokens X\boldsymbol{X} only serve as the keys and values in the Transformer. We consider the language tokens produced by the model at the initialisation step as a deep representation of the instruction which does not need to be further encoded in later steps. Not updating the language features also save a huge amount of computational resources since the instruction and the trajectory can be long in VLN problems.

Vision Processing

At each navigation step (t>0t{>}0), the agent makes new visual observation in the environment and uses the visual clues to assist navigation. To process the visual clues, the network first projects the image features of views at the navigable directions Itv\boldsymbol{I}^{v}_{t} to the same space as the BERT token as V ⁣t=ItvWIv\boldsymbol{V}_{\!t}{=}\boldsymbol{I}^{v}_{t}\boldsymbol{W}^{I^{v}}. Then, the visual tokens will be concatenated with the state token and the language tokens, and fed into the model.

In terms of the remote REF task , we simply consider the object features Ito\boldsymbol{I}^{o}_{t} as additional visual tokens in the input sequence. Similarly, the features will be projected onto the token space as Ot=ItoWIo\boldsymbol{O}_{t}{=}\boldsymbol{I}^{o}_{t}\boldsymbol{W}^{I^{o}} and fed into the model. The object clues can provide valuable information about the important landmarks on the path, which could be very helpful to the navigation with high-level instructions .

State Representation

We formulate the agent’s state at each time step st\boldsymbol{s}_{t} as the summary of all textual and visual clues that the agent collects, as well as all decisions that the agent makes until the current viewpoint. Instead of explicitly defining a memory buffer or implementing an additional recurrent network to store the past experiences, our model relies on BERT’s original architecture to recognise time-dependent inputs, and recurrently updates s0\boldsymbol{s}_{0} from initialisation to represent the state. At each navigation step, the state representation is used as the leading input token of the entire textual-visual sequence. It then performs inter-modal self-attention in VLN ↻\circlearrowright BERT with other tokens to update its content and becomes the leading token of the input at the next step, in an autoregressive way.

State Refinement

Unlike most of the V&L BERT models which apply the output feature of the [CLS] token for classification, our state is not directly used for inferring a decision (see following Decision Making subsection), which means, the vanilla state representation is not explicitly enforced to capture the most important language and visual features. To address this issue, our model matches the raw textual and visual tokens, and feeds the output to the state representation. Formally, let Ql,ks\boldsymbol{Q}^{s}_{l,k} and Kl,kx\boldsymbol{K}^{x}_{l,k} be the state and textual tokens at head kk of the final (l=12l{=}12) layer of VLN ↻\circlearrowright BERT, the attention scores over the textual tokens can be expressed as:

Then, we average the scores over all the attention heads (K=12K{=}12) and apply a Softmax function to get the overall state-language attention weights as:

Similarly, the visual attention scores Al,ks,v\boldsymbol{A}^{s,v}_{l,k} and weights A~ls,v\widetilde{\boldsymbol{A}}^{s,v}_{l} can be obtained. Now, we perform a weighted sum over the input textual tokens and visual tokens respectively to obtain the weighted raw features as:

We then enforce a cross-modal matching between the raw textual and visual features via element-wise product and send such information to the agent’s state as:

where str\boldsymbol{s}^{r}_{t} is the output state features at the final layer. Notice that in REVERIE , only the visual features is sent to the state representation, i.e., Eq. 12 becomes stf=[str;Ftv]Wr\boldsymbol{s}^{f}_{t}=\left[\boldsymbol{s}^{r}_{t};\boldsymbol{F}^{v}_{t}\right]\boldsymbol{W}^{r}. This is because the navigational instructions in REVERIE are high-level, hence performing step-wise matching between the raw textual and visual features is less valuable.

Finally, past decisions are important for the agent to keep track of the navigation progress, our network records the new decision by feeding the directional features of the selected action at\boldsymbol{a}_{t} into the state token as:

where st\boldsymbol{s}_{t} is the new representation of the agent’s state at time step tt.

Decision Making

Many previous VLN agents apply an inner product between the state representation and the visual features at candidate directions to evaluate the state-vision correspondence, and choose a direction with the highest matching score to navigate . We find that the BERT network can nicely perform such scoring because it is fully built upon the inner product based soft-attention. Inspired by the method of predicting alignment between regions and phrases in VisualBERT , we directly apply the mean attention weights of the visual tokens over all the attention heads in the last layer, with respect to the state, as the action probabilities, simply pta=A~ls,v\boldsymbol{p}^{a}_{t}=\widetilde{\boldsymbol{A}}^{s,v}_{l} (as defined in Eq. 10). As for the remote referring expression task , our agent uses the same method to select an object. The selection probabilities can be expressed as pto=A~ls,o\boldsymbol{p}^{o}_{t}=\widetilde{\boldsymbol{A}}^{s,o}_{l}, where A~ls,o\widetilde{\boldsymbol{A}}^{s,o}_{l} is the mean attention weights for all candidate objects. We refer the Appendix §B.2 for more details.

4 Training

We train our network with a mixture of reinforcement learning (RL) and imitation learning (IL) objectives. We apply A2C for RL, in which the agent samples an action according to pta\boldsymbol{p}^{a}_{t} and measures the advantage At\boldsymbol{A}_{t} at each step (we refer the Appendix §B.5 for more details about RL.). In IL, our agent navigates on the ground-truth trajectory by following teacher actions and calculates a cross-entropy loss for each decision. Formally, we minimise the navigation loss function, expressed for each given sample, as

where atsa^{s}_{t} is the sampled action and at∗a^{*}_{t} is the teacher action. Here λ\lambda is a coefficient for weighting the IL loss. In REVERIE , we applied an additional cross-entropy term ∑tot∗log(pto)\sum_{t}o^{*}_{t}\text{log}(p^{o}_{t}) to learn object grounding.

5 Adaptation

We initialise the parameters of VLN ↻\circlearrowright BERT from OSCAR pre-trained without object tags. Although OSCAR is trained on regional features, we find that it is also compatible with the grid features of the entire scene. When adapting to the LXMERT-like model in PREVALENT , we remove the language branch in the cross-modality encoder and concatenate the state token with the visual tokens for self-attention (see Appendix §B.3 for schematics). We also remove the entire downstream network EnvDrop , including the Speaker and the environmental dropout, and then directly fine-tune the model pre-trained by PREVALENT for navigation.

Experiments

All experiments are conducted on a single NVIDIA 2080Ti GPU, the learning rate is fixed to 10−510^{-5} throughout the training and AdamW optimiser is applied. For R2R, we train the agent directly on the mixture of the original training data and the augmented data from PREVALENT , the batch sizeHalf for RL and half for IL in each iteration, corresponding to the first and the second term in Eq. 14, respectively. is set to 16 and the network is trained for 300,000 iterations. For REVERIE, we use batch size\reffootnotebs{}^{\ref{footnote_bs}} 8 and train the agent for 200,000 iterations. Images in the environments are encoded by a ResNet-152 pre-trained on Places365 , and objects are encoded by a Faster-RCNN pre-trained on the Visual Genome . Early stopping is applied when the training saturates, the model which achieves the highest SPL in validation unseen split is adopted for testing.

Evaluation Metrics

We apply the standard metrics employed by previous works to evaluate the performance.

R2R considers Trajectory Length (TL): the average path length in meters, Navigation Error (NE): the average distance between agent’s final position and the target in meters, Success Rate (SR): the ratio of stopping within 3 meters to the target, and Success weighted by the normalised inverse of the Path Length (SPL) .

REVERIE defines Success Rate (SR) as the ratio of stopping at a viewpoint where the target object is visible (in panorama), and considers the corresponding SPL. It also employ Oracle Success Rate (OSR): the ratio of having a viewpoint along the trajectory where the target object is visible, Remote Grounding Success Rate (RGS): the ratio of grounding to the correct objects when stopped, and RGSPL, which weights RGS by the trajectory length.

1 Main Results

Results in Table 1 compare the single-run (greedy search, no pre-exploration ) performance of different agents on the R2R benchmark. Our proposed VLN ↻\circlearrowright BERT initialised from OSCAR (init. OSCAR) performs better than previous methods across all the dataset splits. Comparing to a randomly initialised network (no init. OSCAR), the large performance degeneration suggests that the pre-trained general vision-linguistic knowledge significantly benefits the learning of navigation. The model initialised from PREVALENT , pre-trained especially for VLN, further improves the agent’s performance, achieving 63% SR (+8%) and 57% SPL (+5%) on the test unseen splitR2R Leaderboard: https://evalai.cloudcv.org/web/challenges/challenge-page/97/overview, REVERIE Leaderboard: https://eval.ai/web/challenges/challenge-page/606/overview . Comparing to PRESS and PREVALENT which only fine-tune a pre-trained BERT for extracting language features, adding recurrence into V&L BERT and using the model directly as the navigator network allows the VLN learning to adequately benefit from the pre-trained knowledge. Such performance gain cannot be achieved by using pre-trained V&L BERT only as a feature extractor, as will be shown in §4.2 Ablation Study. Moreover, the large gain in SR with a slight increase in TL suggests that the agent is able to navigate both accurately and efficiently. Comparing to previous methods, we can see that the performance gap between the validation unseen and the test unseen splits is greatly reduced, which means our agent has a stronger generalisation ability to novel instructions and environments.

In REVERIE (Table 2), our VLN ↻\circlearrowright BERT (init. OSCAR) generalises much better to unseen data. On the validation unseen split, the SR of navigation and object grounding has been absolutely improved by 11.13% and 6.36% respectively. On the test unseen split\reffootnoteleaderboard{}^{\ref{footnote_leaderboard}}, our method obtains 24.62% SR and 19.48% SPL for navigation, as well as 12.65% RGS and 10.00% RGSPL for REF, achieving a better performance than the previous best which applies SoTA navigator FAST for navigation and pointer MATTN for object grounding. Compare our model with OSCAR initialisation to without, the navigation results on the validation splits are similar, but the navigation on the test split and the object grounding across all the data splits are largely improved. This result also suggests that it is possible to apply a BERT-based model for VLN and REF multi-task learning. Although the previous method has higher OSR, it is likely due to longer searching (long TL), the lower SR suggests that the agent does not know where to stop correctly. In Table 2, we also present the performance of VLN ↻\circlearrowright BERT initialised from PREVALENT , which achieves the best result across almost all of the metrics in all dataset splits. It is very interesting to see that although PREVALENT is pre-trained on low-level R2R instructions without other V&L knowledge, it significantly boosts the navigation with high-level instructions as well as the object grounding in REVERIE. We hypothesis that the pre-trained knowledge provides the model with some structural priors, while the learning of REF is strongly influenced by navigation, especially at the early training stage where the target object is rarely observable by the agent.

Visualisation of Language Attention

To demonstrate that our VLN ↻\circlearrowright BERT (init. OSCAR) is able to trace the navigation progress, we visualise the changes of language attention weights at the final Transformer layer over all instructions during navigation (Fig. 3). As the agent moves forward, the attention weights with respect to state shifts from the beginning of the instructions to the end. Since the sub-instructions and sub-paths for each sample in R2R is monotonically aligned , our results indicate that the state nicely records the partial instruction that has been completed. In terms of the attention weights with respect to the visual token at the select direction, it follows a similar pattern meaning that the most relevant part of the instruction is used for guiding the action selection.

2 Ablation Study

Table 3 shows comprehensive ablation experiments on the influence of using V&L BERT (init. OSCAR) to replace or to add the key network components in the baseline model (EnvDrop ). The baseline model consists of a language encoder, a visual encoder, a state LSTM and a decision making module, corresponding to the columns of Language, Vision, State and Decision in Table 3, respectively. E.g., Model #3 replaces the language encoder and visual encoder in baseline with a V&L BERT. For fair comparison, all models in the table are trained with the same data and training strategy as our VLN ↻\circlearrowright BERT. As the results suggested, our proposed method is a multi-functional framework, the more network components it covers, the larger performance gain it achieves. Comparing the baseline with Model #1 and #2, we can see that employing a pre-trained BERT as language encoder improves the performance only if the BERT is fine-tuned for navigation. This finding is also supported by using the V&L BERT to encode both the textual and visual signals (Model #3 and #4). However, simply using pre-trained V&L BERT as text and image encoders does not fully utilise its power; model #5 indicates that relying on the original architecture of BERT to learn recurrence is feasible and it is able to achieve better results. Moreover, using the averaged visual attention weights of the final layer of the Transformer as the action probabilities (Model #6) and enhancing the state representation with visual-textual matching as defined in Eq. 12 (full model) further improves the agent’s performance.

Self-Attended Language Features

Due to the long instructions and episodes, high memory cost during training is one of the key issues that prevents previous research to apply BERT for self-attention at every time step. To demonstrate the influence of self-attending textual features during navigation, we compare the agent’s performance and the training time GPU memory consumption (constrained to a single 11GB memory GPU) of re-attending the language at each step. As shown in Table 4, training for Emb-Attn, Init-Attn and Re-Attn consume much more memory for each sample than performing language self-attention only at initialisation (Ours), and their performances are worse than Ours. The results of Re-Attn degenerates significantly because at each time step the output language features aggregate the most relevant visual-textual clues at a certain viewpoint, which suppresses the valuable information in other part of the instruction for the future steps. We refer Appendix §C.2 for experiment on language self-attention with larger batch size by applying gradient accumulation.

Learning Curves

As shown in Fig. 4, we compare the learning curves of VLN ↻\circlearrowright BERT initialised from different models. The training losses of our method initialised from pre-trained OSCAR converges faster than a randomly initialised model, and it reaches much higher SPL in both validation seen and unseen environments. Moreover, our model initialised from PREVALENT learns significantly faster than the other two methods and it is able to achieve a much better performance within much fewer iterations. In terms of training in real time, the model init. OSCAR takes about 7 daysTime includes evaluation on the validation splits every 2,000 iterations. Matterport 3D Simulator v0.1 is applied, whereas the latest version supports batches of agents so it is much more efficient. to complete 600,000 iterations of training (best result achieved in 3.5 days), while the model init. PREVALENT takes about 4.5 days\reffootnotetime{}^{\ref{footnote_time}} (best result achieved in 1 day). Using wall-clock time in Fig. 4 as the x-axis will enhance the discrepancy between the three models. These results suggest that the pre-trained generic visiolinguistic knowledge is beneficial to the learning of VLN, and pre-training especially for navigation skills allows the agent to learn better in fine-tuning.

Conclusion

In this paper, we introduce recurrence into Vision-and-Language BERT and rely on its original architecture to recognise time-dependent inputs. Such innovation allows V&L BERT to address problems with a partially observable Markov decision process, and allows the learning of downstream tasks to adequately benefit from the pre-trained generic V&L knowledge. For VLN, our proposed VLN ↻\circlearrowright BERT applies BERT itself as the navigator network, which achieves SoTA performance in R2R and REVERIE . Moreover, results suggest that V&L BERT with recurrence is capable of VLN and REF multi-tasking.

As suggested by the significant improvement on R2R achieved by our VLN ↻\circlearrowright BERT, we expect that the model can improve the performance of other navigation settings such as street navigation and navigation in continuous environments . In this paper, we only apply our recurrent BERT for VLN. However, we believe that it has a huge potential in addressing other tasks which require sequential interactions/decisions, such as language and visual dialog , dialog navigation and action anticipation for reactive robot response .

References

Appendices

Appendix A Datasets

We evaluate the performance of our proposed model on two distinct datasets for VLN:

Room-to-Room (R2R) : The agent is required to navigate in photo-realistic environments (Matterport3D ) to reach a target following low-level natural language instructions. Most of the previous works apply a panoramic action space for navigation , where the agent jumps among viewpoints pre-defined on the connectivity graph of the environment. The dataset contains 61 scenes for training; 11 and 18 scenes for validation and testing, respectively, in unseen environments.

REVERIE : The agent needs to first navigate to a point where the target object is visible, then, it needs to identify the target object from a list of given candidates. In REVERIE, the navigational instructions are high-level while the instructions for object grounding are very specific. The dataset has in total 4,140 target objects in 489 categories, and each target viewpoint has 7 objects with 50 bounding boxes in average.

Appendix B Implementation Details

We provide the implementation details of preparing visual features (§3.3Link to Section 3.3 in Main Paper. ), decision making (§3.3), adaptation to PREVALENT (§3.3 & 3.5), adaptation to REVERIE (§3.3 & 3.5), critic function and reward shaping in reinforcement learning (§3.4).

Navigation in R2R and REVERIE are conducted in the Matterport3D Simulator . At each navigable viewpoint in the environment, the agent observes a 360∘360^{\circ} panorama, consisting of 36 single-view images at 12 headings (30∘30^{\circ} separation) and 3 elevation angles (±30∘\pm 30^{\circ}).

The directional encoding is formed by replicating vector (cosθti,sinθti,cosϕti,sinϕti)(\text{cos}\theta^{i}_{t},\text{sin}\theta^{i}_{t},\text{cos}\phi^{i}_{t},\text{sin}\phi^{i}_{t}) by 32 times, where θti\theta^{i}_{t} and ϕti\phi^{i}_{t} represents the heading and elevation angles of the image with respect to the agent’s orientation . Moreover, the action features at\boldsymbol{a}_{t} which is fed to the state representation is exactly the directional encoding at the selected direction dtc\boldsymbol{d}^{c}_{t}, where c=atsc=a^{s}_{t}.

Object Features

In REVERIE , the target objects can appear at any single-view of the panorama, we extract the object features (regional features encoded by Faster-RCNN pre-trained on Visual Genome by Anderson et al. ) according to the positions of objects provided in REVERIE . The object features are position- and direction-aware, formulated as

B.2 Decision Making (§bold-§\boldsymbol{\S}3.3)

In R2R , there are two types of decisions that an agent infers during navigation; it either selects a navigable direction to move, or it decides to stop at the current viewpoint. As in most of the previous work, stopping in VLN ↻\circlearrowright BERT is implemented by adding a zero vector vtstop\boldsymbol{v}^{\text{stop}}_{t} to the list of visual features at navigable directions , as

Our VLN ↻\circlearrowright BERT determines to stop by predicting the largest attention score to the stop representation at the final transformer layer. Otherwise, the agent will move to a navigable direction with the largest score.

However, in REVERIE , we directly apply the attention scores over the candidate objects for stopping. To be specific, the visual tokens in REVERIE consists of sequences of scene features and object features:

When the model predicts larger attention scores for at least one of the object token than all of the scene tokens, the agent will stop and will select the object with the largest score as the grounded object for REF. Such formulation has two advantages; First, it relates the object searching process to navigation, \ie, the agent should not stop navigation if it has low confidence of localising the target object. Second, it allows the reinforcement learning to benefit the object grounding, since the action logits for stop is the greatest attention scores over objects.

B.3 Adaptation to PREVALENT (§bold-§\boldsymbol{\S}3.3 & §bold-§\boldsymbol{\S}3.5)

As shown in Fig. 5, we adapt our VLN ↻\circlearrowright BERT to the LXMERT-like architecture in PREVALENT . At initialisation, the transformer TRM-Lang1 encodes the instruction U\boldsymbol{U} and uses the output features of the [CLS] token to represent agent’s initial state. During navigation, the concatenated sequence of the previous state st−1\boldsymbol{s}_{t-1}, the encoded language from initialisation X\boldsymbol{X} and the new visual observation V ⁣t\boldsymbol{V}_{\!t} will be fed to the cross-modality encoder to obtain the language-aware state feature st−1L\boldsymbol{s}^{L}_{t-1} and the language-aware visual features V ⁣tL\boldsymbol{V}^{L}_{\!t}. Finally, TRM-Vis2 will process st−1L\boldsymbol{s}^{L}_{t-1} and V ⁣tL\boldsymbol{V}^{L}_{\!t} to produce a new state st\boldsymbol{s}_{t} and a decision pt\boldsymbol{p}_{t}.

Pre-training of PREVALENT applies the outputs from TRM-Lang3 for attended masked language modelling and action prediction, but fine-tuning on R2R only use the output language features from TRM-Lang1 and rely on a downstream network EnvDrop for navigation. In contrast, our method does not require any downstream network, we leverage the visual transformers TRM-Vis1 and TRM-Vis2 to learn state-language-vision relationship for better decision marking.

The cross-modal matching for State Refinement (see Eq. 12 in §3.3) is applied to enhance the state representation of PREVALENT-based VLN ↻\circlearrowright BERT. Since the final transformer TRM-Vis2 only process the state and visual features, we apply the averaged attention scores over the visual features with respect to st\boldsymbol{s}_{t} to weight the input visual tokens Vt\boldsymbol{V}_{t}, but apply the averaged attention scores over the textual features with respect to st−1L\boldsymbol{s}^{L}_{t-1} to weight the input language tokens X\boldsymbol{X}. Results in Table 5 show that cross-modal matching for state improves the agent’s performance in unseen environments.

B.4 Critic Function (§bold-§\boldsymbol{\S}3.4)

We apply the A2C for reinforcement learning. At each time step, the critic, a multi-layer perceptron, predicts an expected value from the updated state representation as:

where ReLU is the Rectified Linear Unit function, WE1\boldsymbol{W}^{E1} and WE2\boldsymbol{W}^{E2} are learnable linear projections.

B.5 Reward Shaping (§bold-§\boldsymbol{\S}3.4)

In addition to the progress rewards defined in EnvDrop , we apply the normalised dynamic time warping as a part of the reward to encourage the agent to follow the instruction to navigate. Moreover, we introduce a negative reward to penalise the agent if it misses the target.

As formulated in the EnvDrop , we apply the progress rewards as strong supervision signals for directing the agent to approach the target. To be specific, let DtD_{t} to be the distance from agent to target at time step tt, and ΔDt=Dt−Dt−1{\Delta}D_{t}=D_{t}-D_{t-1} to be the change of distance by action ata_{t}, the reward at each step (at≠stopa_{t}{\neq}\texttt{stop}) is defined as:

When the agent decides to stop (at=stopa_{t}{=}\texttt{stop}), a final reward is assigned depending on if the agent successfully complete the task:

Overall, the agent will receive a positive reward if it approaches the target and completes the task (by stopping within 3 meters to the target viewpoint), while it will be penalised with a negative reward if it moves away from the target or stops at a wrong viewpoint.

Path Fidelity Rewards

The Progress Reward encourages the agent to approach the target, but it does not constrain the agent to take the shortest path. As there could be multiple routes to the target, the agent could learn to take a longer path or even a cyclic path to maximise the total reward. To address the problem, we apply the normalised dynamic time warping reward , a measurement of the similarity between the ground-truth path and the predicted path, to urge the agent to follow the instruction accurately. Let PtP_{t} be the normalised dynamic time warping at time tt, and ΔPt=Pt−Pt−1{\Delta}P_{t}=P_{t}-P_{t-1} to be the change of PP caused by action ata_{t}, the reward for at≠stopa_{t}{\neq}\texttt{stop} is defined as:

and the reward for at=stopa_{t}{=}\texttt{stop}:

Moreover, as suggested by many previous works, there exists a large discrepancy between the agent’s oracle success rate and success rate, indicating that it does not learn well to stop accurately. To address this issue, we introduce a negative stopping reward rtSr^{S}_{t} which will be triggered whenever the agent first approaches the target but then departs from it. To be precise, if Dt−1≤1.0D_{t-1}{\leq}1.0 and ΔDt>0.0{\Delta}D_{t}>0.0:

The normalised dynamic time warping reward rtPr^{P}_{t} and the stopping reward rtSr^{S}_{t} together form the path fidelity rewards which can encourage the agent to navigate efficiently. In summary, the overall reward at each step during navigation can be expressed as:

Ablation Study

We perform ablation experiments on training the OSCAR-based and the PREVALENT-based VLN ↻\circlearrowright BERT with and without the path fidelity rewards. As shown in Table 6, the models trained with the path fidelity rewards achieve higher Success Rate (SR) and lower Trajectory Length (TL), leading to higher Success weighted by Path Length (SPL). Despite the improvements in SR, the gap between the Oracle Success Rate (OSR) and SR is kept roughly the same for the PREVALENT-based model and is reduced by about 1.45% the OSCAR-based models. These results suggest that the path-fidelity rewards benefit the agent to navigate more accurate and efficient. Note that, comparing to the results that are shown in Table 1 and Table 3 of our Main Paper, such reward shaping only contributes to a slight gain of the improvement, whereas the structure of the recurrent BERT is much more influential.

Appendix C Language Attention

In Figure 3 of the Main Paper, we visualised the averaged attention weights over all instructions in validation unseen split during navigation. Notice that in this visualisation, we interpolate the instruction to 80 words and the trajectory to 15 steps for each sample to compute the averaged attention weights. Since a large portion of instructions in R2R are less than 80 words, the last few tokens in the plot are corresponding to the full stop punctuation and the [SEP] token, which are less likely to provide valuable information to the state and hence receive very little weight. However, for the last step in the bottom plot, the selected visual token for stopping could have a high correspondence to the full stop and [SEP] which are strong indications of the end of navigation.

C.2 Self-Attended Language Features with Gradient Accumulation (§bold-§\boldsymbol{\S}4.2)

To evaluate the influence of performing different language self-attention under the same batch size, we train Emb-Attn, Init-Attn and Re-Attn with gradient accumulation to achieve batch size of 16 (same as Ours) while keeping learning rate unchanged. Comparing to the results reported in Table 4 of the Main Paper, the performance of all the three methods has been significantly improved. On the validation unseen split, Emb-Attn even outperforms Ours which does not re-attend the language at each time step. However, a downside for accumulating gradient is that the training speed is significantly reduced, which is about 2.5 times slower for accumulating 4 batches of size 4. Nevertheless, this result suggests that agent’s performance can be further improved by training with larger batch size, potentially including Ours. Therefore, whether it is necessary to perform language self-attention at each navigational step when there is sufficient computational power remains a question worth investigating. We will leave it as a future work.

Appendix D Visualisation (§bold-§\boldsymbol{\S}4.1)

As shown in Figure 6, 7 and 8, we visualise the language-to-language and the state/vision-to-state/vision/language attention weights of a sample in the validation unseen split. As shown by the panoramas in Fig. 7 and the given instruction “Exit the bedroom. Walk the opposite way of the picture hanging on the wall through the kitchen. Turn right at the long white countertop. Stop when you get past the two chairs.”, the agent needs to understand the complex contextual clues, including different scene clues (bedroom, kitchen), different object clues (picture, wall, countertop, chair) and various directional clues (forward, opposite, left/right, stop) to complete the task.

Fig. 6 shows the language self-attention attention weights at some selected heads at initialisation (t=0t{=}0); different heads demonstrate different functions and the attentions at different layers behave very differently. Plot (1) and (4) show the general pattern of attentions at shallow (Layer 0) and deep layers (Layer 8); words in shallow layers tend to collect information from the entire sentence, while words at deep layers have higher correspondence to the adjacent words since they are semantically more relevant. It is very interesting to see that the attention head in Plot (2) learns to attend the adjectives and the action-related terms, which describe the objects and scenes, for the initialised state representation ([CLS]). This head also learns about the co-occurrence between different entities, for example, picture has higher correspondence to wall, and countertop has higher correspondence to kitchen. In contrast, the attention head in Plot (3) learns to extract the important landmarks such as bedroom, picture, kitchen and chairs. Attention head in Plot (5) learns about the [SEP] token, which indicates the ending of the instruction. Attention weights in Plot (6) show the most frequent pattern of the attention heads at the final (l=12l{=}12-th) layer, those heads seem to aggregate information from the punctuations in the instruction. This implies the heads could have learnt about breaking the sentence into multiple sub-sentences. Refer to the idea of sub-instruction proposed by Hong et al. , such attention pattern could be beneficial for matching the current observation to a particular and the most relevant part of the instruction.

State/Vision Step-wise Attention

Fig. 7 shows the trajectory of the agent starting from a bedroom, taking a series of actions and eventually stopping at the target location. It also displays the averaged attention weights at the final (l=12l{=}12-th) layer for the state and visual tokens. Note that the averaged attention for state/vision tokens is representative since we apply the averaged attention weights for visual-textual matching to state and use it as the action probabilities (see Eq. 12 and Decision Making in §3.3, respectively). As shown in Plot (1a-6a), the attention shifts from the beginning of the instruction to the end, which agrees with the agent’s navigation progress. It is also interesting to see that at the final layer, the state token is more influential to the predicted action at the first two steps, while the language tokens are more influential at the later steps. The state/vision self-attention in Plot (1b-6b) reflects the action prediction at each time step. The first row in each plot (attention of state with respect to candidate actions) shows the prediction result, we can see that the agent is very confident in each decisions. Moreover, starting from Step 3, the information from state and from different views are aggregated to support choosing the correct direction.

State/Vision Layer-wise Attention

To better understand how the visual and language features are aggregated to support action prediction, we visualise the layer-wise attention at Step 4 of the trajectory (t=4t{=}4). As shown by Fig. 8 Plot (1-12), at the first two layers, candidate views collect information from the entire instruction. But as the signals propagate to deeper layers (Layer 3-6), the visual tokens attend more the middle part of the instruction, which should be more relevant to the current observations. Interestingly, starting from Layer 6, the visual features tend to dominate the attention, and information aggregates toward the visual token at the predicted direction (which is the correct direction). We can see that, at Layer 7-9, the network still has some doubts about the candidate directions that are spatially closer to the correct direction. But after implicitly reasoning in deeper layers, the network becomes very confident about choosing the correct direction.