History Aware Multimodal Transformer for Vision-and-Language Navigation

Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan Laptev

Introduction

Vision-and-language navigation (VLN) has recently received growing attention . VLN requires an agent to understand natural language instructions, perceive the visual world, and perform navigation actions to arrive at a target location. A number of datasets have been proposed to support various VLN tasks such as indoor and outdoor navigation with fine-grained instructions , language-driven remote object finding and navigation in dialogs .

VLN agents are faced with several challenges. First, as opposed to static vision-text grounding , the agent continuously receives new visual observations and should align them with instructions. Most of existing works adopt recurrent neural networks (RNNs) to encode historical observations and actions within a fixed-size state vector to predict the next action. Such condensed states might be sub-optimal for capturing essential information in extended trajectories . For instance, “bring the spoon to me” requires the agent to remember its start location after navigating to the “spoon”, while early memories are prone to fade in the recurrent state. Few endeavors construct external map-like memories for received observations. Nevertheless, these approaches still rely on RNNs to track the navigation state. As the history plays an important role in environment understanding and instruction grounding, we propose to explicitly encode the history as a sequence of previous actions and observations instead of using recurrent states.

Another VLN challenge concerns the generalizations of agents to new environments that have not been observed during training . One direction is to learn more generic text-image representations. The PRESS model improves language representation with a pretrained BERT encoder , and PREVALENT uses pairs of instruction and single-step observations to pretrain a multimodal transformer. Though achieved promising results, these works do not optimize visual representation for the target navigation task. Moreover, lack of history in training makes it hard to learn cross-modal alignment and increases the risk of overfitting to training environments. Another direction towards better generalization is to overcome exposure bias due to discrepancy between training and inference. Different methods have been adopted for VLN including DAgger and scheduled sampling . Reinforcement Learning (RL) is one of the most effective approach among them, but it is considered unstable to directly train large-scale transformers via RL .

To address the above challenges, we propose the History Aware Multimodal Transformer (HAMT), a fully transformer-based architecture for multimodal decision making in VLN tasks. As illustrated in Figure 1, HAMT consists of unimodal transformers for text, history and observation encoding, and a cross-modal transformer to capture long-range dependencies of the history sequence, current observation and instruction. Since our history contains a sequence of all previous observations, its encoding is computationally expensive. To resolve complexity issues, we propose a hierarchical vision transformer as shown in Figure 2, which progressively learns representations for a single view, spatial relationships among views within a panorama and, finally, the temporal dynamics across panoramas of the history. In order to learn better visual representations, we propose auxiliary proxy tasks for end-to-end training. Such tasks include single-step action prediction based on imitation learning, self-supervised spatial relationship reasoning, masked language and image predictions and instruction-trajectory matching. We empirically show that our training facilitates the subsequent fine-tuning of our model with RL . We carry out extensive experiments on various VLN tasks, including VLN with fine-grained instructions (R2R and RxR ), high-level instructions (REVERIE and our proposed R2R-Last), dialogs as well as long-horizon VLN (R4R and our proposed R2R-Back which requires the agent to return back after arriving at the target location). HAMT outperforms state of the art on both seen and unseen environments in all the tasks.

We summarize our contributions as follows: (1) We introduce HAMT to efficiently model long-horizon history of observed panoramas and actions via hierarchical vision transformer; (2) We train HAMT with auxiliary proxy tasks in an end-to-end fashion and use RL to improve the navigation policy; (3) We validate our method and outperform state of the art in a diverse range of VLN tasks, while demonstrating larger gains for long-horizon navigation.

Related work

Vision-and-language navigation. Training instruction-following navigation agents has attracted increasing research attention . Anderson et al. propose a sequence-to-sequence LSTM baseline for the VLN task. Fried et al. extend it with panoramic action space and synthesized instructions. To improve cross-modal alignment, the self-monitoring agent proposes co-grounding and progress estimation, and RelGraph uses graphs to model relationships across scene, objects and directions. Reinforcement learning (RL) is typically used to improve navigation policy. The EnvDrop model mixes imitation learning and A3C . The RCM utilizes intrinsic reward of cross-modal matching in REINFORCE algorithm. Wang et al. propose to learn rewards via soft expert distillation. Due to the success of transformer , recent works explore transformer architectures in VLN. PRESS replaces LSTM instruction encoder with pretrained BERT . SIA uses transformer for single-step multimodal fusion and LSTM for sequential action prediction. PTA is a transformer VLN model using CNNs to extract visual features . Here we propose the first full transformer architecture for VLN and train it end-to-end.

Memory-based policy for navigation. LSTMs have been the dominant approach to encode memories for navigation . Condensing all history into one feature vector, however, is prone to the loss of information. Alternative approaches include topological map memory structures . Deng et al. use graphs to capture environment layout and enable long-term planing. A similar graph is adopted in with frontier-exploration based decision making. But these works still utilize LSTMs for state tracking. To exploit long-term spatio-temporal dependencies, Fang et al. store histories in a sequence encoded with transformer. Recurrent VLN-BERT injects a recurrent unit to encode histories in transformer for VLN. The most similar work to ours is Episodic Transformer (E.T.) . Differently from , we propose a hierarchical encoding of the panoramic observation history and optimize the whole model in end-to-end training.

Multimodal pretraining with transformers. Recent works show significant progress in vision and language tasks using multimodal pretraining. In particular, transformer architectures such as one-stream and dual-stream achieve state of the art for a number of downstream tasks including visual question answering, image-text retrieval and image captioning. While most previous methods rely on CNN to extract image representations, ViLT adopts Vision Transformer (ViT) and trains it with associated texts in an end-to-end manner thanks to the efficiency of ViT. A few endeavors explore multimodal pretraining for VLN. PREVALENT pretrains a transformer using instructions and single-step observations without referring to trajectory history. VLN-BERT measures the compatibility between an instruction and images in a path but does not support action prediction. Our work presents the first end-to-end trainable VLN transformer that jointly encodes text, history and observation, and is able to sequentially predict actions.

Method

Unlike dominant recurrent approaches to condense Ht\mathcal{H}_{t} into a fixed-size vector, in this section, we present the History Aware Multimodal Transformer (HAMT) that jointly encodes text, long-horizon history, and observation for sequential action prediction. The model architecture is described in Section 3.1. We propose end-to-end training for HAMT in Section 3.2 to learn unimodal and multimodal representations, and then use RL to fine-tune the navigation policy in Section 3.3.

1 HAMT: History Aware Multimodal Transformer

Figure 1 illustrates the model architecture of HAMT. The inputs text W\mathcal{W}, history Ht\mathcal{H}_{t} and observation Ot\mathcal{O}_{t} are first encoded via the corresponding unimodal transformers respectively, and then fed into the cross-modal transformer encoder to capture multimodal relationships.

Text Encoding. For each token ii in the instruction W\mathcal{W}, we embed it as the summation of its word embedding wiw_{i}, position embedding EiPE^{P}_{i} and type embedding of text E0TE^{T}_{0}. Then we employ a transformer with NLN_{L} layers to obtain contextual representation xix_{i} following the standard BERT .

Observation Encoding. For each view [vio;aio][v^{o}_{i};a^{o}_{i}] in the panoramic observation Ot\mathcal{O}_{t}, we first represent the relative angle aioa^{o}_{i} as EaioA=(sin⁡θi,cos⁡θi,sin⁡ϕi,cos⁡ϕi)E^{A}_{a^{o}_{i}}=(\sin\theta_{i},\cos\theta_{i},\sin\phi_{i},\cos\phi_{i}) where θi\theta_{i} and ϕi\phi_{i} are the relative heading and elevation angle to the agent’s orientation. Then the observation embedding oio_{i} is as follows:

Hierarchical History Encoding. As Ht\mathcal{H}_{t} consists of all the past panoramic observations Oi\mathcal{O}_{i} and performed actions aiha^{h}_{i} before step tt, it is important to encode Ht\mathcal{H}_{t} efficiently as context. Figures 2(c)-2(c) depict the flattened and temporal-only history encoding approaches used in VLN-BERT and E.T. respectively. The flattened approach treats each view image in Oi\mathcal{O}_{i} as a token, so the history sequence contains tKtK tokens. Though it enables to learn relationships among all image views, the computation cost quadratically increases with the sequence length, making it inefficient for long-horizon tasks. In the temporal-only approach, only the oriented view of the agent in each Oi\mathcal{O}_{i} is taken as inputs instead of the whole panorama, so only tt temporal tokens are encoded. However, this approach can lose critical information in past observations. For example, in the instruction “with the windows on your left, walk through the large room past the sitting areas”, the object “window” does not appear in the oriented view of the agent. Therefore, the encoded history is insufficient to tell whether the agent passed the window or not, making the model confused to take the next action.

In order to balance computational efficiency and information integrity, we propose a hierarchical history encoding approach as illustrated in Figure 2(a). It hierarchically encodes view images within each panorama and then temporal relationships across panoramas, similar to the factorized spatial-temporal video transformer . For each Oi\mathcal{O}_{i}, its constituent view images are first embeded via ViT and Eq (1), and then encoded via a panoramic transformer with NhN_{h} layers to learn spatial relationships within the panorama. We apply average pooling to obtain panorama embedding, and add it with the oriented view image feature in residual connection. The parameters in ViT and panoramic transformer are shared for different steps. In this way, each historical observation Oi\mathcal{O}_{i} is represented as vihv^{h}_{i}, and the final temporal token hih_{i} is computed as:

where EiSE^{S}_{i} denotes the ii-th step embedding, E2TE^{T}_{2} is the type embedding of history. The computational cost is O(tK2+t2)O(tK^{2}+t^{2}), which significantly reduces from O(t2K2)O(t^{2}K^{2}) in the flattened approach. To be noted, we add a special token [cls] to the start of the history sequence to obtain a global representation. The embedding of [cls] is a parameter to learn, which is initialized from a zero vector.

2 End-to-end training with proxy tasks

As it is difficult to train large-scale transformers with RL due to sparse supervision , we propose to first end-to-end train HAMT via several proxy tasks to learn unimodal and multimodal representation.

Table 1 compares our HAMT with previous VLN transformers PREVALENT and VLN-BERT in inputs and proxy tasks. As neither PREVALENT nor VLN-BERT jointly encodes text, history and observation, a limited choice of proxy tasks can be applied in training. Our model instead can take advantage of various proxy tasks to learn cross-modal alignment, spatial and temporal reasoning, and history-aware action prediction. Given the input pair (W,HT)(\mathcal{W},\mathcal{H}_{T}) where TT is the length of full trajectory, we can apply common proxy tasks as in vision-and-language pretraining , including Masked Language Modeling (MLM), Masked Region Modeling (MRM) and Instruction Trajectory Matching (ITM). Details of the three proxy tasks are presented in the supplementary material. In the following, we introduce new proxy tasks given the triplet input (W,Ht,Ot)(\mathcal{W},\mathcal{H}_{t},\mathcal{O}_{t}) specifically for VLN tasks.

Training Strategy. Instead of directly training the whole HAMT model at once, we propose to progressively train HAMT in two stages. In the first stage, we freeze ViT pretrained on ImageNet and train the rest of the modules which are randomly initialized. This aims to avoid catastrophic forgetting of the pretrained weights in ViT. Then we unfreeze ViT and train the whole model end-to-end. The learning rate for ViT is set to be higher than for others modules to avoid vanishing gradients and to speedup convergence. We empirically show that the proposed two-stage training outperforms one-stage training in the supplementary material.

3 Fine-tuning for sequential action prediction

Structure Variants. We present two variants of HAMT for action prediction in the following. 1) MLP action head: we directly reuse the action prediction network fSAPf_{\text{SAP}} in the SAP task to predict navigable views. We use it as default for VLN tasks. 2) MLP action head based on encoder-decoder structure: the original HAMT model applies cross-modal attention for both vision-to-text and text-to-vision, which is computationally expensive when instructions are long. Therefore, we remove the cross-modal attention from text to vision. In this way, we separate the cross-modal transformer into an encoder which only takes instruction as input, and a decoder that inputs history and observation as query and attends over encoded text tokens. Please see supplementary material for details.

where μ\mu is the learning rate, at∗a^{*}_{t} is the expert action at step tt of the expert trajectory of length T∗T^{*}.

Experiments

Datasets. We evaluate our method on four VLN tasks (seven datasets): VLN with fine-grained instructions (R2R , RxR ); VLN with high-level instructions (REVERIE , R2R-Last); vision-and-dialogue navigation (CVDN ); and long-horizon VLN (R4R , R2R-Back).

R2R builds upon Matterport3D and includes 90 photo-realistic houses with 10,567 panoramas. It contains 7,189 shortest-path trajectories, each associated with 3 instructions. The dataset is split into train, val seen, val unseen and test unseen sets with 61, 56, 11 and 18 houses respectively. Houses in val seen split are the same as training, while houses in val unseen and test splits are different from training.

RxR is a large multilingual VLN dataset based on Matterport 3D. The instructions are in three different languages (English, Hindi and Telugu). The dataset emphasizes the role of language in VLN by addressing biases in paths and describing more visible entities than R2R.

R4R extends R2R dataset by concatenating two adjacent tail-to-head trajectories in R2R. Therefore, it has longer instructions and trajectories. The trajectories are also less biased as they are not necessarily the shortest-path from start to end location.

R2R-Back is a new VLN setup proposed in this work. The agent is required to return to its start location after arriving at the destination. The agent needs to remember its navigation histories to solve the task. We add a return command at the end of each instruction in R2R and a reverse path from the end to start locations as expert demonstration.

CVDN defines a navigation from dialog history task, which requires an agent to arrive at goal regions based on multi-turn question-answering dialogs. Such types of instructions are often ambiguous and under-specified. The lengths of instructions and paths are also long.

REVERIE replaces step-by-step instructions in R2R with high-level instructions, which mainly describe the target location and object. The agent, hence, is required to navigate to the goal without detailed guidance and depends on its past experiences.

R2R-Last is our proposed VLN setup similar to REVERIE. It only uses the last sentence from the original R2R instructions describing the final destination.

Evaluation metrics. We adopt standard metrics , including (1) Trajectory Length (TL): the agent’s navigated path in meters; (2) Navigation Error (NE): the average distance in meters between the agent’s final position and the target; (3) Success Rate (SR): the ratio of trajectories reaching the destination with a maximum error of 3 meters to the target; and (4) Success Rate normalized by the ratio between the length of the shortest path and the predicted path (SPL). SPL is more relevant than SR as it balances the navigation accuracy and efficiency. For long-horizon VLN task (R4R and R2R-Back), we further employ three metrics to measure the path fidelity between the predicted path and target path, including (5) Coverage weighted by Length Score (CLS) ; (6) the normalized Dynamic Time Warping (nDTW) ; and (7) the Success weighted by nDTW (SDTW).

Implementation details. For the HAMT model, we set NL=9N_{L}=9 for language transformer, Nh=2N_{h}=2 for panoramic transformer in hierarchical history encoding, and Nx=4N_{x}=4 for cross-modal transformer. There are K=36K=36 view images in each panoramic observation. We use ViT-B/16 for image encoding if not otherwise specified. In training with proxy tasks, we randomly select proxy tasks for each mini-batch with predefined ratio. We train HAMT for 200k iterations with fixed ViT using learning rate of 5e-5 and batch size of 64 on 4 NVIDIA Tesla P100 GPUs (∼\sim1 day). The whole HAMT model is trained end-to-end for 20k iterations on 20 NVIDIA V100 GPUs with learning rate of 5e-5 for ViT and 1e-5 for the others (∼\sim20 hours). We use R2R training set and augmented pairs from for training unless otherwise noted. In fine-tuning with RL+IL, we set λ=0.2\lambda=0.2 in Eq (3) and γ=0.9\gamma=0.9. The model is fine-tuned for 100k iterations with learning rate of 1e-5 and batch size of 8 on a single GPU. Unimodal encoders are fixed by default. The best model is selected according to performance on val unseen split. We use the same augmented data as for R2R for fair comparison, while no augmented data is used for other datasets. Greedy search is applied in inference following the single-run setting. Please see supplementary material for more details.

2 Ablation studies

In this section, we evaluate each component in the HAMT model, including: hierarchical history encoding, end-to-end training with proxy tasks, and fine-tuning objectives.

For fair comparison with the state-of-the-art recurrent architecture RecBERT , we use the same Resnet152 visual features and train all the models from scratch with RL+IL objectives to avoid the influence of different weight initialization. The models are optimized for 300k iterations end-to-end except for the visual feature. Table 2 compares different history encoding approaches on R2R dataset. Our recurrent model slightly differs from RecBERT (no init. OSCAR) in transformer architecture as shown in Figure 1. It achieves slightly better performance on val unseen split. The temporal-only model uses transformer to encode agent’s oriented visual observations in history sequence, and outperforms the recurrent method by relative gains of 1.9% on SR and 2.1% on SPL for val unseen split. Adding panoramic observations in a hierarchical way results in 4.2% (SR) and 3.6% (SLP) relative improvements on the val unseen split compared to the recurrent method. Even larger improvements are achieved on val seen split as the hierarchical model has a larger capacity to fit the seen environments. This evaluation demonstrates the advantage of our hierarchical history representation compared to the recurrent and temporal-only history representation.

How much does training with proxy tasks help?

We next evaluate the advantage of training HAMT end-to-end with proxy tasks. In Table 3(a), the first row uses RL+IL objectives to train HAMT from scratch, while the second row uses proxy tasks for training prior to RL+IL fine-tuning. We can see that it significantly boosts the performance to first train with proxy tasks. It improves on val unseen split with 16.7% and 18.0% relative gains on SR and SPL respectively, indicating that training with auxiliary proxy tasks enables better generalization. In the third row, we replace the visual feature from Resnet152 to ViT. The ViT feature improves the performance on both val seen and val unseen splits, showing that more powerful visual representations matter. Finally, training ViT end-to-end obtains 2.1% gains on SPL on val unseen split. This is the first time to show that optimizing visual representations end-to-end is beneficial for VLN tasks. In Table 3(b), we evaluate the benefit of the two new proxy tasks for frozen ViT features using the other proxy tasks by default. The SAP(R) uses imitation learning to predict actions, which directly influences the navigation policy and improves the performance by a large margin. The SPREL is a self-supervised proxy task that forces the model to learn spatial relationships in panorama and helps generalization in unseen environments. More experiments to ablate contributions from history encoding and proxy tasks, contributions of proxy tasks in end-to-end training etc. are presented in supplementary material.

What is the impact of the fine-tuning objectives?

Table 4 presents results using different objectives in fine-tuning. The first row directly applies HAMT trained by proxy tasks, which achieves lower performance than that after IL fine-tuning, because we mainly use augmented data in proxy task training to increase visual diversity, but such noisy data deteriorates action prediction performance. Previous work has shown that RL alone performs poorly. However, training with proxy tasks stabilizes the followup RL fine-tuning. HAMT optimized by RL achieves much better performance than that when fine-tuning with IL on the SR metric. It indicates that RL is able to learn better exploration strategy on unseen environments. However, as the reward for RL focuses more on shortest paths rather than path fidelity with instructions, the improvement on SPL metric is relatively small compared to SR metric. Moreover, the fluctuation of the pure RL objective is larger than IL. Therefore, mixing the RL and IL achieves the best performance.

3 Comparison to state of the art

Table 5 compares HAMT with previous VLN methods on the R2R benchmark. Our model outperforms state-of-the-art results of RecBERT by relative 5.9% and 7.0% improvements in SPL on val seen and unseen splits respectively. We achieve state-of-the-art performance under the single-run setting on the unseen testing split of the leaderboardWe report the published results on the testing unseen split as shown in https://eval.ai/web/challenges/challenge-page/97/leaderboard/270 (25/10/2021).. It demonstrates the effectiveness and generalization of our model. We further provide computation time in inference for HAMT and RecBERT in the supplementary material to show the efficiency of our HAMT model. We also achieve large improvements on RxR dataset. The full results are presented in supplementary material.

Long-horizon VLN: R4R and R2R-Back.

Table 4.3 shows navigation results on R4R dataset. As R4R contains longer instructions and trajectories compared to R2R, we use the encoder-decoder variant of HAMT for better efficiency. Our method outperforms previous approaches in all metrics and shows particularly large improvements for the path fidelity related metrics. Compared to RecBERT, HAMT achives 8.2% and 9.5% relative improvement in CLS and nDTW respectively. The large improvements on these path fidelity related metrics indicate that HAMT is better to follow the designated path of the fine-grained instruction. Figure 4.3 evaluates the performance of HAMT and RecBERT with respect to instruction length measured by words. Though the nDTW decreases for longer instructions, the relative improvement of HAMT increases with the instruction length.

The navigation performance on R2R-Back dataset is presented in Table 7. We compare with two state-of-the-art recurrent models EnvDrop and RecBERT based on LSTM and transformer respectively (both models are trained on R2R-Back for fair comparison). The improvements are more significant on this task as it requires the agent to remember the way it came to the target to successfully return back. The recurrent state is insufficient to capture such history and leads to inferior performance compared to the HAMT model.

Vision-and-Dialog Navigation: CVDN.

The CVDN dataset contains dialogs as instructions and use Goal Progress (GP) in meters as the primary evaluation metric. GP measures the difference between completed distance and left distance to the goal, so the higher the better. There are two types of demonstrations in the dataset. One is shortest-path trajectory and the other is player’s navigation trajectory. We mix the two types of demonstrations as supervision in training which has shown to be the most effective in previous works . As navigation paths in CVDN dataset are much longer than R2R dataset, we adopt the encoder-decoder variant of HAMT. As shown in Table 8, HAMT outperforms existing recurrent approaches on both seen and unseen environments, and achieves the top position in the leaderboardhttps://eval.ai/web/challenges/challenge-page/463/leaderboard/1292 (25/10/2021). It demonstrates that our HAMT model is generalizable to different types of instructions in new VLN tasks.

VLN with high-level instructions: R2R-Last and REVERIE.

Table 9 shows results on the R2R-Last dataset that specifies the goal location and contains no step-by-step instructions. The HAMT model with the hierarchical history encoding is able to better accumulate the knowledge of the environment and achieves 9.8% and 10.5% relative gains on SPL metric on seen and unseen splits respectively compared to RecBERT . The REVERIE dataset also contains high-level instructions but requires object grounding at the target location besides navigation. We provide results on REVERIE dataset in supplementary material. Our HAMT achieves SPL 30.20 and 26.67 on val unseen and test splits respectively, outperforming the state of the art navigation performance by 5.3% and 2.7%.

Conclusion

This paper presents the first end-to-end transformer for vision-and-language navigation, denoted as History Aware Multimodal Transformer (HAMT). Our method efficiently encodes long-horizon history and combines it with instructions and observations to derive multimodal action prediction. The HAMT is first trained with proxy tasks in an end-to-end manner, and is then fine-tuned with RL to improve the navigation policy. We achieve state-of-the-art navigation performance on a diverse range of challenging VLN tasks, demonstrating improved accuracy and generalization of our approach compared to the dominant recurrent methods. Future work could extend our history-aware transformer to VLN with continuous actions and could benefit from pretraining on larger navigation datasets. This paper has minimal ethical, privacy and safety concerns.

Acknowledgments and Disclosure of Funding

This work was granted access to the HPC resources of IDRIS under the allocation 101002 made by GENCI. It was funded in part by the French government under management of Agence Nationale de la Recherche as part of the “Investissements d’avenir” program, reference ANR19-P3IA-0001 (PRAIRIE 3IA Institute) and by Louis Vuitton ENS Chair on Artificial Intelligence.

References

Supplementary Material for HAMT

Section A provides additional details for the model. The experimental setup is described in Section B, including datasets, metrics and implementation details. Section C presents computation time on R2R dataset and full experimental results on RxR and REVERIE datasets. Section D includes more ablations. Finally, Section E illustrates qualitative results.

Appendix A Model details

We employ five proxy tasks to train HAMT and introduced SAP/SAR and SPREL in Section 3.2. In the following, we present the other three proxy tasks, which are all based on the input pair (W,HT)(\mathcal{W},\mathcal{H}_{T}), where W\mathcal{W} is the textual instruction and HT\mathcal{H}_{T} is the full trajectory with length TT.

A.2 Structure variants in fine-tuning

We present the encoder-decoder variant of HAMT in fine-tuning on the right of Figure 4. Compared to the original cross-modal transformer on the left, the variant removes text-to-vision cross-modal attention. The encoder encodes the texts to obtain textual embeddings. Then the decoder reuses the same text embeddings in vision-to-text attention layer at each navigation step. In this way, the variant is more efficient when instructions are long e.g. in R4R and RxR datasets.

Appendix B Experimental setup

Table 10 summarizes details of the dataset split. The proposed R2R-Back and R2R-Last setups consider exactly the same splits as the R2R dataset. We present details to construct R2R-Back and R2R-Last in the following.

R2R-Back. We append a returning command at the end of annotated instructions in R2R to create new instructions for R2R-Back. The returning command is randomly sampled from the following sentences: “walk back to the start”, “return by the way you came”, “double back to where you start”, “backtrack to the start”, “back the way you came”, “return to the starting point”. The original target location is viewed as a middle stop point. The groundtruth trajectory in R2R-Back is the concatenation of the original and its inverse trajectory.

R2R-Last. We use spacy toolkithttps://spacy.io/ to split sentences for instructions in R2R. We only select the last sentence in each instruction as the new high-level instruction. It mainly describes where the goal location is e.g. “stop in front of the vent”, requiring the agent to explore houses without step-by-step textual guidance. The groundtruth trajectory is the same as R2R.

B.2 Evaluation Metrics

In R2R, RxR, R4R and R2R-Last datasets, a predicted trajectory is considered to be successful if the agent arrives 3 meters near to the final destination. However, such definition would make a motionless agent achieve 100% success rate (SR) on R2R-Back dataset as the final destination is the same as the starting location. Therefore, in R2R-Back evaluation, we define the success as that an agent firstly arrives 3 meters near to the original destination and then returns 3 meters near to its starting location. The groundtruth length in the SPL metric is also modified as the total traversed distance in groundtruth trajectory rather than the shortest distance between start and target location. As the REVERIE task aims for remote object grounding, the success on REVERIE is defined as arriving at a viewpoint where the target object is visible.

B.3 Implementation Details

Training with proxy tasks. We sample proxy tasks for each mini-batch to train the HAMT model. The sampling ratio is MLM:MRM:ITM:SAP:SAR:SPREL=5:2:2:1:1:1. The optimizer is AdamW . In the end-to-end training stage, we use image augmentation and regularization techniques to avoid overfitting of the ViT model, including RandAugment and stochastic depth .

Fine-tuning for sequential action prediction. Due to different goals in various VLN tasks, we design different rewards in reinforcement learning for each downstream VLN dataset. In R2R, RxR and R4R datasets, the reward is introduced in Section 3.3 to take both goal distance and path fidelity into account. In R2R-Last, REVERIE and CVDN datasets where the instruction may not describe detailed navigation path, we only use the reduced distance to the goal viewpoints as rewards. We normalize the reduced distance in the same way as in the R2R dataset. In R2R-Back dataset, we use a different fine-tune strategy to avoid trivial motionless solutions. We require the agent to predict stop actions twice for the original destination (midpoint) and its starting point (final destination) respectively. Before arriving at the midpoint, the RL reward is computed based on distances to the midpoint. If the agent predicts a wrong location to stop for the midpoint, the episode is stopped; otherwise the agent continues its task while receiving rewards based on the distance to the final destination for fine-tuning. We run each experiment twice for ablation study and use the best result on the validation unseen split for the state-of-the-art comparison.

Appendix C Experimental results

To assess the influence of history encoding on the inference time, we compare HAMT with RecBERT . The HAMT and RecBERT use the same number of layers in the language transformer and cross-modal transformer. The main difference of two models is in the history encoding and the attended length of history for action prediction. We run each model on the R2R val unseen split (2349 instructions) and report inference times averaged over two runs using a single Tesla P100 GPU. For our method we compare variants with and without Text-to-Vision Attention (see Section A.2), denoted here as HAMT and HAMT noT2V respectively. We can see that HAMT and its noT2V variant are only 1.5x and 1.1x slower compared to RecBERT, suggesting that attending to the whole history does not increase the inference time significantly. Moreover, while HAMT noT2V is only 10% slower compared to , it still outperforms in SR and SPL on val unseen split.

C.2 RxR dataset

As shown in Table 10, RxR dataset contains much more instructions than R2R dataset. Therefore, we directly use RxR in training proxy tasks rather than R2R with augmented data. As there are three different languages in RxR, we take advantage of pretrained multilingual BERT to initialize the unimodal language encoder, so we are able to deal with multilingual instructions using the same HAMT model. We employ the encoder-decoder variant of HAMT for computational efficiency. For fair comparison with other approaches in RxR testing leaderboardhttps://ai.google.com/research/rxr/competition?active_tab=leaderboard (25/10/2021). which adopt pretrained CLIP features, we use the same visual features without end-to-end optimization. Table 12 presents navigation performances on RxR test split. Our multilingual HAMT model achieves 12.83% and 6.25% gains on SR and nDTW respectively than the second place. Nevertheless, there is still a large gap compared to the human performance. We further present results on val seen and val unseen splits in Table 13.

C.3 REVERIE dataset

The remote object localization task in REVERIE dataset requires both navigation and object grounding. To support the two subtasks in HAMT, we concatenate object features with original view image features for each viewpoint, and add an object grounding head to predict the target object given output embeddings of all object tokens. We fine-tune HAMT that is end-to-end pretrained on R2R dataset, and use the optimized ViT to extract object features given groundtruth object bounding boxes in REVERIE dataset. As shown in Table 14, HAMT achieves better navigation performance (SR and SPL), but the object grounding performance (RGS and RGSPL) on test split is worse than state of the art. Since HAMT can more effectively encode observed visual scenes and actions in the history sequence, it is able to better understand house environments and navigate to target viewpoints more efficiently as shown in the much higher SPL score. However, as we use ViT optimized on R2R dataset to extract object features, the object representation might not be as generalizable as object features used in previous works which are pretrained on large-scale object detection datasets.

Appendix D Additional ablations

We show that the history input plays a critical role for training with proxy tasks. We compare HAMT with history input and PREVALENT without history. For fair comparison, we re-implement PREVALENT which only takes instruction W\mathcal{W} and single-step observation Ot\mathcal{O}_{t} as input and the other architectures are set the same as HAMT. We train PREVALENT with all proxy tasks except the ITM task because there is no trajectory input in PREVALENT for instruction-trajectory matching. ViT features pretrained on ImageNet are used in this experiment.

In Figure 5, we present the single-step action prediction (SAP) accuracy of HAMT and PREVALENT during the training. The SAP accuracies on val seen split are similar for the two models, however, PREVALENT performs much worse on the val unseen split than HAMT. Due to the capacity of large-scale transformer, PREVALENT is likely to memorize the map structure of seen houses, and thus achieves comparable performance to HAMT. However, such knowledge cannot be transferred to unseen houses because the structure and visual observations are distinct for seen and unseen houses. Feeding history as inputs avoids the model simply cramming the structure of seen houses, and enables it to align the history with an instruction to predict actions for better generalization. After fine-tuning the two models on R2R dataset, we obtain SPL 57.5 on val unseen split for HAMT, while 52.7 for PREVALENT without history input. As the same proxy tasks are used in training, the large gains of our HAMT model contribute to the history encoding. Therefore, the proposed history encoding can largely improve the navigation performance on top of training proxy tasks.

D.2 Visual features in training with proxy tasks

Table 15 provides an additional experiment in the third row compared to Table 3(a). It demonstrates that ViT features outperform ResNet152 features with and without training proxy tasks. Comparing the last two rows in Table 15, end-to-end feature optimization improves SPL by 2.1% on val unseen split but decreases SPL by 0.8% on val seen split. Note that we follow previous VLN works to select the best model based on val unseen and use the same model for val seen split. We observe that the performance on val seen split can be improved with longer training time. After optimizing visual representations, HAMT converges faster on val unseen split and achieves the best performance at earlier iterations. Therefore, the performance on val seen split is slightly worse than no end-to-end optimization. If training longer, the performance with optimized ViT features on val seen split can be higher.

D.3 Different proxy tasks in end-to-end training

In Table 3(b) of the main paper, we fix ViT features to ablate contributions of different proxy tasks in training. We further present the ablation results in a fully end-to-end training setup in Table 16, where different proxy tasks are used to train HAMT including the ViT features. The results show the same trend as Table 3(b), where our proposed two new proxy tasks (SAP/R and SPREL) are beneficial. Moreover, we can see that the end-to-end ViT features are superior to fixed ViT features in Table 3(b) on val unseen split for all the three proxy task combinations.

D.4 Two-stage end-to-end (e2e) training strategy

We compare our two-stage e2e training strategy with a single-stage e2e training of HAMT. However, single-stage e2e training achieves inferior performance to the two-stage training or even no e2e training. When trained for 25k iterations and evaluated on the val unseen split, the single-stage e2e training of HAMT results in SPL 53.5 while no e2e training achieves SPL 56.5. We hypothesize that the single-stage e2e training is less effective for VLN given (a) the limited training data available for the VLN task and (b) the higher complexity of VLN compared to common vision and language tasks.

D.5 History encoding in long-horizon VLN task

We compare different history encoding approaches on the R2R-Back dataset to show that the history information is more beneficial for the long-horizon VLN task. Table 17 presents navigation results. All the models are initialized from weights after training with proxy tasks. In order to successfully return back, the agent should remember the way it comes to the targets. The recurrent state is insufficient to capture all the information and achieves the worst navigation performance. Encoding agent’s oriented view at each step in temporal-only model improves over the recurrent approach. However, as the oriented view of the agent in backward trajectory is different from the view in forward trajectory, temporal-only model does not take advantage of the full memory in previous exploration and performs inferior to our hierarchical history encoding model. It demonstrates the effectiveness of our proposed method in long-horizon VLN task that requires long-term dependency. We also show that using the end-to-end trained ViT features further benefits the navigation performance.

D.6 Structure variants in fine-tuning

Our model reuses the fSAP(oi′⊙xcls′)f_{\text{SAP}}(o^{\prime}_{i}\odot x^{\prime}_{\text{cls}}) in training proxy tasks to sequentially predict action in fine-tuning. In Table 18, we compare using different input tokens for the action prediction in fSAPf_{\text{SAP}}, including different combinations of the observation token oi′o^{\prime}_{i}, global history token hcls′h^{\prime}_{\text{cls}} and special text token xcls′x^{\prime}_{\text{cls}}. We can see that the performance varies little on the val unseen split, which indicates that the cross-modal transformer in our model is able to effectively fuse different modalities so that the performance is influenced little by tokens used in prediction.

Appendix E Qualitative results

Figures 6-9 illustrate trajectories obtained by our HAMT model and compare them to results of the state-of-the-art RecBERT model. We can see that HAMT enables to better interpret instructions (Figure 6), recognize the scene (Figure 7), follow the correct direction (Figure 8), and align the current observation with the instruction (Figure 9). We also provide some failure cases in Figures 10-11, where the HAMT model still needs improvements on scene and object recognition.