Sub-Instruction Aware Vision-and-Language Navigation
Yicong Hong, Cristian Rodriguez-Opazo, Qi Wu, Stephen Gould
Introduction
Creating an agent that can navigate through an unknown environment following natural language instructions has been a dream of human-beings for many years. Such an agent needs to possess the ability to perceive its environment, understand the instructions and learn the relationship between these two streams of information. Recently, Anderson et al. (2018b) proposed the vision-and-language navigation (VLN) task that formalized such requirements through an evaluation of an agent’s ability to follow natural language instructions in photo-realistic environments.
Despite the significant progress made by recent approaches, there is little evidence that agents learn the correspondence between observations and instructions. Hu et al. (2019) found that a modified self-monitoring agent Ma et al. (2019a), could achieve similar performance with (success rate 40.5%) and without (success rate 39.7%) visual information. Among other reasons, such as dataset bias, the result suggests that this agent gains little from having the two streams of information.
We argue that one of the main reasons behind this is that current methods are not adequately teaching agents the relationship between perception—things that the robot is observing—and parts of the instructions. Since datasets do not provide such information agents can only use the ground-truth trajectory as a learning signal. Moreover, given the lack of fine-grained annotation, current methods cannot evaluate the (perceptual or linguistic) grounding process at each step as there is no ground truth signal to indicate which part of the instruction has been completed.
To address this problem, we enhance the R2R dataset Anderson et al. (2018a) to acquire intermediate supervision for the agents, providing a fine-grained matching between sub-instructions and the agent’s visual perception, as illustrated in Figure 1, to produce our Fine-Grained Room-to-Room dataset (FGR2R). We argue that the granularity of the navigation task should be at the level of these sub-instructions, rather than attempting to ground a specific part of the original long and complex instruction without any direct supervision or measure navigation progress at word level.
Our work aims to make the navigation process traceable and encourage the agent to run precisely on the described path rather than just focusing on reaching the target. We hypothesize that the agent should reach the target with higher success rate by following a detailed instruction with richer information, and in practice, the agent could complete some additional tasks on its way to the target.
In light of this, we propose a novel sub-instruction attention mechanism to better learn the correspondence between visual features and language features. Our agent first segments the long and complicated instruction into short and easier-to-understand sub-instructions using a heuristic method based on the grammatical relations provided by the Stanford NLP Parser Qi et al. (2018). Moreover, we propose a shifting module that infers whether the current sub-instruction has been completed. Hence, only one sub-instruction is available to the agent at each time step for textual grounding. These modules can be easily applied to previous VLN models.
We conduct experiments to compare the performance of four state-of-the-art agents to evaluate with or without our sub-instruction module, for agents based on imitation learning Anderson et al. (2018b); Fried et al. (2018); Ma et al. (2019a) and reinforcement learning Tan et al. (2019). Analyzing the results we find that the intermediate supervision and our proposed modules help the agents to better follow the instructions. Furthermore, we demonstrate the traceability of the navigation process through qualitative and quantitative analysis.
Related Work
Visual and textual grounding. Visual grounding aims to infer the relationship between a text description and a spatial or temporal region in an image or video, respectively. It is an essential component for a variety of tasks in vision-and-language research such as visual question answering (VQA) Schwartz et al. (2017); Anderson et al. (2018a); Hudson and Manning (2019), image captioning Xu et al. (2015); Anderson et al. (2018a); Cornia et al. (2019); Ma et al. (2020), video understanding Gao et al. (2017); Ma et al. (2018); Rodriguez et al. (2020) and phrase localization Engilberge et al. (2018); Yu et al. (2018). In the case of vision-and-language navigation, at each navigational step, the agent attends to the relevant part of the instruction according to visual clues to direct the future action. Meanwhile, the agent attends the visual inputs at different directions as described by text to perceive the environment Fried et al. (2018); Ma et al. (2019a).
Vision and language navigation. Anderson et al. (2018b) formalized the vision-and-language navigation task in a photo-realistic environment, and proposed a benchmark Room-to-Room (R2R) dataset and a sequence-to-sequence agent as a baseline model. Other datasets in real environments, such as R4R Jain et al. (2019), which is an extended version of R2R with longer instruction-path pairs, and Touchdown Chen et al. (2019) for navigation on streets have also been proposed for study.
Researchers have addressed the R2R task through a great variety of approaches. Wang et al. (2018) propose a look-ahead model that combines model-based and model-free reinforcement learning, predicts the agent’s next state and reward during navigation. Fried et al. (2018) proposed the Speaker-Follower model which generates augmented samples for training and makes use of the panoramic action space to ground and navigate efficiently. Later, Ma et al. (2019a) introduced the Self-Monitoring agent which includes a vision and language co-grounding network and a progress monitor. The progress monitor estimates a normalized distance to the target and guides the transition of the textual attention. Wang et al. (2019) applied the REINFORCE algorithm Williams (1992) to improve the agent’s generalizability and proposed a Self-Supervised Imitation Learning (SIL) method to facilitate lifelong learning in a new environment. The Back Translation agent Tan et al. (2019) applied the A2C algorithm Mnih et al. (2016) and made use of a speaker module with environmental dropout for data augmentation. Landi et al. (2019) applied dynamic convolutional filters for image feature extraction for low-level grounding of visual inputs and Hu et al. (2019) grounded multiple modalities using a mixture-of-experts approach and applied joint training strategy. Besides, the Regretful agent Ma et al. (2019b) and the Tactical Rewind agent Ke et al. (2019) are models which focus on path-scoring and backtracking methods. Very recently, Zhu et al. (2020a) introduces multiple auxiliary losses in training to help exploring the semantic meaning of visual features, Huang et al. (2019) and Hao et al. (2020) apply pre-trained encoders to generate generic visual and textual representations for the agent.
In contrast to all previously mentioned methods that ground the complete instruction, we propose to divide the instruction into meaningful semantic sub-instructions, and teach the agent to complete each one at a time before reading the next sub-instruction. In that spirit, our method is similar to the image captioning work by Cornia et al. (2019). They design a shifting gate over the image regions to control the visual features that feed into each time step of the caption module. We differ from their work in the modality that is attended. Our method works in the language domain, and the shifting depends only on local context rather than looking over all the sub-instructions. BabyWalk Zhu et al. (2020b) is a concurrent work to ours, it uses sub-instructions for curriculum learning which trains the agent to complete shorter navigation tasks before trying to solve the longer ones. Comparing the sub-instruction and sub-path pairs in FGR2R and BabyWalk, BabyWalk aligns the textual and visual sequences by solving a dynamic programming problem, whereas FGR2R employs human annotation, which is more fine-grained and accurate.
Sub-instruction Aware VLN
In this section, we first introduce the VLN problem and the general architecture of the agent. Then, we discuss about the proposed chunking function for producing the sub-instructions and the novel sub-instruction module for enabling sub-instruction attention and transition.
The VLN task requires the agent to navigate through a real environment to a target location following a natural language instruction. Formally, an instruction is a sequence of words provided to the agent at the beginning of its navigation, where denotes the -th word in the sequence. The environment is defined as set of viewpoints denoting all the navigable locations. At time step , the agent at viewpoint receives a panoramic view composed of single view images . Using the given instruction and the current observation , the agent needs to infer an action which triggers a transition signal from to . The episode ends when the agent output a action or the maximum number of steps allowed is reached.
We build our sub-instruction module based on the current state-of-the-art VLN agents, as shown in Figure 2. Those agents share a similar pipeline, a sequence-to-sequence architecture with textual and visual attentions. In this section, we refer to the Self-Monitoring Agent Ma et al. (2019a) to present the flow of information in the network.
Visual and textual encoding. Before the start of navigation, the agent first encodes the given instruction, using an LSTM with a learned embedding as Embed and LSTM, where is the hidden state of word in the instruction. In the case of the panoramic view, the agent encodes the images using a ResNet-152 model He et al. (2016) pre-trained on ImageNet Russakovsky et al. (2015) for each navigable direction. A 4-dimensional vector is concatenated with the image encoding to represent the direction of visual features, where and are the heading and elevation angles, respectively.
Policy module and co-grounding. We define the agent’s state at time as a combination of the attended textual representation , the attended visual representation and the previous selected action , encoded by an LSTM as
We refer to and as the agent’s state and memory, respectively.
The attended textual representation is obtained by performing soft-attention over the language features with the agent’s state at the previous time step. The attention weights over all the words are calculated as and Softmax, obtaining the attended textual representation by . Similarly, we perform soft-attention over the single-view visual features as where is a multi-layer perceptron (MLP), and the attention weight Softmax. The attended visual representation is . The previous selected action is represented by the visual features at the previously selected action direction. Finally, the agent decides an action by finding the visual features at a navigable direction with the highest correspondence to the attended language features and the agent’s current state . The probability at each navigable direction is computed as:
where is the same MLP as in visual attention for feature projection. Then, the agent moves in a panoramic action space Fried et al. (2018), so that it jumps directly to an adjacent viewpoint in the selected direction.
All baseline agents in our experiments are variants of this pipeline. For instance, the Speaker-Follower agent Fried et al. (2018) encodes the agent’s state with only the previous action and the attended visual features. In the case of the Back-Translation agent Tan et al. (2019), it attends the language features by the agent’s current state.
2 Chunking
To encourage the learning of vision and language correspondences, we provide short and easier-to-learn sub-instructions to the agent at each time step. Formally, for each instruction , there exists a set of sub-instructions , where and is the total number of sub-instructions. The sub-instructions are ordered, mutually exclusive and cover the entire .
We propose a chunking function to break the original instruction into several sub-instructions, where each sub-instruction is an independent navigation task and usually requires the agent to perform one or two actions to complete. To achieve this automatically, we design chunking rules based on the grammatical relations between words in the instruction, where the relations are produced by the Stanford NLP Parser Qi et al. (2018), a pre-trained natural language analysis tool. First, we pass the entire instruction into the StanfordNLP Parser for extracting the and the of each word, denoted as and , respectively. Then, using the two attributes, we formulate a heuristic as shown in Algorithm 1.
The chunking function considers words in the instruction that meet one of the following three conditions as the beginning of a new sub-instruction: (1) its dependency is root and all the words before belong to the previous chunk, (2) its dependency is conj and its governor is the previous root, (3) its dependency is parataxis and all the words before belong to the previous chunk. If any one of the three conditions is met, a Check function will be performed on the temporary chunk to decide whether to save the temporary chunk into the final sub-instruction list . Here, the Check function examines if the temporary chunk meets two conditions: (1) the chunk length should exceed the minimum length of two words, and (2) the temporary chunk should not only contains a single action-related phrase which is following the previous chunk or is leading the next chunk (e.g. “go straight then …”), if it happens, then the temporary chunk should be appended to the previous chunk or added to the next chunk respectively.
We provide an illustrative example here. Our chunking function breaks the given instruction “Enter through the glass door. Go up the wooden plank stairs on the right. Enter the doorway next to the bear head and wait there.” into 1 “Enter through the glass door”, 2 “Go up the wooden plank stair on the right”, 3 “Enter the doorway next to the bear head” and 4 “And wait there”, as shown in Figure 1. In the third and the fourth sub-instructions, the words “Enter” and “Wait” satisfy the conditions (1) and (2), respectively. Notice that the of conjunction word “And” is “Wait”, so it has been assigned to the fourth sub-instruction.
3 Sub-Instruction Module
To encourage the agent to learn the correspondence between visual and language features in a sub-instruction, we modify the base agents to include a sub-instruction module, which enables the agent to focus on a particular sub-instruction at each time step, as shown in Figure 2. It contains two main components: the sub-instruction attention and the sub-instruction shifting module.
Sub-instruction attention. The module attends the words inside the selected sub-instruction through a soft-attention mechanism. Formally, at each time step, we calculate the distribution of weights over each word in as:
where is the previous state of the agent and is the learned weights. The grounded representation of the sub-instruction is hence .
With the sub-instruction attention, the agent is forced to attend the most relevant part of the instruction and prevent the agent from “getting distracted” by the other part of the instruction that has been completed or to be completed in the further steps.
Sub-instruction shifting. At each time step, the agent needs to decide whether the current sub-instruction will be completed by the next action or not. We enable this functionality by designing a shifting module that estimates the probability of proceeding to the next sub-instruction.
The module uses a recurrent neural architecture to encode a representation that reflect the vision and language co-grounded features:
where and is the agent’s current state and memory, is the visual feature at the selected action direction, represents a sigmoid function, and are the learned weights and denotes the Hadamard product.
The module then computes the shifting probability from and a one-hot encoding of the number of sub-instructions left to be completed, as:
where and are the learned parameters. Here, introduces a learnable prior on when to shift before viewing the scene. This prior is then modified by taking into account the visual evidence, which is essential for efficient navigation. If the shifting probability exceed a certain threshold, a shift signal () of reading the next sub-instruction will be produced. We only enable the module to do a single step uni-directional shifting, which agrees with the fact that instructions and trajectory in the R2R dataset are monotonically aligned.
4 Training
In the training stage, for each instruction , there exists a corresponding ground-truth path . In the case of sub-instructions, we partition the path into sub-paths, one for each sub-instruction.
The binary cross-entropy loss compares the estimated shifting probabilities to the target shifting signals , where the target is either 1 or 0 depending on the distance between the agent’s current position and the ending viewpoint of the current sub-path. In summary, the agent’s parameters are learned to optimized
where is the predicted action, and are the ground-truth action and shifting signal respectively at time step .
During training, we apply student-forcing supervision to the action to encourage exploration, but use teacher-forcing for the sub-instruction shifting Williams and Zipser (1989); Anderson et al. (2018b). In early stages of training, the ground-truth shifting signal will have a large number of zeros since the agent has a high probability of deviating from the desired path. We prevent the sub-instruction shifting module from converging to an undesirable local minimum by forcing the shifting loss to consider an equal number of randomly selected shift and do-not-shift samples in each time step.
The FGR2R Dataset
To acquire the matching between vision and language sub-sequences, we introduce a Fine-Grained Room-to-Room (FGR2R) dataset which enriches the benchmark Room-to-Room dataset by dividing the instructions into sub-instructions and pairing each of those with their corresponding viewpoints in the path.
Dataset collection. We first apply the chunking function introduced in Section 3.2 to generate the sub-instructions automatically from the original R2R data. We demonstrate the quality of the generated sub-instructions by comparing the output sub-instructions against a manually annotated subset of 300 samples, obtaining a smoothed BLEU-4 score of . Then, we add annotations of sub-path corresponding to each sub-instruction using the Amazon Mechanical Turk (AMT)Amazon Mechanical Turk: https://www.mturk.com/. We refer the readers to Appendix A.1 for more information about the data collection interface and the qualification process of the annotators that we designed to ensure the quality of the collected data.
Dataset statistics. The original R2R possesses 21,567 navigation instructions and 7,189 paths in 91 real-world environments, where 3 or 4 different natural language instructions describe each path. The R2R data has been split for learning proposes, with 4,675 paths for training and 340 paths for seen validation in 61 scenes, 783 paths in 11 scenes for unseen validation and the remaining 1,391 paths in 18 scenes for testingMore information about R2R can be found in the Matterport3D datasetChang et al. (2017) and the R2R dataset Anderson et al. (2018b). Based on the original R2R data, FGR2R divides the instructions for the training and validation set in an average of sub-instructions. Each sub-instruction has words on average. Sub-instructions are paired on average viewpoints, and with a minimum and maximum of and viewpoints, respectively. We refer the readers to Appendix A.1 for more dataset statistics.
Experiments
We experiment with four state-of-the-art VLN agents with and without our sub-instruction module and compare their performance on the original R2R validation unseen split.
The agents are chosen to include the most common network architectures, training strategies and inference methods among the previous VLN agents. They include the Sequence-to-Sequence (Seq2Seq) Anderson et al. (2018b) model which does not apply panoramic action space, two visual-textual co-grounding models, the Speaker-Follower Fried et al. (2018) and the Self-Monitoring agent Ma et al. (2019a), as well as the Back-Translation model Tan et al. (2019) which applies reinforcement learning. For all agents, we implement our sub-instruction module in their network based on their officially released code. For the self-monitoring agent, we remove the progress monitor since it requires the attention weight over the entire instruction for estimating the navigation progress.
Implementation details. To obtain the word representations in each sub-instruction, the entire instruction is first passed to a unidirectional LSTM, then we implement chunking on the language hidden states to obtain the word representations of the selected sub-instructions. The ground-truth shifting signal at each time-step is dependent on the distance between the agent’s current position and the end viewpoint of the selected sub-instruction. If the distance is smaller than or equal to 0.5 meters, the ground-truth shift signal will be , and otherwise. For the Back-Translation model Tan et al. (2019), we only apply chunk shifting loss to the teacher-forcing imitation learning branch, so that the agent navigates on the ground-truth path and learns the chunk-shifting with less noise. We train all agents on a single NVIDIA Tesla K80 GPU, using the same hyperparameters as the baselines.
Evaluation metrics. We follow the standard metrics that previous work employed for evaluating the agent’s performance on the R2R dataset Anderson et al. (2018b), which include Path Length (PL) of the agent’s trajectory, average Navigation Error (NE) for the distance between agent’s final position and the target, Oracle Success Rate (OSR) for the ratio of agents which the shortest distance between the target and the trajectory is within , Success Rate (SR) for the ratio of agents which the distance between agent’s final position and the target is within , and Success Rate Weighted by Path Length (SPL). Furthermore, we also consider the normalized Dynamic Time Warping (nDTW) score Magalhaes et al. (2019), which is a metric that measure the overall performance of the agent with a focus on the similarity between the ground-truth and the actual trajectories.
Results and Analysis
We compare the performance of the four agents on the R2R unseen validation set. We also present the traceability of the navigation process resulting from our FGR2R data.
Quantitative results. Table 1 shows the results of the four agents in unseen environments. The performance of the imitation learning agents (Row 1–3) with our sub-instruction attention module outperforms the base agents. In terms of the success rate, the Seq2Seq, Speaker-Follower and the Self-Monitoring agents achieve an absolute increase of 1.1%, 4.9% and 1.7% respectively. The improvement is consistent in most of the other metrics, e.g. for the Self-Monitoring agent, its SPL improves from 0.30 to 0.32 and its nDTW score grows from 0.58 to 0.61. The overall improvement on Path Length and nDTW score for the first three agents indicates that using sub-instructions improves the agent’s ability to navigate on the described path. As for the Back-Translation agent (Row 4), the performance with sub-instruction attention is very similar to the baseline, one possible reason could be that the introduction of sub-instruction shifting perturbs the learning of action during for the reinforcement learning scheme which the agent could deviate far from the ground-truth path.
Learning when the agent needs to read a new sub-instruction is a difficult task, the same viewpoint in a specific environment can be considered as a shifting point or not depending on the sub-instruction that the agent follows. In Table 2, we show the confusion matrix of the shifting signals and we compute accuracy, precision, recall and F1-score to evaluate the performance of our proposed shifting module. Results show that all the agents have huge room for improvement for shifting, since the best F1-Score obtained is only 0.331. But we can see from the four agents that, as the success rate increases, the precision, recall and F1-score also improve. We propose to consider these results to be useful baselines for future methods that apply sub-instructions. Notice that agents visit a different number of viewpoints due to the maximum number of steps allowed, the use of panoramic action space and the ability to stop. In the case of Seq2Seq model, since the agent is not using a panoramic view, it performs many actions to change the camera orientation.
Qualitative performance. We illustrate a qualitative example in Figure 3 to show how the sub-instruction module works in the agent. In the example, both the baseline model and the model with the sub-instruction module completes the task successfully. However, unlike the baseline model which fails to follow the instruction and stops within 3 meters of the target by chance, our model correctly identifies the completeness of each sub-instruction, guides the agent to walk on the described path and eventually stops right at the target position. We refer the readers to Supplementary Materials for visualization of more trajectories.
2 Traceability
With the FGR2R data, we reveal the navigation process of the agent working on specific sub-instructions. For each sub-instruction, we measure the similarity between the ground-truth path and the actual trajectory using nDTW as well as the distance between the end viewpoint of the sub-instruction and the predicted shift viewpoint. As a result, we can estimate the performance of the agent in each sub-task.
We cluster the sub-instructions into 100 clusters using complete-linkage hierarchical agglomerative clustering algorithm. Instead of using a standard metric of distance such as the Euclidean distance, we compute a similarity matrix of sub-instructions using the BLEU-4 metric. We experiment with the Self-Monitoring agent on validation unseen split and present a summary of the top five and the bottom five clusters ranked by the mean distance, as shown in Table 3.
We can see from the table that the clusters which the agent performs better consist of simple and direct sub-instructions which refer to a single action, such as “head down the stair” and “exit the bedroom”. On the other hand, with sub-instructions that refer to specific objects such as “walk past the sink, fridge, oven” or express an action which is conditioned on the completion of another action, such as “walk along the grass until you reach …”, the agent deviates far from the described path. Moreover, the ranking does not show a strong correlation with the frequency or the number of viewpoints of each sub-instruction. These results suggest that agent is incapable of understanding complex natural language instructions or ground to specific objects with a high accuracy.
Conclusion
In this paper we introduce a novel sub-instruction module and the Fine-Grained R2R Dataset to encourage the learning of correspondences between vision and language. The sub-instruction module enables the agent to attend to one particular sub-instruction at each time-step and decides whether the agent needs to proceed to the next sub-instruction. Our experiments show that by implementing the sub-instruction module in state-of-the-art agents, most of the agents are able to follow the given instruction more closely and achieve better performance. We also show that, with the sub-instruction annotations, the entire navigation trajectory is trackable. We believe that the idea of sub-instruction module and a sub-instruction annotated dataset can benefit future studies in the VLN task as well as other vision-and-language problems.
References
Appendix A Appendices
Data collection. We build a web interface to collect FGR2R data using Amazon Mechanical Turk (AMT), as shown in Figure 5. In the interactive window, each viewpoint on the ground-truth path is highlighted with a large cylinder and an index of the viewpoint. Besides each sub-instruction, there is a drop-down list for assigning the start and end viewpoints of the corresponding sub-path. The annotators can click in the interactive window to freely move on the ground-truth path and freely rotate the camera to observe its surroundings. Before the start of labelling, we first ask the annotators to watch the automatic trajectory run-through to get familiar with the environment. Then, we ask them to partition the ground-truth path and assign a sub-instruction to those partitions. Once the labelling is completed, a function will automatically check if the annotation disobeys any rules (e.g., the start viewpoint of a sub-path should be the same as the end viewpoint of the previous sub-path) before approval for submission.
Annotator qualification. To ensure the quality of the annotation returned by the annotators, we annotated a subset of 300 samples as ground-truths and we exam each annotator with 15 ground-truth samples before approval for labelling. In total, there are 126 participants. We reject workers with a low agreement to the ground-truth. The qualification process leaves us 58 qualified annotators to complete the annotation task.
Dataset statistics. Apart from the FGR2R statistics mentioned in the paper, we present the distribution of sub-instructions in an instruction and the distribution of viewpoints for a sub-instruction in Figure 4. As we can see, most of the instructions are broken down into more than one sub-instruction and the frequency of more than seven sub-instructions is very low. Also, notice that about 15% of the sub-instructions are paired with only one viewpoint, as a result of the sub-instructions that only refer to camera rotation such as “rotate slightly to the left” or stopping command such as “wait by the sink”.
Training with FGR2R. During training, consider that more coherent motion could be beneficial for the agent to learn the textual-visual correspondence. We combine the sub-instructions which are only paired with one viewpoint to the next sub-instruction (and combine with the previous sub-instruction if it is the last one). The sub-instructions in validation sets remain in their original format so that the ground-truth trajectories are kept unknown. In this work, we only enable the sub-instruction module with a single step uni-directional shifting, which agrees with the observation that instructions and trajectory in the R2R dataset are monotonically aligned. However, different rules could be designed. For example, one can allow the agent to shift for more than one step or enable the agent to read the previous sub-instructions once it backtracks to the visited viewpoint. Our proposed FGR2R make all these research directions possible. We leave these ideas to future research.
A.2 Extension to Fine-Grained R4R
R2R to R4R The R4R dataset is created by concatenating two trajectories in R2R, which the first path ends within three meters from the start of the second path Jain et al. (2019). We enrich the R4R data with sub-instructions annotations by joining two sequences of sub-instructions corresponding to the two trajectories. However, for some trajectories in R4R, there exist several additional viewpoints for connecting the two paths, which has no sub-instruction annotation. Therefore, we assign those additional viewpoints to the first sub-instruction of the second path.
Evaluation We further experimented the four agents on the R4R dataset, with and without sub-instruction modules. As shown in Table 4, the performance of the first three agents are very similar. For agents with sub-instruction modules, the SR of Seq2Seq and Speaker-Follower are slightly lower, whereas the SR of Self-Monitoring agent is 1.6% higher. As for the Back-Translation, the agent experiences a large OSR and SR drop after applying sub-instructions, but the SPL and nDTW are increased by 3% and 9%. This result indicates that although the agent with sub-instruction modules has a lower chance to reach the target (stop within 3m), it follows the instruction much better.
However, we argue that a large performance gain has not been obtained in R4R mainly for two reasons: (1) The additional viewpoints created for linking the two trajectories have no corresponding sub-instructions. Hence, agents trained to follow each sub-instructions strictly have no guidance for those steps. (2) The last sub-instruction of the first trajectory is very confusing to the agent, as it usually refers to the action, but the navigation does not end. This prevents the agent from learning a good stopping policy since the ground-truth action requires the agent to keep moving.
In conclusion, we believe that it is inappropriate to apply FGR2R data directly for FGR4R task. To obtain FGR4R data, our suggestion is to remove the final sub-instruction about the action from the first trajectory, and use a Speaker module Fried et al. (2018) to generate a new sub-instruction for the additional viewpoints for linking the two trajectories. We will leave this idea for future work.
A.3 Visualization of Navigation
We visualize the navigation trajectories of the Self-Monitoring agent with and without our proposed sub-instruction module in the following pages.