Look Wide and Interpret Twice: Improving Performance on Interactive Instruction-following Tasks
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani
Introduction
There is a growing interest in the community in making an embodied AI agent perform a complicated task following natural language directives. Recent studies of vision-language navigation tasks (VLN) have made significant progress Anderson et al. (2018b); Fried et al. (2018); Zhu et al. (2020). However, these studies consider navigation in static environments, where the action space is simplified, and there is no interaction with objects in the environment.
To consider more complex tasks, a benchmark named ALFRED was developed recently Shridhar et al. (2020). It requires an agent to accomplish a household task in interactive environments following given language directives. Compared with VLN, ALFRED is more challenging as the agent needs to (1) reason over a greater number of instructions and (2) predict actions from larger action space to perform a task in longer action horizons. The agent also needs to (3) localize the objects to manipulate by predicting the pixel-wise masks. Previous studies (e.g., Shridhar et al. (2020)) employ a Seq2Seq model, which performs well on the VLN tasks Ma et al. (2019). However, it works poorly on ALFRED. Overall, existing methods only show limited performance; there is a huge gap with human performance.
In this paper, we propose a new method that leads to significant performance improvements. It is based on several ideas. Firstly, we propose to choose a single instruction to process at each timestep from the given series of instructions. This approach contrasts with previous methods that encode them into a single long sequence of word features and use soft attention to specify which instruction to consider at each timestep implicitly Shridhar et al. (2020); Yeung et al. (2020); Singh et al. (2020a). Our method chooses individual instructions explicitly by learning to predict when the agent completes an instruction. This makes it possible to utilize constraints on parsing instructions, leading to a more accurate alignment of instructions and action prediction.
Secondly, we propose a two-stage approach to the interpretation of the selected instruction. In its first stage, the method interprets the instruction without using visual inputs from the environment, yielding a tentative prediction of an action-object sequence. In the second stage, the prediction is integrated with the visual inputs to predict the action to do and the object to manipulate. The tentative interpretation makes it clear to interact with what class of objects, contributing to an accurate selection of objects to interact with.
Moreover, we acquire multiple agent egocentric views of a scene as visual inputs and integrate them using a hierarchical attention mechanism. This allows the agent to have a wider field of views, leading to more accurate navigation. To be specific, converting each view into an object-centric representation, we integrate those for the multiple views into a single feature vector using hierarchical attention conditioned on the current instruction.
Besides, we propose a module for predicting precise pixel-wise masks of objects to interact with, referred to as the mask decoder. It employs the object-centric representation of the center view, i.e., multiple object masks detected by the object detector. The module selects one of these candidate masks to specify the object to interact with. In the selection, self-attention is applied to the candidate masks to weight them; they are multiplied with the tentative prediction of the pairs of action and an object class and the detector’s confidence scores for the candidate masks.
The experimental results show that the proposed method outperforms all the existing methods by a large margin and ranks first in the challenge leaderboard as of the time of submission. A preliminary version of the method won the ALFRED Challenge 2020 The ALFRED Challenge 2020 https://askforalfred.com/EVAL. The present version further improved the task success rate in unseen and seen environments to 8.37% and 29.16%, respectively, which are significantly higher than the previously published SOTA (0.39% and 3.98%, respectively) Shridhar et al. (2020).
Related Work
Many studies have been recently conducted on the problem of making an embodied AI agent follow natural language directives and accomplish the specified tasks in a three-dimensional environment while properly interacting with it. Vision-language navigation (VLN) tasks have been the most extensively studied, which require an agent to follow navigation directions in an environment.
Several frameworks and datasets for simulating real-world environments have been developed to study the VLN tasks. The early ones lack photo-realism and/or natural language directions Kempka et al. (2016); Kolve et al. (2017); Wu et al. (2018). Recent studies consider perceptually-rich simulated environments and natural language navigation directions Anderson et al. (2018b); Chen et al. (2019); Hermann et al. (2020). In particular, since the release of the Room-to-Room (R2R) dataset Anderson et al. (2018b) that is based on real imagery Chang et al. (2017), VLN has attracted increasing attention, leading to the development of many methods Fried et al. (2018); Wang et al. (2019); Ma et al. (2019); Tan et al. (2019); Majumdar et al. (2020).
Several variants of VLN tasks have been proposed. A study Nguyen et al. (2019) allows the agent to communicate with an adviser using natural language to accomplish a given goal. In a study Thomason et al. (2020), the agent placed in an environment attempts to find a specified object by communicating with a human by natural language dialog. A recent study Suhr et al. (2019) proposes the interactive environments where users can collaborate with an agent by not only instructing it to complete tasks, but also acting alongside it. Another study Krantz et al. (2020) introduces a continuous environment based on the R2R dataset that enables an agent to take more fine-grained navigation actions. A number of other embodied vision-language tasks have been proposed such as visual semantic planning Zhu et al. (2017); Gordon et al. (2019) and embodied question answering Das et al. (2018); Gordon et al. (2018); Wijmans et al. (2019); Puig et al. (2018).
2 Existing Methods for ALFRED
As mentioned earlier, ALFRED was developed to consider more complicated interactions with environments, which are missing in the above tasks, such as manipulating objects. Several methods for it have been proposed so far. A baseline method Shridhar et al. (2020) employs a Seq2Seq model with an attention mechanism and a progress monitor Ma et al. (2019), which is prior art for the VLN tasks. In Singh et al. (2020a), a pre-trained Mask R-CNN is employed to generate object masks. It is proposed in Yeung et al. (2020) to train the agent to follow instructions and reconstruct them. In Corona et al. (2020), a modular architecture is proposed to exploit the compositionality of instructions. These methods have brought about only modest performance improvements over the baseline. A concurrent study Singh et al. (2020b) proposes a modular architecture design in which the prediction of actions and object masks are treated separately, as with ours. Although it achieves notable performance improvements, the study’s ablation test indicates that the separation of the two is not the primary source of the improvements. Closely related to ALFRED, ALFWorld Shridhar et al. (2021) has been recently proposed to combine TextWorld Côté et al. (2018) and ALFRED for creating aligned environments, which enable transferring high-level policies learned in the text world to the embodied world.
Proposed Method
The proposed model consists of three decoders (i.e., instruction, mask, and action decoders) with the modules extracting features from the inputs, i.e., the visual observations of the environment and the language directives. We first summarize ALFRED and then explain the components one by one.
ALFRED is built upon AI2Thor Kolve et al. (2017), a simulation environment for embodied AI. An agent performs seven types of tasks in 120 indoor scenes that require interaction with 84 classes of objects, including 26 receptacle object classes. For each object class, there are multiple visual instances with different shapes, textures, and colors.
The dataset contains 8,055 expert demonstration episodes of task instances. They are sequences of actions, whose average length is 50, and they are used as a ground truth action sequence at training time. For each of them, language directives annotated by AMT workers are provided, which consist of a goal statement and a set of step-by-step instructions, . The alignment between each instruction and a segment of the action sequence is known. As multiple AMT workers annotate the same demonstrations, there are 25,743 language directives in total.
We wish to predict the sequence of agent’s actions, given and of a task instance. There are two types of actions, navigation actions and manipulation actions. There are five navigation actions (e.g., MoveAhead and RotateRight) and seven manipulation actions (e.g., Pickup and ToggleOn). The manipulation actions accompany an object. The agent specifies it using a pixel-wise mask in the egocentric input image. Thus, the outputs are a sequence of actions with, if necessary, the object masks.
2 Feature Representations
Unlike previous studies Shridhar et al. (2020); Singh et al. (2020a); Yeung et al. (2020), we employ the object-centric representations of a scene Devin et al. (2018), which are extracted from a pretrained object detector (i.e., Mask R-CNN He et al. (2017)). It provides richer spatial information about the scene at a more fine-grained level and thus allows the agent to localize the target objects better. Moreover, we make the agent look wider by capturing the images of its surroundings, aiming to enhance its navigation ability.
2.2 Language Representations
3 Instruction Decoder
Previous studies Shridhar et al. (2020); Singh et al. (2020a); Yeung et al. (2020) employ a Seq2Seq model in which all the language directives are represented as a single sequence of word features, and soft attention is generated over it to specify the portion to deal with at each timestep. We think this method could fail to correctly segment instructions with time, even with the employment of progress monitoring Ma et al. (2019). This method does not use a few constraints on parsing the step-by-step instructions that they should be processed in the given order and when dealing with one of them, the other instructions, especially the future ones, will be of little importance.
We propose a simple method that can take the above constraints into account, which explicitly represents which instruction to consider at the current timestep . The method introduces an integer variable storing the index of the instruction to deal with at .
To update properly, we introduce a virtual action representing the completion of a single instruction, which we treat equally to the original twelve actions defined in ALFRED. Defining a new token COMPLETE to represent this virtual action, we augment each instruction’s action sequence provided in the expert demonstrations always end with COMPLETE. At training time, we train the action decoder to predict the augmented sequences. At test time, the same decoder predicts an action at each timestep; if it predicts COMPLETE, this means completing the current instruction. The instruction index is updated as follows:
where is the predicted probability distribution over all the actions at time , which will be explained in Sec. 3.4. The encoded feature of the selected instruction is used in all the subsequent components, as shown in Fig. 1.
3.2 Decoder Design
As explained earlier, our method employs a two-stage approach for interpreting the instructions. The instruction decoder (see Fig. 1) runs the first stage, where it interprets the instruction encoded as without any visual input. To be specific, it transforms into the sequence of action-object pairs without additional input. In this stage, objects mean the classes of objects.
As it is not based on visual inputs, the predicted action-object sequence has to be tentative. The downstream components in the model (i.e., the mask decoder and the action decoder) interpret again, yielding the final prediction of an action-object sequence, which are grounded on the visual inputs. Our intention of this two-stage approach is to increase prediction accuracy; we expect that using a prior prediction of (action, object class) pairs helps more accurate grounding.
In fact, many instructions in the dataset, particularly those about interactions with objects, are sufficiently specific so that they are uniquely translated into (action, object class) sequences with a perfect accuracy, even without visual inputs. For instance, “Wash the mug in the sink” can be translated into (Put, Sink), (TurnOn, Faucet), (TurnOff, Faucet), (PickUp, Mug). However, this is not the case with navigation instructions. For instance, “Go straight to the sink” may be translated into a variable number of repetition of MoveAhead; it is also hard to translate “Walk into the drawers” when it requires to navigate to the left/right. Therefore, we separately deal with the manipulation actions and the navigation actions. In what follows, we first explain the common part and then the different parts.
These probabilities and are predicted separately by two LSTMs in an autoregressive fashion. The two LSTMs are initialized whenever a new instruction is selected; to be precise, we reset their internal states as for when we increment as (see the example in Fig. 2). Then, and are predicted as follows:
Now, as they do not need visual inputs, we can train the two LSTMs in a supervised fashion using the pairs of instructions and the corresponding ground truth action-object sequences. We denote this supervised loss, i.e., the sum of the losses for the two LSTMs, by . Although it is independent of the environment and we can train the LSTMs offline, we simultaneously train them along with other components in the model by adding to the overall loss. We think this contributes to better learning of instruction representation , which is also used by the mask decoder and the action decoder.
4 Action Decoder
As explained in Sec. 3.2, we use the multi-view object-centric representation of visual inputs. To be specific, we aggregate outputs of Mask R-CNN from ego-centric images, obtaining a single vector . The Mask R-CNN outputs for view are the visual features and the confidence scores of detected objects.
where is the confidence score associated with .
As shown in the ablation test in the appendix, the performance drops significantly when replacing the above gated-attention by soft-attention, indicating the necessity for merging observations of different views, not selecting one of them.
4.2 Decoder Design
where denotes concatenation operation. We initialize the LSTM by setting the initial hidden state to , the encoded feature of the goal statement; see Sec. 3.2. The updated state is fed into a fully-connected layer to yield the probabilities over the actions including COMPLETE as follows:
5 Mask Decoder
To predict the mask specifying an object to interact with, we utilize the object-centric representations of the visual inputs of the central view (). Namely, we have only to select one of the detected objects. This enables more accurate specification of an object mask than predicting a class-agnostic binary mask as in the prior work Shridhar et al. (2020).
where is -vector with all ones.
We then compute the probability of selecting -th object from the candidates using the above self-attended object features along with other inputs , , and . We concatenate the latter three inputs into a vector and then compute the probability as
Experiments
We follow the standard procedure of ALFRED; 25,743 language directives over 8,055 expert demonstration episodes are split into the training, validation, and test sets. The latter two are further divided into two splits, called seen and unseen, depending on whether the scenes are included in the training set.
Following Shridhar et al. (2020), we report the standard metrics, i.e., the scores of Task Success Rate, denoted by Task and Goal Condition Success Rate, denoted by Goal-Cond. The Goal-Cond score is the ratio of goal conditions being completed at the end of an episode. The Task score is defined to be one if all the goal conditions are completed, and otherwise 0. Besides, each metric is accompanied by a path-length-weighted (PLW) score Anderson et al. (2018a), which measures the agent’s efficiency by penalizing scores with the length of the action sequence.
We use views: the center view, up and down views with the elevation degrees of , and left and right views with the angles of . We employ a Mask R-CNN model with ResNet-50 backbone that receives a image and outputs object candidates. We train it before training the proposed model with 800K frames and corresponding instance segmentation masks collected by replaying the expert demonstrations of the training set. We set the feature dimensionality . We train the model using imitation learning on the expert demonstrations by minimizing the following loss:
We use the Adam optimizer with an initial learning rate of , which is halved at epoch 5, 8, and 10, and a batch size of 32 for 15 epochs in total. We use a dropout with the dropout probability 0.2 for the both visual features and LSTM decoder hidden states.
2 Experimental Results
Table 1 shows the results. It is seen that our method shows significant improvement over the previous methods Shridhar et al. (2020); Yeung et al. (2020); Singh et al. (2020a, b) on all metrics. Our method also achieves better PLW (path length weighted) scores in all the metrics (indicated in the parentheses), showing its efficiency. Notably, our method attains 8.37% success rate on the unseen test split, improving approximately 20 times compared with the published result in Shridhar et al. (2020). The higher success rate in the unseen scenes indicates its ability to generalize in novel environments. Detailed results for each of the seven task types are shown in the appendix.
The preliminary version of our method won an international competition, whose performance is lower than the present version. It differs in that are not forwarded to the mask decoder and the action decoder and the number of Mask R-CNN’s outputs is set to . It is noted that even with a single view (i.e., ), our model still outperforms Shridhar et al. (2020); Yeung et al. (2020); Singh et al. (2020a) in all the metrics.
Following Shridhar et al. (2020), we evaluate the performance on individual sub-goals. Table 2 shows the results. It is seen that our method shows higher success rates in almost all of the sub-goal categories.
3 Ablation Study
We conduct an ablation test to validate the effectiveness of the components by incrementally adding each component to the proposed model. The results are shown in Table 3.
The model variants 1-4 use a single-view input (); they do not use multi-view inputs and the hierarchical attention method. Model 1 further discards the instruction decoder by replacing it with the soft-attention-based approach Shridhar et al. (2020), which yields a different language feature at each timestep. Accordingly, and are not fed to the mask/action decoders; we use . These changes will make the method almost unworkable. Model 2 retains only the instruction selection module, yielding . It performs much better than Model 1. Model 3 has the instruction decoder, which feeds and to the subsequent decoders. It performs better than Model 2 by a large margin, showing the effectiveness of the two-stage method.
Model 4 replaces the mask decoder with the counterpart of the baseline method Shridhar et al. (2020), which upsamples a concatenated vector by deconvolution layers. This change results in inaccurate mask prediction, yielding a considerable performance drop. Model 5 is the full model. The difference from Model 3 is the use of multi-view inputs with the hierarchical attention mechanism. It contributes to a notable performance improvement, validating its effectiveness.
4 Qualitative Results
Figure 3 shows the visualization of how the agent completes one of the seven types of tasks. These are the results for the unseen environment of the validation set. Each panel shows the agent’s center view with the predicted action and object mask (if existing) at different time-steps. See the appendix for more results.
4.2 Mask Prediction for Sub-goal Completion
Figure 4 shows an example of the mask prediction by the baseline Shridhar et al. (2020) and the proposed method. It shows our method can predict a more accurate object mask when performing Slice sub-goal. More examples are shown in the appendix. Overall, our method shows better results, especially for difficult sub-goals like Pickup, Put, and Clean, for which a target object needs to be chosen from a wide range of candidates.
Conclusion
This paper has presented a new method for interactive instruction following tasks and applied it to ALFRED. The method is built upon several new ideas, including the explicit selection of one of the provided instructions, the two-stage approach to the interpretation of each instruction (i.e., the instruction decoder), the employment of the object-centric representation of visual inputs obtained by hierarchical attention from multiple surrounding views (i.e., the action decoder), and the precise specification of objects to interact with based on the object-centric representation (i.e., the mask decoder). The experimental results have shown that the proposed method achieves superior performances in both seen and unseen environments compared with all the existing methods. We believe this study provides a useful baseline framework for future studies.
This work was partly supported by JSPS KAKENHI Grant Number 20H05952 and JP19H01110.
References
Appendix
Appendix A Additional Experimental Results
Table 4 shows the success rates across the 7 task types achieved by the existing methods including ours on the validation set of ALFRED. It is seen that our method outperforms others by a large margin in both seen and unseen environments.
A.2 Full Results of Ablation Tests
Table 5 shows the full results of the ablation test reported in the main paper. We also provide additional results in Table 6 with different activation functions (i.e. sigmoid or softmax) in the second step of the proposed hierarchical attention mechanism, and with different ’s; is selected from 1 (only ‘center’ view), 3 (‘center’, ‘left’, and ‘right’ views), or 5 (‘center’, ‘left’, ‘right’, ‘up’, and ‘down’ views)). The results show that the use of gated-attention in Eq.(5) (of the main paper) is essential. We also confirm the number of views also affect the success rate.
Appendix B Qualitative Results
Figure 3 shows examples of mask predictions by the baseline Shridhar et al. (2020) and the proposed method for different sub-goals. Overall, it is seen that our method can predict more accurate object masks. It shows better results especially for difficult sub-goals like Pickup, Put, and Clean, where a target object needs to be chosen from a wide range of candidates.
B.2 Entire Task Completion
Figures 4-10 show example visualization of how the agent completes the seven types of tasks. These are the results for the unseen environments of the validation set. Each panel shows the agent’s center view with the predicted action and object mask (if existing) at different timesteps.
We also provide seven video clips as independent files, which contain several examples of the agent’s entire task completion for seven above task instances in unseen environments.
Appendix C Analyses of Failure Cases
We analyze the failure cases of our method using the results on the validation splits. We categorize them into navigation failures and manipulation failures.
It is seen from the sub-goal results of Table 2 in the main paper that the Goto sub-goal is the most challenging. Failures with it tend to make it hard to complete the entire goal, since they will inevitably affect the subsequent actions to take. We think there are three major cases for the navigation failures.
The first case, which occurs most frequently, is that the agent follows a navigation instruction and reaches a position that should be fine as far as the instruction goes; nevertheless, it is not the right position for the next manipulation action to take. For instance, following the instruction “Go to the table,” the agent goes to the table. The next instruction is “ Pickup the remote control at the table,” but the remote lies on the other side of the table. This is counted as a failure of completing the Goto sub-goal.
The second case is when the instructions are either abstract or misleading. An example is that when the agent has to take several left and right turns together with multiple MoveAhead steps to reach the destination, e.g., a drawer, the provided instruction is simply“Go to the drawer.”
The third case, which occurs less frequently, is that while there is an obstacle in front of the agent, e.g., wall, it attempts to take the MoveAhead action. This occurs because of the lack of proper visual inputs. This is demonstrated by the fact that when we reduce the number of views, the task success rate drops significantly, as shown in the second block of Table 6.
C.2 Manipulation Failures
As shown in Table 2, after it has moved to the ideal position right before performing any interaction sub-goals (i.e., all the sub-goals but Goto), the agent can manipulates objects with high success rates of 91% and 69% in the seen and unseen environments, respectively.
However, the success rates for completing the Goto sub-goal in the seen and unseen environments are only 59% and 39%, respectively. Therefore, the primary cause of the manipulation failures is that the agent cannot find the target object becuase it fails to reach the right destination due to a navigation failure.
Even if the agent has successfully navigated to the right destination, it can fail to detect the target object. This seem to happen mostly because the object is either too small or indistinguishable from the surroundings. The agent tends to fail to detect, for example, a small knife placed on the steel/metal-made sink of the same color.
The agent also fails to detect an object that has not seen in the training. This is confirmed by the fact that the performance drops considerably in unseen environments for some interaction sub-goals (including Put, Clean, and Toggle). There are also a small number of cases where failures are attributable to bad instructions, e.g., incorrect statement of objects.
Appendix D Further Details of Implementation
Table 7 summarizes the hyperparameters used in our experiments. We perform all the experiments on a GPU server that has four Tesla V100-SXM2 of 16GB memory. It has Intel(R) Xeon(R) Gold 6148 CPU @ 2.40GHz of 20 cores with the RAM of 240GB memory. We use Pytorch version 1.6 Paszke et al. (2019). As for the Mask R-CNN model, we train it in advance, separately from the main model, for 15 epochs with the learning rate and halved at the epoch 5, 8, and 12. We train models with batch size of 32 and 4 workers per GPU.
Appendix E Entry Submission to the ALFRED Embodied AI Challenge 2021
In this section, we describe the details of our entry submission to the ALFRED Embodied AI Challenge 2021https://askforalfred.com/EAI21, which is organized in conjuction with the CVPR 2021.
To further improve the performance of our model, we create an ensemble of 4 models with some differences, from initializations with different random seeds to whether to use self-attention mechanism in the mask decoder. We train these 4 models independently with the same pretrained Mask R-CNN, whose weights fixed during the training phase. Refer to these above sections for training and implementation details.
During the evaluation, a shared detector Mask R-CNN is to extract the features from visual inputs, which are forwarded into all the models. Each model will ouput the probability distributions over the actions and objects of interest if any. We then take average these probabilities and select the actions and objects of highest probabilities. Figure 2 illustrates the use of ensemble during the inference.
Table 8 shows the performance of our single model and its corresponding ensemble. It is seen that ensembling increases the success rate of the agent on both seen and unseen environments in comparison with a single model.
It is noted that the evaluation speed is slowed down proportionally when capturing multiple ego-centric views from the agent’s camera. It is due to the fact that the current version of AI2Thor does not support direct acquisition of multiple/panoramic views; we thus have to move agents into different view points to capture such corresponding views. Our best-performing results from multi-view models suggest that it be of importance to improve the benchmark which is capable of obtaining multiple/panoramic views directly. Also noted that using ensemble during evaluation requires running all of its models simultaneously. Therefore, it leads to a proportional increase in memory and compute.