Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion

Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, Gaurav Sukhatme

Introduction

Language is grounded in agent experience based on interactions with the world Bisk et al. (2020); Bender and Koller (2020). Task-oriented, instructional language focuses on objects and interactions between objects and actors, as seen in instructional datasets Damen et al. (2020); Koupaee and Wang (2018), as a function of the inextricable relationship between language and objects Quine (1960). That focus yields language descriptions of object targets for manipulation such as put the strawberries on the cutting board and slice them into pieces Chai et al. (2018). We demonstrate that predicting navigational object landmarks in addition to manipulation object targets improves the performance of an instruction following agent in a rich, 3D simulated home environment. We posit that object-centric navigation is a key piece of semantic and topological navigation Kuipers and Byun (1991) for Embodied AI (EAI) agents generally.

Substantial modeling Majumdar et al. (2020) and benchmark Qi et al. (2020b) efforts in EAI navigation focus on identifying object landmarks Blukis et al. (2018) and destinations Batra et al. (2020b). However, for agent task completion, where agents must navigate an environment and manipulate objects towards a specified goal Gordon et al. (2017); Shridhar et al. (2020), most predict movement actions without explicitly identifying navigation object targets Singh et al. (2020); Pashevich et al. (2021); Nguyen et al. (2021); Abramson et al. (2020). We address this gap, grounding navigation instructions like Head to the sink in the corner by predicting the spatial locations of the goal sink object at each timestep (Figure 1).

Transformer-based models in EAI score the alignment between a language instruction and an already-completed path Majumdar et al. (2020) or introduce recurrence by propagating part of the hidden state to the next timestep Hong et al. (2020). The former requires beam search over sequences of environment actions, which is not feasible when actions cannot be undone, such as slicing an apple. The latter introduces a heavy memory requirement, and is feasible only with short trajectories of four to six steps. We overcome both limitations by decoupling the embedding of language and visual features from the prediction of what action to take next in the environment. We first embed language and visual observations at single timesteps using a multi-modal transformer architecture, then train a transformer decoder model to consume sequences of such embeddings to decode actions (Figure 3).

We introduce Embodied BERT (EmBERT), which implements these two key insights:

Object-centric Navigation unifies the disjoint navigation and interaction action sequences in ALFRED, giving navigation actions per-step object landmarks.

Decoupled Multimodal Transformers enable extending transformer based multimodal embeddings and sequence-to-sequence prediction to the fifty average steps present in ALFRED trajectories.

Related Work

Natural language guidance of robots Tellex et al. (2020) has been explored in contexts from furniture assembly Tellex et al. (2011) to quadcoptor flight control Blukis et al. (2019).

Embodied AI. For task completion benchmarks, actions like pickup must be coupled with object targets in the visual world, with specification ranging from mask prediction only Shridhar et al. (2020) to proposals for full low level gripper control Batra et al. (2020a). Similarly, navigation benchmarks incorporate objects as targets in tasks like object navigation Qi et al. (2020b); Batra et al. (2020b); Kurenkov et al. (2020), and explicitly modeling those objects assists generally at navigation success Shrivastava et al. (2021); Qi et al. (2020a, 2021). Many successful modeling approaches for navigation benchmarks incorporate multimodal transformer models that require large memory from recurrence Hong et al. (2020), beam search over potential action sequences Majumdar et al. (2020), or shallow layers without large-scale pretraining to encode long histories Pashevich et al. (2021); Magassouba et al. (2021). In this work, we incorporate navigation object targets into the ALFRED task completion benchmark Shridhar et al. (2020), and decouple transformer-based multimodal state embedding from transformer-based translation of state embeddings to action and object target predictions. In addition, differently from other approaches that train from scratch their language encoder, we successfully exploit the BERT stack in our multi-modal architecture. In this way, EmBERT can be applied to other language-guided tasks such as VLN and Cooperative Vision-and-Dialog Navigation Thomason et al. (2019).

Language-Guided Task Completion. Table 1 summarizes how EmBERT compares to current ALFRED modeling approaches. ALFRED language instructions are given as both a single high level goal and a sequence of step-by-step instructions (Figure 2). At each timestep, we encode the goal instruction and a predicted current step-by-step instruction. We train EmBERT to predict when to advance to the next instruction, a technique introduced by LWIT Nguyen et al. (2021).

EmBERT uses a panoramic view space to see all around the agent. Rather than processing dense, single vector representations Shridhar et al. (2020); Singh et al. (2020); Pashevich et al. (2021); Kim et al. (2021); Blukis et al. (2021), EmBERT attends directly over object bounding box predictions embedded with their spatial relations to the agent, inspired by LWIT Nguyen et al. (2021) and a recurrent VLN BERT model Hong et al. (2020). We similarly follow prior work Singh et al. (2020); Pashevich et al. (2021); Nguyen et al. (2021); Kim et al. (2021); Zhang and Chai (2021) in predicting these bounding boxes as object targets for actions like Pickup, rather than directly predicting a dense object segmentation mask Shridhar et al. (2020).

Consider the step heat the mug of water in the microwave, where the visual observation before turning the microwave on and after turning the microwave off are identical. Transformer encodings of ALFRED’s large observation history are possible only with shallow networks Pashevich et al. (2021) that cannot take advantage of large scale, pretrained language models used on shorter horizons Hong et al. (2020). We decouple multimodal transformer state encoding from sequence to sequence state to action prediction, drawing inspiration from the AllenNLP SQuAD Rajpurkar et al. (2016) training procedure Gardner et al. (2017).

Our EmBERT model is the first to utilize an auxiliary, object-centric navigation prediction loss during joint navigation and manipulation tasks, building on prior work that predicted only the direction of the target object Storks et al. (2021) or honed in on landmarks during navigation-only tasks Shrivastava et al. (2021). While mapping environments during inference has shown promise on both VLN Fang et al. (2019); Chen et al. (2021) and ALFRED Blukis et al. (2021), we leave the incorporation of mapping to future work.

The ALFRED Benchmark

The ALFRED benchmark Shridhar et al. (2020) pairs household task demonstrations with written English instructions in 3d simulated rooms Kolve et al. (2017). ALFRED tasks are from seven categories: pick & place, stack & place, pick two & place, clean & place, heat & place, cool & place, and examine in light. Each task involves one or more objects that need to be manipulated, for example an apple, and a final receptacle on which they should come to rest, for example a plate. Many tasks involve intermediate state changes, for example heat & place requires cooking the target object in a microwave.

Supervision Data. Each ALFRED episode comprises an initial state for a simulated room, language instructions, planning goals, and an expert demonstration trajectory. The language instructions are given as a high-level goal instruction Ig\mathcal{I}_{g}, for example Put a cooked egg in the sink, together with a sequence of step-by-step instructions I⃗\vec{\mathcal{I}}, for example Turn right and go to the sink, Pick up the egg on the counter to the right of the sink, …\dots The planning goals P\mathcal{P} (or sub-goals) are tuples of goals and arguments, such as (SliceObject, Apple) that unpack to low-level sequences of actions like picking up a knife, performing a slice action on an apple, and putting the knife down on a countertop. The expert demonstration trajectory T\mathcal{T} is a sequence of action and object mask pairs, where Tj=(aj,Mj)\mathcal{T}_{j}=(a_{j},M_{j}). Each step-by-step instruction Ii\mathcal{I}_{i} corresponds to a sub-sequence of the expert demonstration, Tj:k\mathcal{T}_{j:k} given by alignment lookup ma(i)=(j,k)m_{a}(i)=(j,k) and to a planning goal Pb\mathcal{P}_{b} by alignment lookup mp(i)=bm_{p}(i)=b. For example, in Figure 2, instruction I0\mathcal{I}_{0} corresponds to a GotoLocation navigation goal, as well as a sequence of turning and movement API actions that a model must predict.

Model Observations. At the beginning of each episode in timestep t=0t=0, an ALFRED agent receives the high-level and step-by-step language instructions Ig,I⃗\mathcal{I}_{g},\vec{\mathcal{I}}. At every timestep tt, the agent receives a 2d, RGB visual observation representing the front-facing agent camera view, VF\mathcal{V^{F}}. ALFRED models produce an action ata_{t} from among 5 navigation (e.g., Turn Left, Move Forward, Look Up) and 7 manipulation actions (e.g., Pickup, ToggleOn, Slice), as well as an object mask MtM_{t}. Predicted action ata_{t} and mask MtM_{t} are executed in the ALFRED environment to yield the next visual observation. For navigation actions, prediction MtM_{t} is ignored, and there is no training supervision for objects associated with navigation actions.

EmBERT Predictions. EmBERT gathers additional visual data (Figure 2). After every navigation action, we turn the agent in place to obtain left, backwards, and right visual frames VL\mathcal{V^{L}}, VB\mathcal{V^{B}}, VR\mathcal{V^{R}}. Following prior work Singh et al. (2020), we run a pretrained Mask-RCNN He et al. (2017) model to extract bounding boxes from our visual observations at each view. We train EmBERT to select the bounding box which has the highest intersection-over-union with MtM_{t} (more details in Section 4).

We define a navigation object target for navigation actions. For navigation actions taken during language instruction Ii\mathcal{I}_{i}, we examine the frame VFk\mathcal{V^{F}}_{k} at time kk for Tk\mathcal{T}_{k}; ma(i)=(j,k)m_{a}(i)=(j,k). We identify the object instance OO of the class specified in the planning goal Pmb(i)\mathcal{P}_{m_{b}(i)} in VFk\mathcal{V^{F}}_{k}. We define this object OO as the navigation object target for all navigation actions in Tj:k\mathcal{T}_{j:k} by pairing those actions with object mask MOM^{O} to be predicted during training. We also add a training objective to predict the parent receptacle P(O)P(O) of OO. Parent prediction enables navigating to landmarks such as the table for instructions like Turn around and head to the box on the table, where the box compared to the table on which it rests (Figure 2).

Embodied BERT

EmBERT uses a transformer encoder for jointly embedding language and visual tokens and an transformer decoder for long-horizon planning and object-centric navigation predictions (Figure 3).

We use OSCAR Li et al. (2020) as a backbone transformer module to fuse language and visual features at each ALFRED trajectory step. We obtain subword tokens for the goal instruction Ig={g1,g2,…,gn}\mathcal{I}_{g}=\{g_{1},g_{2},\dots,g_{n}\} and the step-by-step instruction Ij={i1,i2,…,im}\mathcal{I}_{j}=\{i_{1},i_{2},\dots,i_{m}\} using the WordPiece tokenizer Wu et al. (2016) and process the sequence as: [CLS] Ig\mathcal{I}_{g} [SEP] Ij\mathcal{I}_{j} [SEP], using token type ids to distinguish the goal and step instructions. We derive token embeddings L∈R(m+n+3)×de\mathbf{L}\in\mathcal{R}^{(m+n+3)\times d_{e}} using the BERT Devlin et al. (2019) embedding layer, where ded_{e} is the embedding dimensionality.

2 Segment-Level Recurrent Action Decoder

The ALFRED challenge requires models to learn to complete action sequences averaging 50 steps and spanning multiple navigation and manipulation sub-goals. However, due to the quadratic complexity of the self-attention mechanism, feeding long sequences to transformers is computationally expensive Beltagy et al. (2020). Inspired by the TransformerXL model Dai et al. (2019), we design the Segment-Level Recurrent Action Decoder architecture that models long trajectories with recurrent segment-level state reuse. At training time we divide trajectories into temporal segments of size ss. Given two consecutive segments, si\mathbf{s}_{i} and si+1\mathbf{s}_{i+1}, EmBERT caches the representations generated for segment si\mathbf{s}_{i}. The computed gradient does not flow from si+1\mathbf{s}_{i+1} to si\mathbf{s}_{i}, but cached representations are used as extended context. When predicting the next action, the model can still perform self-attention over the previous segment representations, effectively incorporating additional contextual information that spans an high number of previous timesteps.

3 Auxiliary tasks

During the EmBERT training, we jointly optimize LA\mathcal{L}_{A}, LO\mathcal{L}_{O}, and several auxiliary tasks.

Next Instruction Prediction. Several existing models for ALFRED encode the sequence of language instructions I\mathcal{I} together with the goal (Table 1), or concatenate step-by-step instructions. These simplifications can prevent the model from carefully attending to relevant parts of the visual scene. EmBERT takes the first instruction at time t=0t=0, and performs add an auxiliary prediction task to advance from instruction Ij\mathcal{I}_{j} to instruction Ij+1\mathcal{I}_{j+1}. To supervise the next-instruction decision, we create a binary label for each step of the trajectory that indicates whether that step is the last step for a specific sub-goal, as obtained by ma(i)m_{a}(i). We use a similar FNN as Equation 1to model a Bernoulli variable used to decide when to advance to the next instruction. We denote the binary cross-entropy loss used to supervise this task as LINST\mathcal{L}_{INST}.

Object Target Predictions. EmBERT predicts a target object for navigation actions, together with the receptacle object containing the target, for example a table on which a box sits (Figure 2). For these tasks, we use an equivalent prediction layer to the one used for object prediction. We denote the cross-entropy loss associated with these task by LNAV\mathcal{L}_{NAV} and LRECP\mathcal{L}_{RECP}.

Experiments and Results

EmBERT achieves competitive performance with state of the art models on the ALFRED leaderboard test sets (Table 2), surpassing all but ET Pashevich et al. (2021) and ABP Kim et al. (2021) on Seen test fold performance (Table 3) at the time of writing. Notably, EmBERT achieves this performance without augmenting ALFRED data with additional language instructions, as is done in ET Pashevich et al. (2021), or visual distortion as used in ABP Kim et al. (2021).

Implementation Details. EmBERT is implemented using AllenNLP Gardner et al. (2017), PyTorch-Lightning,https://www.pytorchlightning.ai/ and Huggingface-Transformers Wolf et al. (2019). We train using the Adam optimizer with weight fix Loshchilov and Hutter (2017), learning rate 2e−52e^{-5}, and linear rate scheduler without warmup steps. We use dropout of 0.10.1 for the hidden layers of the FFN modules and gradient clipping of 1.01.0 for the overall model weights. Our TransformerXL-based decoder is composed of 22 layers, 88 attention heads, and uses a memory cache of 200200 slots. At training time, we segment the trajectory into 1010 timesteps. In order to optimize memory consumption, we use bucketing based on the trajectory length. We use teacher forcing Williams and Zipser (1989) to supervise EmBERT during the training process. To decide when to stop training, we monitor the average between action and object selection accuracy for every timestep based on gold trajectories. The best epoch according to that metric computed on the validation seen set is used for evaluation. The total time for each epoch is about 1 hour for a total of 20 hours for each model configuration using EC2 instances p3.8xlarge using 1 GPU.

Action Recovery Module. For obstacle avoidance, if a navigation action fails, for example the agent choosing MoveAhead when facing a wall, we take the next most confident navigation action at the following timestep, as in MOCA Singh et al. (2020). We introduce an analogous object interaction recovery procedure. When the agent chooses an interaction action such as Slice, we first select the bounding box of highest confidence to retrieve an object interaction mask. If the resulting API action fails, for example if the agent attempts to Slice a Kettle object, we choose the next highest confidence bounding box at the following timestep. The ALFRED challenge ends an episode when an agent causes 10 such API action failures.

Comparison to Other Models. Table 2 gives EmBERT performance against top and baseline models on the ALFRED leaderboard at the time of writing. Seen and Unseen sets refer to tasks in rooms that were or were not seen by the agent at training time. We report Task success rate and Goal Conditioned (GC) success rate. Task success rate is the average number of episodes completed successfully. Goal conditioned success rate is more forgiving; each episode is scored in $based on the number of subgoals satisfied, for example, in a stack & place task if one of two mugs are put on a table, the GC score is0.5$ Shridhar et al. (2020). Path weighted success penalizes taking more than the number of expert actions necessary for the task.

EmBERT outperforms MOCA Singh et al. (2020) on Unseen scenes, and several models on Seen scenes. The primary leaderboard metric is Unseen success rate, measuring models’ generalization abilities. Among competitive models, EmBERT outperforms only MOCA at Unseen generalization success. Notably, EmBERT remains competitive on Unseen path-weighted metrics, because it does not perform any kind of exploration or mapping as in HLSM Blukis et al. (2021) and ABP Kim et al. (2021).

We do not utilize the MOCA Instance Association in Time module Singh et al. (2020) that is mimicked by ET Pashevich et al. (2021). That module is conditioned based on the object class of the target object selected across timesteps. Because we directly predict object instances without conditioning on a predicted object class, our model must learn instance associations temporally in an implicit manner, rather than using such an inference time “fix”.

EmBERT Ablations. Removing the object-centric navigation prediction unique to EmBERT decreases performance on all metrics (Table 3). We show that limiting memory for the action decoder to a single previous timestep, initializing with BERT rather than OSCAR weights, and limiting vision to the front view all decrease performance in both Seen and Unseen folds.

We find that our parent prediction and visual region classification losses, however, do not improve performance. To investigate whether a smaller model would benefit more from these two auxiliary losses, we ran EmBERT with only 99 bounding boxes per side view, which enables fitting longer training segments in memory (we use 1414 timesteps, rather than 1010). We found that those losses improved EmBERT performance on the Unseen environment via both success rate and goal conditions metrics, and improved success rate alone in Seen environments when the non-frontal views were limited to 9, rather than 18, bounding boxes. Given the similar performance of EmBERT with all three auxiliary losses at 18 and 9 side views, we believe EmBERT is over-parameterized with the additional losses and 18 side view bounding boxes. It is possible that data augmentation efforts to increase the volume of ALFRED training data, such as those in ET Pashevich et al. (2021), would enable us to take advantage of the larger EmBERT configuration.

Conclusions

We apply the insight that object-centric navigation is helpful for language-guided Embodied AI to a benchmark of tasks in home environments. Our proposed Embodied BERT (EmBERT) model adapts the pretrained language model transformer OSCAR Li et al. (2020), and we introduce a decoupled transformer embedding and decoder step to enable attending over many features per timestep as well as a history of previous embedded states (Figure 1). EmBERT is the first to bring object-centric navigation to bear on language-guided, manipulation and navigation-based task completion. We find that EmBERT’s object-centric navigation and ability to attend across a long time horizon both contribute to its competitive performance with state-of-the-art ALFRED models (Table 3).

Moving forward, we will apply EmBERT to other benchmarks involving multimodal input through time, such as vision and audio data Chen et al. (2020a), as well as wider arrays of tasks to accomplish Puig et al. (2018). To further improve performance on the ALFRED benchmark, we could conceivably continue training the Mask RCNN model from MOCA Singh et al. (2020) forever by randomizing scenes in AI2THOR Kolve et al. (2017) and having the agent view the scene from randomized vantage points with gold-standard segmentation masks available from the simulator. For language supervision, we could train and apply a speaker model for ALFRED to generate additional training data for new expert demonstrations, providing an initial multimodal alignment for EmBERT, a strategy shown effective in VLN tasks Fried et al. (2018).

Implications and Impact

We evaluated EmBERT only on ALFRED, whose language directives are provided as a one-sided “recipe” accomplishing a task. The EmBERT architecture is applicable to single-instruction tasks like VLN, as long as auxiliary navigation object targets can be derived from the data as we have done here for ALFRED, by treating the “recipe” of step-by-step instructions as empty. In future work, we would like to incorporate our model on navigation tasks involving dialogue Thomason et al. (2019); de Vries et al. (2018) and real robot platforms Banerjee et al. (2020) where lifelong learning is possible Thomason et al. (2015); Johansen et al. (2020). Low-level physical robot control is more difficult than the abstract locomotion used in ALFRED, and poses a separate set of challenges Blukis et al. (2019); Anderson et al. (2020). By operating only in simulation, our model also misses the full range of experience that can ground language in the world Bisk et al. (2020), such as haptic feedback during object manipulation Thomason et al. (2020, 2016); Sinapov et al. (2014), and audio Chen et al. (2020a) and speech Harwath et al. (2019); Ku et al. (2020) features of the environment. Further, in ALFRED an agent never encounters novel object classes at inference time, which represent an additional challenge for successful task completion Suglia et al. (2020).

The ALFRED benchmark, and consequently the EmBERT model, only evaluates and considers written English. EmBERT inherently excludes people who cannot use typed communication. By training and evaluating only on English, we can only speculate whether the object-centric navigation methods introduced for EmBERT will generalize to other languages. We are cautiously optimistic that, with the success of massively multi-lingual language models Pires et al. (2019), EmBERT would be able to train with non-English language data. At the same time, we acknowledge the possibility of pernicious, inscrutable priors and behavior Bender et al. (2021) and the possibility for targeted, language prompt-based attacks Song et al. (2021) in such large-scale networks.

References

Appendix A Appendix

In this section we describe alternative auxiliary losses that we designed for EmBERT training using ALFRED data. After validation, these configurations did not produce results comparable with the best performing model. This calls for a more detailed analysis of how to adequately design and combine such losses in the complex training regime of the ALFRED benchmark.

Masked Region Modeling

This is analogous to the Visual Region Classification (VCR) loss that we integrated in the model. The main difference is that 15%15\% of the visual features are entirely masked (i.e., replaced with zero values) and we ask the model to predict them given the time-dependent representations generated by EmBERT for them.

Image-text Matching

A.2 EmBERT Asset Licenses

AI2THOR Kolve et al. (2017) is released under the Apache-2.0 License, while the ALFRED benchmark Shridhar et al. (2020) is released under the MIT License.