A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution

Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, Yoav Artzi

Introduction

Mobile manipulation in a home environment requires addressing multiple challenges, including exploration and making long-term inference about actions to perform. In addition to reasoning, robots require an accessible, yet sufficiently expressive interface to specify their tasks. Natural Language provides an intuitive mechanism for task specification, and coupled with advances in automated language understanding, is increasingly applied to embodied agents [e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11].

In this paper, we study the problem of learning to map high-level natural language instructions to low-level mobile manipulation actions in an interactive 3D environment . Existing work largely studies language tightly aligned to the robot actions, either using single-sentence instructions [e.g., 1, 2, 5, 9] or sequences of instructions . In contrast, we focus on high-level instructions, which provide more efficient human-robot communication, but require long-horizon reasoning across layers of abstraction to generate actions not explicitly specified in the instruction.

Robust reasoning about manipulation goals from unrestricted high-level natural language instructions has a variety of open challenges. Consider the instruction secure two discs in a bedroom safe (Figure 1). The robot must first locate the safe in the bedroom. It then needs to distribute the actions entailed by secure to two objects (two discs), each requiring a distinct sequence of actions, but targeting the same safe. It is also required to map the verb secure to its action space. In parallel, the robot must address mobile manipulation challenges, and often can only identify required actions as it observes and manipulates the world (e.g., if the safe needs to be opened).

We propose to construct and continually update a spatial semantic representation of the world from robot observations (Figure 2). Similar to widely used map representations , we retain the spatial properties of the environment, allowing the robot to navigate and reason about relations between objects, as required to accomplish its task. We propose the Hierarchical Language-conditioned Spatial Model (HLSM), a hierarchical approach that uses our spatial representation as a long-term memory to solve long-horizon tasks. HLSM consists of a high-level controller that generates subgoals, and a low-level controller that generates sequences of actions to accomplish them. In our example (Figure 1), the sequence of subgoals is ⟨\langlepick up a CD, open the safe, put the CD in the safe, …⟩\rangle, each requiring a sequence of actions. The spatial representation allows selecting subgoals that use previously observed objects outside of the agent’s view, or to decide about needed exploration.

We evaluate our approach on the ALFRED benchmark and achieve state-of-the-art results without using the low-level instructions used by previous work , neither during training nor at test-time. This paper makes three key contributions: (a) a modular representation learning approach for the problem of mapping high-level natural language task descriptions to actions in a 3D environment; (b) a method for utilizing a spatial semantic representation within a hierarchical model for solving mobile manipulation tasks; and (c) state-of-the-art performance on the ALFRED benchmark, even outperforming all approaches that use detailed sequential instructions.

Related Work

Natural language has been extensively studied in robotics research, including with focus on instruction , reference resolution , question generation , and dialogue . Most work in this area has considered either synthetic instructions of relatively simple goals , or natural language instructions where all intermediate steps are explained in detail . In contrast, we focus on high-level instructions, which are more likely in home environments .

Representation of world state, action history, and language semantics plays a central role in robot systems and their algorithm design. Symbolic representations have been extensively studied for instruction following agents . While they simplify the symbol grounding problem and enable robustness, the ontologies on which they rely on are laborious to scale to new, unstructured environments and language. Representation learning presents an alternative by learning to map observations and language directly to actions . World state and language semantics are represented with vectors or by memorizing past observations . Modelling improvements have enabled these approaches to achieve good performance on complex navigation tasks , a success that has not yet translated to mobile manipulation .

We propose integrating a semantic voxel map state representation within a hierarchical representation learning system. Similar semantic 2D maps have been successfully used in navigation and more recently even in mobile manipulation instruction-following tasks . We extend these maps to 3D and show state-of-the-art results on a challenging mobile manipulation benchmark. Our map design is related to sparse metric, topological and semantic maps that have enabled grounding symbolic instruction representations. Our map does not impose a topological structure or require reasoning about object instances, instead modelling a distribution over semantic classes for every voxel.

Problem Definition

Let A\mathcal{A} be the set of agent actions, and S\mathcal{S} the set of world states. Given a natural language instruction LL and an initial state s0∈Ss_{0}\in\mathcal{S}, the agent’s goal is to generate an execution Ξ=⟨s0,a0,s1,a1,…,sT,aT⟩\Xi=\langle s_{0},a_{0},s_{1},a_{1},\dots,s_{T},a_{T}\rangle, where at∈Aa_{t}\in\mathcal{A} is an action taken by the agent at time tt, st∈Ss_{t}\in\mathcal{S} is the state before taking ata_{t}, and st+1=T(st,at)s_{t+1}=\mathcal{T}(s_{t},a_{t}) under environment dynamics T:S×A→S\mathcal{T}:\mathcal{S}\times\mathcal{A}\rightarrow\mathcal{S}. The state sts_{t} is defined by the environment layout and the poses and states of all objects and the agent. The agent does not have access to the state sts_{t}, but only to an observation oto_{t}. An observation ot=(It,Pt,vtS,L)o_{t}=(I_{t},P_{t},v^{S}_{t},L) includes a first-person RGB camera image ItI_{t}, the agent’s pose PtP_{t}, a one-hot encoding of the object class the agent is holding vtSv^{S}_{t}, and the instruction LL.The task is considered successful if all goal-conditions corresponding to the task LL are true at the final state sTs_{T}. Partial success is measured as the fraction of goal-conditions that have been achieved.

The ALFRED dataset includes sets of seen and unseen environments. The set of actions A=Anav∪Aint\mathcal{A}=\mathcal{A}_{\rm nav}\cup\mathcal{A}_{\rm int} includes parameter-free navigation actions Anav={\textscMoveAhead,\textscRotateLeft,\textscRotateRight}\mathcal{A}_{\rm nav}=\{\textsc{MoveAhead},\textsc{RotateLeft},\textsc{RotateRight}\} and interaction actions Aint={\textscPickup,\textscPut,\textscToggleOn,\textscToggleOff,\textscOpen,\textscClose,\textscSlice}\mathcal{A}_{\rm int}=\{\textsc{Pickup},\textsc{Put},\textsc{ToggleOn},\textsc{ToggleOff},\textsc{Open},\textsc{Close},\textsc{Slice}\} parameterized by a binary mask that identifies the object of the interaction in the agent’s current first-person view. We compute PtP_{t} and vtSv^{S}_{t} using dead-reckoning from RGB observations and actions.

Hierarchical Model with a Persistent Spatial Semantic Representation

We model the agent behavior with a policy π\pi that maps an instruction LL and the observation oto_{t} at time tt to an action ata_{t}. The policy π\pi is made of an observation model FF and two controllers: a high-level controller πH\pi^{H} and a low-level controller πL\pi^{L}. The observation model builds a spatial state representation s^t\hat{s}_{t} that captures the cumulative agent knowledge of the world at time tt. s^t\hat{s}_{t} is used by both πH\pi^{H} for high-level long-horizon task planning, and πL\pi^{L} for near-term reasoning, such as object search, navigation, collision avoidance, and manipulation. Figure 2 illustrates the policy.

The high-level controller πH\pi^{H} computes a probability over subgoals. A subgoal gg is a tuple (type,argC,argM)(\texttt{type},\texttt{arg}^{C},\texttt{arg}^{M}), where type∈Aint\texttt{type}\in\mathcal{A}_{\rm int} is an interaction type (e.g., Open, Pickup), argC\texttt{arg}^{C} is the semantic class of the interaction argument (e.g., Safe, CD), and argM\texttt{arg}^{M} is a 3D mask identifying the location of the argument instance. In ALFRED, each interaction action in the set Aint\mathcal{A}_{\rm int} corresponds to a subgoal type. When predicting the kk-th subgoal at time tt, πH\pi^{H} considers the instruction LL, the current state representation s^t\hat{s}_{t}, and the sequence of past subgoals ⟨gi,⟩i<k\langle g_{i},\rangle_{i<k}. During inference, we sample from πH\pi^{H}. Unlike arg⁡max⁡\arg\max, sampling allows the agent to re-try the same or different subgoal incase of a potentially random failure (e.g., if a Mug was not found, pick up a Cup).

The low-level controller πL\pi^{L} is given the subgoal gkg_{k} as its goal specification at time tt. At every timestep j>tj>t, πL\pi^{L} maps the state representation s^j\hat{s}_{j} and subgoal gkg_{k} to an action aja_{j}, until it outputs one of the stop actions: aPASSa_{\rm PASS} or aFAILa_{\rm FAIL} to indicate successful or failed subgoal completion.

The execution flow is as follows. At time t=0t=0 the initial observation o0o_{0} is received. At each timestep, we update the state representation s^t\hat{s}_{t} using the observation model. If there is no currently active subgoal, we sample a new subgoal gkg_{k} from πH\pi^{H}, and then sample an action ata_{t} from πL\pi^{L}. If ata_{t} is aPASSa_{\rm PASS}, we increment subgoal counter kk. If it is aFAILa_{\rm FAIL}, we discard the current subgoal kk. We repeat sampling subgoals and actions until an executable action ata_{t} is sampled. We execute ata_{t}, increment the timestep tt, and receive the next observation oto_{t}. The episode ends when the subgoal gSTOPg_{\rm STOP} is sampled or the horizon TmaxT_{max} is exceeded. Algorithm 1 in Appendix A.4 describes this process.

The state representation s^t\hat{s}_{t} at time tt captures the agent’s current understanding of the state of the world, including the locations of objects observed and the agent’s relation to them. The state representation is a tuple (VtS,VtO,vtS,Pt)(V^{S}_{t},V^{O}_{t},v^{S}_{t},P_{t}). The semantic map VtS∈X×Y×Z×CV^{S}_{t}\in^{X\times Y\times Z\times C} is a 3D voxel map that for every position indicates which of the c∈[1,C]c\in[1,C] object classes are present in the voxel. The observability map VtO∈{0,1}X×Y×ZV^{O}_{t}\in\{0,1\}^{X\times Y\times Z} is a 3D voxel map that indicates whether the corresponding position has been observed. The inventory vector vtS∈{0,1}Cv^{S}_{t}\in\{0,1\}^{C} indicates which of the CC object classes the agent is currently holding. The agent pose Pt=(x,y,ωp,ωy)P_{t}=(x,y,\omega_{p},\omega_{y}) is specified by the 2D position (x,y)(x,y), pitch angle ωp\omega_{p}, and yaw angle ωy\omega_{y}.

We also compute 2D state affordance features \textscAfford(s^t)∈7×X×Y\textsc{Afford}(\hat{s}_{t})\in^{7\times X\times Y} in a top-down view that represent each position with one or more of seven affordance classes {pickable, receptacle, togglable, openable, ground, obstacle, observed}. Each [\textscAfford(s^t)](τ,x,y)=1.0[\textsc{Afford}(\hat{s}_{t})]_{(\tau,x,y)}=1.0 if at least one of the voxels at position (x,y)(x,y) has affordance class τ\tau, otherwise it is zero. \textscAfford(s^t)\textsc{Afford}(\hat{s}_{t}) is suited for object class agnostic reasoning, for example predicting a pose to pick up an object. We assume a known mapping between object semantic classes and affordance classes.

2 Observation Model

The observation model F(s^t−1,ot,gk)F(\hat{s}_{t-1},o_{t},g_{k}) updates the state representation with new observations. It considers the current subgoal gkg_{k} to actively acquire information relevant to gkg_{k}. The computation of FF consists of three steps: perception, projection, accumulation.

We predict semantic segmentation ItSI^{S}_{t} and depth map 0ptt0pt_{t} from the RGB observation ItI_{t}. We use neural networks pre-trained in the ALFRED environment. The semantic segmentation [ItS](u,v)[I^{S}_{t}]_{(u,v)} is a distribution over CC object classes at pixel (u,v)(u,v). The depth map [0ptt](u,v)[0pt_{t}]_{(u,v)} is a binned distribution over BB bins.We use BB uniformly spaced depth bins {0,ΔD,2ΔD,…,(B−1)ΔD}\{0,\Delta_{D},2\Delta_{D},\dots,(B-1)\Delta_{D}\}, where ΔD\Delta_{D} is a depth resolution. We suggest ΔD\Delta_{D} should be less than 50% of the voxel size. We used voxels with edge length 0.25m. We also heuristically compute a binary mask MtDM^{D}_{t} that indicates which pixels have confident depth readings. We allow more confidence slack in pixels that correspond to the current subgoal argument argtC\texttt{arg}^{C}_{t} according to ItSI^{S}_{t}. Appendix A.3 provides further details. We use perception models based on the U-Net architecture, but our framework supports other, potentially more powerful models as well (e.g. ).

We integrate V^tS\hat{V}^{S}_{t} and V^tO\hat{V}^{O}_{t} into a persistent state representation:

This operation updates each voxel with the most recent semantic distribution, while retaining the values of all voxels not visible at time tt. The output of the observation model is the spatial state representation s^t=(VtS,VtO,vtS,Pt)\hat{s}_{t}=(V^{S}_{t},V^{O}_{t},v^{S}_{t},P_{t}). The inventory vtSv^{S}_{t} and pose PtP_{t} are taken directly from oto_{t}.

At timestep tt, when invoked for the kk-th time, the input to πH\pi^{H} is the instruction LL, the sequence of past subgoals ⟨gi⟩i<k\langle g_{i}\rangle_{i<k}, and the current state representation s^t\hat{s}_{t}. The output is the next subgoal gk=(typek,argkC,argkM)g_{k}=(\texttt{type}_{k},\texttt{arg}^{C}_{k},\texttt{arg}^{M}_{k}). Figure 3 illustrates the high-level controller architecture.

Subgoal Prediction We concatenate the three representations h(t,k)=[ϕL;ϕts;ϕk−1g]\mathbf{h}_{(t,k)}=[\phi^{L};\phi^{s}_{t};\phi^{g}_{k-1}]. We use a densely connected multi-layer perceptron to predict two distributions P(typek∣h(t,k))P(\texttt{type}_{k}\mid\mathbf{h}_{(t,k)}) and P(argkC∣typek,h(t,k))P(\texttt{arg}^{C}_{k}\mid\texttt{type}_{k},\mathbf{h}_{(t,k)}), from which we sample a subgoal type typek\texttt{type}_{k} and argument class argkC\texttt{arg}^{C}_{k}.

The remaining component of the subgoal is the action argument mask argkM\texttt{arg}^{M}_{k}. Let [VtS](argkC)[V^{S}_{t}]_{(\texttt{arg}^{C}_{k})} be a voxel map that only retains the object information for objects of class argkC\texttt{arg}^{C}_{k} in the semantic map VtSV^{S}_{t}. We refine it to identify a single object instance. We compute a birds-eye view representation:

where \textscEgoTransform(x,Pt)\textsc{EgoTransform}(\mathbf{x},P_{t}) transforms the map x\mathbf{x} to the agent egocentric pose PtP_{t}, Refiner is a neural network based on the LingUNet architecture , and ϕL\phi^{L} is the language embedding. The refined argkM\texttt{arg}^{M}_{k} is a $−valued3Dmaskthatidentifiestheinstanceoftheinteractionargumentobject.Iftheobjectisbelievedtobeunobserved,then-valued 3D mask that identifies the instance of the interaction argument object. If the object is believed to be unobserved, then\texttt{arg}^{M}_{k}containsallzeroes.Thecontrolleroutputisthesubgoalcontains all zeroes. The controller output is the subgoalg_{k}=(\texttt{type}_{k},\texttt{arg}^{C}_{k},\texttt{arg}^{M}_{k})$.

The low-level controller πL\pi^{L} is conditioned on the most recent subgoal gk=(typek,argkC,argkM)g_{k}=(\texttt{type}_{k},\texttt{arg}^{C}_{k},\texttt{arg}^{M}_{k}). At time tt, it maps the state representation s^t\hat{s}_{t} to an action ata_{t}. It combines engineered and learned components. Appendix A.6 provides the implementation details. The controller πL\pi^{L} invokes a set of procedures: NavigateTo, SampleExplorationPosition, SampleInteractionPose, and InteractMask. Their invokation follows a pre-specified execution flow across multiple timesteps. First, we perform a 360° rotation to observe the nearby environment. If no objects of type argkC\texttt{arg}^{C}_{k} are observed, we explore the environment by sampling a position (x,y)=SampleExplorationPosition(s^t)(x,y)=\texttt{SampleExplorationPosition}(\hat{s}_{t}), navigating there using the procedure NavigateTo(x,y,s^t)\texttt{NavigateTo}(x,y,\hat{s}_{t}), and performing a 360° rotation. We repeat exploration until a voxel in VtSV^{S}_{t} contains the class argkC\texttt{arg}^{C}_{k} with >>50% probability. To interact with an object, we sample an interaction pose (x,y,ωy,ωp)=SampleInteractionPose(s^t,gk)(x,y,\omega_{y},\omega_{p})=\texttt{SampleInteractionPose}(\hat{s}_{t},g_{k}), invoke NavigateTo(x,y,s^t)\texttt{NavigateTo}(x,y,\hat{s}_{t}) to reach the position (x,y), and then rotate according to yaw and pitch angles (ωy,ωp)(\omega_{y},\omega_{p}). Finally, we generate the egocentric interaction mask maskt=InteractMask(s^t,argkM)\texttt{mask}_{t}=\texttt{InteractMask}(\hat{s}_{t},\texttt{arg}^{M}_{k}), and output the interaction action (typek,maskt)(\texttt{type}_{k},\texttt{mask}_{t}).

All procedures use the spatial representation s^t\hat{s}_{t}. NavigateTo navigates to a goal position using a value iteration network (VIN) that reasons over obstacle and observability maps from s^t\hat{s}_{t}. SampleExplorationPosition samples positions on the boundary of observed space in s^t\hat{s}_{t}. SampleInteractionPose uses a learned neural network NavModel to predict a distributon of poses from which the interaction gkg_{k} will likely succeed. InteractMask uses the segmentation image ItSI^{S}_{t} and the 3D argument mask argtM\texttt{arg}^{M}_{t} to compute the first-person mask of the target object.

Learning

The policy contains four learned models: the segmentation and depth networks, πH\pi^{H}, and the navigation model NavModel used by πL\pi^{L}. We train all four networks independently using supervised learning. We assume access to a training dataset D={(L(j),Ξ(j))}j=1ND\mathcal{D}=\{(L^{(j)},\Xi^{(j)})\}_{j=1}^{N_{D}} of high-level natural language instructions L(j)L^{(j)} paired with demonstration execution Ξ(j)\Xi^{(j)} in a set of seen environments. Each execution Ξ(j)\Xi^{(j)} is a sequence of states and actions ⟨s0(j),a0(j),…,sT(j),aT(j)⟩\langle s_{0}^{(j)},a_{0}^{(j)},\dots,s_{T}^{(j)},a_{T}^{(j)}\rangle. We denote NPN_{P} the total number of states in dataset D\mathcal{D}, and NGN_{G} the total number of subgoals.

We process D\mathcal{D} into three datasets. The perception dataset DP={([I](i),[0pt](i),[IS](i)}i=1NP\mathcal{D}^{P}=\{([I]^{(i)},[0pt]^{(i)},[I^{S}]^{(i)}\}_{i=1}^{N_{P}} includes RGB images [I](i)[I]^{(i)} with ground truth depth [0pt](i)[0pt]^{(i)} and segmentation [IS](i)[I^{S}]^{(i)}. The subgoal dataset Dg={(L(i),s^t(i),⟨gj(i)⟩j=0k)}i=1NG\mathcal{D}^{g}=\{(L^{(i)},\hat{s}_{t}^{(i)},\langle g_{j}^{(i)}\rangle_{j=0}^{k})\}_{i=1}^{N_{G}} contains natural language instructions L(i)L^{(i)}, state representations s^t(i)\hat{s}_{t}^{(i)} at the start of kk-th subgoal execution, and sequences of the first kk subgoals ⟨gj(i)⟩j=0k\langle g_{j}^{(i)}\rangle_{j=0}^{k} extracted from Ξ(j)\Xi^{(j)}. The navigation dataset DN={(s^(i),g(i),P(i))}i=1NP\mathcal{D}^{N}=\{(\hat{s}^{(i)},g^{(i)},P^{(i)})\}_{i=1}^{N_{P}} consists of state representations s^(i)\hat{s}^{(i)}, subgoals g(i)g^{(i)}, and agent poses P(i)P^{(i)} at the time of taking the interaction action corresponding to subgoal g(i)g^{(i)}. The state representations s^(⋅)\hat{s}^{(\cdot)} in datasets Dg\mathcal{D}^{g} and DN\mathcal{D}^{N} are constructed using the observation model (Section 4.2), but using ground-truth depth and segmentation images.

We train the perception models on DP\mathcal{D}^{P} and the πH\pi^{H} on Dg\mathcal{D}^{g} to predict the kk-th subgoal by optimizing cross-entropy losses. We use DN\mathcal{D}^{N} to train the navigation model NavModel by optimizing a cross-entropy loss for positions and yaw angles, and an L2 loss for the pitch angle.

Experimental Setup

We evaluate our approach on the ALFRED benchmark. It contains 108 training scenes, 88/4 validation seen/unseen scenes, and 107/8 test seen/unseen scenes. There are 21,023 training tasks, 820/821 validation seen/unseen tasks, and 1533/1529 test seen/unseen tasks. Each task is specified with a high-level natural language instruction. The goal of the agent is to map raw RGB observations to actions to complete the task. ALFRED also provides detailed low-level step-by-step instructions, which simplify the reasoning process. We do not use these instructions for training or evaluation. We collect a training dataset of language-demonstration pairs for learning (Section 5). To extract subgoal sequences, we label each interaction action at=(typet,maskt)a_{t}=(\texttt{type}_{t},\texttt{mask}_{t}) and any preceding navigation actions with a single subgoal of type=typet\texttt{type}=\texttt{type}_{t}. We compute the subgoal argument class argC\texttt{arg}^{C} and 3D mask argM\texttt{arg}^{M} labels from the first-person mask maskt\texttt{mask}_{t}, and ground truth segmentation and depth. Completing a task requires satisfying several goal conditions. Following the common evaluation , we report two metrics. Success rate (SR) is the fraction of tasks for which all goal conditions were satisfied. Goal condition rate (GC) is the fraction of goal-conditions satisfied across all tasks.

We compare our approach, the Hierarchical Language-conditioned Spatial Model (HLSM) to others on the ALFRED leaderboard that only use the high-level instructions. At the time of writing, the only such published approach is HiTUT , an approach that uses a flat BERT architecture to model a hierarchical task structure without using a spatial representation. See Appendix A.2 for a detailed comparison. We also compare to approaches that use the step-by-step instructions, which puts our method at a disadvantage. Of these, LAV also imposes a hierarchical task structure and uses pre-trained depth and segmentation models, but without using a spatial state representation.

Additionally, we perform ablations and study sensory oracles. To study the observation model, we compare to using sensory oracles for ground truth depth, ground truth segmentation, and both. We report high-level controller ablations that remove the subgoal encoder, language encoder, and state representation encoder as used for predicting subgoal type typek\texttt{type}_{k} and argument class argkC\texttt{arg}^{C}_{k}, while still using the state representation s^t\hat{s}_{t} to predict the subgoal argument mask argkM\texttt{arg}^{M}_{k}. We also study a low-level controller ablation that removes the exploration procedure.

Results

Table 1 shows test and validation results. Our approach achieves state-of-the-art performance across both seen and unseen environments in the setting with only high-level instructions. We achieve 10.04% absolute (98.1% relative) improvement in SR on the test unseen split, and 11.53% absolute (62.6% relative) improvement in SR on the test seen split compared to HiTUT G-only.

Our approach performs competitively even when compared to approaches that also use the low-level step-by-step instructions. We achieve 4.84% absolute (31.4% relative) improvement in SR on the test unseen split compared to ABP . On the test seen split, our approach performs reasonably well, however ABP and LWIT perform better, reflecting potentially stronger scene overfitting.

Tables 2 and 3 show development results. We performed five runs of the full HLSM model on the validation unseen data and found the sample standard deviation of the success rate is 1.1% (absolute). All other results are from a single-evaluation runs. Ground truth depth alone (+ gt depth) does not significantly affect performance. Ground truth segmentation (+ gt seg) provides 6.6%/16.4% absolute improvement in seen/unseen scenes. Using both (+ gt depth, gt seg) provides 11.1%/21.9% absolute improvement and narrows the seen/unseen gap from 11.3% to 0.5%. This points to perception being the main bottleneck in generalization to unseen scenes.

We report high-level controller πH\pi^{H} input encoder ablations. The poor performance without the language encoder reflects task difficulty. Zeroing the input to the subgoal history encoder (but keeping position encodings) does not significantly affect performance, showing that knowing the index of the current subgoal in addition to the state representation is often sufficient. Not using the state representation for predicting subgoal type and argument class gives mixed results in seen and unseen scenes, but without a significant difference in performance. Therefore, predicting the sequence of subgoal types and argument classes (i.e., what to do) is at times possible without spatial reasoning, while grounding the subgoal (i.e., where to do it) requires spatial information. Removing random exploration from πL\pi^{L} does not significantly affect unseen performance.

Figure 4 illustrates the model behavior, showing both successes and common failures. The main failures in valid unseen scenes are due to (1) perception errors that result in missing or extraneous obstacles or picking up wrong objects; (2) insufficiency of random exploration (e.g., not searching inside cabinets); (3) navigation model errors (e.g., blocking objects from opening); (4) subgoal prediction errors (e.g., picking up wrong objects); and (5) lack of state-aware multi-step planning and backtracking. More qualitative results are available in Appendix A.10.

Discussion and Limitations

We showed that a persistent spatial semantic representation enables a hierarchical model to achieve state-of-the-art performance on a challenging instruction-following mobile manipulation task. The main performance bottlenecks include long-horizon exploration, perception generalization to unseen environments, and low-level motion planning for continuous collision avoidance. In terms of learning, incorporating reinforcement learning to train πH\pi^{H}, πL\pi^{L}, and observation model FF jointly could improve robustness. We defined the interface to πL\pi^{L} to be faithful to skills available on physical robots, but the exact implementation of πL\pi^{L} is not the focus of our work. Physical deployment would require changes to πL\pi^{L}, and study on robustness to errors in continuous environments, such as localization or motion uncertainty.

Acknowledgements

This research was supported by ARO W911NF-21-1-0106, a Google Focused Award, and NSF under grant No. 1750499. Animesh Garg is supported in part by CIFAR AI Chair and NSERC Discovery Grant. A significant part of the work was done during the first author’s internship at Nvidia. We thank the authors of ALFRED for maintaining the benchmark. We thank Mohit Shridhar and Jesse Thomason for their help answering our questions, and the anonymous reviewers for their helpful comments.

References

Appendix A Appendix

Are the ALFRED sequential instructions needed during training? The sequential step-by-step instructions are not needed neither during training, nor at test-time.

What has to be done to apply this approach to a real robot? The observation model, high-level controller, state representation, and the interface to the low-level controller together constitute our contribution and are intended to generalize to physical robots. Deployment on a real robot would require an implementation of the low-level controller designed for continuous motion in cluttered environments, and an implementation of the ALFRED interface to enable execution of manipulation actions such as Pickup and ToggleOn. Such physical robot capabilities are subject of ongoing research .

Does this simulated environment result constitute progress towards real-world capabilities? Real-robot operation is the long-term motivation of this work and has been carefully considered in the design of the representation and the approach. However, we do not claim to execute high-level natural language mobile manipulation instructions from raw vision on real robots in unseen environments. To date, such capabilities haven’t been demonstrated even in simulated environments, such as ALFRED. Even in this scenario, though our method achieves better results than existing work, it can still only solve 18.28% of problems in unseen environments.

Would the system scale to physically larger environments? The main bottleneck towards scaling to larger environments is the memory constraint of the semantic memory. While our implementation is likely restricted to interior scenes when using commodity hardware, follow-up work could address this, perhaps using multi-scale representations such as Octress .

How are the state dynamics modeled? Are they assumed to be known or are they learned? The GoTo procedure in the low-level controller is based on a value-iteration network that utilizes a deterministic grid-navigation dynamics model on the internal representation, which is a crude approximation of the dynamics of the RotateLeft, RotateRight, MoveAhead navigation commands. Other than that, the dynamics of the environment are assumed to be completely unknown to the agent, and are not explicitly learned or modeled.

How would localization uncertainty affect the approach? Our representation approach assumes a reliable robot pose estimate. Precisely studying the effects of pose errors would require integration into a system for continuous environments. Intuitively, voxels further away are affected by pose errors more, but may better tolerate it due to being used mainly to decide navigation goals. Voxels close to the agent require more precision as they are used for object instance mask generation, but would be less affected by pose errors. Our voxel map uses a relatively coarse 25cm resolution.

Which model was used to obtain test results? The full HLSM model was evaluated on the test set, even though the model without state representation encoding input to the high-level controller performed better in unseen environments on the validation set.

Why does the ablation without state representation encodings perform better in unseen environments? In unseen environments, the semantic segmentation is erroneous due to the generalization gap, resulting in state encodings that contain errors. This may affect perfromance of the high-level controller that was trained on data with perfect segmentation, and thus with perfect state encodings.

What is the benefit of sampling the subgoals instead of attempting execution from most to least likely in order? There are two types of subgoal execution failures: systematic and random. An example of a systematic failure is the selection of an incorrect subgoal. For example, ToggleOn(FloorLamp) would fail if a FloorLamp does not exist in the environment. An example of a random failure is the low-level controller sampling an interaction pose for which the interaction fails (e.g., Figure 4, row 1, timestep 272). A next-best approach would alleviate a systematic failure, but a sampling approach alleviates both: the systematic failures by trying different subgoals, and random failures by potentially sampling the same subgoal multiple times.

A.2 Extended Related Work

In order for natural language human-robot interfaces to be useful and widely adopted in practice, they should support instructions that are as brief as possible while still being informative of the task, i.e., that adhere to Grice’s maxim of quantity . Following such high-level instructions requires bridging the gap from high-level language to long sequences of low-level actions. This is commonly achieved using temporal abstraction, where subgoals or options abstract over sequences of low-level actions, reducing the effective time horizon of the problem. Most work on instruction following in robotics utilizes temporal abstraction .

Various methods explicitly model correspondences between linguistic constituents in a symbolic instruction representation, environment percepts in the world model, and subgoals (behavior primitives) . This requires the instruction to at least mention each subgoal, and precludes instructions that omit intermediate goals that are expected to be inferred. This limitation can be overcome by directly mapping from language to reward specifications or post-conditions , and then using a planner or learning a task-specific policy to solve for the sequence of actions. Both are difficult in practice. Planning requires a compact, symbolic environment representation with an underlying ontology that is hard to construct for unstructured environments, such as the household environment studied in this work. Policy learning is computationally expensive, and poorly adapts to novel tasks specified in natural language in real-time.

Recently, methods that map language and observations directly to actions using neural networks have seen rising popularity and success on simulated and real-robot navigation, as well as simulated manipulation tasks. Simulated mobile manipulation is a promising next frontier . Representation learning approaches avoid planning, by using a direct sequence-to-sequence formulation and a data-driven approach that theoretically permits mapping arbitrarily terse input text to arbitrarily long action sequences that potentially include any necessary intermediate steps not explicitly mentioned in the text. In practice, however, most research has focuses on relatively detailed step-by-step instructions, sometimes using modelling tools such as attention and progress monitoring to leverage the sequential nature of the instructions.

We learn to follow high-level instructions in an interactive mobile manipulation 3D environment. To bridge the gap between language and actions, we use temporal abstraction, where the high-level controller predicts subgoals that abstract over sequences of actions, and the low-level controller generates actions to fulfil each subgoal. The controllers rely on a spatial-semantic state representation to enable reasoning about what subgoals make progress towards the high-level task, and what actions make progress towards the specific subgoal, given all past sensory observations. The persistent representation enables operation over long time horizons. Using a shared world representation for both the high- and low-level controllers reduces representation engineering effort and error accumulation typically associated with pipeline approaches.

The idea of building maps that combine spatial and semantic information and using them for following natural language instructions has a long history in robotics. Common approaches can be classified into sparse topological and dense grid-based maps.

Walter et al. introduced a sparse semantic graph that combines pose, semantic, and topological information, extracted from sensory observations and speech descriptions along a route. Hemachandra et al. added a spatial map layer, and fused language with other sensory modalities. Hemachandra et al. used these representations for grounding natural language route instructions. More recently, Patki et al. extended this framework to build compact world models specific to the input instruction, and Patki et al. enabled supporting previously unseen environments. This class of sparse topological maps are well suited for probabilistic language grounding from symbolic representations.

Dense grid-based 2D semantic maps are suited for downstream processing using learned neural network modules, and have been used in modular neural network approaches for language grounding . Saha et al. used a grid-based spatial representation and a map filtering method, showing promising early results on a subset of the ALFRED dataset. We extend this line of work to 3D voxel maps, add explicit tracking of occupancy and observability, and maintain the representation through time to facilitate grounding high-level language over long time horizons. Our dense representation has a number of advantages. First, it is easy to build in real-time from RGBD data using segmentation models and geometric operations. Second, it captures structures found in indoor environments, such as L-shaped countertops or kitchen islands with sinks that are hard to represent topologically. Third, it encodes spatial object relationships without requiring an ontology of spatial relations, or even tracking of object instances. The main limitation of our approach is a memory footprint that scales with the physical size of the environment, making it less suited for outdoor or field applications. Follow-up work could address this limitation, for example by using multi-scale representations such as octrees .

We provide a detailed technical comparison between our approach and HiTUT [Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring; 47], our main point of comparison.

Both our approach and HiTUT use a hierarchical task decomposition of goals into sequences of subgoals, and subgoals into sequences of actions. The set of subgoals assumed by our approach and HiTUT have differences. HiTUT has an additional subgoal GoTo(Location), while we view any navigation as a means to an end of a manipulation subgoal, and therefore do not have an explicit GoTo subgoal. HiTUT additionally has subgoals for Clean and Heat, (e.g., Clean(Obj) usually abstracts over the sequence Put(Sink), ToggleOn(Faucet), ToggleOff(Faucet), Pickup(Obj)), while our high-level policy would have to predict this entire sequence.

In terms of the model architecture, we use a hierarchical model with high-level and low-level controllers to mimic the task structure. In contrast, HiTUT uses a flat transformer model to jointly solve high-level subgoal planning and low-level action prediction. One of their main contributions is showing how a flat transformer model can be used to model a hierarchical task structure. The benefit of our hierarchical model decomposition in combination with a shared spatial state representation is its ability to solve low-level navigation and manipulation problems with specialized modules, while avoiding the representational error accumulation and representation engineering issues typically associated with modular pipeline approaches.

In terms of inference, HiTUT and our approach both sample subgoals one at a time, dynamically responding to changes in environment and execution. Both approaches perform backtracking to previous subgoals upon subgoal failure.

In terms of perception, our approach requires a pre-trained segmentation model, while HiTUT requires a pre-trained object detection model to generate object entity information that is fed into the transformer.

A.3 Observation Model Details

At time tt, during the perception step, we predict first-person semantic segmentation ItSI^{S}_{t} and depth 0ptt0pt_{t} from the observation ot=(It,Pt,vtS,L)o_{t}=(I_{t},P_{t},v^{S}_{t},L), from the RGB image ItI_{t} with neural network models pre-trained in the ALFRED environment. Each pixel [ItS](u,v)[I^{S}_{t}]_{(u,v)} at coordinates (u,v)(u,v) is a distribution over CC object classes. Likewise, [0ptt](u,v)[0pt_{t}]_{(u,v)} is a distribution over BB uniformly spaced depth bins {0,ΔD,2ΔD,…,(B−1)ΔD}\{0,\Delta_{D},2\Delta_{D},\dots,(B-1)\Delta_{D}\}, where ΔD\Delta_{D} is a depth resolution. In early experiments, we observed that ΔD\Delta_{D} should be less than 50% of the voxel size. We use ΔD=0.1m\Delta_{D}=0.1m, B=50B=50, and voxel size of 0.25m0.25m. We also heuristically compute a binary mask MtDM^{D}_{t} that indicates which pixels have confident depth readings. We allow more confidence slack in pixels that correspond to the current subgoal argument argtC\texttt{arg}^{C}_{t} according to ItSI^{S}_{t}. The mask MtDM^{D}_{t} is used in the projection step to discard points (x,y,z)(x,y,z) that correspond to pixels (u,v)(u,v) for which [MD](u,v)=0[M^{D}]_{(u,v)}=0. The mask computation is:

If the agent is currently holding an object (i.e. ∑i[[vtS](i)]\sum_{i}[[v^{S}_{t}]_{(i)}] > 0), we also discard points closer than 0.7m0.7m to the camera to make sure that the object in the agent inventory does not get added to the voxel map.

We use custom models based on the U-Net architecture for depth and segmentation networks. The architecture is illustrated in Figure 5. It consists of a cascade of five downscale blocks followed by five upscale blocks with skip-connections. Each block includes two convolutions, two leakyReLU activations, and an instance normalization layer. The upscale blocks contain a 2x spatial upscaling operation. We found that training a separate network for depth and segmentation worked better than sharing one network for both tasks.

A.4 Model Execution Flow

Algorithm 1 describes the execution flow. At time t=0t=0 the initial observation o0o_{0} is received. At each timesep, we update the state representation s^t\hat{s}_{t} (Line 6). If needed, we sample a new subgoal gkg_{k} from πH\pi^{H} (Line 9), and then sample an action ata_{t} from πL\pi^{L}. If ata_{t} is aPASSa_{\rm PASS}, we increment subgoal counter kk (Line 14). If it is aFAILa_{\rm FAIL}, we discard the current subgoal kk (Line 16). We repeat Lines 9–16 until an executable action ata_{t} is sampled. We execute ata_{t}, receive the next observation (Line 18), and proceed to the next timestep. The episode ends when the subgoal gSTOPg_{\rm STOP} is sampled (Line 11) or the horizon TmaxT_{max} is exceeded (Line 19).

A.5 High-Level Controller Details

Subgoals are predicted periodically. Let gk=(typek,argkC,argkM)g_{k}=(\texttt{type}_{k},\texttt{arg}^{C}_{k},\texttt{arg}^{M}_{k}) be the kk-th subgoal predicted at time tt. Predicting the subgoal type typek\texttt{type}_{k} and the argument class argkC\texttt{arg}^{C}_{k} is described in the main paper (Section 4.3). This section provides further details of Refiner, the model we use to generate argkM\texttt{arg}^{M}_{k}. The mask refiner Refiner has four inputs:Errata: Equation 4 in the main paper is missing [VtS](argkC)[V^{S}_{t}]_{(\texttt{arg}^{C}_{k})} and PtP_{t} arguments to the Refiner. (a) a spatial feature map xtego∈N×W×L\mathbf{x}^{ego}_{t}\in^{N\times W\times L} oriented in the agent egocentric reference frame; (b) [VtS](argkC)∈W×L×H[V^{S}_{t}]_{(\texttt{arg}^{C}_{k})}\in^{W\times L\times H}, a 3D mask indicating all voxels that contain objects of class argkC\texttt{arg}^{C}_{k} in the voxel map VtSV^{S}_{t}; (c) the agent’s pose PtP_{t}; and (d) a vector representation of the instruction ϕL\phi^{L}. It outputs a 3D mask argkM∈W×L×H\texttt{arg}^{M}_{k}\in^{W\times L\times H} that identifies the subgoal argument object. Formally, the computation is:⊗\otimes is an operation that multiplies a W×LW\times L matrix by a W×L×HW\times L\times H tensor to obtain a W×L×HW\times L\times H tensor

where AlloTransform transforms a spatial 2D map from an egocentric to the global reference frame, and \textscLingUNetm\textsc{LingUNet}_{m} is a language-conditioned image-to-image encoder-decoder . The architecture of \textscLingUNetm\textsc{LingUNet}_{m} is illustrated in Figure 6.

A.6 Low-Level Controller Details

We describe the implementation of each of the low-level controller procedures. This implementation is not the focus of this paper, and could be improved or replaced with other algorithms. Some of the procedures cause actions in the AI2Thor environment, others simply process data to pass between procedures.

The procedures are NavigateTo, SampleExplorationPosition, SampleInteractionPose, and InteractMask. The low-level controller receives the subgoal gkg_{k}, and follows a pre-specified execution flow across multiple timesteps to complete it. The execution flow (Figure 7) consists of an exploration and interaction phase. In the exploration phase, we perform a 360° rotation by generating a sequence of three RotateLeft actions to observe the environment and add information to the semantic map. If the semantic map indicates that no object of type argkC\texttt{arg}^{C}_{k}, the action argument, is observed, we explore the environment by sampling a position (x,y)=SampleExplorationPosition(s^t)(x,y)=\texttt{SampleExplorationPosition}(\hat{s}_{t}), navigating there using NavigateTo(x,y,s^t)\texttt{NavigateTo}(x,y,\hat{s}_{t}), and performing another 360° rotation. We repeat this process until a voxel in VtSV^{S}_{t} contains the class argkC\texttt{arg}^{C}_{k} with >>50% probability, at which point we move on to the interaction phase. In the interaction phase, we sample an interaction pose (x,y,ωy,ωp)=SampleInteractionPose(s^t,gk)(x,y,\omega_{y},\omega_{p})=\texttt{SampleInteractionPose}(\hat{s}_{t},g_{k}), invoke NavigateTo(x,y,s^t)\texttt{NavigateTo}(x,y,\hat{s}_{t}) to reach the position (x,y), and rotate according to yaw and pitch angles (ωy,ωp)(\omega_{y},\omega_{p}). Finally, we generate the egocentric interaction action mask maskt=InteractMask(s^t,argkM)\texttt{mask}_{t}=\texttt{InteractMask}(\hat{s}_{t},\texttt{arg}^{M}_{k}), and execute the interaction action (typek,maskt)(\texttt{type}_{k},\texttt{mask}_{t}) in the ALFRED environment. We output aPASSa_{\rm PASS} or aFAILa_{\rm FAIL} depending if the interaction action has succeeded, and pass control back to the high-level controller to sample the next subgoal.

At time tt, the NavigateTo procedure maps a 2D navigation goal position (x,y)(x,y) and the state representation s^t\hat{s}_{t} to one of the actions: {\textscRotateLeft,\textscRotateRight,\textscMoveAhead,aSTOP}\{\textsc{RotateLeft},\textsc{RotateRight},\textsc{MoveAhead},a_{\rm STOP}\}. We implement it with a Value Iteration Network [VIN; 57] that solves a 2D grid-MDP to predict navigation actions using fast GPU-accelerated convolution and max-pooling operations. The VIN parameters are pre-defined, and not learned. Other motion planners such as A∗ could be used as well.

The reward function assigns different rewards for visiting states with different attributes:

Obstacle states receive reward −0.9-0.9, Goal states receive reward 1.01.0, and Unobserved states receive reward −0.02-0.02. Taking the Stop action in any state gives reward 0.0010.001, which has the effect of the agent stopping in unsolvable cases. We use the VIN iteratively for NvinN^{vin} iterations, and predict an action avin=arg⁡max⁡avin∈Avin(Qvin(stvin,avin)a^{vin}=\arg\max_{a^{vin}\in\mathcal{A}^{vin}}(Q^{vin}(s^{vin}_{t},a^{vin}). We map from the VIN action avina^{vin} to a single valid AI2Thor navigation action using a deterministic mapping (Table 4).

A.6.2 SampleExplorationPosition

The SampleExplorationPosition procedure maps a state representation s^t\hat{s}_{t} to a discrete 2D position pexplore=(x,y)p^{\rm explore}=(x,y). Let Ps\mathcal{P}_{s} be the set of 2D positions corresponding to voxel centroids in the voxel map along the horizontal axes, and the ground set Pg\mathcal{P}_{g} as the set of all unoccupied positions that have the class Floor or Rug in at least one voxel. A position is unoccupied if all voxels in the height range [0,1.75m][0,1.75m] are free of obstacles. We define a frontier set Pf\mathcal{P}_{f} as the set of all positions Pg\mathcal{P}_{g} for which at least one immediately neighboring position contains zero observed voxels. If Pf\mathcal{P}_{f} is non-empty, we sample the position pexplorep^{\rm explore} uniformly at random from Pf\mathcal{P}_{f}. Otherwise, we sample pexplorep^{\rm explore} uniformly at random from Pg\mathcal{P}_{g}.

A.6.3 SampleInteractionPose

The SampleInteractionPose procedure maps the state representation s^t\hat{s}_{t} and subgoal gk=(typek,argkC,argkM)g_{k}=(\texttt{type}_{k},\texttt{arg}^{C}_{k},\texttt{arg}^{M}_{k}) to a pose P=(x,y,ωy,ωp)P=(x,y,\omega_{y},\omega_{p}), where (x,y)(x,y) is a discrete 2D position, ωy\omega_{y} is the agent yaw angle, and ωp\omega_{p} is the agent camera pitch angle. The pose is predicted such that upon reaching it, the interaction action of type typek\texttt{type}_{k} is likely to succeed on the object of class argkC\texttt{arg}^{C}_{k} at location identified by the mask argkM\texttt{arg}^{M}_{k}.

The network NavModel is based on the LingUNet architecture (Figure 6):

where Afford is an affordance feature map (Section 4.1), Linear is a linear layer with bias, \textscLutT\textsc{Lut}_{T} and \textscLutC\textsc{Lut}_{C} are embedding lookup tables, and [⋅;⋅][\cdot;\cdot] is a vector concatenation.

A.6.4 InteractionMask

The InteractionMask procedure maps a state representation s^t=(VtS,VtO,vtS,Pt)\hat{s}_{t}=(V^{S}_{t},V^{O}_{t},v^{S}_{t},P_{t}), the most recent RGB observation ItI_{t}, the most recent predicted segmentation ItSI^{S}_{t}, and a subgoal gk=(typek,argkC,argkM)g_{k}=(\texttt{type}_{k},\texttt{arg}^{C}_{k},\texttt{arg}^{M}_{k}) to a 0-1 valued mask maskt∈H×W\texttt{mask}_{t}\in^{H\times W} that identifies the interaction object in the first-person view observation. The interaction mask maskt\texttt{mask}_{t} is in the format expected by ALFRED. Formally, it is computed in three steps:

where PinholeCam projects the 0-1 valued 3D voxel map argkM\texttt{arg}^{M}_{k} to the agent’s camera plane according to the pose PtP_{t}. The mask masktA\texttt{mask}_{t}^{A} is an egocentric 0-1 valued mask that identifies all objects of class argkC\texttt{arg}^{C}_{k} in the image ItI_{t}. The masktB\texttt{mask}_{t}^{B} is an egocentric 0-1 valued mask that identifies the voxels argkM\texttt{arg}^{M}_{k}. For each pixel (u,v)(u,v), the value [masktB](u,v)[\texttt{mask}_{t}^{B}]_{(u,v)} is the maximum of all values [argkM](x,y,z)[\texttt{arg}^{M}_{k}]_{(x,y,z)} over voxels with coordinates (x,y,z)(x,y,z) that the ray cast from the camera through the pixel (u,v)(u,v) intersects with. The final mask maskt\texttt{mask}_{t} is a 0-1 valued mask that identifies not only the correct object class, but also the correct instance according to the voxel mask argkM\texttt{arg}^{M}_{k}.

A.7 Additional Learning Details

As described in Section 5, we use a perception dataset DP\mathcal{D}^{P} for training depth and segmentation models. The dataset DP={([I](i),[0pt](i),[IS](i)}i=1NP\mathcal{D}^{P}=\{([I]^{(i)},[0pt]^{(i)},[I^{S}]^{(i)}\}_{i=1}^{N_{P}} includes RGB images [I](i)[I]^{(i)} with ground truth depth [0pt](i)[0pt]^{(i)} and segmentation [IS](i)[I^{S}]^{(i)}. The ground truth depth [0pt](i)[0pt]^{(i)} at each pixel (u,v)(u,v) is a distribution [0pt]((u,v))(i)[0pt]^{(i)}_{((u,v))} over BB depth bins, where 100% of the probability mass is assigned to the bin containing the reference depth value. The ground truth segmentation [IS](i)[I^{S}]^{(i)} is likewise at each pixel (u,v)(u,v) a one-hot vector indicating the object class that pixel belongs to.

The ALFRED dataset consists of 108 different training scenes, where each scene has a fixed furniture and light fixtures. Observations are highly correlated within each scene, which greatly reduces the effective size of the perception dataset and hurts generalization to unseen scenes. We use a custom segmentation-aware data augmentation strategy that increases the diversity of RGB observations.

During training, we apply Augment with 50% probability to each training example. Additionally, with 50% probability we perform a horizontal flip.

A.8 Additional Experimental Details

We collect a training dataset of language-demonstration pairs as described in Section 5. The demonstrations in ALFRED typically navigate while looking down at the floor, likely a side-effect of the PDDL planner that had access to the world state during data generation, and as such has no need to explore or observe the visual environment. We modify the demonstration trajectories to get more informative first-person observations. First, we insert four RotateLeft actions at the start of each trajectory. Second, we maintain a nominal camera pitch angle of 30°during navigation, by inserting LookDown and LookUp actions before and after every interaction action. We discard trajectories for which these modifications cause failures. These modifications result in observations that are more useful for learning and constructing our persistent spatial representation.

A.9 Hyperparameters

Table 5 shows hyperparameter values. The hyperparameters were hand-tuned on the validation unseen split.

A.10 Additional Results

Additional qualitative results are available at: https://hlsm-alfred.github.io/.

A successful example of task execution is available at: https://drive.google.com/file/d/1APKe3cR_-vliyU2elT5Un30w7PvkEdYs/view?usp=sharing

A failed example of task execution is available at: https://drive.google.com/file/d/1j8BJ_ALoXGyf8a-IOkmQAg38awSWYt6f/view?usp=sharing