LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Dhruv Shah, Blazej Osinski, Brian Ichter, Sergey Levine
Introduction
One of the central challenges in robotic learning is to enable robots to perform a wide variety of tasks on command, following high-level instructions from humans. This requires robots that can understand human instructions, and are equipped with a large repertoire of diverse behaviors to execute such instructions in the real world. Prior work on instruction following in navigation has largely focused on learning from trajectories annotated with textual instructions . This enables understanding of textual instructions, but the cost of data annotation impedes wide adoption. On the other hand, recent work has shown that learning robust navigation is possible through goal-conditioned policies trained with self-supervision. These utilize large, unlabeled datasets to train vision-based controllers via hindsight relabeling . They provide scalability, generalizability, and robustness, but usually involve a clunky mechanism for goal specification, using locations or images. In this work, we aim to combine the strengths of both approaches, enabling a self-supervised system for robotic navigation to execute natural language instructions by leveraging the capabilities of pre-trained models without any user-annotated navigational data. Our method uses these models to construct an “interface” that humans can use to communicate desired tasks to robots. This system enjoys the impressive generalization capabilities of the pre-trained language and vision-language models, enabling the robotic system to accept complex high-level instructions.
Our main observation is that we can utilize off-the-shelf pre-trained models trained on large corpora of visual and language datasets — that are widely available and show great few-shot generalization capabilities — to create this interface for embodied instruction following. To achieve this, we combine the strengths of two such robot-agnostic pre-trained models with a pre-trained navigation model. We use a visual navigation model (VNM: ViNG ) to create a topological “mental map” of the environment using the robot’s observations. Given free-form textual instructions, we use a pre-trained large language model (LLM: GPT-3 ) to decode the instructions into a sequence of textual landmarks. We then use a vision-language model (VLM: CLIP ) for grounding these textual landmarks in the topological map, by inferring a joint likelihood over the landmarks and nodes. A novel search algorithm is then used to maximize a probabilistic objective, and find a plan for the robot, which is then executed by VNM.
Our primary contribution is Large Model Navigation, or LM-Nav, an embodied instruction following system that combines three large independently pre-trained models — a self-supervised robotic control model that utilizes visual observations and physical actions (VNM), a vision-language model that grounds images in text but has no context of embodiment (VLM), and a large language model that can parse and translate text but has no sense of visual grounding or embodiment (LLM) — to enable long-horizon instruction following in complex, real-world environments. We present the first instantiation of a robotic system that combines the confluence of pre-trained vision-and-language models with a goal-conditioned controller, to derive actionable plans without any fine-tuning in the target environment. Notably, all three models are trained on large-scale datasets, with self-supervised objectives, and used off-the-shelf with no fine-tuning — no human annotations of the robot navigation data are necessary to train LM-Nav. We show that LM-Nav is able to successfully follow natural language instructions in new environments over the course of 100s of meters of complex, suburban navigation, while disambiguating paths with fine-grained commands.
Related Work
Early works in augmenting navigation policies with natural language commands use statistical machine translation to discover data-driven patterns to map free-form commands to a formal language defined by a grammar . However, these approaches tend to operate on structured state spaces. Our work is closely inspired by methods that instead reduce this task to a sequence prediction problem . Notably, our goal is similar to the task of VLN — leveraging fine-grained instructions to control a mobile robot solely from visual observations .
However, most recent approaches to VLN use a large dataset of simulated trajectories — over 1M demonstrations — annotated with fine-grained language labels in indoor and driving scenarios , and rely on sim-to-real transfer for deployment in simple indoor environments . However, this necessitates building a photo-realistic simulator resembling the target environment, which can be challenging for unstructured environments, especially for the task of outdoor navigation. Instead, LM-Nav leverages free-form textual instructions to navigate a robot in complex, outdoor environments without access to any simulation or any trajectory-level annotations.
Recent progress in using large-scale models of natural language and images trained on diverse data has enabled applications in a wide variety of textual , visual , and embodied domains . In the latter category, Shridhar et al. , Khandelwal et al. and Jang et al. fine-tune embeddings from pre-trained models on robot data with language labels, Huang et al. assume that the low-level agent can execute textual instructions (without addressing control), and Ahn et al. assumes that the robot has a set of text-conditioned skills that can follow atomic textual commands. All of these approaches require access to low-level skills that can follow rudimentary textual commands, which in turn requires language annotations for robotic experience and a strong assumption on the robot’s capabilities. In contrast, we combine these pre-trained vision and language models with pre-trained visual policies that do not use any language annotations without fine-tuning these models in the target environment or for the task of VLN.
Data-driven approaches to vision-based mobile robot navigation often use photorealistic simulators or supervised data collection to learn goal-reaching policies directly from raw observations. Self-supervised methods for navigation instead can use unlabeled datasets of trajectories by automatically generating labels using onboard sensors and hindsight relabeling. Notably, such a policy can be trained on large, diverse datasets and generalize to previously unseen environments . Being self-supervised, such policies are adept at navigating to desired goals specified by GPS locations or images, but are unable to parse high-level instructions such as free-form text. LM-Nav uses self-supervised policies trained in a large number of prior environments, augmented with pre-trained vision and language models for parsing natural language instructions, and deploys them in novel real-world environments without any fine-tuning.
Preliminaries
LM-Nav consists of three large, pre-trained models for processing language, associating images with language, and visual navigation.
Large language models are generative models based on the Transformer architecture , trained on large corpora of internet text. LM-Nav uses the GPT-3 LLM , to parse textual instructions into a sequence of landmarks.
Vision-and-language models refer to models that can associate images and text, e.g. image captioning, visual question-answering, etc. . We use the CLIP VLM , a model that jointly encodes images and text into an embedding space that allows it to determine how likely some string is to be associated with a given image. We can jointly encode a set of landmark descriptions obtained from the LLM and a set of images to obtain their VLM embeddings (see Fig. 3). Computing the cosine similarity between these embeddings, followed by a softmax operation results in probabilities , corresponding to the likelihood that image corresponds to the string . LM-Nav uses this probability to align landmark descriptions with images.
Visual navigation models learn navigation behavior and navigational affordances directly from visual observations , associating images and actions through time. We use the ViNG VNM , a goal-conditioned model that predicts temporal distances between pairs of images and the corresponding actions to execute (see Fig. 3). This provides an interface between images and embodiment. The VNM serves two purposes: (i) given a set of observations in the target environment, the distance predictions from the VNM can be used to construct a topological graph that represents a “mental map” of the environment; (ii) given a “walk”, comprising of a sequence of connected subgoals to a goal node, the VNM can navigate the robot along this plan. The topological graph is an important abstraction that allows a simple interface for planning over past experience in the environment and has been successfully used in prior work to perform long-horizon navigation . To deduce connectivity in , we use a combination of learned distance estimates, temporal proximity (during data collection), and spatial proximity (using GPS measurements). For every connected pair of vertices , we assign this distance estimate to the corresponding edge weight . For more details on the construction of this graph, see Appendix B.
LM-Nav: Instruction Following with Pre-Trained Models
While we can use a variety of traversability likelihood functions, a simple choice is to use a discounted Markovian model, where the discount models the probability of exiting at each time step, leading to a termination probability of at each step, and a probability of reaching given by , where is the estimated number of time steps the robot needs to travel from to , which is predicted by the VNM. While other traversability likelihoods could also be used, this choice is a convenient consequence of goal-conditioned reinforcement learning formulations , and thus, the log-likelihood corresponds to . We can use these likelihoods to derive the probability that a given sequence can be traversed successfully, which we denote with the auxiliary Bernoulli random variable (i.e., implies that was traversed successfully):
The full likelihood used for planning is then given by:
2 Parsing Free-Form Textual Instructions
The user specifies the route they want the robot to take using natural language, while the objective above is defined in terms of a sequence of desired landmarks. To extract this sequence from the user’s natural language instruction we employ a standard large language model, which in our prototype is GPT-3 . We used a prompt with 3 examples of correct landmarks’ extractions, followed up by the description to be translated by the LLM. Such an approach worked for the instructions that we tested it on. Examples of instructions together with landmarks extracted by the model can be found in Fig. 4. The appropriate selection of the prompt, including those 3 examples, was required for more nuanced cases. For details of the “prompt engineering” please see Appendix A.
3 Visually Grounding Landmark Descriptions
4 Graph Search for the Optimal Walk
represents the maximal value of for a walk ending in that visited the landmarks up to index . The base case visits none of the landmarks, and its value of is simply equal to minus the length of shortest path from the starting node . For we have:
The base case for DP is to compute . Then, in each step of DP we compute . This computation resembles the Dijkstra algorithm (). In each iteration, we pick the node with the largest value of and update its neighbors based on the Eqn. 5. Algorithm 1 summarizes this search process. The result of this algorithm is a walk that maximizes the probability of successfully carrying out the instruction. Such a walk can be executed by VNM, using its action estimates to sequentially navigate to these nodes.
System Evaluation
We now describe our experiments deploying LM-Nav in a variety of outdoor settings to follow high-level natural language instructions with a small ground robot. For all experiments, the weights of LLM, VLM, and VNM are frozen — there is no fine-tuning or annotation in the target environment. We evaluate the complete system, as well as the individual components of LM-Nav, to understand its strengths and limitations. Our experiments demonstrate the ability of LM-Nav to follow high-level instructions, disambiguate paths, and reach goals that are up to 800m away.
We implement LM-Nav on a Clearpath Jackal UGV platform (see Fig. 1(right)). The sensor suite consists of a 6-DoF IMU, a GPS unit for approximate localization, wheel encoders for local odometry, and front- and rear-facing RGB cameras with a field-of-view for capturing visual observations and localization in the topological graph. The LLM and VLM queries are pre-computed on a remote workstation and the computed path is commanded to the robot wirelessly. The VNM runs on-board and only uses forward RGB images and unfiltered GPS measurements.
2 Following Instructions with LM-Nav
In each evaluation environment, we first construct the graph by manually driving the robot and collecting image and GPS observations. The graph is constructed automatically using the VNM from this data, and in principle such data could also be obtained from past traversals, or even with autonomous exploration methods . Once the graph is constructed, the robot can carry out instructions in that environment. We tested our system on 20 queries, in environments of varying difficulty, corresponding to a total combined length of over 6 km. Instructions include a set of prominent landmarks in the environment that can be identified from the robot’s observations, e.g. traffic cones, buildings, stop signs, etc.
Fig. 4 shows qualitative examples of the path taken by the robot. Note that the overhead image and spatial localization of the landmarks is not available to the robot and is shown for visualization only. In Fig. 4(a), LM-Nav is able to successfully localize the simple landmarks from its prior traversal and find a short path to the goal. While there are multiple stop signs in the environment, the objective in Eqn. 2 causes the robot to pick the correct stop sign in context, so as to minimize overall travel distance. Fig. 4(b) highlights LM-Nav’s ability to parse complex instructions with multiple landmarks specifying the route — despite the possibility of a shorter route directly to the final landmark that ignores instructions, the robot finds a path that visits all of the landmarks in the correct order.
Disambiguation with instructions. Since the objective of LM-Nav is to follow instructions, and not merely to reach the final goal, different instructions may lead to different traversals. Fig. 5 shows an example where modifying the instruction can disambiguate multiple paths to the goal. Given the shorter prompt (blue), LM-Nav prefers the more direct path. On specifying a more fine-grained route (magenta), LM-Nav takes an alternate path that passes a different set of landmarks.
Missing landmarks. While LM-Nav is effective at parsing landmarks from instructions, localizing them on the graph, and finding a path to the goal, it relies on the assumption that the landmarks (i) exist in the environment, and (ii) can be identified by the VLM. Fig. 4(c) illustrates a case where the executed path fails to visit one of the landmarks — a fire hydrant — and takes a path that goes around the top of the building rather than the bottom. This failure mode is attributed to the the inability of the VLM to detect a fire hydrant from the robot’s observations. On independently evaluating the efficacy of our the VLM at retrieving landmarks (see Sec. 5.4), we find that despite being the best off-the-shelf model for our task, CLIP is unable to retrieve a small number of “hard” landmarks, including fire hydrants and cement mixers. In many practical cases, the robot is still successful in finding a path that visits the remaining landmarks.
3 Quantitative Analysis
To quantify the performance of LM-Nav, we introduce some performance metrics. A walk produced by the graph search is considered successful, if (1) it matches the path intended by the user or (2) if the landmark images extracted by the search algorithm indeed contain said landmarks (i.e. if the produced path is valid, if not identical). The fraction of successful walks produced by the search algorithm is defined as planning success. For a successfully executed plan in the real world, we define efficiency as:
The second term — corresponding to the optimality of the executed route — is clipped at a maximum of 1 to account for occasional cases when the VNM executes a shorter, more direct path than the user intended. For a set of queries, we report the average efficiency over successful experiments. The planning efficiency is analogously defined as:
Yet another metric — number of disengagements — counts the average number of human interventions required per experiment, due to unsafe maneuvers like collisions or falling off a curb, etc.
Table 1 summarizes the quantitative performance of the system over 20 instructions. LM-Nav can consistently follow the instructions in 85% of the experiments, without collisions or disengagements (an average of 1 intervention per 6.4km of traversals). Comparing to baselines where the navigation model has been ablated (described in Sec. 5.4), LM-Nav performs consistently better in executing efficient, collision-free paths to the goal. In all the unsuccessful experiments, the failure can be attributed to the inability of the planning stage — the search algorithm is unable to visually localize certain “hard” landmarks in the graph — leading to incomplete execution of the instructions. Investigating these failure modes suggests that the most critical component of our system is the ability of VLM to detect unfamiliar landmarks, e.g. a fire hydrant, and in challenging lighting conditions, e.g. underexposed images.
4 Dissecting LM-Nav
To understand the influence of each of the components of LM-Nav, we conduct experiments to evaluate these components in isolation. For more details about these experiments, see Appendix C.
To evaluate the performance of LLM candidates in parsing instructions into an ordered list of landmarks, we compare GPT-3 (used by LM-Nav) to other state-of-the-art pre-trained language models — fairseq , GPT-J-6B , and GPT-NeoX-20B — as well as a simple baseline using spaCy NLP library that extracts base noun phrases, followed by filtering. In Table 2 we report the average extraction success for all the methods on the prompts used in Section 5.3. GPT-3 significantly outperforms other models, owing to its superior representation capabilities and in-context learning . The noun chunking performs surprisingly reliably, correctly solving many simple prompts. For further details on these experiments, see Appendix C.2.
To evaluate the VLM’s ability to ground these textual landmarks in visual observations, we set up an object detection experiment. Given an unlabeled image from the robot’s on-board camera and a set of textual landmarks, the task is to retrieve the corresponding label. We run this experiment on a set of 100 images from the environments discussed earlier, and a set of 30 commonly-occurring landmarks. These landmarks are a combination of the landmarks retrieved by the LLM in our experiments from Sec. 5.2 and manually curated ones. We report the detection successful if any of the top 3 predictions adhere to the contents of the image. We compare the retrieval success of our VLM (CLIP) with some credible object detection alternatives — Faster-RCNN-FPN , a state-of-the-art object detection model pre-trained on MS-COCO , and ViLD , an open-vocabulary object detector based on CLIP and Mask-RCNN . To evaluate against the closed-vocabulary baseline, we modify the setup by projecting the landmarks onto the set of MS-COCO class labels. We find that CLIP outperforms baselines by a wide margin, suggesting that its visual model transfers very well to robot observations (see Table 3). Despite deriving from CLIP, ViLD struggles with detecting complex landmarks like “manhole cover” and “glass building”. Faster-RCNN is unable to detect common MS-COCO objects like “traffic light”, “person” and ”stop sign”, likely due to the on-board images being out-of-distribution for the model.
To understand the importance of the VNM, we run an ablation experiment of LM-Nav without the navigation model. Using GPS-based distance estimates and a naïve straight line controller between nodes of the topological graph. Table 1 summarizes these results — without VNM’s ability to reason about obstacles and traversability, the system frequently runs into small obstacles such as trees and curbs, resulting in failure. Fig. 6 illustrates such a case — while such a controller works well on open roads, it cannot reason about connectivity around buildings or obstacles, and results in collisions with a curb, a tree, and a wall in 3 individual attempts. This illustrates that using a learned policy and distance function from the VNM is critical for enabling LM-Nav to navigate in complex environments without collisions.
Lastly, to understand the importance of the two components of the graph search objective (Eqn. 3), we ran a set of ablations where the graph search only depends on , i.e. Max Likelihood Planning, which only picks the most likely landmark without reasoning about topological connectivity or traversability. Table 4 shows that such a planner suffers greatly in the form of efficiency, because it does not utilize the spatial organization of nodes and their connectivity. For more details on these experiments, and qualitative examples, see Appendix C.
Discussion
We presented Large Model Navigation, or LM-Nav, a system for robotic navigation from textual instructions that can control a mobile robot without requiring any user annotations for navigational data. LM-Nav combines three pre-trained models: the LLM, which parses user instructions into a list of landmarks, the VLM, which estimates the probability that each observation in a “mental map” constructed from prior exploration of the environment matches these landmarks, and the VNM, which estimates navigational affordances (distances between landmarks) and robot actions. Each model is pre-trained on its own dataset, and we show that the complete system can execute a variety of user-specified instructions in real-world outdoor environments — choosing the correct sequence of landmarks through a combination of language and spatial context — and handle mistakes (such as missing landmarks). We also analyze the impact of each pre-trained model on the full system.
Limitations and future work. The most prominent limitation of LM-Nav is its reliance on landmarks: while the user can specify any instruction they want, LM-Nav only focuses on the landmarks and disregards any verbs or other commands (e.g., “go straight for three blocks” or “drive past the dog very slowly”). Grounding verbs and other nuanced commands is an important direction for future work. Additionally, LM-Nav uses a VNM that is specific to outdoor navigation with the Clearpath Jackal robot. An exciting direction for future work would be to design a more general “large navigation model” that can be utilized broadly on any robot, analogous to how the LLM and VLM handle any text or image. However, we believe that in its current form, LM-Nav provides a simple and attractive prototype for how pre-trained models can be combined to solve complex robotic tasks, and illustrates that these models can serve as an “interface” to robotic controllers that are trained without any language annotations. One of the implications of this result is that further progress on self-supervised robotic policies (e.g., goal-conditioned policies) can directly benefit instruction following systems. More broadly, understanding how modern pre-trained models enable effective decomposition of robotic control may enable broadly generalizable systems in the future, and we hope that our work will serve as a step in this direction.
This research was supported by DARPA Assured Autonomy, DARPA RACER, Toyota Research Institute, ARL DCIST CRA W911NF-17-2-0181, and AFOSR. BO was supported by the Fulbright Junior Research Award granted by the Polish-U.S. Fulbright Commission. We would like to thank Alexander Toshev for pivotal discussions in early stages of the project. We would also like to thank Kuan Fang, Siddharth Karamcheti, and Albertyna Osińska for useful discussions and feedback.
References
Appendix
To use large language models for a particular task, as opposed to a general text completion, one needs to encode the task as a part of the text input to the model. There exist many ways to create such encoding and the process of the representation optimization is sometimes referred to as prompt engineering . In this section, we discuss the prompts we used for LLM and VLM.
All our experiments use GPT-3 as the LLM, accessible via OpenAI’s API: https://openai.com/api/. We used this model to extract a list of landmarks from free-form instructions. The model outputs were very reliable and robust to small changes in the input prompts. For parsing simple queries, GPT-3 was surprisingly effective with a single, zero-shot prompt. See the example below, where the model output is highlighted:
First, you need to find a stop sign. Then take left and right and continue until you reach a square with a tree. Continue first straight, then right, until you find a white truck. The final destination is a white building. Landmarks: 1. Stop sign 2. Square with a tree 3. White truck 4. White building
While this prompt is sufficient for simple instructions, more complex instructions require the model to reason about occurrences such as re-orderings, e.g. Look for a glass building after after you pass by a white car. We leverage GPT-3 ability to perform in-context learning by adding three examples in the prompt:
Look for a library, after taking a right turn next to a statue. Landmarks: 1. a statue 2. a library Look for a statue. Then look for a library. Then go towards a pink house. Landmarks: 1. a statue 2. a library 3. a pink house [Instructions] Landmarks: 1.
We use the above prompt in all our experiments (Section 5.2, 5.3), and GPT-3 was successfully able to extract all landmarks. The comparison to other extraction methods is described in Section 5.4 and Appendix C.2.
A.2 VLM Prompt Engineering
In the case of our VLM— CLIP — we use a simple family of prompts: This is a photo of ___, appended with the landmark description. This simple prompt was sufficient to detect over of the landmarks encountered in our experiments. While our experiments did not require more careful prompt engineering, Radford et al. and Zeng et al. report improved robustness by using an ensemble of slightly varying prompts.
Appendix B Building the Topological Graph with VNM
This section outlines finer details regarding how the topological graph is constructed using VNM. We use a combination of learned distance estimates (from VNM), spatial proximity (from GPS), and temporal proximity (during data collection), to deduce edge connectivity. If the corresponding timestamps of two nodes are close (), suggesting that they were captured in quick succession, then the corresponding nodes are connected — adding edges that were physically traversed. If the VNM estimates of the images at two nodes are close, suggesting that they are reachable, then the corresponding nodes are also connected — adding edges between distant nodes along the same route and giving us a mechanism to connect nodes that were collected in different trajectories or at different times of day but correspond to the nearby locations. To avoid cases of underestimated distances by the model due to aliased observations, e.g. green open fields or a white wall, we filter out prospective edges that are significantly further away as per their GPS estimates — thus, if two nodes are nearby as per their GPS, e.g. nodes on different sides of a wall, they may not be disconnected if the VNM does not estimate a small distance; but two similar-looking nodes 100s of meters away, that may be facing a white wall, may have a small VNM estimate but are not added to the graph to avoid wormholes. Algorithm 2 summarizes this process — the timestamp threshold is 1 second, the learned distance threshold is 80 time steps (corresponding to meters), and the spatial threshold is 100 meters.
Since a graph obtained by such an analysis may be quite dense, we perform a transitive reduction operation on the graph to remove redundant edges.
Appendix C Miscellaneous Ablation Experiments
The graph search objective described in Section 4.4 can be factored into two components: visiting the required landmarks (denoted by ) and minimizing distance traveled (denoted by ). To analyze the importance of these two components, we ran a set of experiments where the nodes to be visited are selected based only on . This corresponds to a Max Likelihood planner, which only picks the most likely node for each landmark, without reasoning about their relative topological positions and traversability. This approach leads to a simpler algorithm: for each of the landmark descriptions, the algorithm selects the node with the highest CLIP score and connects it via the shortest path to the current node. The shortest path between each pair of nodes is computed using the Floyd–Warshall algorithm.
Table 4 summarizes the performance metrics for the two planners. Unsurprisingly, the max likelihood planner suffers greatly in the form of efficiency, because it does not incentivize shorter paths (see Figure 7 for an example). Interestingly, the planning success suffers as well, especially in complex environments. Further analysis of these failure modes reveals cases where VLM returns erroneous detections for some landmarks, likely due to the contrastive objective struggling with variable binding (see Figure 8 for an example). While LM-Nav suffers from these failures as well, the second factor in the search objective imposes a soft constraint on the search space of the landmarks, eliminating most of these cases and resulting in a significantly higher planning success rate.
C.2 Ablating the LLM
As described in Section 5.4 we run experiments comparing performance of different methods on extracting landmarks. Here we provide more details on the experiments. The source code to run this experiments is available in the file ablation_text_to_landmark.ipynb in the repository (see Appendix D).
As the metric of performance we used average extraction success. For a query with a ground truth list of landmarks , where a method extracts list , we define the methods extraction success as:
where LCS is longest common subsequence and denotes a length of a sequence or a list. This metric is measuring not only if correct landmarks were extracted, but also whether they are in the same order as in the ground truth sequence. When comparing landmarks we ignore articles, as we don’t expect them to have impact on the downstream tasks.
All the experiments were run using APIs serving models. We used OpenAI’s API (https://beta.openai.com/) for GPT-3 and GooseAI (https://goose.ai) for the other open-source models. Both providers conveniently share the same API. We used the same, default parameters, apart from setting temperature to : we don’t expect that landmark extractions to require creativity and model’s determinism improves reputability. For all the reported experiments, we used the same prompt as described in Appendix A. Please check out the released code for the exact prompts used.
Appendix D Code Release
We released the code corresponding to the LLM interface, VLM scoring, and graph search algorithm — along with a user-friendly Colab notebook capable of running quantitative experiments from Section 5.3. The links to the code and pickled graph objects can be found at our project page: sites.google.com/view/lmnav.
Appendix E Experiment Videos
We are sharing experiment videos of LM-Nav deployed on a Clearpath Jackal mobile robotic platform — please see sites.google.com/view/lmnav. The videos highlight the behavior learned by LM-Nav for the task of following free-form textual instructions and its ability to navigate complex environments and disambiguate between fine-grained commands.