Open-vocabulary Queryable Scene Representations for Real World Planning
Boyuan Chen, Fei Xia, Brian Ichter, Kanishka Rao, Keerthana Gopalakrishnan, Michael S. Ryoo, Austin Stone, Daniel Kappler
I Introduction
For robots to perform varied, real-world tasks, they must be able to comprehend diverse human commands and then act on these commands in the context of their environment. Imagine a robot in a home environment tasked with “water the plants in the living room”. It has to first identify relevant objects and locations within the scene (e.g., the watering can, the sink, and each potential plant) and then plan over these objects in sequential order (get the watering can, then go the sink, and then fill it up), conditioning on its affordances (e.g., can it carry a full watering can), and conditioning on the scene (e.g., how many plants there are, and where are they). Semantic representation and downstream mobile manipulation planners capable of accessing this representation emerge as critical challenges in such a pipeline.
Semantic understanding is crucial for a robot to achieve long-horizon tasks in unstructured environments. Though a robot can avoid building a semantic representation by finding objects each time they are required, e.g., with Object Goal Navigation , this repeated exploration can be inefficient. A persistent scene representation on the other hand avoids this exploration, but past works are generally limited to locating object categories known during the construction of the representation and may not encode the open-vocabulary objects that arise from human queries, such as in “bring me the purple unicorn plush toy”. Recent progress in contrastively trained visual language models offers a promising solution to open-ended scene presentation. Contrastive Language-Image Pre-training (CLIP) models are trained on image-language associations and can provide open-vocabulary image understanding and object detection . They have demonstrated impressive zero-shot classification performance and thus might be used to build a semantic representation in a zero-shot manner.
Another challenge lies in connecting the semantic scene representation to a planning algorithm that is capable of acting upon it. Recent progress in large language models (LLMs), has shown impressive few-shot performance in language comprehension, semantic understanding, and reasoning, as well as application to robotics problems like planning and instruction following . Using such models in embodied settings can provide significant challenges, most critically because LLMs are not grounded in the physical world. For example, pioneers in using LLMs for planning, but it has no grounding in environmental context. In contrast, SayCan showed how value functions of learned skills can provide such a grounding through selecting options scored highly by a language model and an affordance model. However, this is limited by the options provided and hardcoded knowledge of where objects exist.
In this work, we introduce Natural-Language Map (NLMap), a flexible and language-queryable spatial semantic representation based on visual-language models including ViLD and CLIP and integrate with SayCan. We show that NLMap grounds LLM-based planners in their environments, significantly improves long-horizon planning via natural language instructions in the open-world domain, and enables new tasks prior state-of-the-art algorithms failed to address. To summarize, we make the following contributions:
We propose an open-vocabulary, queryable semantic representation based on ViLD and CLIP.
We integrate NLMap into a language-based planner to enable grounding on the context.
We benchmark NLMap + SayCan in a real-world kitchen, showing it is capable of performing tasks at success rate. Notably, of these tasks are impossible with previous state-of-the-art planners that do not have access to NLMap.
II Related Work
Semantic Scene Representations. Scene representation is a central theme in robot perception and planning. Semantic SLAM is an augmentation over traditional SLAM, it assigns semantic features over geometric features provided by SLAM (points, lines, planes). Many representations are proposed, ranging from a faithful 3D recontruction of the environment, to more object-centric ones , such as object detection bounding boxes and 3D bounding boxes . Recently, topological maps and scene graphs emerge as an effective discrete representation of scenes.
One issue with those representations is that they cannot be queried with natural language. Interfacing with those scene representations requires reducing the object set to a closed set, indicating that they are not as useful for LLM-based planners and that they are limited in an open-vocabulary setting. In contrast, our work allows the scene representation to be queried at test time with natural language. Concurrent work VLMaps also explores this concept, by fusing pretrained visual-language model features into a geometric reconstruction of the scene. The representation is then used for visual-language navigation tasks via program synthesis.
Object Goal Navigation. There is also a significant body of related work on object navigation, which focuses on flexible exploration to find objects in unknown scenes. A few of these algorithms construct a semantic map of the current region before planning in that region . Map-based methods are modular and interpretable and hence easier to deploy in the real world. Other algorithms do not require a map and can decide where to go based directly on the current observations and memories, without maintaining a global representation of the environment. Recently, methods that leverage pre-trained image-text models can do zero-shot Object Goal Navigation . CoW performs zero-shot object goal navigation by leveraging CLIP. LM-Nav uses three pretrained model to perform visual language navigation. Our work differs from Object Goal Navigation since the eventual goal is not purely finding objects, but using object presence and location information for planning. Our work can use the representation from a single exploration for many downstream planning tasks without the need to run Object Goal Navigation every time.
Planning with Scene Representations. In task and motion planning, scene representations are often composed of predicates compatible with symbolic planners . Recent progresses attempt to build a symbolic and geometric scene graph to facilitate task and motion planning . However, they still require defining the objects in the scene. Recently LLM-based planners are more flexible and do not require handcrafting predicates, however, they do not handle the complexity of open-vocabulary object proposal and require defining a set of objects involved in planning. They also fail to integrate perception in real robot experiments due to the difficulty of connecting unstructured natural language instruction to perception algorithms that need structured inputs.
III Problem Statement
In this work, we aim to efficiently fulfill high-level, natural-language instructions, such as “Bring me a snack” or “I spilled my coffee, can you help?”. This requires a robotic system to solve problems at the intersection of natural language comprehension, scene understanding, task planning, navigation, and manipulation. Recent work, SayCan , has shown how large language models can be applied to such problems through world-grounding affordance functions, allowing LLMs to understand what a robot can do from a state. However, SayCan did not provide scene-scale affordance grounding, and thus cannot reason over what a robot can do in a scene. To that end, we address two core problems (i) how to maintain open-vocabulary scene representations that are capable of locating arbitrary objects and (ii) how to merge such representations within long-horizon LLM planners to imbue them with scene understanding.
IV NLMap + SayCan
We provide a high-level description of our algorithm in Listing LABEL:lst:algobox. The design of each component is described below:
The scene representation is generated from an exploration phase of the unstructured scene, which our approach is agnostic to, but could be for example frontier exploration or pre-determined waypoints. During this exploration, NLMap runs a class agnostic region proposal network as in ViLD on all the observed RGB images. For each proposed region of interest (ROI) , our method uses an ensemble of VLM image encoders to extract image embeddings . As shown in Fig. 2, such embedding can be queried with text at plan time since VLMs are capable of estimating the correlation between texts and images. In our setup we leverage CLIP and ViLD as visual encoders , where image-text-alignment is scored with inner product of image feature and CLIP text feature. We also extract the estimated location using depth at the center of the image as well as estimated size of the object in . Defining the tuple as a context element, the collection forms our scene representation.
IV-B Querying the Representation
To complete a task specified by human instruction, the robot will query the scene representation for relevant information. This is achieved by first parsing natural language instruction into a list of relevant object names, then using the names as keys to query object locations and availability. Finally, we generate executable options based on what’s found in the scene, then plan and execute as instructed.
The core challenge of querying scene information is bridging unstructured natural language input and structured representations. In order to decide what objects to look up in the scene representation, we use few-shot prompting to let LLM actively propose required objects given an instruction. Different from previous work that uses LLM to extract names from a sentence, our object proposal is much more demanding in four different ways as we will discuss in Sec. V-B.
In order to achieve a reliable object proposal that addresses four requirements, we introduce example prompts for each case and use the few-shot prompting technique of LLMs to propose them. The few-shot examples can be found on our project website.
IV-B2 Object Query
Here we use both CLIP embedding and ViLD embedding because the former detects out-of-distribution objects better while the latter is more robust to common objects as shown in Fig. 5. We can directly take the maximum over the two inner products because both of them are normalized vectors designed to be queried by the inner product CLIP text encoder. Given metric , the top k nearest neighbor elements for object name can be found in the scene representation . We note that based on the value of , we can impose a threshold to filter out low-confidence detections. These top context elements are associated with ROIs, multiple of which may correspond to the same real-world 3D object instance. We then run a multi-view fusion algorithm to aggregate these context elements into 3d object locations and filter out objects that don’t exist according to an aggregated score. Details of the algorithm can be found in Sec. VI-B.
IV-C Combining NLMap and SayCan
Our method constructs a scene representation queryable by natural language. Such representation can be connected with LLM-based planners to enable robots to operate in a truly uncontrolled environment. Previously, SayCan presents a framework that allows robots to plan and execute in the real world following human instructions. We highlight the difference between our work and SayCan in Fig. 3. SayCan work as follows: with few-shot prompting, SayCan uses the scoring of a language model to break down a high-level instruction like “Bring me an apple” to “1. Find the apple, 2. Pick up the apple, 3. Bring it to you, 4. Put down the apple”. Each option from a pre-defined list is scored by an LLM and an affordance prediction module. However, SayCan relies on a hard-coded list of object names, locations, and executable options so its capability is largely limited by the lack of contextual grounding.
NLMap makes up this missing component in SayCan. Our object proposal, combined with the object query, generates the relevant object names and locations conditioned on the instruction and the scene. There are two major remaining challenges.
Vanilla SayCan provides a list of skills associated with either 1) navigation policies to hard-coded locations 2) manipulation policies (pick and place) of objects, specified by object names. Given a detected object and its location, we can create a new skill “find the [object name]” bound to a navigation policy to that location. This means we can expand a small fixed set of navigation options to infinitely many options. On the other hand, although training manipulation policies for infinitely many objects is beyond the scope of our work, we can still augment the manipulation capability of SayCan by binding all possible references to a manipulable object with the available manipulation policies. This is achieved by finding CLIP nearest neighbor of object names. For example, given discovered objects, we can generate executable options like “pick up the red can” and “pick up a tin of coke”. Our method will bind both of them to the closest manipulation policy “pick up coke can” with CLIP. This nearest neighbor query is similar to that used with BERT in .
IV-C2 Ground LLM planner with context
Unlike the setup in SayCan, which assumes all objects in the hard-coded list are present, our method is expected to tackle infeasible instructions, such as instructions involving objects that aren’t present. SayCan weakly addresses this problem by grounding plans with local affordance, which is only conditioned on what’s directly visible in the field of view rather than what’s available in the entire scene. NLMap gives us a list of available objects so we can add the missing global contextual grounding to SayCan. This is achieved by modifying the original few shot prompts in SayCan to also condition the plan on discovered objects, expressed in templates like “Scene: apple, coke can.” We include both positive examples when necessary objects are all present and negative examples when available objects cannot fulfill the instruction. In the former case, LLM is prompted to plan just like in vanilla SayCan; In the latter case, LLM is prompted to output the terminate signal “done” directly, indicating the task is infeasible.
With these components, we can ground SayCan with context awareness. After exploring the scene, when a human gives the robot an instruction, the robot will propose potentially involved objects in the scene and query the gathered scene representation for their locations and availability. NLMap then generates executable options, plans with LLM conditioned on what’s found and finally executes the plan in the real world under the SayCan framework.
V Experiments
In this section, we evaluate NLMap and its individual components with real-world robotics tasks. We test a robot running NLMap in a real office kitchen, as shown in Fig. 4. We test the entire system in an end-to-end setting such that the robot attempts to accomplish tasks specified by humans with natural language. We list a subset of the manipulable objects in Fig. 4(a) receptacle locations in Fig. 4(b). The robot is a mobile manipulator from Everyday Robot, which has a mobile base and a 7-degree-of-freedom arm, as shown in Fig. 4(c). The main sensor is an RGBD camera, which returns RGBD images. Similar to SayCan, we use a set of manipulation policies trained from imitation learning and PaLM 540B as the LLM for all experiments, due to its good performance on new tasks with few-shot prompting. Throughout this section, all experiments share the same set of hyper-parameters and LLM prompts unless specified otherwise. A full list of test instructions can be found on the project website.
In this section, we demonstrate our natural language queryable representation can be combined with LLM planners to significantly augment the capability of real robot operation. We choose to combine NLMap with SayCan, a recent work that uses LLM planners to let robots plan and execute according to natural language instructions. One of the biggest limitations of SayCan, as stated in Sec. III, is that it has no global context awareness. By combing our method with SayCan using the method described in Sec. IV-C, we free SayCan from a fixed, hard-coded set of objects, locations, or executable options. With NLMap, SayCan can now perform a great number of previously unachievable tasks. In addition, we demonstrate that our method allows SayCan to plan with the global context to identify infeasible tasks. We quantitatively evaluate the real robot performance of NLMap + SayCan in Table I with three sets of benchmarks. We compare our method with a privileged version of SayCan, which uses ground truth perception results in the scene.
We hope to understand how much performance will be lost compared to SayCan due to the addition of perception and context-aware planning. Therefore, we benchmark tasks adopted from of the task families from the original SayCan paper with random tasks from each family (except for Embodiment family). Our method achieves a success rate of among these tasks compared to the of privileged SayCan. We also tried tasks with deliberate typos ‘ppsi” ‘chpis”. Our method failed in both instructions with typos, with one failure during object proposal and one failure due to policy binding. With these two typo experiments included, our method achieves an overall success rate of compared to in real robot experiments compared to privileged SayCan that has hard-coded object locations. This shows our NLMap maintains a reasonable overall success rate even if multiple components like object proposal, perception, and context-conditioned planning are added.
V-A2 Novel objects
SayCan relies on a hard-coded list of object names, locations, and executable options. Since the hard-coded set of objects and executable options are finite, SayCan is incapable of performing tasks that involve objects or skills outside these small sets. However, since NLMap can propose and detect objects, and generate executable options itself, NLMap can be combined with SayCan to execute infinitely many tasks that involve such novel objects as described in Sec. IV-C. As shown in Table I, SayCan fails to plan nor execute any of these tasks while our method achieves a success rate of in the end-to-end execution experiment. It even succeeds in some very out-of-distribution instructions such as “I want to watch TV, can you get a bottle of tea and put it there” or “Show me where is the first aid station”. We note that manipulation policies used in this project are still limited to be with the objects that are visually similar to training objects in and rely on the generalization to slightly out-of-distribution data. Therefore, the novel object names in this experiment are either used for navigation only, or for describing objects that are visually similar to training objects in . Such constraint can be lifted in the future when a general text-conditioned manipulation policy is available but lifting it is beyond the scope of the project.
V-A3 Missing Objects
Vanilla SayCan isn’t grounded by what’s available in the scene. If a necessary object is removed from the scene, there is no way for SayCan’s LLM planner to tell the task is infeasible. With NLMap, we can use the method in Sec. IV-C to condition SayCan planning on what’s actually detected. In this benchmark, we ask NLMap + SayCan to perform tasks that require objects not present in the scene. Instructions in the benchmark consist of size subset of all instructions in the “novel object” benchmark since we cannot remove objects like “first aid station” from the wall. In a successful run, the robot is expected to not detect an object doesn’t exist and output a termination signal immediately in its plan. Our method achieves a success rate of in the missing object setting, where of the total failure cases are due to false positive detections. Although vanilla SayCan will achieve a success rate of zero in comparison, this benchmark still indicates false positive detection is a challenge for context-aware planning.
V-B Benchmarking Object Proposal
Object proposal is a foundational component in our framework to parse unstructured instructions into structured object names. We investigate the robustness and generalization capability of object proposal from four perspectives:
Infer objects from implication of the instruction: e.g. “Heat up the taco” (taco, microwave)
Unstructured crowd-sourced instructions: e.g. “Redbull is my favorite drink, can I have a one please?” (redbull, human)
Objects with fine-grained description: e.g. “turn off the macbook with yellow stickers” (macbook with yellow stickers)
Decomposition to proper granularity: e.g. “check out what types of ingredients are available to cook a luxurious breakfast” (milk,eggs,bacon,bread,butter,cheese,ham,sausage…)
A summary of result of each perspective can be found in Table II.
In previous work that use LLM to extract object names from language, all object names are nouns that are directly present in the language input. However, in the real world, humans frequently give instructions that involve objects that have to be inferred from the implication of the task. We test object proposal on such instructions and evaluate whether proposed objects would complete the task. Object proposal achieved a success rate of in test cases including “season the steak (salt, pepper)”, “fillet the fish (fish, knife)”.
V-B2 Unstructured crowd-sourced instructions
Object proposal module is expected to take in instructions from a variety of highly unstructured formats. We evaluate the robustness of our object proposal on a set of test instructions adopted from crowd-sourced instructions for SayCan. Object proposal achieved a success rate of in this study, including multi-step tasks like “Move an multigrain chips to the table and an apple to the far counter”. Object proposal succeeded in all out of multi-step tasks in this study.
V-B3 Reference to objects with fine-grained description
Human instructions often involve reference to objects with fine-grained descriptions. Such descriptions are often important to visually identify a particular instance in the scene. Thus it’s important for the object proposal to keep these fine-grained descriptions in its output. We evaluate object proposal on test instructions that involve fine-grained descriptions by adjectives or clauses. The model attains a success rate of in this experiment. The model even succeeded in some complicated descriptions like “mug in the shape of a donut”.
V-B4 Decomposition to proper granularity
Many instructions require a different level of object proposal granularity. Certain tasks can only be accomplished if the object proposal is more fine-grained. We evaluate object proposal on tasks that require expanding a category mentioned in the instruction. Overall, the object proposal achieves a success rate of in this set, indicating that proper granularity is still a hard challenge for LLM due to its multi-modality nature.
V-C Benchmarking Object Queries to NLMap
In this section, we evaluate the open-vocabulary object query module on a list of common objects in our testing kitchens. We run robot exploration and object query in two different kitchen scenes, each with some object deliberately missing. Our method uses both maximum ensemble metric and multi-view fusion described in Sec. IV-B with . We compare this choice with alternative embeddings and metrics like or . Maximum ensemble metric without multi-view fusion is also evaluated as a baseline. We have in the above three baselines since no multi-view fusion is happening. As shown in Table III, ViLD and CLIP embedding alone achieves a very low success rate in both environments. As illustrated in Fig. 5, we observe that ViLD embedding detects common objects like cans or apples more reliably while suffering from false negative detection of out-of-distribution objects such as “first aid station”. On the other hand, CLIP embedding gives us better results on uncommon objects but is less robust for basic objects. Additionally CLIP embeddings better captures features of text and signs. Our method uses multi-view fusion in addition to the maximum ensemble. Multiview fusion leads to a slight accuracy increase in scene but a significant increase in the second scene. This shows that multi-view fusion can help remove outlier observations that produce high likelihood scores but are actually noise by noticing a lack of detection of it from different views. Overall, the perception success rate for our method is and respectively in the two kitchens. Such accuracy is limited by the low resolution and exposure of our robot camera. However, since instructions don’t always contain visually ambiguous objects like many in these test queries, perception is still reliable enough as we see in the real robot experiments Sec. V-A.
V-D Benchmarking Context Grounded Planing
Failures from perception or object proposal are coupled with planning in real robot experiments. In this section, we ablate context-aware LLM planning as a standalone component, assuming correct object proposal and detection. We test LLM planning in a generative way. A generated plan is considered correct if it will accomplish the instruction, is consistent with the available objects, and is executable. We benchmark generative planning with test cases consisting of instructions with set of available objects for each. One set is a positive set that contains all needed objects for the task while the other set is a negative set with some necessary objects missing. To be considered successful, the planner should behave like Vanilla SayCan in the positive set while outputting the terminal signal immediately in the negative set. Our LLM planner, conditioned on available objects using the method described in Sec. IV-C, achieves a success rate of and on the instructions with positive object set and negative set respectively. The performance gap is expected because negation is known to be a hard problem for LLM.
VI Conclusions
We integrate NLMap, a flexible and queryable spatial semantic representation based on visual-language models including ViLD and CLIP with SayCan. We show that NLMap is a flexible scene representation that grounds LLM-based planners in their environments, significantly improving long-horizon planning via natural language instructions in open-worlded domain, enabling new tasks prior state-of-the-art algorithms failed to address.
Future work. Currently, NLMap only handles a static scene representation without dynamic objects and human, which we will leave this for future work. All the modules used in NLMap + SayCan is pre-trained and deployed zero-shot. It is a great advantage but we hope to fine-tune them for better performance. Additionally, we will look into efficient exploration algorithms to speed up the creation of scene representation.
Acknowledgements
Special thanks to Arjun Majumdar, Andy Zeng and Karol Hausman for helpful discussions; we thank Peng Xu and Xiran Liu for helpful feedbacks on writing.
References
APPENDIX
Our context-aware SayCan algorithm is similar to , it expands the last line LLM.plan(instruction, scene_objects) in Listing LABEL:lst:algobox. Compared to the original SayCan , our context-aware version needs a list of detected object names , along with a list of template functions as extra input. A template function maps an object name to an option name such as . We note that the template function is used here because training manipulation policies beyond pick-and-place are beyond the scope of our project. If we have a language-conditioned policy in the future, we don’t need to use template functions anymore. Trusting LLM for new options will suffice in that case. A full pseudo-code can be found in Algo 1.
VI-B Multi-view fusion algorithm
In this section, we describe details of the multi-view fusion algorithm mentioned in Sec. IV-B. In the gathered scene representation , multiple context elements may be associated with the same object. Each context element contains an estimation of object centroid and along with a object width . To simplify formulation, we use cylindrical bounding volumes to model 3d objects. We create such bounding boxes with center and radius in an upright position. Given each queried object name , we can quickly narrow down bounding box candidates by finding the top k nearest neighbors with metric . We now have a problem similar to post-processing in object detection - for each real object instance, we may have overlapping bounding box predictions, which are supposed to be aggregated together. In computer vision, this is achieved by the NMS algorithm that group predictions based on the intersection over union(IOU) of the bounding box followed by keeping only the bounding box with the highest confidence in each group. We made three major changes to the NMS algorithm by noticing the special structure of our problem.
First, since our bounding volumes are not cubes, IOU is hard to compute. We instead use KL divergence of Gaussian distributions to model. For each cylindrical bounding box with a circular projection on the 2d plane, define Gaussian distribution . The 2d Gaussian will have its center at the estimated centroid and standard deviation proportional to the width of the object. KL divergence measures how different two distributions are so it acts like the IOU for gaussian distributions. When estimations have very different centers or sizes, they will be considered to correspond to two different object instances by our algorithm. Second, different from the setup in 2d object detection, different estimations of the same object in our problem are considered valid, independent data points that contribute to a better estimation of object location. Therefore, we don’t discard non-maximum estimations in each clustered group, but rather use their score as importance weights to derive the final estimation through weighted average. Third, bounding boxes are directly filtered out based on a threshold on confidence score in 2d detection. In our setup, we give confidence scores a bonus based on how many elements there are by noticing available objects should be detected from multiple view points.
We then offer a formal algorithm box for multi-view fusion in Algo 2. Given object name , we can use metric to score each context element in and find the top k ones. Denote the indices of top k context elements as , sorted in descending order by score. For each context element , define Gaussian distribution . In our experiments, we choose the monotonic increasing function to be in the form where is some hyper-parameter.
The algorithm then outputs clustered locations for objects queried by name .
VI-C Prompt used for object proposal and for planning
VI-D Object proposal experiment task list
VI-E Robot experiment task list
VI-F Additional qualitative experiment results
We show additional qualitative experiment results in Fig. 7, Fig. 8 and Fig. 9.