VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, Li Fei-Fei

Introduction

Language is a compressed medium through which humans distill and communicate their knowledge and experience of the world. Large language models (LLMs) have emerged as a promising approach to capture this abstraction, learning to represent the world through projection into language space . While these models are believed to internalize generalizable knowledge as text, it remains a question about how to use it to enable embodied agents to physically act in the real world.

We look at the problem of grounding abstract language instructions (e.g., “set up the table”) in robot actions . Prior works have leveraged lexical analysis to parse the instructions , while more recently language models have been used to decompose the instructions into a textual sequence of steps . However, to enable physical interactions with the environment, existing approaches typically rely on a repertoire of pre-defined motion primitives (i.e., skills) that may be invoked by an LLM or a planner, and this reliance on individual skill acquisition is often considered a major bottleneck of the system due to the lack of large-scale robotic data. The question then arises: how can we leverage the wealth of internalized knowledge of LLMs at the even fine-grained action level for robots, without requiring laborious data collection or manual designs for each individual primitive?

In addressing this challenge, we first note that it is impractical for LLMs to directly output control actions in text, which are typically driven by high-frequency control signals in high-dimensional space. However, we find that LLMs excel at inferring language-conditioned affordances and constraints, and by leveraging their code-writing capabilities, they can compose dense 3D voxel maps that ground them in the visual space by orchestrating perception calls (e.g., via CLIP or open-vocabulary detectors ) and array operations (e.g., via NumPy ). For example, given an instruction “open the top drawer and watch out for the vase”, LLMs can be prompted to infer: 1) the top drawer handle should be grasped, 2) the handle needs to be translated outwards, and 3) the robot should stay away from the vase. By generating Python code to invoke perception APIs, LLMs can obtain spatial-geometric information of relevant objects or parts and then manipulate the 3D voxels to prescribe reward or cost at relevant locations in observation space (e.g., the handle region is assigned high values while the surrounding of the vase is assigned low values). Finally, the composed value maps can serve as objective functions for motion planners to directly synthesize robot trajectories that achieve the given instruction The approach also bears resemblance and connections to potential field methods in path planning and constrained optimization methods in manipulation planning . , without requiring additional training data for each task or for the LLM. An illustration diagram and a subset of tasks we considered are shown in Fig. 1.

We term this approach VoxPoser , a formulation that extracts affordances and constraints from LLMs to compose 3D value maps in observation space for guiding robotic interactions. Rather than relying on robotic data that are often of limited amount or variability, the method leverages LLMs for open-world reasoning and VLMs for generalizable visual grounding in a model-based planning framework that directly enables physical robot actions. We demonstrate its zero-shot generalization for open-set instructions with open-set objects for various everyday manipulation tasks. We further showcase how VoxPoser can also benefit from limited online interactions to efficiently learn a dynamics model that involves contact-rich interactions.

Related Works

Grounding Language Instructions. Language grounding has been studied extensively both in terms of intelligent agents and of robotics , where language can be used as a tool for compositional goal specification , semantic anchor for training multi-modal representation , or as an intermediate substrate for planning and reasoning . Prior works have looked at using classical tools such as lexical analysis, formal logic, and graphical models to interpret language instructions . More recently, end-to-end approaches, popularized by successful applications to offline domains , have been applied to directly ground language instructions in robot interactions by learning from data with language annotations, spanning from model learning , imitation learning , to reinforcement learning . Most closely related to our work is Sharma et al. , where an end-to-end cost predictor is optimized via supervised learning to map language instructions to 2D costmaps, which are used to steer a motion planner to generate preferred trajectories in a collision-free manner. In contrast, we rely on pre-trained language models for their open-world knowledge and tackle the more challenging robotic manipulation in 3D.

Language Models for Robotics. Leveraging pre-trained language models for embodied applications is an active area of research, where a large body of works focus on planning and reasoning with language models . To allow language models to perceive the physical environments, textual descriptions of the scene or perception APIs can be given, vision can be used during decoding or can be directly taken as input by multi-modal language models . In addition to perception, to truly bridge the perception-action loop, an embodied language model must also know how to act, which typically is achieved by a library of pre-defined primitives. Liang et al. showed that LLMs exhibit behavioral commonsense that can be useful for low-level control. Despite the promising signs, hand-designed motion primitives are still required, and while LLMs are shown to be capable of composing sequential policy logic, it remains unclear whether composition can happen at spatial level. A related line of works has also explored using LLMs for reward specification in the context of reward design , exploration , and preference learning . For robotic applications, concurrent works explored LLM-based reward generation , among which Yu et al. use MuJoCo as a high-fidelity physics model for model predictive control. In contrast, we focus exclusively on grounding the reward generated by LLMs in the 3D observation space of the robot.

Learning-based Trajectory Optimization. Many works have explored leveraging learning-based approaches for trajectory optimization. While the literature is vast, they can be broadly categorized into those that learn the models and those that learn the cost/reward or constraints , where data are typically collected from in-domain interactions. To enable generalization in the wild, a parallel line of works has explored learning task specification from large-scale offline data , particularly egocentric videos , or leveraging pre-trained foundation models . The learned cost functions are then used by reinforcement learning , imitation learning , or trajectory optimization to generate robot actions. In this work, we leverage LLMs for zero-shot in-the-wild cost specification with superior generalization. Compared to prior works that leverage foundation models, we ground the cost directly in 3D observation space with real-time visual feedback, which makes VoxPoser amenable to closed-loop MPC that’s robust in execution.

Method

We first provide the formulation of VoxPoser as an optimization problem (Sec. 3.1). Then we describe how VoxPoser can be used as a general zero-shot framework to map language instructions to 3D value maps (Sec. 3.2). We subsequently demonstrate how trajectories can be synthesized in closed-loop for robotic manipulation (Sec. 3.3). While zero-shot in nature, we demonstrate how VoxPoser can learn from online interactions to efficiently solve contact-rich tasks (Sec. 3.4).

2 Grounding Language Instruction via VoxPoser

3 Zero-Shot Trajectory Synthesis with VoxPoser

4 Efficient Dynamics Learning with Online Experiences

While Sec. 3.3 presents a zero-shot framework for synthesizing trajectories for robot manipulation, VoxPoser can also benefit from online experiences by efficiently learning a dynamics model. Consider the standard setup where a robot interleaves between 1) collecting environment transition data (ot,at,ot+1)(\mathbf{o}_{t},\mathbf{a}_{t},\mathbf{o}_{t+1}), where ot\mathbf{o}_{t} is the environment observation at time tt and at=MPC(ot)\mathbf{a}_{t}=\text{MPC}(\mathbf{o}_{t}), and 2) training a dynamics model gθ\mathbf{g}_{\theta} parametrized by θ\theta by minimizing the L2 loss between predicted next observation o^t+1\hat{\mathbf{o}}_{t+1} and ot+1\mathbf{o}_{t+1}. A critical component that determines the learning efficiency is the action sampling distribution P(at∣ot)P(\mathbf{a}_{t}|\mathbf{o}_{t}) in MPC, which typically is a random distribution over the full action space A\mathbf{A}. This is often inefficient when the goal is to solve a particular task, such as opening a door, because most actions do not interact with the relevant objects in the scene (i.e., the door handle) nor do they necessarily interact with the objects in a meaningful way (i.e., pressing down the door handle). Since VoxPoser synthesizes robot trajectories with LLMs, which have a wealth of commonsense knowledge, the zero-shot synthesized trajectory τ0r\tau^{\mathbf{r}}_{0} can serve as a useful prior to bias the action sampling distribution P(at∣ot,τ0r)P(\mathbf{a}_{t}|\mathbf{o}_{t},\tau^{\mathbf{r}}_{0}), which can significantly speed up the learning process. In practice, this can be implemented by only sampling actions in the vicinity of τ0r\tau^{\mathbf{r}}_{0} by adding small noise ε\varepsilon to encourage local exploration instead of exploring in the full action space A\mathbf{A}.

Experiments and Analysis

We first discuss our implementation details. Then we validate VoxPoser for real-world everyday manipulation (Sec. 4.1). We also study its generalization in simulation (Sec. 4.2). We further demonstrate how VoxPoser enables efficient learning of more challenging tasks (Sec. 4.3). Finally, we analyze its source of errors and discuss how improvement can be made (Sec. 4.4).

LLMs and Prompting. We follow prompting structure by Liang et al. , which recursively calls LLMs using their own generated code, where each language model program (LMP) is responsible for a unique functionality (e.g., processing perception calls). We use GPT-4 from OpenAI API. For each LMP, we include 5-20 example queries and corresponding responses as part of the prompt. An example can be found in Fig. 2 (simplified for clarity). Full prompts are in Appendix.

VLMs and Perception. Given an object/part query from LLMs, we first invoke open-vocab detector OWL-ViT to obtain a bounding box, then feed it into Segment Anything to obtain a mask, and finally track the mask using video tracker XMEM . The tracked mask is used with RGB-D observation to reconstruct the object/part point cloud.

Dynamics Model. We use the known robot dynamics model in all tasks, where it is used in motion planning for the end-effector to follow the waypoints. For the majority of our considered tasks where the “entity of interest” is the robot, no environment dynamics model is used (i.e., scene is assumed to be static), but we replan at every step to account for the latest observation. For tasks in which the “entity of interest” is an object, we study only a planar pushing model parametrized by contact point, push direction, and push distance. We use a heuristic-based dynamics model that translates an input point cloud along the push direction by the push distance. We use MPC with random shooting to optimize for the action parameters. Then a pre-defined pushing primitive is executed based on the action parameters. However, we note that a primitive is not necessary when action parameters are defined over the end-effector or joint space of the robot, which would likely yield smoother trajectories but takes more time for optimization. We also explore the use of a learning-based dynamics model in Section 4.3, which enables VoxPoser to benefit from online experiences.

We study whether VoxPoser can zero-shot synthesize robot trajectories to perform everyday manipulation tasks in the real world. Details of the environment setup can be found in Appendix A.4. While the proposed method can generalize to an open-set of instructions and an open-set of objects as shown in Fig. 1, we pick 5 representative tasks to provide quantitative evaluations in Table 1. Qualitative results including environment rollouts and value map visualizations are shown in Fig. 3. We find that VoxPoser can effectively synthesize robot trajectories for everyday manipulation tasks with a high average success rate. Due to fast replanning capabilities, it is also robust to external disturbances, such as moving targets/obstacles and pulling the drawer open after it has been closed by the robot. We further compare to a variant of Code as Policies that uses LLMs to parameterize a pre-defined list of simple primitives (e.g., move_to_pose, open_gripper). We find that compared to chaining sequential policy logic, the ability to compose spatially while considering other constraints under a joint optimization scheme is a more flexible formulation, unlocking the possibility for more manipulation tasks and leading to more robust execution.

2 Generalization to Unseen Instructions and Attributes

To provide rigorous quantitative evaluations on generalization, we set up a simulated block-world environment that mirrors our real-world robot setup but features 13 highly-randomizable tasks with 2766 unique instructions. Eash task comes with a templated instruction (e.g., “push [obj] to [pos]”) that contains randomizable attributes chosen from a pre-defined list. Details are in Appendix A.5. Seen instructions/attributes may appear in the prompt (or in the training data for supervised baselines). The tasks are grouped into 2 categories, where “Object Interactions” are tasks that require interactions with objects, and “Spatial Composition” are tasks involving spatial constraints (e.g., moving slower near a particular object). For baselines, we ablate the two components of VoxPoser, LLM and motion planner, by comparing to a variant of that combines an LLM with primitives and to a variant of that learns a U-Net to synthesize costmaps for motion planning. Table 2 shows the success rates averaged across 20 episodes per task. We find VoxPoser exhibits superior generalization in all scenarios. Compared to learned cost specification, LLMs generalize better by explicitly reasoning about affordances and constraints. On the other hand, grounding LLM knowledge in robot perception through value map composition rather than directly specifying primitive parameters offers more flexibility and better generalization.

3 Efficient Dynamics Learning with Online Experiences

As discussed in Sec. 3.4, we investigate how VoxPoser can optionally benefit from online experiences for tasks that involve more intricacies of contact, such as opening doors, fridges, and windows, in a simulated environment. Specifically, we first synthesize kk zero-shot trajectories using VoxPoser, each represented as a sequence of end-effector waypoints, that act as priors for exploration (e.g., “handle needs to be pressed down first in order to open a door”). Then an MLP dynamics model is learned through an iterative procedure where the agent alternates between data collection and model learning. During data collection, we add ε∼N(0,σ2)\varepsilon\sim\mathcal{N}(0,\sigma^{2}) to each waypoint in τ0r\tau^{\mathbf{r}}_{0} to encourage local exploration. As shown in Tab. 3, we find zero-shot synthesized trajectories are typically meaningful but insufficient. However, we can learn an effective dynamics model with less than 3 minutes of online interactions by using these trajectories as exploration prior, leading to high eventual success rates. In comparison, exploring without prior all exceed the maximum 12-hour limit.

4 Error Breakdown

In this section, we analyze the errors resulting from each component of VoxPoser and how the overall system can be further improved. We conduct experiments in simulation where we have access to ground-truth perception and dynamics model (i.e., the simulator). . “Dynamics error” refers to errors made by the dynamics modelLLM + Primitives does not use model-based planning, thus not having a dynamics module.. “Perception error” refers to errors made by the perception moduleU-Net + MP maps RGB-D to costmaps using U-Net , thus not having perception module. Errors by which are attributed to “specification error”.. “Specification error” refers to errors made by the module specifying cost or parameters for the low-level motion planner or primitives. Examples for each method include 1) noisy prediction by the U-Net, 2) incorrect parameters specified by the LLM, and 3) incorrect value maps specified by the LLM. As shown in Fig. 4, VoxPoser achieves lowest “specification error” due to its generalization and flexibility. We also find that having access to a more robust perception pipeline and a physically-realistic dynamics model can contribute to better overall performance. This observation aligns with our real-world experiment, where most errors are from perception. For example, we find that the detector is sensitive to initial poses of objects and is less robust when detecting object parts.

Conclusion, Limitations, & Future Works

In this work, we present VoxPoser, a general framework for extracting affordances and constraints, grounded in 3D perceptual space, from LLMs and VLMs for everyday manipulation tasks in the real world, offering significant generalization advantages for open-set instructions and objects. Despite compelling results, VoxPoser has several limitations. First, it relies on external perception modules, which is limiting in tasks that require holistic visual reasoning or understanding of fine-grained object geometries. Second, while applicable to efficient dynamics learning, a general-purpose dynamics model is still required to achieve contact-rich tasks with the same level of generalization. Third, our motion planner considers only end-effector trajectories while whole-arm planning is also feasible and likely a better design choice . Finally, manual prompt engineering is required for LLMs. We also see several exciting venues for future work. For instance, recent success of multi-modal LLMs can be directly translated into VoxPoser for direct visual grounding. Methods developed for alignment and prompting may be used to alleviate prompt engineering effort. Finally, more advanced trajectory optimization methods can be developed that best interface with value maps synthesized by VoxPoser.

We would like to thank Andy Zeng, Igor Mordatch, and the members of the Stanford Vision and Learning Lab for the fruitful discussions. This work was in part supported by AFOSR YIP FA9550-23-1-0127, ONR MURI N00014-22-1-2740, ONR MURI N00014-21-1-2801, ONR N00014-23-1-2355, the Stanford Institute for Human-Centered AI (HAI), JPMC, and Analog Devices. Wenlong Huang is partially supported by Stanford School of Engineering Fellowship. Ruohan Zhang is partially supported by Wu Tsai Human Performance Alliance Fellowship.

References

Appendix A Appendix

We provide an open-sourced implementation of VoxPoser at github.com/huangwl18/VoxPoser based on RLBench , as its diversity of tasks and scenes best resembles our real-world setup.

A.2 Emergent Behavioral Capabilities

Emergent capabilities refer to unpredictable phenomenons that are only present in large models . As VoxPoser uses pre-trained LLMs as backbone, we observe similar embodied emergent capabilities driven by the rich world knowledge of LLMs. In particular, we focus our study on the behavioral capabilities that are unique to VoxPoser. We observe the following capabilities:

Behavioral Commonsense Reasoning: During a task where robot is setting the table, the user can specify behavioral preferences such as “I am left-handed”, which requires the robot to comprehend its meaning in the context of the task. VoxPoser decides that it should move the fork from the right side of the bowl to the left side.

Fine-grained Language Correction: For tasks that require high precision such as “covering the teapot with the lid”, the user can give precise instructions to the robot such as “you’re off by 1cm”. VoxPoser similarly adjusts its action based on the feedback.

Multi-step Visual Program : Given a task “open the drawer precisely by half” where there is insufficient information because object models are not available, VoxPoser can come up with multi-step manipulation strategies based on visual feedback that first opens the drawer fully while recording handle displacement, then close it back to the mid-point to satisfy the requirement.

Estimating Physical Properties: Given two blocks of unknown mass, the robot is tasked to conduct physics experiments using an existing ramp to determine which block is heavier. VoxPoser decides to push both blocks off the ramp and choose the block traveling the farthest as the heavier block. Interestingly, this mirrors a common human oversight: in an ideal, frictionless world, both blocks would traverse the same distance under the influence of gravity. This serves as a lighthearted example that language models can exhibit limitations similar to human reasoning.

A.3 APIs for VoxPoser

Central to VoxPoser is an LLM generating Python code that is executed by a Python interpreter. Besides exposing NumPy and the Transforms3d library to the LLM, we provide the following environment APIs that LLMs can choose to invoke:

detect(obj_name): Takes in an object name and returns a list of dictionaries, where each dictionary corresponds to one instance of the matching object, containing center position, occupancy grid, and mean normal vector.

execute(movable,affordance_map,avoidance_map,rotation_map,velocity_map,gripper_map): Takes in an “entity of interest” as “movable” (a dictionary returned by detect) and (optionally) a list of value maps and invokes the motion planner to execute the trajectory. Note that in MPC settings, “movable” and the input value maps are functions that can be re-evaluated to reflect the latest environment observation.

cm2index(cm,direction): Takes in a desired offset distance in centimeters along direction and returns 3-dim vector reflecting displacement in voxel coordinates.

index2cm(index,direction): Inverse of cm2index. Takes in an integer “index” and a “direction” vector and returns the distance in centimeters in world coordinates displaced by the “integer” in voxel coordinates.

pointat2quat(vector): Takes in a desired pointing direction for the end-effector and returns a satisfying target quaternion.

set_voxel_by_radius(voxel_map,voxel_xyz,radius_cm,value): Assigns “value” to voxels within “radious_cm” from “voxel_xyz” in “voxel_map”.

get_empty_affordance_map(): Returns a default affordance map initialized with 0, where a high value attracts the entity.

get_empty_avoidance_map(): Returns a default avoidance map initialized with 0, where a high value repulses the entity.

get_empty_rotation_map(): Returns a default rotation map initialized with current end-effector quaternion.

get_empty_gripper_map(): Returns a default gripper map initialized with current gripper action, where 1 indicates “closed” and 0 indicates “open”.

get_empty_velocity_map(): Returns a default affordance map initialized with 1, where the number represents scale factor (e.g., 0.5 for half of the default velocity).

reset_to_default_pose(): Reset to robot rest pose.

A.4 Real-World Environment Setup

We use a Franka Emika Panda robot with a tabletop setup. We use Operational Space Controller with impedance from Deoxys . We mount two RGB-D cameras (Azure Kinect) at two opposite ends of the table: bottom right and top left from the top down view. At the start of each rollout, both cameras start recording and return the real-time RGB-D observations at 20 Hz.

For each task, we evaluate each method on two settings: without and with disturbances. For tasks with disturbances, we apply three kinds of disturbances to the environment, which we pre-select a sequence of them at the start of the evaluation: 1) random forces applied to the robot, 2) random displacement of task-relevant and distractor objects, and 3) reverting task progress (e.g., pull drawer open while it’s being closed by the robot). We only apply the third disturbances to tasks where “entity of interest” is an object or object part.

We compare to a variant of Code as Policies as a baseline that uses an LLM with action primitives. The primitives include: move_to_pos, rotate_by_quat, set_vel, open_gripper, close_gripper. We do not provide primitives such as pick-and-place as they would be tailored for a particular suite of tasks that we do not constrain to in our study (similar to the control APIs for VoxPoser specified in Sec. A.3).

Move & Avoid: “Move to the top of [obj1] while staying away from [obj2]”, where [obj1] and [obj2] are randomized everyday objects selected from the list: apple, banana, yellow bowl, headphones, mug, wood block.

Set Up Table: “Please set up the table by placing utensils for my pasta”.

Close Drawer: “Close the [deixis] drawer”, where [deixis] can be “top” or “bottom”.

Open Bottle: “Turn open the vitamin bottle”.

Sweep Trash: “Please sweep the paper trash into the blue dustpan”.

A.5 Simulated Environment Setup

We implement a tabletop manipulation environment with a Franka Emika Panda robot in SAPIEN . The controller takes as input a desired end-effector 6-DoF pose, calculates a sequence of interpolated waypoints using inverse kinematics, and finally follows the waypoints using a PD controller. We use a set of 10 colored blocks and 10 colored lines in addition to an articulated cabinet with 3 drawers. They are initialized differently depending on the specific task. The lines are used as visual landmarks and are not interactable. For perception, a total of 4 RGB-D cameras are mounted at each end of the table pointing at the center of the workspace.

We create a custom suite of 13 tasks shown in Table 4. Each task comes with a templated instruction (shown in Table 4) where there may be one or multiple attributes randomized from the pre-defined list below. At reset time, a number of objects are selected (depending on the specific task) and are randomized across the workspace while making sure that task is not completed at reset and that task completion is feasible. A complete list of attributes can be found below, divided into “seen” and “unseen” categories:

[pos]: [“back left corner of the table”, “front right corner of the table”, “right side of the table”, “back side of the table”]

[obj]: [“blue block”, “green block”, “yellow block”, “pink block”, “brown block”]

[preposition]: [“left of”, “front side of”, “top of”]

[deixis]: [“topmost”, “second to the bottom”]

[region]: [“right side of the table”, “back side of the table”]

[velocity]: [“faster speed”, “a quarter of the speed”]

[line]: [“blue line”, “green line”, “yellow line”, “pink line”, “brown line”]

[pos]: [“back right corner of the table”, “front left corner of the table”, “left side of the table”, “front side of the table”]

[obj]: [“red block”, “orange block”, “purple block”, “cyan block”, “gray block”]

[preposition]: [“right of”, “back side of”]

[deixis]: [“bottommost”, “second to the top”]

[region]: [“left side of the table”, “front side of the table”]

[line]: [“red line”, “orange line”, “purple line”, “cyan line”, “gray line”]

A.5.2 Full Results on Simulated Environments

A.6 Prompts

Prompts used in Sec. 4.1 and Sec. 4.2 can be found below.

real-world: voxposer.github.io/prompts/real_planner_prompt.txt.

simulation: voxposer.github.io/prompts/sim_composer_prompt.txt.

real-world: voxposer.github.io/prompts/real_composer_prompt.txt.

parse_query_obj: Takes in a text query of object/part name and returns a list of dictionaries, where each dictionary corresponds to one instance of the matching object containing center position, occupancy grid, and mean normal vector.

simulation: voxposer.github.io/prompts/sim_parse_query_obj_prompt.txt.

real-world: voxposer.github.io/prompts/real_parse_query_obj_prompt.txt.

get_affordance_map: Takes in natural language parametrization from composer and returns a NumPy array for task affordance map.

simulation: voxposer.github.io/prompts/sim_get_affordance_map_prompt.txt.

real-world: voxposer.github.io/prompts/real_get_affordance_map_prompt.txt.

get_avoidance_map: Takes in natural language parametrization from composer and returns a NumPy array for task avoidance map.

simulation: voxposer.github.io/prompts/sim_get_avoidance_map_prompt.txt.

real-world: voxposer.github.io/prompts/real_get_avoidance_map_prompt.txt.

get_rotation_map: Takes in natural language parametrization from composer and returns a NumPy array for end-effector rotation map.

simulation: voxposer.github.io/prompts/sim_get_rotation_map_prompt.txt.

real-world: voxposer.github.io/prompts/real_get_rotation_map_prompt.txt.

get_gripper_map: Takes in natural language parametrization from composer and returns a NumPy array for gripper action map.

simulation: voxposer.github.io/prompts/sim_get_gripper_map_prompt.txt.

real-world: voxposer.github.io/prompts/real_get_gripper_map_prompt.txt.

get_velocity_map: Takes in natural language parametrization from composer and returns a NumPy array for end-effector velocity map.

simulation: voxposer.github.io/prompts/sim_get_velocity_map_prompt.txt.

real-world: voxposer.github.io/prompts/real_get_velocity_map_prompt.txt.