RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
Zhao Mandi, Shreeya Jain, Shuran Song
Introduction
Multi-robot systems are intriguing for their promise of enhancing task productivity, but are faced with various challenges. For robots to effectively split and allocate the work, it requires high-level understanding of a task, and consideration of each robot’s capabilities such as reach range or payload. Another challenge lies in low-level motion planning: as the configuration space grows with the number of robots, finding collision-free motion plans becomes exponentially difficult. Finally, traditional multi-robot systems typically require task-specific engineering, hence compromise generalization: with much of the task structures pre-defined, these systems are incapable of adapting to new scenarios or variations in a task.
We propose RoCo, a zero-shot multi-robot collaboration method to address the above challenges. Our approach includes three key components:
Dialogue-style task-coordination: To facilitate information exchange and task reasoning, we let robots ‘talk’ among themselves by delegating each robot to an LLM agent in a dialog, which allows robots to discuss the task in natural language, with high interpretability for supervision.
Feedback-improved Sub-task Plan Generated by LLMs: The multi-agent dialog ends with a sub-task plan for each agent (e.g. pick up object). We provide a set of environment validations and feedback (e.g. IK failures or collision) to the LLM agents until a valid plan is proposed.
LLM-informed Motion-Planning in Joint Space: From the validated sub-task plan, we extract goal configurations in the robots’ joint space, and use a centralized RRT-sampler to plan motion trajectories. We explore a less-studied capability of LLM: 3D spatial reasoning. Given the start, goal, and obstacle locations in task space, we show LLMs can generate waypoint paths that incorporate high-level task semantics and environmental constraints, and significantly reduce the motion planner’s sample complexity.
We next introduce RoCoBench, a benchmark with 6 multi-robot manipulation tasks. We experimentally demonstrate the effectiveness of RoCo on the benchmark tasks: by leveraging the commonsense knowledge captured by large language models (LLMs), RoCo is flexible in handling a variety of collaboration scenarios without any task-specific training.
In summary, we propose a novel approach to multi-robot collaboration, supported by two technical contributions: 1) An LLM-based multi-robot framework (RoCo) that is flexible in handling a large variety of tasks with improved task-level coordination and action-level motion planning; 2) A new benchmark (RoCoBench) for multi-robot manipulation to systematically evaluate these capabilities. It includes a suite of tasks that are designed to examine the flexibility and generality of the algorithm in handling different task semantics (e.g., sequential or concurrent), different levels of workspace overlaps, and varying agent capabilities (e.g., reach range and end-effector types) and embodiments (e.g., 6DoF UR5, 7DoF Franka, 20DoF Humanoid).
Preliminaries
Task Assumption. We consider a cooperative multi-agent task environment with robots, a finite time horizon , full observation space . Each agent has observation space . Agents may have asymmetric observation spaces and capabilities, which stresses the need for communication. We manually define description functions that translate task semantics and observations at a time-step into natural language prompts: . We also define parsing functions that map LLM outputs (e.g. text string “PICK object”) to the corresponding sub-task, which can be described by one or more gripper goal configurations.
Multi-Robot Collaboration with LLMs
We present RoCo, a novel method for multi-robot collaboration that leverages LLMs for robot communication and motion planning. The three key components in our method are demonstrated in Fig. 2 and described below:
We assume multi-agent task environments with asymmetric observation space and skill capabilities, which means agents can’t coordinate meaningfully without first communicating with each other. We leverage pre-trained LLMs to facilitate this communication. Specifically, before each environment interaction, we set up one round of dialog where each robot is delegated an LLM-generated agent, which receives information that is unique to this robot and must respond strictly according to its role (e.g. “I am Alice, I can reach …”).
For each agent’s LLM prompt, we use a shared overall structure but with agent-specific content, varied with each robot’s individual status. The prompt is composed of the following key components:
Task Context: describes overall objective of the task.
Round History: past dialog and executed actions from previous rounds.
Agent Capability: the agent’s available skills and constraints.
Communication Instruction: how to respond to other agents and properly format outputs.
Current Observation: unique to each agent’s status, plus previous responses in the current dialog.
Plan Feedback: (optional) reasons why a previous sub-task plan failed.
We use the following communication protocol to move the dialog forward: each agent is asked (instructions are given in Communication Instruction part of the prompt) to end their response by deciding between two options: 1) indicate others to proceed the discussion; 2) summarize everyone’s actions and make a final action proposal; this is allowed only if each agent has responded at least once in the current dialog round. This protocol is designed to allow the agents to freely discuss, while guaranteeing one sub-task plan will be proposed within a finite number of exchanges.
2 LLM-Generated Sub-task Plan
Once a round of dialog ends, the last speaking agent summarizes a ‘sub-task plan’, where each agent gets one sub-task (e.g. pick up object) and optionally a path of 3D waypoints in the task space. This sub-task plan is first passed through a set of validations before going into execution. If any of the checks fail, the feedback is appended to each agent’s prompt and another round of dialog begins. The validations are conducted in the following order, and each one assumes the previous check has passed (e.g. a plan must be parsed before checking for task and agent constraints):
Text Parsing ensures the plan follows desired format and contains all required keywords
Task Constraints checks whether each action complies with the task and agent constraints
IK checks whether each robot arm’s target pose is feasible via inverse kinematics
Collision Checking checks if the IK-solved arm join configurations cause collision
Valid Waypoints optionally, if a task requires path planning, each intermediate waypoint must be IK-solvable and collision-free, and all steps should be evenly spaced
The agents are allowed to re-plan until reaching a maximum number of attempts, after which the current round ends without any execution and the next round begins. The episode is considered failed if the task is not completed within a finite number of rounds.
3 LLM-informed Motion Planning in Joint Space
Once a sub-task plan passes all validations, we combine it with IK to produce a goal configuration jointly over all robot arms, and optionally, each step of the task space waypoint paths produces an intermediate goal configuration. The goal configuration(s) are passed to an RRT-based multi-arm motion planner that jointly plans across all robot arms, and outputs motion trajectories for each robot to execute in the environment, then the task moves on to the next round.
Toy Example: 3D Spatial Reasoning Capability in LLMs. As a motivating example, we ask an LLM (GPT-4) to plan multi-agent paths in a 3D grid. We randomly sample 3 agents’ (start, goal) coordinates and a set of obstacles, and prompt GPT-4 to plan collision-free paths. Given feedback on failed plans, allowing up to 5 attempts, GPT-4 obtains success rate over 30 runs, using on average attempts. See Fig. 3 for an example output. See Appendix 13 for further details and results. While this capability is encouraging, we found it to also be limited as the grid size and number of obstacles increase.
Benchmark
RoCoBench is a suite of 6 multi-robot collaboration tasks in a tabletop manipulation setting. The tasks involve common-sense objects that are semantically easy to understand for LLMs, and span a repertoire of collaboration scenarios that require different robot communication and coordination behaviors. See Appendix 10 for detailed documentation on the benchmark. We remark three key properties that define each task, summarized in Table 1:
Task decomposition: whether a task can be decomposed into sub-parts that can be completed in parallel or in certain order. Three tasks in RoCoBench have a sequential nature (e.g. Make Sandwich task requires food items to be stacked in correct order), while the other three tasks can be executed in parallel (e.g. objects in Pack Grocery task can be put into bin in any order).
Observation space: how much of the task and environment information each robot agent receives. Three tasks provide shared observation of the task workspace, while the other three have a more asymmetric setup and robots must inquire each other to exchange knowledge.
Workspace overlap: proximity between operating robots; we rank each task from low, medium or high, where higher overlap calls for more careful low-level coordination (e.g. Move Rope task requires manipulating the same object together).
Experiments
Overview. We design a series of experiments using RoCoBench to validate our approach. In Section 5.1, we evaluate the task performance of RoCo compared to an oracle LLM-planner that does not use dialog, and ablate on different components of the dialog prompting in RoCo. Section 5.2 shows empirically the benefits of LLM-proposed 3D waypoints in multi-arm motion planning. Section 5.3 contains qualitative results that demonstrate the flexibility and adaptability of RoCo. Additional experiment results, such as failure analysis, are provided in Appendix 12.
Experiment Setup. We use GPT-4 for all our main results. In addition to our method ‘Dialog’, we set up an oracle LLM-planner ‘Central Plan’, which is given full environment observation, information on all robots’ capabilities and the same plan feedback, and prompts an LLM to plan actions for all robots at once. We also evaluate two ablations on the prompt components of RoCo: first removes dialog and action history from past rounds, i.e. ‘Dialog w/o History’. Second removes environment feedback, i.e. ‘Dialog w/o Feedback’, where a failed action plan is discarded and agents are prompted to continue discussion without detailed failure reasons. To offset the lack of re-plan rounds, each episode is given twice the budget of episode length.
Evaluation Metric. Provided with a finite number of rounds per episode and a maximum number of re-plan attempts per round, we evaluate the task performance on three metrics: 1) task success rate within the finite rounds; 2) number of environment steps the agents took to succeed an episode, which measures the efficiency of the robots’ strategy; 3) average number of re-plan attempts at each round before an environment action is executed – this reflects the agents’ ability to understand and use environment feedback to improve their plans. Overall, a method is considered better if the task success rate is higher, it takes fewer environment steps, and requires fewer number of re-plans.
Results. The evaluation results are reported in Table 2. We remark that, despite receiving less information, dialog agents sometimes achieve comparable performance to the oracle planner. Particularly in Sort Cubes task, agents are able to find a strategy to help each other through dialog, but the oracle makes mistakes in trying to satisfy all agents’ constraints at once. While removing history information or plan feedback rounds does not negatively impact performance on some tasks, full prompt that includes both achieves the best overall results. Lastly, on Pack Gocery task, the oracle planner shows better capability in waypoint planning, displaying better capability at incorporating feedback and improve on individual coordinate steps.
2 Effect of LLM-proposed 3D Waypoints
We demonstrate the utility of LLM-proposed task space waypoints. We use two tasks that were designed to have high workspace overlap, i.e. Pack Grocery and Move Rope, which require both picking and placing to complete the task. For comparison, we define a hard-coded waypoint path that performs top-down pick or place, i.e. always hovers over a gripper atop a certain height before picking an object, and moves an object up to a height above the table before moving and placing. We single-out one-step pick or place snapshots, and run multi-arm motion planning using the compared waypoints, under a maximum of 300 second planning time budget. As shown in Fig. 4, LLM-proposed waypoints show no clear benefits for picking sub-tasks, but significantly accelerate planning for placing, where collisions are more likely to happen between the arms and the desktop objects.
3 Zero-shot Adaptation to Task Variations
Leveraging the zero- and few-shot ability of LLMs, RoCo demonstrates strong adaptation ability to varying task semantics, which traditionally would require modification or re-programming of a system, e.g. fine-tuning a learning-based policy. We showcase 3 main variation categories, all using Make Sandwich task in RoCoBench. 1. Object Initialization: The locations of the food items are randomized, and we show the dialog agents’ reasoning is robust to this variation. 2. Task Goal: The agents must stack food items in the correct order given in the sandwich recipe, and are able to coordinate sub-task strategies accordingly. 3. Robot Capability: The agents are able to exchange information on items that are within their respective reach and coordinate their plans accordingly.
4 Real-world Experiments: Human-Robot Collaboration
We validate RoCo in a real world setup, where a robot arm collaborates with a human to complete a sorting blocks task (Fig. 6). We run RoCo with the modification that only the robot agent is controlled by GPT-4, and it discusses with a human user that interacts with part of the task workspace. For perception, we use a pre-trained object detection model, OWL-ViT , to generate scene description from top-down RGB-D camera images. The task constrains the human to only move blocks from cups to the table, then the robot only picks blocks from table into wooden bin. We evaluate 2 main variation categories: 1) object initialization, i.e. initial block locations are randomized for each run (Fig. 6.1); 2) task order specification, where the agents are asked to follow a fixed order to move the blocks (Fig. 6.2). We also evaluate two types of human behaviors: first is an oracle human that corrects mistakes in the OWL-ViT-guided scene descriptions and the robot’s responses; second is an imperfect human that provides no feedback to those errors.
We evaluate 10 runs for each setup, see Table 3 for results. We report task success rate within the finite rounds, and number of steps the agents took to succeed an episode. We remark that task performance is primarily bottle-necked by incorrect object detection from OWL-ViT, which leads to either an incorrect object being picked up and resulting in failure or no object being picked up and resulting in higher steps. See Appendix 12.2 for further details on the real world experiments.
Multi-Agent Representation and Reasoning Dataset
In addition to our main experimental results, we curate a text-based dataset, RoCoBench-Text, to evaluate an LLM’s agent representation and task reasoning ability. This dataset aligns LLM with desirable capabilities in multi-agent collaboration, without requiring robotic environment interaction. It builds on data from our evaluation on RoCoBench, and contains a series of additional questions that are more open-ended and go beyond simply finding the next best action plan.
This dataset contains yes/no, multiple-choice or short question-answering questions, spanning a range of different reasoning abilities (see Appendix 14 for more details): Self-knowledge evaluates how well the agent establishes its identity under a given task context, divided into two categories: 1) understanding an agent’s own capability (e.g. which objects/area are not reachable); 2) memory retrieval, i.e. inferring information from past dialog and actions. Communication Skills evaluates an agent’s ability to effectively exchange information and drive a discussion into an agreeable plan. The questions ask an LLM to 1) choose appropriate response to other agents’ questions; 2) choose appropriate inquiries to other agents. Adaptation evaluates adaptation to unexpected situations that were not specified in context. We use a subset of RoCoBench tasks to design unexpected occurrences, either regarding task state (e.g. a missing object) or a response from another agent, and ask an LLM agent to choose the best response. See below for an example question: two agents make a sandwich together; one agent is informed of a broken gripper and must infer that the sandwich can actually be completed without any item from its side of the table.
You are a robot Chad collaborating with Dave …[task context] Your gripper is not working. What should you say to Dave? Select exactly one option from below. A: Sorry Dave, we can’t complete the task anymore, my gripper is broke. B: Let’s stop. The recipe needs ham but Dave can’t reach my side and my gripper is not functioning. C: Dave, go ahead and finish the sandwich without me, there isn’t anything we need on my side anyway.
2 Evaluation Results
Setup. All questions are designed to have only one correct answer, hence we measure the average accuracy in each category. We evaluate GPT-4 (OpenAI), GPT-3.5-turbo (OpenAI), and Claude-v1 (Anthropic). For GPT-4, we use two models marked with different time-stamps, i.e. 03/14/2023 and 06/13/2023. Results are summarized in Table 4: we observe that, with small performance variations between the two versions, GPT-4 leads the performance across all categories. We remark that there is still a considerable gap from fully accurate, and hope this dataset will be useful for improving and evaluating language models in future work. Qualitative Results. We observe GPT-4 is better at following the instruction to formulate output, whereas GPT-3.5-turbo is more prone to confident and elongated wrong answers. See below for an example response from an agent capability question (the prompt is redacted for readability). You are robot Chad .. [cube-on-panel locations…]. You can reach: [panels] Which cube(s) can you reach? […] Answer with a list of cube names, answer None if you can’t reach any. Solution: None GPT-4: None Claude-v1: yellow_trapezoid GPT-3.5-turbo: At the current round, I can reach the yellow_trapezoid cube on panel3.
Limitation
Oracle state information in simulation. RoCo assumes perception (e.g., object detection, pose estimation and collision-checking) is accurate. This assumption makes our method prone to failure in cases where perfect perception is not available: this is reflected in our real-world experiments, where the pre-trained object detection produces errors that can cause planning mistakes.
Open-loop execution. The motion trajectories from our planner are executed by robots in an open-loop fashion and lead to potential errors. Due to the layer of abstraction in scene and action descriptions, LLMs can’t recognize or find means to handle such execution-level errors.
LLM-query Efficiency. We rely on querying pre-trained LLMs for generating every single response in an agent’s dialog, which can be cost expensive and dependant on the LLM’s reaction time. Response delay from LLM querying is not desirable for tasks that are dynamic or speed sensitive.
Related Work
LLMs for Robotics. An initial line of prior work uses LLMs to select skill primitives and complete robotic tasks, such as SayCan , Inner Monologue , which, similarly to ours, uses environment feedback to in-context improve planning. Later work leverages the code-generation abilities of LLMs to generate robot policies in code format, such as CaP , ProgGPT and Demo2Code ; or longer programs for robot execution such as TidyBot and Instruct2Act . Related to our use of motion-planner, prior work such as Text2Motion , AutoTAMP and LLM-GROP studies combining LLMs with traditional Task and Motion Planning (TAMP). Other work explores using LLMs to facilitate human-robot collaboration , to design rewards for reinforcement learning (RL) , and for real-time motion-planning control in robotic tasks . While prior work uses single-robot setups and single-thread LLM planning, we consider multi-robot settings that can achieve more complex tasks, and use dialog prompting for task reasoning and coordination.
Multi-modal Pre-training for Robotics. LLMs’ lack of perception ability bottlenecks its combination with robotic applications. One solution is to pre-train new models with both vision, language and large-scale robot data: the multi-modal pre-trained PALM-E achieves both perception and task planning with a single model; Interactive Language and DIAL builds a large dataset of language-annotated robot trajectories for training generalizable imitation policies. Another solution is to introduce other pre-trained models, mainly vision-language models (VLMs) such as CLIP ). In works such as Socratic Models , Matcha , and Kwon et al. , LLMs are used to repeatedly query and synthesize information from other models to improve reasoning about the environment. While most use zero-shot LLMs and VLMs, works such as CogLoop also explores fine-tuning adaptation layers to better bridge different frozen models. Our work takes advantage of simulation to extract perceptual information, and our real world experiments follow prior work in using pre-trained object detection models to generate scene description.
Dialogue, Debate, and Role-play LLMs. Outside of robotics, LLMs have been shown to possess the capability of representing agentic intentions and behaviors, which enables multi-agent interactions in simulated environments such as text-based games and social sandbox . Recent work also shows a dialog or debate style prompting can improve LLMs’ performance on human alignment and a broad range of goal-oriented tasks . While prior work focuses more on understanding LLM behaviors or improve solution to a single question, our setup requires planning separate actions for each agent, thus adding to the complexity of discussion and the difficulty in achieving consensus.
Multi-Robot Collaboration and Motion Planning. Research on multi-robot manipulation has a long line of history . A first cluster of work studies the low-level problem of finding collision-free motion trajectories. Sampling-based methods are a popular approach , where various algorithmic improvements have been proposed . Recent work also explored learning-based methods as alternative. While our tasks are mainly set in static scenes, much work has also studied more challenging scenarios that require more complex low-level control, such as involving dynamic objects or closed-chain kinematics . A second cluster of work focuses more on high-level planning to allocate and coordinate sub-tasks, which our work is more relevant to. While most prior work tailor their systems to a small set of tasks, such as furniture assembly , we highlight the generality of our approach to the variety of tasks it enables in few-shot fashion.
Conclusion
We present RoCo, a new framework for multi-robot collaboration that leverages large language models (LLMs) for robot coordination and planning. We introduce RoCoBench, a 6-task benchmark for multi-robot manipulation to be open-sourced to the broader research community. We empirically demonstrate the generality of our approach and many desirable properties such as few-shot adaptation to varying task semantics, while identifying limitations and room for improvement. Our work falls in line with recent literature that explores harnessing the power of LLMs for robotic applications, and points to many exciting opportunities for future research in this direction.
This work was supported in part by NSF Award #2143601, #2037101, and #2132519. We would like to thank Google for the UR5 robot hardware. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of the sponsors. The authors would like to thank Zeyi Liu, Zhenjia Xu, Huy Ha, Cheng Chi, Samir Gadre, Mengda Xu, and Dominik Bauer for their fruitful discussions throughout the project and for providing helpful feedback on initial drafts of the manuscript.
References
RoCoBench
RoCoBench is built with MuJoCo physics engine. The authors would like to thank the various related open-source efforts that greatly assisted the development of RoCoBench tasks: DMControl , Menagerie, and MuJoCo object assets from Dasari et al. . The sections below provide a detailed documentation for each of the 6 simulated collaboration tasks.
2 Task: Sweep Floor
Task Description. 2 Robots bring a dustpan and a broom to opposite sides of each cube to sweep it up, then the robot holding dustpan dumps cubes into a trash bin.
Agent Capability. Two robots stand on opposite sides of the table:
UR5E with robotiq gripper (‘Alice’): holds a dustpan
Observation Space. 1) cube locations: a. on table; b. inside dustpan; c. inside trash bin; 2) robot status: 3D gripper locations
Available Robot Skills. 1) MOVE [target]: target can only be a cube; 2) SWEEP [target]: moves the groom so it pushes the target into dustpan; 3) WAIT; 4) DUMP: dump dustpan over the top of trash bin.
3 Task: Make Sandwich
Task Description. 2 Robots make a sandwich together, each having access to a different set of ingredients. They must select the required items and take turns to stack them in the correct order.
Agent Capability. Two robots stand on opposite sides of the table:
UR5E with suction gripper (‘Chad’): can only reach right side
Humanoid robot with suction gripper (‘Dave’): can only reach left side
Observation Space 1) the robot’s own gripper state (either empty or holding an object); 2) food items on the robot’s own side of the table and on the cutting board.
Available Robot Skills. 1) PICK [object]; 2) PUT [object] on [target]; WAIT
4 Task: Sort Cubes
Task Description. 3 Robots sort 3 cubes onto their corresponding panels. The robots must stay within their respective reach range, and help each other to move a cube closer.
Agent Capability. Three robots each responsible for one area on the table
UR5E with robotiq gripper (‘Alice’): must put blue square on panel2, can only reach: panel1, panel2, panel3.
Franka Panda (‘Bob’): must put pink polygon on panel4, can only reach: panel3, panel4, panel5.
UR5E with suction gripper (‘Chad’): must put yellow trapezoid on panel6, can only reach: panel5, panel6, panel7.
Observation Space 1) the robot’s own goal, 2) locations of each cube.
Available Robot Skills. 1) PICK [object] PLACE [panelX]; 2) WAIT
5 Task: Pack Grocery
Task Description. 2 Robots pack a set of grocery items from the table into a bin. The objects are in close proximity and robots must coordinate their paths to avoid collision.
Agent Capability. Two robots on opposite sides of table
UR5E with robotiq gripper (‘Alice’): can pick and place any object on the table
Franka Panda (‘Bob’): can pick and place any object on the table
Observation Space 1) robots’ gripper locations, 2) locations of each object, 3) locations of all slots in the bin.
Available Robot Skills. (must include task-space waypoints) 1) PICK [object] PATH [path]; 2) PLACE [object] [target] PATH [path]
6 Task: Move Rope
Task Description. 2 robots lift a rope together over a wall and place it into a groove. They must coordinate their grippers to avoid collision.
Agent Capability. Two robots on opposite sides of table
UR5E with robotiq gripper (‘Alice’): can pick and place any end of the rope within its reach
Franka Panda (‘Bob’): can pick and place any end of the rope within its reach
Observation Space 1) robots’ gripper locations, 2) locations of rope’s front and end back ends; 3) locations of corners of the obstacle wall; 4) locations of left and right ends of the groove slot.
Available Robot Skills. (must include task-space waypoints) 1) PICK [object] PATH [path]; 2) PLACE [object] [target] PATH [path]
7 Task: Arrange Cabinet
Task Description. 3 robots, two of them each hold one side of the cabinet door open, while the third robot takes the cups out and place them onto the correct coasters.
Agent Capability. Three robots, one on left side of the table, two on right side of table
UR5E with robotiq gripper (‘Alice’): stands on left side, can only reach left cabinet door
Franka Panda (‘Bob’): : stands on right side, can only reach right cabinet door
UR5E with suction gripper (‘Chad’):: stands on right side, can reach right cabinet door and cups and mugs inside the cabinet.
Observation Space 1) locations cabinet door handles; 2) each robot’s reachable objects, unaware of other robot’s reach range.
Available Robot Skills. 1) PICK [object]; 2) OPEN [one side of door handle]; 3) WAIT; 3) PLACE [object] [target]
Details on LLM Prompting
We use a separate query call for every agent’s individual response in a dialog, see the text box below for a redacted example of an agent’s prompt:
You are robot [agent name], collaborating with [other agent(s)] to [task context]. You can [agent capability]. Previously: [round history] At current round: [current observation] Discuss with [other agent(s)] to coordinate and complete the task together.[communication instruction]. Never forget you are [agent name]! Respond very concisely by [response format] Previous chat: [dialog from previous re-plan rounds] This proposed plan failed: [plan feedback] Current chat: [dialog from current round] Your response is: I am Alice, …
Additional Experiment Results
We provide example failure scenarios observed in the agent dialog and environment interaction.
Factual Errors in Dialog. We observe that when one LLM-generated agent makes a mistake, it could sometimes cause subsequent dialog to build on the wrong reasoning. See an example below: one agent (Alice) wrongfully decided the task is complete, and other agents repeat this wrong claim and choose to wait for multiple rounds while the task is, in fact, not finished. [Alice]: […] we have successfully completed our task. […], let’s wait for further instructions. [Bob]: I agree that we have successfully completed our task. Great teamwork, everyone! Let’s wait […] [Chad]: I concur that we have accomplished our goal. Excellent teamwork! Let’s wait […].
Errors in Motion Trajectory Execution Due to the open-loop nature of the execution, small errors in a motion trajectory could lead to unexpected errors, e.g. knocking of an object by accident.
2 Real World Experiment Setup
The robot agent is a 6DoF UR5E arm with suction gripper, and dialog is enabled by querying a GPT-4 model to respond as agent ‘Bob’, who is discussing with a human collaborator ‘Alice’. The human user provides text input to engage in the dialog, and arranges cubes on the same tabletop. For perception, we use top-down RGB-D image from an Azure Kinect sensor.
See the text below for an example of the robot’s prompt:
==== System Prompt ==== [Action Options] 1) PICK
==== User Prompt ==== You are Bob, your response is: response from GPT-4: EXECUTE NAME Bob ACTION …
Toy Example on LLM for 3D Path Planning
We use the chatCompletion mode in GPT-4. At each evaluation run, we randomly sample a new set of obstacles and agents’ start-goal locations. Each run is given up to 5 attempts: using the same system prompt, which describes the task and output instructions, and a user prompt, which describes the current grid layout. If a first attempt fails, the feedback from previous plans are appended to the user prompt at later attempts, until reaching the max number of attempts. See below for the prompts and GPT-4’s response (some coordinate tuples are omitted (marked as “omitted”) due to limited space).
==== System Prompt ==== Plan paths for agents to navigate a 3D grid to reach their respective goals and avoid collision. You are given: 1) a list of obstacle coordinates (x, y, z): locations of the obstacle grid cells, agents must avoid them. 2) a list of [([name], [init], [goal]) tuples], [init] and [goal] are 3D coordinates of the initial position and goal position of agent named [name]. 3) a previous plan, if any, and why it failed. Analyze this information and re-plan a collision-free path. How to plan a path: 1) Make sure each path does not touch any obstacle, … [planning instructions] [instruction on output format]
==== User Prompt ==== At the current step: Grid size: 10 x 10 x 10 Obstacles:(4, 9, 6) (0, 2, 0)… [list of all obstacle 3D locations] Agent Alice init: (7, 6, 2) goal: (5, 3, 2) Agent Bob init: (9, 9, 5) goal: (7, 1, 4) Agent Chad init: (1, 0, 0) goal: (9, 3, 6) Agent Dave init: (9, 1, 1) goal: (7, 1, 0) Feedback: this previous plan failed: [past plan, omitted] Some steps in this plan are not exactly 1 step away from each other: Bob: (7, 4, 5), (7, 1, 5); Use this information to try again, update this plan so it has collision-free, strictly one-step-apart paths. Feedback: this previous plan failed: …[another feedback, same format as above] Your reasoning and plan is: response from GPT-4: PLAN NAME Alice PATH [(7, 6, 2), (6, 6, 2), …[omitted] ] NAME Bob PATH [(9, 9, 5), (8, 9, 5), …[omitted] ] NAME Chad PATH [(1, 0, 0), (2, 0, 0), …omitted ] NAME Dave PATH [(9, 1, 1), (8, 1, 1), (7, 1, 1), (7, 1, 0)]
2 Full Experiment Results
We run 30 evaluation runs, each with a randomly sampled grid layout. See the left plot in Figure 13 for evaluation results. We make the following observations: first, the planning success rate varies moderately with the temperature parameter of GPT-4, which controls how deterministic its output generation is. A moderate temperature around 0.6, i.e. partially deterministic, shows the best success rate, and notably is able to 1-shot succeed the most times. Next, we observe the model’s ability to read from the feedback of a previous plan, and eventually succeed after a few more attempts. We also provide a visualization of an example output from one of the runs, where the path from first attempt has collision, and the model is able to correct this plan.
Multi-Agent Representation and Reasoning Dataset
57 total questions. Sim. Task: sequential transport
- Context (system prompt): 7 panels on the table, ordered left to right: panel1,…,panel7. They form a straight assembly line, panel1 is closed to panel2 and farthest from panel7. You are robot Alice in front of panel2. You are collaborating with Bob, Chad to sort cubes into their target panels. The task is NOT done until all three cubes are sorted. At current round: blue_square is on panel5 pink_polygon is on panel1 yellow_trapezoid is on panel3 Your goal is to place blue_square on panel2, but you can only reach panel1, panel2, panel3: this means you can only pick cubes from these panels, and can only place cubes on these panels. Never forget you are Alice! Never forget you can only reach panel1, panel2, panel3! - Question (user prompt): You are Alice. List all panels that are out of your reach. Think step-by-step. Answer with a list of panel numbers, e.g. means you can’t reach panel 1 and 2. - Solution: panels
1.2 Memory Retrieval
44 total questions. Task: Make Sandwich, Sweep Floor
- Context (system prompt): [History] Round#0: [Chat History] [Chad]: … [Dave]:… [Chad]: … … [Executed Action]… Round#1: …… - Current Round You are a robot Chad, collaborating with Dave to make a vegetarian_sandwich [……] You can see these food items are on your reachable side: … - Question (user prompt) You are Chad. Based on your [Chat History] with Dave and [Executed Action] from previous rounds in [History], what food items were initially on Dave’s side of the table? Only list items that Dave explicitly told you about and Dave actually picked up. Don’t list items that you are unsure about. Output the item names as a list. Think step-by-step. - Solution: bread_slice1
2 Challenge: Effective Communication
3 Inquiry
41 multiple-choice questions, using Arrange Cabinet task.
- Context (system prompt): You are Bob, collaborating with Alice, Chad to pick a mug and a cup out of cabinet, and place them on correct coasters. Both left and right cabinet doors should be OPENed and held open, while anything inside can be PICKed. You must coordinate to complete the task. At current round: left door is closed, right door is closed, mug is inside cabinet; cup is inside cabinet; Alice’s gripper is holding nothing, Your gripper is holding nothing, Chad’s gripper is holding nothing, Never forget you are Bob! Never forget you can only reach right door handle! - Question (user prompt): You are thinking about picking right door handle. Who and what should you ask to confirm this action? Think step-by-step, then choose exactly one option from below. [A] tell others about this plan because you are free and right door handle is within your reach. [B] ask if Alice and Chad can reach right door handle because it’s not within your reach. [C] ask if Alice and Chad can help, because you can reach right door handle, but you are busy and they are free. [D] all three of you are busy, so it’s better to wait until later. - Solution: [A]
4 Responsiveness
7 panels on the table, ordered left to right: panel1,…,panel7. They form a straight assembly line, panel1 is closed to panel2 and farthest from panel7. You are robot Alice in front of panel2. You are collaborating with Bob, Chad to sort cubes into their target panels. The task is NOT done until all three cubes are sorted. At current round: blue_square is on panel5 pink_polygon is on panel3 yellow_trapezoid is on panel3 Your goal is to place blue_square on panel2, but you can only reach panel1, panel2, panel3: this means you can only pick cubes from these panels, and can only place cubes on these panels. You can help others by moving other cubes that are within your reachable panels. Never forget you are Alice! Never forget you can only reach panel1, panel2, panel3! You are Alice. Chad needs to put yellow_trapezoid on panel6, and asked you for help. How should you respond? Think step-by-step. If you can help, your answer must start with YES and propose a plan to help. If you can’t help, must answer NO.
5 Challenge: Adaptation to Unexpected Scenarios
You are a robot Chad, collaborating with Dave to make a [vegetarian_sandwich]. Food items must be stacked following this order: bread_slice1, tomato, cheese, cucumber, bread_slice2, where bread_slice1 must be PUT on cutting_board. You must stay on right side of the table! This means you can only PICK food from right side, and Dave can only PICK from the other side. Only one robot can PUT at a time, so you must coordinate with Dave. At the current round: You can see these food items are on your reachable side: bread_slice1: on cutting_board cheese: atop tomato tomato: atop bread_slice1 cucumber: atop cheese ham: on your side beef_patty: on your side Your gripper is empty You are Chad. Your gripper is not working right now. What should you say to Dave? Select exactly one option from below. You must first output a single option number (e.g. A), then give a very short, one-line reason for why you choose it. Options: A: Sorry Dave, we can’t complete the task anymore, my gripper is broke. B: Let’s stop. The recipe needs ham but Dave can’t reach my side and my gripper is not functioning. C: Dave, go ahead and finish the sandwich without me, there isn’t anything we need on my side anyway.