Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, Andy Zeng

Introduction

Recent progress in training large language models (LLMs) has led to systems that can generate complex text based on prompts, answer questions, or even engage in dialogue on a wide range of topics. These models absorb vast quantities of knowledge from text corpora mined from the web, and we might wonder whether knowledge of everyday tasks that is encoded in such models can be used by robots to perform complex tasks in the real world. But how can embodied agents extract and harness the knowledge of LLMs for physically grounded tasks?

This question poses a major challenge. LLMs are not grounded in the physical world and they do not observe the consequences of their generations on any physical process . This can lead LLMs to not only make mistakes that seem unreasonable or humorous to people, but also to interpret instructions in ways that are nonsensical or unsafe for a particular physical situation. Figure 1 shows an example – a kitchen robot capable of executing skills such as “pick up the sponge” or ”go to the table” may be asked for help cleaning up a spill (“I spilled my drink, can you help?”). A language model may respond with a reasonable narrative that is not feasible or useful for the robot. “You could try using a vacuum cleaner” is impossible if there is no vacuum in the scene or if the robot is incapable of using one. With prompt engineering, a LLM may be capable of splitting the high-level instruction into sub-tasks, but it cannot do so without the context of what the robot is capable of given its abilities and the current state of the robot and the environment.

Motivated by this example, we study the problem of how to extract the knowledge in LLMs for enabling an embodied agent, such as a robot, to follow high-level textual instructions. The robot is equipped with a repertoire of learned skills for “atomic” behaviors that are capable of low-level visuomotor control. We make use of the fact that, in addition to asking the LLM to simply interpret an instruction, we can use it to score the likelihood that an individual skill makes progress towards completing the high-level instruction. Then, if each skill has an affordance function that quantifies how likely it is to succeed from the current state (such as a learned value function), its value can be used to weight the skill’s likelihood. In this way, the LLM describes the probability that each skill contributes to completing the instruction, and the affordance function describes the probability that each skill will succeed – combining the two provides the probability that each skill will perform the instruction successfully. The affordance functions make the LLM aware of the current scene, and constraining the completions to the skill descriptions makes the LLM aware of the robot’s capabilities. Furthermore, this combination results in a fully explainable sequence of steps that the robot will execute to accomplish an instruction – an interpretable plan that is expressed through language.

Our method, SayCan, extracts and leverages the knowledge within LLMs in physically-grounded tasks. The LLM (Say) provides a task-grounding to determine useful actions for a high-level goal and the learned affordance functions (Can) provide a world-grounding to determine what is possible to execute upon the plan. We use reinforcement learning (RL) as a way to learn language-conditioned value functions that provide affordances of what is possible in the world. We evaluate the proposed approach on 101 real-world robotic tasks that involve a mobile robot accomplishing a large set of language instructions in a real kitchen in a zero-shot fashion. Our experiments validate that SayCan can execute temporally-extended, abstract instructions. Grounding the LLM in the real-world via affordances nearly doubles the performance over the non-grounded baselines. Additionally, by evaluating the performance of the system with different LLMs, we show that a robot’s performance can be improved simply by enhancing the underlying language model.

Preliminaries

Language models seek to model the probability p(W)p(W) of a text W={w0,w1,w2,...,wn}W=\{w_{0},w_{1},w_{2},...,w_{n}\}, a sequence of strings ww. This is generally done through factorizing the probability via the chain rule to be p(W)=Πj=0np(wj∣w<j)p(W)=\Pi^{n}_{j=0}p(w_{j}|w_{<j}), such that each successive string is predicted from the previous. Recent breakthroughs initiated by neural network-based Attention architectures have enabled efficient scaling of so-called Large Language Models (LLMs). Such models include Transformers , BERT , T5 , GPT-3 , Gopher , LAMDA , FLAN , and PaLM , each showing increasingly large capacity (billions of parameters and terabytes of text) and subsequent ability to generalize across tasks.

In this work, we utilize the vast semantic knowledge contained in LLMs to determine useful tasks for solving high-level instructions.

In this work, we utilize TD-based methods to learn said value function that is additionally conditioned on the language command and utilize those to determine whether a given command is feasible from the given state. It is worth noting that in the undiscounted, sparse reward case, where the agent receives the reward of 1.01.0 at the end of the episode if it was successful and 0.00.0 otherwise, the value function trained via RL corresponds to an affordance function that specifies whether a skill is possible in a given state. We leverage that intuition in our setup and express affordances via value functions of sparse reward tasks.

SayCan: Do As I Can, Not As I Say

Connecting Large Language Models to Robots. While large language models can draw on a wealth of knowledge learned from copious amounts of text, they will not necessarily break down high-level commands into low-level instructions that are suitable for robotic execution. If a language model were asked “how would a robot bring me an apple”, it may respond “a robot could go to a nearby store and purchase an apple for you”. Though this response is a reasonable completion for the prompt, it is not necessarily actionable to an embodied agent, which may have a narrow and fixed set of abilities. Therefore, to adapt language models to our problem statement, we must somehow inform them that we specifically want the high-level instruction to be broken down into sequences of available low-level skills. One approach is careful prompt engineering , a technique to coax a language model to a specific response structure. Prompt engineering provides examples in the context text (“prompt”) for the model that specify the task and the response structure which the model will emulate; the prompt used in this work is shown in Appendix D.3 along with experiments ablating it. However, this is not enough to fully constrain the output to admissible primitive skills for an embodied agent, and indeed at times it can produce inadmissible actions or language that is not formatted in a way that is easy to parse into individual steps.

This has the added benefit of interpretability, as the model not only outputs generative responses, but also gives a notion of likelihood across many possible responses. Figure 3 (and Appendix Figure 12 in more detail) shows this process of forcing the LLM into a language pattern, where the set of tasks are the skills the low-level policy is capable of and prompt engineering shows plan examples and dialog between the user and the robot. With this approach, we are able to effectively extract knowledge from the language model, but it leaves a major issue: while the decoding of the instruction obtained in this way always consists of skills that are available to the robot, these skills may not necessarily be appropriate for executing the desired high-level task in the specific situation that the robot is currently in. For example, if I ask a robot to “bring me an apple”, the optimal set of skills changes if there is no apple in view or if it already has one in its hand.

Implementing SayCan in a Robotic System

Language-Conditioned Robotic Control Policies. To instantiate SayCan, we must provide it with a set of skills, each of which has a policy, a value function, and a short language description (e.g., “pick up the can”). These skills, value functions, and descriptions can be obtained in a variety of different ways. In our implementation, we train the individual skills either with image-based behavioral cloning, following the BC-Z method , or reinforcement learning, following MT-Opt . Regardless of how the skill’s policy is obtained, we utilize value functions trained via TD backups as the affordance model for that skill. While we find that the BC policies achieve higher success rates at the current stage of our data collection process, the value functions provided by the RL policies are crucial as an abstraction to translate control capabilities to a semantic understanding of the scene. In order to amortize the cost of training many skills, we utilize multi-task BC and multi-task RL, respectively, where instead of training a separate policy and value function per skill, we train multi-task policies and models that are conditioned on the language description. Note, however, that this description only corresponds to low level skills – it is still the role of the LLM in SayCan to interpret the high-level instruction and break it up into individual low level skill descriptions.

To condition the policies on language, we utilize a pre-trained large sentence encoder language model . We freeze the language model parameters during training and use the embeddings generated by passing in text descriptions of each skill. These text embeddings are used as the input to the policy and value function that specify which skill should be performed (see the details of the architectures used in the Appendix C.1). Since the language model used to generate the text embeddings is not necessarily the same as the language model used for planning, SayCan is able to utilize different language models well suited for different abstraction levels – understanding planning with respect to many skills as opposed to expressing specific skills more granularly.

Training the Low-Level Skills. We utilize both BC and RL policy training procedures to obtain the language-conditioned policies and value functions, respectively. To complete the description of the underlying MDP that we consider, we provide the reward function as well as the skill specification that is used by the policies and value functions. As mentioned previously, for skill specification we use a set of short, natural language descriptions that are represented as language model embeddings. We utilize sparse reward functions with reward values of 1.01.0 at the end of an episode if the language command was executed successfully, and 0.00.0 otherwise. The success of language command execution is rated by humans where the raters are given a video of the robot performing the skill, together with the given instruction. If two out of the three raters agree that the skill was accomplished successfully, the episode is labelled with a positive reward.

To learn language-conditioned BC policies at scale in the real world, we build on top of BC-Z and use a similar policy-network architecture (shown in Fig. 10). To learn a language-conditioned RL policy, we use MT-Opt in the Everyday Robots simulator using RetinaGAN sim-to-real transfer . We bootstrap the performance of simulation policies by utilizing simulation demonstrations to provide initial successes, and then continuously improve the RL performance with online data collection. We use a network architecture similar to MT-Opt (shown in Fig. 9). The action space of our policies includes the six degrees of freedom of the end-effector pose as well as gripper open and close commands, x-y position and yaw orientation delta of the mobile base of the robot, and the terminate action. Additional details on data collection and training are in Appendix Section C.2.

Robotic System and Skills. For the control policies, we study a diverse set of manipulation and navigation skills using a mobile manipulator robot. Inspired by common skills one might pose to a robot in a kitchen environment, we propose 551 skills that span seven skill families and 17 objects, which include picking, placing and rearranging objects, opening and closing drawers, navigating to various locations, and placing objects in a specific configurations. In this study we utilize the skills that are most amenable to more complex behaviors via composition and planning as well as those that have high performance at the current stage of data collection; for more details, see Appendix D.

Experimental Evaluation

Experimental Setup. We evaluate SayCan with a mobile manipulator and a set of object manipulation and navigation skills in two office kitchen environments. Figure 4 shows the environment setup and the robot. We use 15 objects commonly found in an office kitchen and 5 known locations with semantic meaning (two counters, a table, a trash can, and the user location). We test our method in two environments: a real office kitchen and a mock environment mirroring the kitchen, which is also the environment in which the robot’s skills were trained. The robot used is a mobile manipulator from Everyday Robots https://everydayrobots.com/ with a 7 degree-of-freedom arm and a two-fingered gripper. The LLM used is 540B PaLM unless stated otherwise for LLM ablations. We refer to SayCan with PaLM as PaLM-SayCan.

Instructions. To evaluate PaLM-SayCan, we test across 101 instructions from 7 instruction families, summarized in Table 1 and enumerated in Appendix E.1. These were developed to test various aspects of SayCan and were inspired by crowd sourcing via Amazon Mechanical Turk and in-person kitchen users, as well as benchmarks such as ALFRED and BEHAVIOR . The instructions span multiple axes of variation: time-horizon (from single primitives to 10+ in a row), language complexity (from structured language to fully crowd-sourced requests), and embodiment (variations over the robot and environment state). Table 1 details examples for each family.

Metrics. To understand the performance of the proposed method we measure two main metrics. The first is plan success rate, which measures whether the skills selected by the model are correct for the instruction, regardless of whether or not they actually successfully executed. We ask 3 human raters to indicate whether the plan generated by the model can achieve the instruction, and if 2 out of 3 raters agree that the plan is valid, it is marked a success. Note that for many instructions there may be multiple valid solutions. For example if the instruction is to “bring a sponge and throw away the soda can”, the plan can choose to bring sponge first or throw away the soda can first.

The second metric is execution success rate, which measures whether the full PaLM-SayCan system actually performs the desired instruction successfully. We ask 3 human raters to watch the robot execution. The raters are asked to answer the question “whether the robot achieves the task specified by the task string?” We mark an execution successful if 2 out of 3 raters agree that it is successful.

Table 2 shows the performance of PaLM-SayCan across 101 tasks. In the mock kitchen, PaLM-SayCan achieved a planning success rate of 84% and an execution rate of 74%. We also investigate PaLM-SayCan out of the lab setting and in the real kitchen to verify the performance of the policies and value functions in this setting. We find a reduction of planning performance by 3% and execution by 14%, indicating PaLM-SayCan and the underlying policies generalize reasonably well to the full kitchen. The full task list and results can be found in the Appendix Table 6, and videos of experiment rollouts and the decision making process can be found on the project website: say-can.github.io.

Figure 5 shows two long-horizon queries and the resulting rollouts. These tasks require PaLM-SayCan to plan many steps without error and for the robot to navigate and interact with a significant portion of the kitchen. Each query requires PaLM-SayCan to understand context implicit within the instruction. In Figure 5(a), the algorithm must understand the operator has asked for something to “recover from a workout”, i.e. something healthy, and thus it brings water and an apple rather than, e.g., a soda and chips. Furthermore, the algorithm must understand ordering and history, that it has already brought a drink and now must bring a snack before terminating. In Figure 5(b), PaLM-SayCan must track which objects are the “them” that need to be disposed of and where the sponge should be brought.

Figure 6 highlights PaLM-SayCan’s decision making, along with its interpretability. The decision making process can be understood as it solves instructions by visualizing what the two sides of the algorithm output. This allows a user to understand what options PaLM-SayCan is considering as language completions and what it believes is possible. We find that sequence order is understood (approaching objects before picking them up and picking them up before bringing them). Figure 6 shows that though the query mentions a coke, PaLM-SayCan understands that the important object is something to clean and brings a sponge. Appendix E.6 shows additional rollouts with complex decisions, embodiment grounding, and long-horizon tasks in Figures 21-24 as well as failures in Figure 23. We believe such real-time and clear interpretability opens avenues to more interactive operation.

When comparing the performance of different instruction families in Table 2 (see Table 1 for an explanation of families), we see that the natural language nouns performed worse than natural language verbs, due to the number of nouns possible (15 objects and 5 locations) versus number of verbs (6). The structured language tasks (created to ablate the performance loss of spelling out the solution versus understanding the query) were planned correctly 93% of the time, while their natural language verb counterparts were planned correctly 100%. This indicates the language model effectively parses the queries. The embodiment tasks were planned correctly 64% of the time, generally with failures as a result of affordance function misclassification. PaLM-SayCan planned and executed crowd-sourced natural queries with performance on par with other instruction families. PaLM-SayCan performed worst on the most challenging long-horizon tasks, where most failures were a result of early termination by the LLM (e.g., bringing one object but not the second). We also find that PaLM-SayCan struggles with negation (e.g., “bring me a snack that isn’t an apple”) and ambiguous references (e.g. asking for drinks with caffeine), which is a known issue inherited from underlying language models . Overall, 65% of the errors were LLM failures and 35% were affordance failures.

Returning to our initial example, “I spilled something, can you help?”, an ungrounded language model would respond with statements like “I can call you a cleaner” or “I can vacuum that up for you”, which given our robot are unreasonable. We have shown that PaLM-SayCan responds “I would: 1. find a sponge, 2. pick up the sponge, 3. bring it to you, 4. done” and is able execute this sequence on the robot in a real kitchen. This requires long-horizon reasoning over a required order, an abstract understanding of the instruction, and knowledge of both the environment and robot’s capabilities.

Ablating Language. To study the importance of the LLM, we conduct two ablation experiments using the language-conditioned policy (see Sections 4-4). In BC NL we feed the full instruction ii into the policy – this approach is representative of standard RL or BC-based instruction following methods . In BC USE we project the high-level instruction into the set of known language commands via the Universal Sentence Encoder (USE) embeddings by embedding the instruction, all the tasks, and the combinatorial set of sequences tasks (i.e., we consider “pick coke can” as well as “1. find coke can, 2. pick coke can” and so on), and selecting the highest cosine similarity instruction. The results in Table 2 illustrate the necessity of the language grounding where BC NL achieves 0% in all tasks and BC USE achieves 60% for single primitives, but 0% otherwise.

Ablating Value Functions. Table 2 illustrates the necessity of the affordance grounding. We compare PaLM-SayCan to (1) No VF, which removes the value function grounding (i.e., choosing the maximum language score skill) and to (2) Generative, which uses the generative output of the LLM and then projects each planned skill to its maximal cosine similarity skill via USE embeddings. The latter in effect compares to , which loses the explicit option probabilities, and thus is less interpretable and cannot be combined with affordance probabilities. For Generative we also tried BERT embeddings , but found poor performance. The No VF and Generative approaches performed similarly, achieving 67% and 74% planning success rate respectively, and worse than PaLM-SayCan’s 84%.

Ablating the Language Model. SayCan is able to improve with improved language models. The LLM used herein was PaLM , a 540B parameter model. In this section we ablate over 8B, 62B, and 540B parameter models as well as the 137B parameter FLAN model which is finetuned on a “instruction answering” dataset. Appendix Table 7 shows each model on a set of generative problems, where we find that generally larger models perform better, though the difference between the 62B and 540B model is small. Results in other works, such as Chain of Thought Prompting , indicate this difference may be more pronounced on more challenging problems – this is shown in Section 5.2. We also find that PaLM outperforms FLAN. Though FLAN was fine-tuned on instruction answering, the broader and improved dataset for PaLM may make up for this difference in training.

While it is expected that the generative performance of the language model will improve with better language models, it is unclear how the LLM size influences the final robotics success rate. Table 3 shows PaLM 540B and FLAN on robot running the full SayCan algorithm. The results show that the system using PaLM with affordance grounding (PaLM-SayCan) chooses the correct sequence of skills 84% of the time and executes them successfully 74% of the time, reducing errors by half compared to FLAN. This is particularly exciting because it represents the first time we can see how an improvement in language models translates to a similar improvement in robotics. This result indicates a potential future where the fields of language processing and robotics can collaboratively improve each other and scale together.

2 Case Studies of New Capabilities of PaLM-SayCan

PaLM-SayCan enables new capabilities. First, we show that it is very easy to incorporate new skills into the system, and use drawer manipulation as an example. Second, we show by leveraging chain of thought reasoning, we are able to solve complex tasks that require reasoning. Finally we show the system can work with multilingual queries, without explicitly being designed to.

Adding Skills: Drawer Manipulation (Appendix E.3). SayCan is capable of integrating new skills by simply adding the new skills as options for the LLM and providing accompanying value functions and add an example in the prompt with that skill. For example, with the skills open, close, and go to the drawer, SayCan is capable of solving tasks such as “restock the coke and pepsi into the drawer”. Over 21 queries we found a planning rate of 100% and an execution rate of 33% (due to failures of the chained manipulation policy), with no loss in performance for other instructions.

Chain of Thought Reasoning. SayCan can be integrated with recent work improving LLM reasoning, such as Chain of Thought . One limitation of vanilla SayCan is that it doesn’t handle tasks that involves negation. This is inherited from underline language models, and studied in the NLP community . However, we found by using chain-of-thought prompting we can improve SayCan on this front.

For chain-of-thought prompting-based SayCan, we need to modify the prompt to include a part called “Explanation”. We also slightly change how we use the language model. Instead of directly using the scoring interface to rank possible options, we first use the generative decoding of LLM to create an explanation, and then use the scoring mode, by including the explanation into the prompt. The full prompt is shown in Appendix E.4 Listing 3.

A few successful rollouts of the model at evaluation time is shown in Table 4. As we can see, with chain of thought prompting, the model can handle negations and tasks that require reasoning.

Multilingual Queries (Appendix E.5). While not explicitly designed to work with multilingual queries, PaLM-SayCan is able to handle them. The LLM was trained on multilingual corpora and thus SayCan can handle multilingual queries other than English. The results of SayCan on multilingual queries are summarized in Table. 9, and there is almost no performance drop in planning success rate when changing the queries from English to Chinese, French and Spanish.

Closed-Loop Planning. As presented herein, SayCan only receives environmental feedback through value functions at the current decision step, meaning if a skill fails or the environment changes, the necessary feedback may not be available. Owing to the extendibility and the natural language interface, Huang et al. builds upon SayCan to enable closed-loop planning by leveraging environment feedback (from e.g., success detectors, scene descriptors, or even human feedback) through an inner monologue.

Open Source Environment

We have open-sourced an implementation of SayCan in a Google Colab notebook at say-can.github.io/#open-source. The environment is shown in Figure 8 and is a tabletop with a UR5 robot and randomly generated sets of colored blocks and bowls. The low-level policy is implemented with CLIPort , which is trained to output a pick and place location. Due to the lack of a value function for this policy, the affordances are implemented with a ViLD object detector . GPT-3 is used as the open source language model . Steps are output in the form “pick up the object and place it in location”, leveraging the ability of LLMs to output code structures.

Related Work

Grounding Language Models. A significant body of work has focused on grounding language . Recent works have studied how to ground modern language models, by training them to accept additional environment inputs or to directly output actions . Others grounded language in an environment through prompt engineering . Concurrently with SayCan, Huang et al. use prompt engineering to extract temporally extended plans, but without any additional mechanism to ensure grounding, roughly corresponding to the “Generative” baseline in our experiments. The above methods are all trained without interaction with a physical environment, thus limiting their ability to reason over embodied interactions. One approach to grounding language models in interaction is by learning downstream networks with pre-trained LLM representations . Another approach finetunes language models with interactive data, such as rewards or ranking feedback of the interaction . In our work, SayCan is able to ground language models in the given environment through previously-trained value functions, enabling general, long-horizon behaviors in a zero-shot manner, i.e., without additional training.

Learning Language-Conditioned Behavior. There is a long history of research studying how to connect language and behavior . A large number of prior works have learned language-conditioned behavior via imitation learning or reinforcement learning . Most of these prior works focus on following low-level instructions, such as for pick-and-place tasks and other robotic manipulation primitives , though some methods address long-horizon, compound tasks in simulated domains . Like these latter works, we focus on completing temporally extended tasks. However, a central aspect of our work is to solve such tasks by extracting and leveraging the knowledge in large language models. While prior works have studied how pre-trained language embeddings can improve generalization to new instructions and to new low-level tasks , we extract much more substantial knowledge from LLMs by grounding them within the robot’s affordances. This allows robots to use language models for planning.

Task Planning and Motion Planning. Task and motion planning is a problem of sequencing tasks to solve a high-level problem, while ensuring the feasibility given an embodiment (task planning ; motion planning ). Classically, this problem has been solved through symbolic planning or optimization , but these require explicit primitives and constraints. Machine learning has recently been applied to enable abstract task specification, allow general primitives, or relax constraints . Others learn to hierarchically solve such long-horizon problems . SayCan leverages an LLM’s semantic knowledge about the world for interpreting instructions and understanding how to execute them. The use of LLMs and generality of learned low-level policies enables long-horizon, abstract tasks that scale effectively to the real world, as demonstrated in our robot experiments.

Conclusions, Limitations and Future Work

We presented SayCan, a method that enables leveraging and grounding the rich knowledge in large language models to complete embodied tasks. For real-world grounding, we leverage pre-trained skills, which are then used to condition the model to choose natural language actions that are both feasible and contextually appropriate. More specifically, we use reinforcement learning as a way to learn value functions for the individual skills that provide affordances of what is possible in the world, and then use textual labels for these skills as potential responses that are scored by a language model. This combination results in a symbiotic relationship where the skills and their value functions can act as the language model’s “hands and eyes,” while the language model supplies high-level semantic knowledge about how to complete a task. We evaluated the proposed approach on a number of real-world robotic tasks that involve a mobile manipulator robot accomplishing a large set of long-horizon natural language instructions in a real kitchen. We also demonstrated an exciting property of our method where a robot’s performance can be improved simply by enhancing the underlying language model.

While SayCan presents a viable way to ground language models in agents’ affordances, it has a number of limitations. First, we expect this method to inherit the limitations and biases of LLMs , including the dependence on the training data. Secondly, we observe that even though SayCan allows the users to interact with the agents using natural language commands, the primary bottleneck of the system is in the range and capabilities of the underlying skills. To illustrate this, we present representative failure cases in Appendix E. Future work that extends the repertoire of skills and improves their robustness would mitigate this limitation. In addition, at the current stage, the system is not easily able to react to situations where individual skills fail despite reporting a high value, though this could potentially be addressed by appropriate prompting of the language model for a correction.

There are many other potential avenues for future work. A natural question that this work raises is how the information gained through grounding the LLM via real-world robotic experience can be leveraged to improve the language model itself, both in terms of its factuality and its ability to perform common-sense reasoning about real-world environments and physics. Furthermore, since our method uses generic value functions to score affordances, it is intriguing to consider what other sources of grounding could be incorporated in the same manner, such as non-robotic contexts.

In the future, it is also interesting to examine whether natural language is the right ontology to use to program robots: natural language naturally incorporates contextual and semantic cues from the environment, and provides a level of abstraction which enables robots do decide on how to execute a strategy based on its own perception and affordances. At the same time, as opposed to, e.g., hindsight goal images , it requires supervision and it might not be the most descriptive medium for certain tasks.

Lastly, SayCan presents a particular way of connecting and factorizing the challenges of language understanding and robotics, and many further extensions can be proposed. Ideas such as combining robot planning and language , using language models as a pre-training mechanism for policies and many other ways of combining language and interaction are exciting avenues for future research.

The authors would like to thank Fred Alcober, Yunfei Bai, Matt Bennice, Maarten Bosma, Justin Boyd, Bill Byrne, Kendra Byrne, Noah Constant, Pete Florence, Laura Graesser, Rico Jonschkowski, Daniel Kappler, Hugo Larochelle, Benjamin Lee, Adrian Li, Maysam Moussalem, Suraj Nair, Jane Park, Evan Rapoport, Krista Reymann, Jeff Seto, Dhruv Shah, Ian Storz, Razvan Surdulescu, Tom Small, and Vincent Zhao for their help and support in various aspects of the project.

References

Appendix A Version Control

v1 →\rightarrow v2: Added PaLM results. Added study about new capabilities (drawer manipulation, chain of thought prompting, multilingual instructions). Added an ablation study of language model size. Added an open-source version of SayCan on a simulated tabletop environment. Improved readability.

Appendix B Contributions

Designed and built distributed robot learning infrastructure: Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Byron David, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Alex Irpan, Eric Jang, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Yao Lu, Peter Pastor, Kanishka Rao, Nicolas Sievers, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, and Mengyuan Yan.

Designed, implemented, or trained the underlying manipulation policies: Yevgen Chebotar, Keerthana Gopalakrishnan, Karol Hausman, Julian Ibarz, Alex Irpan, Eric Jang, Nikhil Joshi, Ryan Julian, Kuang-Huei Lee, Yao Lu, Kanishka Rao, and Ted Xiao.

Designed or implemented the data generation and curation or collected data: Noah Brown, Omar Cortes, Jasmine Hsu, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Linda Luu, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Clayton Tan, and Sichun Xu.

Designed or implemented SayCan: Karol Hausman, Brian Ichter, Sergey Levine, Alexander Toshev, and Fei Xia.

Managed or advised on the project: Chelsea Finn, Karol Hausman, Eric Jang, Sally Jesmonth, Sergey Levine, Yao Lu, Carolina Parada, Kanishka Rao, Alexander Toshev, and Vincent Vanhoucke.

Ran evaluations or experiments: Noah Brown, Omar Cortes, Brian Ichter, Rosario Jauregui Ruano, Kyle Jeffrey, Linda Luu, Jornell Quiambao, Jarek Rettinghouse, Diego Reyes, Clayton Tan, and Fei Xia.

Scaled simulation infrastructure: Nikhil J Joshi, Yao Lu, Kanishka Rao, and Ted Xiao.

Wrote the paper: Chelsea Finn, Karol Hausman, Brian Ichter, Alex Irpan, Sergey Levine, Fei Xia, and Ted Xiao.

Created open-source environment: Brian Ichter, Andy Zeng.

B.2 By Person

Michael Ahn developed the deployment system that enabled the ability to scale up data collection on real robots. Anthony Brohan implemented the logging system for the project and designed and implemented the data labeling pipelines. Noah Brown led and coordinated the real-robot operations including data collection with teleoperators, evaluations and the real-world setup. Yevgen Chebotar designed and implemented multiple offline RL methods allowing the manipulation policies to process data coming from different sources. Omar Cortes collected data on the robots and ran and supervised real-world evaluations. Byron David developed simulation assets and performed system identification. Chelsea Finn advised on the project, helped set the research direction and wrote parts of the paper. Chuyuan Fu developed deformable objects and experimented with sim-to-real techniques. Keerthana Gopalakrishnan provided multiple infrastructure contributions that allowed for scalable learning of manipulation policies. Karol Hausman co-led the project as well as developed SayCan, helped set the research direction, trained the underlying manipulation policies, and wrote the paper. Alex Herzog developed the teleoperation tools and implemented multiple infrastructure tools that allowed for continuous robot operation. Daniel Ho helped develop sim-to-real pipelines for manipulation policies. Jasmine Hsu provided logging and monitoring infrastructure tools as well as data labeling pipelines. Julian Ibarz provided multiple contributions that enabled scaling learning algorithms for manipulation policies, and helped set the research direction. Brian Ichter initiated and led the SayCan algorithm, combined the manipulation and navigation skills, ran experiments for the paper, created the open-sourced SayCan Colab, and wrote the paper. Alex Irpan set up and led the autonomous data collection effort as well as verified the data collected by the robots, and wrote parts of the paper. Eric Jang helped set the research and team direction, managed the data for learning, developed the behavioral cloning manipulation policies, and wrote parts of the paper. Rosario Jauregui Ruano collected data on the robots and ran and supervised real-world evaluations. Kyle Jeffrey collected data on the robots and ran and supervised real-world evaluations. Sally Jesmonth was the program manager for the project. Nikhil J Joshi developed a number of simulation and infrastructure tools that allowed to scale up simulation training. Ryan Julian developed multi-modal network architectures and trained manipulation policies. Dmitry Kalashnikov contributed infrastructure pieces that enabled training from logged data. Yuheng Kuang implemented the logging system for the project and designed and implemented the data labeling pipelines Kuang-Huei Lee made improvements to training algorithms for manipulation policies. Sergey Levine advised on the project, helped set the research direction, developed SayCan, and wrote parts of the paper. Yao Lu led and designed the robot learning infrastructure for the project providing most of the tools and improving manipulation policies. Linda Luu ran multiple evaluations, collected data and helped establish real-robot operations. Carolina Parada advised on the project, managed the team, helped write the paper, and helped set the research direction. Peter Pastor provided infrastructure tools that allowed for continuous robot operations. Jornell Quiambao collected data on the robots and ran and supervised real-world evaluations. Kanishka Rao co-led the project, managed the team, helped set the research direction and contributed to training manipulation policies. Jarek Rettinghouse collected data on the robots and ran and supervised real-world evaluations. Diego Reyes collected data on the robots and ran and supervised real-world evaluations. Pierre Sermanet set up the crowd compute rating pipeline. Nicolas Sievers provided simulation assets and environments used for simulation training. Clayton Tan collected data on the robots and ran and supervised real-world evaluations and helped establish real-robot operations. Alexander Toshev advised on the project, developed SayCan, helped write the paper, and helped set research direction. Vincent Vanhoucke advised on the project, managed the team, and helped write the paper. Fei Xia developed, implemented, and led on-robot SayCan, ran the experiments for the paper, created the demos, and wrote the paper. Ted Xiao led the scaling of manipulation skills, designed and developed learning from simulation for manipulation skills, and developed multi-modal network architectures. Peng Xu was the engineering lead for integrating manipulation and navigation and developed the underlying infrastructure for SayCan. Sichun Xu developed remote teleoperation tools that allowed scaling up data collection in simulation. Mengyuan Yan implemented infrastructure and learning tools that allowed for learning manipulation policies from different data sources. Andy Zeng provided codebase and simulation assets for the open-source SayCan Colab.

B.3 Corresponding Emails:

Appendix C RL and BC Policies

The RL models use an architecture similar to MT-Opt , with slight changes to support natural language inputs (see Fig. 9 for the network diagram). The camera image is first processed by 7 convolutional layers. The language instruction is embedded by the LLM, then concatenated with the robot action and non-image parts of the state, such as the gripper height. To support asynchronous control, inference occurs while the robot is still moving from the previous action. The model is given how much of the previous action is left to execute . The conditioning input goes through FC layers, then tiled spatially and added to the conv. volume, before going through 11 more convolutional layers. The output is gated through a sigmoid, so the Q-value is always in $$.

The BC models use an architecture similar to BC-Z (see Fig. 10 for the network diagram). The language instruction is embedded by a universal sentence encoder , then used to FiLM condition a Resnet-18 based architecture. Unlike the RL model, we do not provide the previous action or gripper height, since this was not necessary to learn the policy. Multiple FC layers are applied to the final visual features, to output each action component (arm position, arm orientation, gripper, and the termination action).

C.2 RL and BC Policy Training

In addition to using demonstrations in the BC setup, we also learn language-conditioned value functions with RL. For this purpose, we complement our real robot fleet with a simulated version of the skills and environment. To reduce the simulation-to-real gap we transform robot images via RetinaGAN to look more realistic while preserving genera object structure. In order to learn a language-conditioned RL policy, we utilize MT-Opt in the Everyday Robots simulator using said simulation-to-real transfer. We bootstrap the performance of simulation policies by utilizing simulation demonstrations to provide initial successes, and then continuously improve the RL performance with online data collection in simulation. Standard image augmentations (random brightness and contrast) as well as random cropping were applied. The 640 x 512 input image was padded by 100 pixels left-right and 40 pixels top-down, then cropped back down to a 640 x 512 image, so as to allow for random spatial shifts without limiting the field of view. We use a network architecture similar to MT-Opt (shown in Fig. 9).

The RL model is trained using 16 TPUv3 chips and for about 100 hours, as well as a pool of 3000 CPU workers to collect episodes and another 3000 CPU workers to compute target Q-values. Computing target Q-values outside the TPU allows the TPU to be used solely for computing gradient updates. Episode rewards are sparse and always 0 or 1, so the Q-function is updated using a log loss. Models were trained using prioritized experience replay , where episode priority was tuned to encourage replay buffer training data for each skill to be close to 50% success. Episodes were sampled proportionally to their priority, defined as 1+10⋅∣p−0.5∣{1+10\cdot|p-0.5|}, where pp is the average success rate of episodes in the replay buffer.

We use 6800068000 teleoperated demonstrations that were collected over the course of 11 months using a fleet of 10 robots. The operators use VR headset controllers to track the motion of their hand, which is then mapped onto the robot’s end-effector pose. The operators can also use a joystick to move the robot’s base. We expand the demonstration dataset with 276000276000 autonomous episodes of learned policies which are later success-filtered and included in BC training, resulting in an additional 1200012000 successful episodes. To additionally process the data, we also ask the raters to mark the episodes as unsafe (i.e., if the robot collided with the environment), undesirable (i.e., if the robot perturbed objects that were not relevant to the skill) or infeasible (i.e., if the skill cannot be done or is already accomplished). If any of these conditions are met, the episode is excluded from training.

C.3 RL and BC Policy Evaluations

In order to obtain the best possible manipulation capabilities for use in SayCan, we use a separate evaluation protocol for iterating on the RL and BC policies in the Mock Office Kitchen stations. Evaluations are divided by skill (pick up, knock over, place upright, open/close drawers, move object close to another one), and within each skill, 18-48 skills are sampled from a predetermined set of three objects. Object positions are randomized on each episode, with one or two objects serving as a distractor.

The episode ends when 50 actions have been taken or the policy samples a terminate action. A human operator supervises multiple robots performing evaluation and performs scene resets as needed, and records each episode as a success or failure. Models whose per-skill performance outperforms prior models are ”graduated” to the same evaluation protocol in the real kitchen, and then integrated into SayCan. We found that despite the domain shift from Mock Office Kitchen stations to the actual kitchen counter and drawers, higher success rates on mock stations usually corresponded to higher success rates in the real kitchen setting.

Figure 11 shows the development of the manipulation skills over time. It reports the per-skill success rate, the average success rate across all skills, and the number of instructions the policy was trained on. Over the course of the project, we increased the number of skills evaluated, from 1 instruction in April 2021 to hundreds of instructions at time of publication over the course of 366 real-world model evaluations.

Appendix D SayCan Details and Parameters

Figure 12 shows scoring approach used and prompt engineering for the LLM side of SayCan. Figure 2 shows how robotic affordances are computed with value functions and real value function computations at different states. These two components are combined to form SayCan, as detailed in Algorithm 1 and in Figure 13.

D.2 Policies and Affordance Functions

We also note a few practical considerations for setting up our affordance functions and policies. The flexibility of our approach allows us to mix and match policies and affordances from different methods. For the pick manipulation skills we use a single multi-task, language-conditioned policy, for the place manipulation skills we use a scripted policy with an affordance based on the gripper state, and for navigation policies we use a planning-based approach which is aware of the locations where specific objects can be found and a distance measure. In order to avoid a situation where a skill is chosen but has already been performed or will have no effect, we set a cap for the affordances indicating that the skill has been completed and the reward received.

SayCan is capable of incorporating many different policies and affordance functions through its probability interface. Though in principle each type of skill has been trained with the pipeline described in Appendix C, to the success rates seen in Figure 11, we wish to show the generality of SayCan to different policies and affordance functions as well as the robustness of other functions (e.g. distance for navigation). Furthermore, some skills (such as the manipulation skill “move object near object” and “knock object over”) are not naturally part of long-horizon tasks and thus we do not utilize them. Other skills, such as drawer opening, were not consistent enough for long-horizon planning and thus unused. However, we note that as skills become performant or as new skills are learned, it is straightforward to incorporate these skills by adding them as options for LLM scoring and as examples in the prompt. We use the following for each skill family:

Pick. For pick we use the learned policies in Appendix C and Section 4 with actions from BC and value functions from RL trained on the same skill. In natural language these are specified as “pick up the object”.

Go to. Since the focus of this work is mainly on planning, we assume the location of objects are known. Thus any navigation skill maps to the coordinate of the object with a classical planning-based navigation stack. In natural language these are specified as “go to location” and “find object”.

Place. Though our manipulation policies have a “place upright” skill, this skill only applies to objects that have a canonical upright direction, e.g., a water bottle but not a bag of chips. One could also train a universal “place” command, but our current policies are trained in a setup-free environment and thus are not amenable to an initial pick. Thus to have a consistent place policy across all objects we use a classical motion planning policy. We use Cartesian space motion planning to plan a path from pre-grasp pose shown in Figure 4 to a gripper release pose. The robot executes that path until the gripper is in contact with a supporting surface, and then the gripper opens and releases the object. In natural language these are specified as “put down the object”.

Pick. We find the trained value functions generally have a minimum value for when a skill is not possible and a maximum when the skill is successful and thus we normalize the value function to get a affordance function with

Go to. The affordance function of go to skills are based on the distance dd (in meters) to the location. We use

D.3 LLM Prompt

The LLM uses prompt engineering and a strict response structure to score skills. But, as SayCan as a whole requires affordances from a world embodiment, it is not straightforward to optimize this structure and tune parameters quickly. Thus we built a language-based simulator which, given a query and a solution sequence of skills, outputs affordances consistent with the query and solution. It also generates consistent distractor affordances to ensure robustness. The simulator then verifies that SayCan recovers the correct solution and tests how confident SayCan is in the correct solutions. In Table 5 we test the effect of the number of examples in the prompt on the planning success rate in the language-based simulator (over 50 demonstrative instructions). We show a success rate with and without requiring the plan to terminate; without examples we found the LLM was unlikely to issue a “done” phase. With no examples SayCan is able to successfully plan 54% without the done condition, but only 10% with the done condition. Though it makes mistakes, clearly some information is already imbued within the language model. It is able to correctly solve “Can I have a redbull please?” and “Move the chips bag from the table to the counter.”. With only one example the LLM quickly improves in both planning rates, though still fails to terminate the plan occasionally. After only four examples the LLM is performant, planning 82% of the queries correctly, though the remaining errors are largely within a single instruction family: Long-Horizon. Finally, the prompt used in this work, Listing 1, involved 17 examples and recovered 88% of the solutions correctly.

We note here briefly a few lessons learned in prompt engineering and structuring the final prompt. Providing explicit numbers between steps (e.g., 1., 2., instead of combining skills with “and then” or other phrases) improved performance, as did breaking each step into a separate line (e.g. adding a “\n” between steps). Examples which overly include objects used in the actual planning tend to bias results to those objects (e.g., if every example is about apples then the apple scoring will be off in planning). Phrasing of the natural language names of skills and objects is important due to the auto-regressive nature of the LLM scoring – skills and objects should be naturally named and errors such as misspellings or mismatches in “a” vs “an” can be problematic. Notably, since user generated instructions are taken as given such fragility is not issues for the input, allowing a robustness to user queries. For our language model, PaLM , structuring the interaction as dialog (How would you - I would) was both more natural and performant. Although dialog is used as prompt, the model generalized to imperative sentences at deployment time.

Appendix E Experiments

Below we include every instruction run, which environment it was run in, and its planning and execution success rate. Table 6 shows all instructions as broken down by instruction family, listed below and initially defined in Section 5 Table 1.

Natural Language (NL) Single Primitive. Given a natural language command corresponding to performing a single primitive, can SayCan recover that primitive skill and terminate?

NL Noun. Given a natural language query that replaces a noun (typically an object or location) with a synonym, can SayCan execute the appropriate sequence?

NL Verbs. Given a natural language query that replaces a verb (typically an action) with a synonym, can SayCan execute an appropriate sequence?

Structured Language. Given a structure language query that mirrors the NL Verbs and spells out the sequence of commands, how well can SayCan plan compared to NL Verbs? This acts as an ablation to see the performance loss of understanding a natural language query over an explicit solution.

Embodiment. Given a query with different environment and robot states, can SayCan still execute at a high rate? This tests the performance of SayCan’s affordance model and the LLM’s ability to reason within it.

Crowd-Sourced. These queries were crowd sourced from Mechanical Turk by giving humans a description of what occurred (e.g., an apple was moved in front of you) and asking them what they would ask the robot to do. They were also crowd sourced by asking humans in a real office kitchen to command the robot to perform tasks (given knowledge of the robot’s abilities). This tests SayCan’s performance with natural requests.

Long-Horizon. These challenging queries require SayCan to reason over temporally extended instructions to investigate how well it scales to such regimes.

E.2 Ablating Over Language Model Size

To test how SayCan scales with LLM size, we ablate over varying PaLM sizes (8B, 62B, and 540B parameter models) as well as the 137B parameter FLAN model. The results are discussed in Section 5.1 and shown in Table 7.

E.3 Adding Skills: Drawer Manipulation

In order to support drawer manipulation we added another category of skills in SayCan.

Drawer Manipulation. For drawer manipulation we use the learned policies in Appendix C and Section 4 with actions from BC and value functions from heuristics (If the robot is next to the drawer, all drawer tasks are possible). In natural language these are specified as “open the drawer”, “close the drawer”, “put the object in the drawer”, “take the object out of the drawer”.

A few drawer-specific prompts also need to be added to teach the robot how to chain the drawer skills together. The prompts are shown in Listing 2.

The results of the drawer tasks are shown in Table. 8. SayCan achieved an overall planning success rate of 100% and execution success rate of 33%. The main failure cases are manipulation failures, where the robot fails to open the drawer wide enough to put objects in it, or fails to completely close the drawer.

E.4 Chain of Thought Reasoning

Here we show the chain of thought prompt that generates the rollout in Table 4.

E.5 Multilingual Queries

Since the underlying LM we used has been trained on multilingual corpora, SayCan can handle multilingual queries out of the box. The results of SayCan on multilingual queries are summarized in Table. 9, and there is almost no performance drop on planning success rate when changing the queries from English to Chinese, French and Spanish.

E.6 Additional Results

Additional results are shown in Figure 21 and Figure 24 and some failure cases in Figure 23. For videos of the rollouts, please visit the our website https://say-can.github.io