Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, Hongsheng Li

Introduction

Recently, Large language models (LLMs) such as GPT-3 brown2020language , LLaMA touvron2023llama , and ChatGPT have made unprecedented progress in human-like text generation and understanding of natural language instructions. These models demonstrate remarkable zero-shot generalization abilities after being trained on large collected corpora and absorbing human feedback. The follow-up work Visual ChatGPT wu2023visual incorporates a variety of visual foundation models to achieve visual drawing and editing using a prompt manager. Additionally, VISPROG gupta2022visual proposes a neuro-symbolic approach for complex visual tasks, including image understanding, manipulation, and knowledge retrieval. Inspired by the tremendous potential of combining LLMs and multi-modality foundation models, we aim to develop general robotic manipulation systems. Could we build a ChatGPT-like robotic systems that support robotic manipulation, visual goal-reaching, and visual reasoning?

Developing a general-purpose robotic system capable of executing complex tasks in dynamic environments poses a significant challenge in robotics research. Such a system must possess the ability to perceive the surroundings, select relevant robotic skills, and sequence them accordingly to accomplish long-term goals. Achieving such functionalities requires the integration of various technologies, including perception, planning, and control, to enable the robot to operate autonomously in unstructured environments. Inspired by the strong capability of synthesizing simple Python programs from docstrings of LLMs, CaP liang2022code directly generates the robot-centric policy code based on several in-context example language commands. However, it is restricted to what the perception APIs can provide and struggle to interpret longer and more complex commands due to the high precision requirements of the code.

To address these challenges, we propose a novel approach that utilizes multi-modality foundation models and LLMs to simultaneously implement perceptual recognition, task planning, and low-level control modules. Unlike existing methods such as CaP liang2022code , which directly generates policy codes, we generate decision-making actions that can help reduce the error rate of executing complex tasks. Specifically, we use various foundation models such as SAM and CLIP to accurately locate and classify objects in the environment. We then combine the information with robotic skills to generate decision-making actions by LLMs.

We evaluate our proposed approach on multiple domains and scenarios, including simple object manipulation, visual goal-reaching, and visual reasoning. Our framework provides an easy-to-use general-purpose robotic system and shows strong competitive performance on six representative meta-tasks from VIMABench jiang2022vima . Our proposed approach can serve as a strong baseline method in the field of robotic research and contribute to the development of more intelligent and capable robots.

The contributions of our papers can be summarized as follows:

General-purpose robotic system. We introduce a general-function robotic system, Instruct2Act, that leverages the in-context learning ability of LLMs and multi-modality instructions to generate middle-level decision-making actions from both natural language and visual instructions.

Flexible modality inputs. This paper investigates unified modality instruction inputs on robotic tasks, such as manipulation and reasoning, and presents a flexible retrieval architecture to handle varying instruction types.

Strong zero-shot performance with minimal code overhead. The proposed Instrcut2Act has shown superior performance in comparison to state-of-the-art learning-based policies, even without fine-tuning. Additionally, the adaptation effect of using foundation models is comparatively minor, as opposed to the training-from-scratch methods.

Related Works

Language in robotics offers not only a user-friendly interface but also the potential for cross-task skill generalization and long-horizon task reasoning. As a result, instruction-based policies have been a popular area of research in robotics brohan2023can ; shao2021concept2robot ; shridhar2018interactive . Recently, with the emergence of multi-modality models, CLIPORT shridhar2022cliport has given Transporter zeng2021transporter the ability to understand semantics and manipulate objects by encoding text input through CLIP radford2021learning . shridhar2023perceiver extended the CLIPORT to the 3D domain by employing voxelized observation and action spaces in their Perceiver-Actor model. brohan2023can employed the 540B PaLM chowdhery2022palm to accomplish zero-shot concept grounding in their SayCan model. huang2022language utilized two LLMs in their approach, where one was used for zero-shot planning generation and the other one was used for admissible action mapping. huang2022inner enhanced their method by integrating closed-loop feedback, such as scene descriptors and success detectors, for performing robotic tasks. VIMA jiang2022vima developed a large-scale benchmarking dataset by designing a multimodal prompts-conditioned framework. CaP liang2022code directly generates policy codes with detailed comments and context-specific examples to guide LLM output. cui2023no conducted a further investigation of the shared autonomy regime. Their approach involves a fusion of the correction signal from human instruction and the static controller with the original policy during inference. PaLM-Edriess2023palm built a large VL model for embodied agents by integrating the power of 540B PaLM chowdhery2022palm and 22B ViT dehghani2023scaling . Text2Motion lin2023text2motion predicted the goal state and selected feasible actions using LLM while considering geometric constraints. Socratic Models zeng2022socratic generated prompts in their approach by incorporating perceptual information into LLM using VL models. The concurrent work wu2023tidybot uses LLM to summarize the human’s preferences given a few examples. Our Instruct2Act framework achieves great flexibility while retaining expert domain knowledge by combining robotic primitive skills with LLM.

2 Foundation Models on Computer Vision Tasks

Several studies chen2020uniter ; zhang2021vinvl have utilized frozen pre-trained image encoders to improve the visual features extracted from images. And wang2022internimage adopted deformable convolutions in their large-scale visual foundation models for better image analysis. In addition, leveraging self-training and a massive dataset of 27M image-text pairs, GLIP li2022grounded achieved strong zero-shot transfer ability. Segment Anything Model (SAM) kirillov2023segment , a segmentation foundation model trained on more than one billion mask samples, allows for zero-shot transfer to diverse tasks through prompt engineering. Furthermore, pre-trained LLMs have shown significant progress in text understanding and generation vaswani2017attention ; brown2020language ; gao2023llama ; touvron2023llama . And a such breakthrough in LLMs also benefits the VL tasks fu2021violet ; zhang2021vinvl . Recently, Recently, there have been explorations to combine the reasoning capacity of LLMs with the visual understanding ability of visual foundation models. VISPROG gupta2022visual utilizes in-context learning in GPT-3 to generate a program for new instruction and demonstrates the system’s compositional visual reasoning ability. ViperGPT suris2023vipergpt leverages code-generation models to produce the results of language queries by means of composing foundation models into subroutines. Visual ChatGPT wu2023visual incorporates multiple visual foundation models and allows users to interact with ChatGPT through the proposed prompt manager, which allows multiple AI models reasoning ability with multi-steps. Similarly, the proposed Instruct2Act aims at endowing robotics the perception ability by incorporating advanced foundation models through the reasoning ability of LLM.

3 Foundation Models in Robotics

In addition to language-conditioned robotic manipulation, the use of foundation models has also led to significant advancements in robotics. LID li2022pre proposes a general approach to sequential decision-making that uses a pre-trained language model (LM) to initialize a policy network where goals and observations are embedded. R3M nair2022r3m explores how visual representations obtained by training on diverse human video data grauman2022ego4d using time-contrastive learning and video-language radosavovic2023real can enable data-efficient learning of downstream robotic manipulation tasks. Meanwhile, CACTI mandi2022cacti suggests a scalable framework for visual imitation learning that utilizes pre-trained models to map pixel values to low-dimensional latent embeddings for improved generalization ability. DALL-E-Bot kapelyukh2022dall uses Stable Diffusion rombach2022high to generate goal scene images that function as guides for robot actions, providing a distinct approach compared to the aforementioned works. In contrast, our Instruct2Act framework employs visual foundation models as modular tools that can be invoked with APIs without the need for fine-tuning, thereby eliminating the need for data collection and training costs.

Methods

Generally, Instruct2Act allows a robot to execute a sequence of actions based on an instruction from the user and an observation image captured by a top-view camera. It aims to change the object state in the environment to match the configuration in the instruction descriptions. Instruct2Act is a language-based robotic system that generates perception-to-action codes using an LLM. Task-related variables, including image crops used in the task instruction and image-to-robot coordinate transformations, are stored in an environment cache C\mathbf{C} that can be accessed through an API. The system gains a visual understanding of its manipulation environment by using perception models. Based on this information, the system generates executable action codes that the robot executes with the help of available low-level controllers. To guide the LLM’s output, we provide API definitions and in-context examples to the LLM.

To facilitate LLMs in completing robotic tasks, a designed prompt that guides the LLMs’ generation is provided together with specific task instructions. The prompt includes application programming interfaces (APIs) and in-context examples to demonstrate their usage, which are introduced in the following section. The LLM’s final input is a sequence of code APIs, usage examples, and task instructions. The LLM’s output is a Python function in string format that can be executed by the Python interpreter to drive the robot’s action.

Using LLM as the robotic driver has several advantages. Firstly, the robotic policies generated by LLM’s API are highly flexible, as in-context examples can be adjusted to guide LLM’s behavior and adapt to new tasks. Secondly, by utilizing visual foundation models directly, there is no need to gather training data or to conduct training processes, and any improvements in foundation models can improve action accuracy without incurring additional costs. Lastly, the simple API naming and readable Python code make the generated policy code highly interpretable.

After acquiring the input image’s semantic information, including the target object’s location and semantic class, we need a mapping between the image space and action space to generate executable actions. We use a pre-defined transformation matrix to transfer object location to robot coordinates and apply boundary clamping to prevent unintended actions. The LLM identifies appropriate actions based on instructions and in-context examples.

2 Prompts for Instruct2Act

Figure 3 illustrates that a complete prompt should include essential information about third-party libraries, API definitions, and in-context examples. The third-party library import information enables the LLM to understand how APIs use the parameter types defined by these libraries to perform calculations and to even create new functions. And we demonstrate its effectiveness in Section 4.4. We also provide API definitions and descriptions, along with a few in-context examples to demonstrate their usage, similar to gupta2022visual . However, in our approach, we distribute and organize APIs based on robotic system information, as shown in Figure 1(b). Specifically, these APIs are classified according to their functionality within the robot system hierarchy, and this categorized information is provided in the prompt. An example is the SAM() function belonging to Perception module which is the second level core module in the robotic system. So we will add # Second Level: Core Modules and ## Perception Modules before introducing the SAM() API in the prompt. Moreover, unlike ViperGPT suris2023vipergpt , which only offers function-level usage examples, we provide full-logical code examples invoking different modules similarly to the approach of gupta2022visual . This choice is based on the notion that robotic tasks tend to be more intricate yet organized. In contrast to gupta2022visual , we design a prompt that provides coverage for all tasks and has fewer in-context examples. To encourage chain-of-thought reasoning, as done in kojima2022large , we add the prompt Think step by step to carry out the instruction before the inserted specific task instruction. To avoid the LLM from generating too many redundant lines, we explicitly instruct it to only implement the main() function. Examples of complete prompts are available in Appendix.

3 Perception with off-the-shelf Foundation Models

Instruct2Act leverages off-the-shelf visual foundation models, specifically, the Segment Anything Model (SAM) and CLIP models, accessed by LLM via two designated APIs: SAM() and CLIPRetrieval(). These Python functions load and call the models to conduct the visual analysis. SAM outputs masks for all potential objects in the input image, based on which object crops IiI_{i} are extracted correspondingly. These crops are then encoded into FIiF_{I_{i}} using the CLIP image encoder and later utilized for the classification task. However, pre-trained visual models used directly on downstream tasks without any fine-tuning often suffer from incompleteness or incorrectness. To address these issues caused by the zero-shot paradigm, we insert processing modules between the output of the large model and the downstream tasks. To mitigate the effect of shadows, we apply a gray threshold filter followed by a morphological closing operation to fill up small holes before sending the image to the SAM. After SAM’s segmentation operation, we perform a morphological opening operation to eliminate overly small holes or unconnected gaps. We also filter out masks with unreasonable sizes and reduce redundant mask output using Non-Maximum Suppression (NMS). For detailed discussions and visualizations of the processing steps, please refer to Appendix.

4 Flexible Instruction Modality Manager

Instruct2Act is flexible and can handle inputs of multiple modalities, such as pure language and language-visual instructions, as depicted in Figure 2. We design a unified retrieval system that utilizes different types of queries to ensure the use of a unified architecture for both types of inputs.

Pure-Language Instruction. For pure language inputs, descriptive sentences are utilized to specify the target object and action. For example, a sentence such as Put the green and purple polka dot block into the green container can be employed. Instruct2Act utilizes the LLM to deduce that the robot needs to fetch the polka dot block from the environment. Therefore, the phrase the green and purple polka dot block acts as the query and is inputted into the CLIP text encoder to obtain the feature vector FTF_{T}. Finally, the similarity between the query embedding FTF_{T} and the image crop features FIiF_{Ii} can localize the intended object precisely.

Language-Visual Instruction. For the multimodal inputs, the instruction uses an image to describe the target object or the target state. An example instruction is Put <dragged_objdragged\_obj> into <base_objbase\_obj>. The placeholders in curly braces represent the corresponding images of individual objects, aligning with the LLM model’s input format. Given the instruction, the LLM determines the placeholder strings to complete the query which is used to fetch the corresponding object image crop II from the cache C\mathbf{C}. Then the image crop II is sent to the CLIP image encoder to obtain the feature vector FIF_{I}. We use FIF_{I} to calculate the similarity with observation image feature vectors FIiF_{I_{i}}.

Certain tasks require scene-level understanding, as demonstrated by the task of Rearrange to this . To fulfill this instruction, we first obtain every possible object and their corresponding feature vectors in the target scene image. We then use the Hungarian algorithm to determine the correspondences between the target scene image and the currently observed images.

Pointing-Language Enhanced Instruction. Pointing-language instructions are an effective alternative when the target object cannot be described using pure language instructions and providing image crops is impractical. Specifically, we adopt the cursor movement method from 2023interngpt and use cursor clicks to generate point prompts that guide the SAM’s segmentation. Additional details can be found in Appendix.

Experiments

We select several representative meta tasks from VIMABench jiang2022vima (17 tasks in total), ranging from simple object manipulation to visual reasoning to evaluate the proposed methods in the tabletop manipulation domain, as shown in Fig 4. The evaluation benchmark uses Pybullet coumans2016pybullet as the backend and the default render. In addition to the original multimodal prompt instruction in the VIMABench, we extract object descriptions from the simulator and create a task prompt utilizing natural language that is more commonly used by actual users in their day-to-day lives. Furthermore, the VIMABench contains a 4-level generalization ability evaluation protocol, e.g. L1 placement, L2 combinatorial, L3 novel object, and L4 novel task generalization. And we provide a more detailed task description in Appendix. Each level differs more from the training distribution, and we used the first three levels to evaluate our methods. For more details regarding the evaluation setting, please refer to jiang2022vima .

2 Large Language Model

For our experiments, we used two language models: (i) the text-davinci-003 language model via the OpenAI API, which is a fine-tuned variant of the InstructGPT ouyang2022training language model optimized by using human feedback, and (ii) the LLaMA-Adapter gao2023llama ; zhang2023llama via the user interface, which is a lightweight adaptation of the original LLaMA models touvron2023llama . Our approach provided limited prompts to influence the output behavior of the language models without any training or fine-tuning. Although the language models may occasionally generate incomplete or incorrect code with a low rate of errors, such as missing brackets, punctuation, and mismatched cases, the Python Interpreter can detect these errors and inform us to generate new code.

The LLaMA-Adapter language model has a slow inference speed and lacks an API interface due to its local hosting. Therefore, this language model is exclusively used to verify the output. In contrast, the ChatGPT language model is used for large-scale verification experiments as an alternative. A concise comparison of LLaMA-Adapter and ChatGPT is provided in the below experimental section of this paper.

3 Experiment Results

To perform open-vocabulary segmentation, we used open-sourced models such as SAM ViT-H, while for classification, we used CLIP ViT-H-14. Code generation was performed using ChatGPT with the text-davinci-003 engine. Task success rates were evaluated on 150 instances for each of the six meta-tasks, with three random seeds selected per meta-task to obtain average success rates. The VIMABench simulator determined task success if the final states matched the configurations outlined in the instructions. We conducted our experiments on an NVIDIA 3090Ti GPU. Additionally, we directly used the experiment results of different baselines from jiang2022vima .

Table 1 presents the experimental results obtained using the proposed Instruct2Act approach on the VIMABench, using the aforementioned foundation models and all processing techniques. The results show that our method achieves comparable average performance with the state-of-the-art (SOTA) learning-based approach, VIMA jiang2022vima , specifically designed for task applications in the multimodal instruction version. Notably, for tasks that require multiple steps to complete, such as Task 05 and Task 17, our Instruct2Act outperforms the previous SOTA by a significant margin, regardless of whether the instructions are single-modal or multimodal. We attribute this improved performance to the strong generalization ability of the large visual model and the powerful reasoning capacity of the LLM. The results also highlight that the performance of multimodal instructions is generally better than that of single-modal instructions. We believe that this is because the former provides a more comprehensive range of information to the model, thereby reducing difficulties for the robot when attempting to reason about the current execution scenario. It is important to note that our method is entirely zero-shot and relies solely on basic task information without using any auxiliary information.

To further validate the generalization ability, we assessed the effectiveness of our methods on L2 and L3 generalization tasks. The results of different methods are shown in Fig. 5. The results from (a-c) show that our method consistently achieved positive results across all levels and various tasks. Additionally, incorporating foundation models enabled our method to experience less distribution shifting, resulting in minimal performance fluctuations, as shown in Fig. 5-(d).

4 Further Analysis

Ablation Studies on Prompt Elements. We examine the efficacy of each prompt component in the code generation process, including the import information of third-party libraries, API definitions, and contextual examples. Generating executable robotic codes solely based on third-party libraries and without contextual information is nearly impossible. Therefore, in this paper, we consider the import information of third-party libraries as the default input and provide a brief discussion in the following section. For further reference, the complete generated codes is available in Appendix.

The LLM is able to produce reasonable logic code with detailed explanations by relying solely on API definitions, as demonstrated in Listing 2-4. Unfortunately, in the absence of usage examples, the model can generate unnecessary (actions like DistractorActions and RerrangeActions), as seen in Listing 2, does not return the executed information in Listing 4. Nevertheless, we are pleased to discover that, the LLM can account for execution failures through the invocation of the SaveFailureImage function, which was not present in our original in-context examples.

In contrast, the LLM output results are much more structured and similar to the examples given, when only in-context ones are available. However, due to the deficiency of functional information, the LLM fails to generate any comments or ambiguously produces incorrect ones, as seen in Listing 5 and Listing 7, where SAM was mistakenly inferred as the abbreviation for Semantic Affinity Module. In addition, variable naming lacks semantic information, and the LLM is unable to add new logic such as failure case handling, beyond the descriptions provided as examples. By providing both the API definitions and in-context examples in the prompt, the LLM can generate accurate and human-readable Python codes, as shown in Listing 8.

After analyzing the observations, we have derived the following conclusions: 1) API style prompts provide more flexibility and enable the LLM to exhibit better reasoning abilities; and 2) the in-context examples style prompts used by the LLM could efficiently sum up and deduce expert information from examples, making them more compatible with structured tasks.

Ablation Studies on Processing Modules. Table 2 presents ablation studies that validate the effectiveness of the proposed processing modules, including image pre-processing and mask post-processing. The absence of any processing methods leads to significant performance degradation, with only 51%51\% SR, in line with the findings of the analysis in Section 3.3. Mask post-processing directly boosts the success rate to 83.0%83.0\%, while further incorporation of image pre-processing achieves the peak success rate of 84.1%84.1\%. It is worth noting that using only image pre-processing leads to no improvement in performance and may even cause degradation in some tasks. This could be attributed to the sub-optimality of the pre-defined parameters, which lack tuning for better generalization.

Additional Foundational Models. We experimented with different visual base models to analyze their impact on performance. Specifically, we replaced the original models with SAM-Base and SAM-Large for the semantic segmentation module, and Base-16 and Large-14 for the CLIP model. The results, presented in Table 4 and Table 4, consistently show that larger foundation models lead to better performance for Instruct2Act. This suggests that our approach can benefit from even stronger foundation models in the future. Additionally, since the visual models are accessed via APIs, they can be easily replaced with other visual foundation models. However, this exploration is currently not our main focus and is left for future work.

Comparison between Different LLMs. In Section 4.2, we demonstrate that our proposed Instruct2Act is potentially effective with open-source LLM, in addition to the commercial ChatGPT. Specifically, we use the LLaMA-Adapter gao2023llama and choose Task 01, reporting the success rate on 40 instances in Table 5.

Remarkably, our Instruct2Act already achieves plausible performance with the LLaMA-Adapter’s original output, as shown in Table 5. Furthermore, our performance improves to 77.5%77.5\% when we increase the number of generation trials when the Python interpreter raises an exception. And comparable results can be achieved by using basic filtering, such as re-generating when the environment cache usage is missed in the code.

Flexibility and Robustness of Instruct2Act. We demonstrate the flexibility and robustness of Instruct2Act by leveraging LLMs’ powerful reasoning abilities. Specifically, we evaluate our approach in three scenarios: Human Intervention, Missing Characteristics, and Synonym Replacement. In the Human Intervention scenario, we add sentences to the original instruction to allow human intervention. For example, by appending I cancel this task. Stop!, the LLM can infer that no executable code should be generated. Moreover, the LLM can understand That instruction is wrong, use this one. and generate the correct code for the latter instruction. In the Missing Characteristic scenario, we randomly remove some characteristics or misspell words to test the LLM’s grounding ability. Remarkably, the LLM can still ground correctly, demonstrating its robustness. Finally, in the Synonym Replacement scenario, we test the LLM’s flexibility by allowing synonym replacement, such as substituting Rotate with Spin. The LLM’s flexibility shines as it can handle such variations effortlessly.

Extension with The Third-party Library. As described in Section 3.2, Instruct2Act can use third-party libraries to perform simple calculations. To demonstrate this, we evaluate the system with the degrees to radians scenario. We offer examples of rotations in degrees for context and inform the system of the numpy library’s availability in the prompt, as shown in Fig. 3. Then, we instruct the system to rotate objects to specific angles in radians, such as 0.5 radians.

As demonstrated in the code block above, Instruct2Act uses the provided numpy module to convert degrees to radians before passing the argument to the action function.

Limitations. A notable drawback of Instruct2Act is its high computational cost, as it employs several foundation models to accomplish robotic tasks, which is nearly unacceptable for a real-time robotic system with limited computation resources. Moreover, our method is presently limited by the basic action primitives provided in the chosen VIMABench, such as Pick and Place. However, we are confident that our approach can be readily expanded to include more intricate actions with API extension. Additionally, we have only tested our method in a simulation environment so far, but we plan to investigate real-world applications in the near future. We do not foresee any negative social impact from the proposed work.

Conclusion

We proposed a Instruct2Act framework to utilize LLM to map multi-modality instructions to sequential actions in the robotics domain. With the LLM-generated policy codes, various visual foundation models are invoked with APIs to gain a visual understanding of the task sets. To mitigate the gaps in the zero-shot setting, some processing modules are plugged in. Extensive experiments verify that Instruct2Act is effective and flexible in robotic manipulation tasks.

References

Appendix A Appendix

A.2 Processing Module in Instruct2Act

When pre-trained models are utilized directly on downstream tasks without any fine-tuning, they inevitably suffer from problems like incompleteness or incorrectness. In light of this, it is advisable to insert processing modules or adapters between the output of the large model and the downstream tasks which can effectively tackle the issues caused by this zero-shot paradigm. The entire loop of execution for robot tasks comprises three modules, namely, perception, planning, and execution. We have curated various processing programs for each of these modules designed for tabletop manipulation domains.

Image Pre-Processing In the case of zero-shot SAM outputs, a significant challenge is distinguishing the target object from shadow regions that might appear within the image. As object shadows cannot be grasped, their presence poses additional difficulty for robot grasping tasks. Furthermore, tabletop manipulation domains typically involve camera placement above the robotic arm, leading to large shadows cast onto the operating table. To account for this, we adopted a simple yet efficient image preprocessing methodology: to mitigate the effect of shadows, we employed the gray threshold filter followed by the close morphological operation to fill in small gaps that the filter might produce.

Mask Post-Processing In the zero-shot setting, it is observed that the SAM produces multiple segments (e.g. masks) that may be discontinuous or discrete. There may also be detected objects with missing parts, or holes inside, as shown in Fig 6-C. These misleading semantic mask outputs will inevitably confuse the subsequent modules and greatly challenge the robot grasping task. In order to address the given problem, we have developed a set of processing modules to work with the SAM output. The modules include the following methods:

We apply a filtering process on the output based on the mask’s size. This process removes any output that is clearly not part of the target object. Such output may include objects that cannot be moved, like tables or patterns on the target object.

A dilation operation is used to effectively eliminate unrealistic small holes or unconnected gaps. To avoid significant changes to the mask’s size due to dilation, we then use an erosion operation. These two procedures combine to create what is known as the opening morphological operation.

In some instances where there are multiple segmentation outputs for a single object, we employ the Non-Maximum Suppression (NMS) operator to reduce redundant mask output.

A.3 Pointing-Language Enhanced Instruction

When utilizing the pointing-language mode of Instruct2Act, the system will display the initial task instruction and observation image. Next, the user will select the target objects by clicking on them with the cursor. These click points will then act as point prompts to guide the SAM’s segmentation process. We evaluated this mode on Task 01 and Task 03, averaging the results over 150 instances with one seed.

Table 6 shows that our Instruct2Act achieves better results with the pointing-language enhanced mode than with the pure-language instruction mode. This could be due to the stronger prior information provided by the user’s click operations.

Instruct2Act, as discussed in Section 3.4, can manage instructions with different modalities. For audio instructions, the Whisper model API can be invoked to convert the audio input into text, after which the process described in Section 3.4 can be followed to guide the LLM and generate policy codes.

A.4 Evaluation Task Description

Simple object manipulation. The agent is asked to follow the basic instructions to take action.

Visual Manipulation. The agent is required to pick a specific object and place it into a specified container. The agent needs to first recognize the target objects aligning with the task instruction specified by the natural language description or by the image pattern.

Scene Understanding. The agent must first identify the target object with the described texture by grounding the natural language description and the scene image simultaneously and then put the target object into the container with a specific color.

Rotation. The agent is required to rotate a specific object by certain degrees along the z-axis.

Visual goal-reaching. The agent is asked to manipulate the objects to match the goal states described by the goal scene image,

Rearrange. The agent is required to rearrange the target objects to reach the goal configuration. The agent needs to identify the possible existing distractors and move them away to avoid position conflicts.

Rearrange then restore. The agent is required to restore the object placements after the rearrangement operations.

Visual Reasoning The agent is asked to make decisions and take actions where reasoning and memory ability are required.

Pick in order then restore. The agent is required to pick and place the target object sequentially into different containers and finally restore it to the initial container.

VIMABench presents a 4-level evaluation protocol that progressively increases in difficulty for trained agents. Our experiments utilize the first 3 levels of this protocol, which test the generalization abilities of our agents. Level 1 (L1) placement generalization randomly arranges the placement of target objects, whereas L2 combinatorial generalization generates new combinations of target materials and object descriptions. Finally, L3 novel object generalization tests our agents’ ability to generalize to novel materials and objects.

A.5 Ablation Studies on Prompt Element

We use the same task instruction Put the polka dot block into the green container. for all experiments here. Because the outputs of LLMs are somewhat random, we will display the results of three consecutive outputs.