JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models

Zihao Wang, Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang, Haowei Lin, Zhaofeng He, Zilong Zheng, Yaodong Yang, Xiaojian Ma, Yitao Liang

Introduction

Creating sophisticated agents that can accomplish myriad of tasks in complex domains remains a pivotal milestone towards generally capable artificial intelligence (Reed et al., 2022; Brown et al., 2020; Alayrac et al., 2022; Brohan et al., 2022a; Zhao et al., 2023). Recent advancements have shown a trend towards employing a hierarchical goal execution architecture (Wang et al., 2023b; Huang et al., 2022b, a), and leveraging large language models (LLMs) as the high-level planner to generate action plans that will be ultimately executed by low-level instruction-following controllers. Albeit the fruitful progress they have yielded in many robotics (Huang et al., 2022b) and even open-world environments like Minecraft (Fan et al., 2022; Guss et al., 2019b), today’s agents built with these approaches are still struggling with three major issues: 1) perceive the world from multimodal sensory observations, such as images, videos in addition to natural language instructions and feedback for planning; This is mostly due to the inability of LLM-based planners on processing multimodal data (Huang et al., 2022a; Yao et al., 2022); 2) perform consistent and accurate long-term planning. This requires multi-round, knowledge, and reasoning-intensive dialogues, which remain great challenges to LLMs (Huang et al., 2022b); 3) learn and evolve in a life-long fashion. This calls out the need for agents to propose their own tasks and self-improve. Addressing these issues will unleash the full planning potential of LLM-based agents, and expedite the development of more generalist agents.

In this work, we introduce JARVIS-1, a brand new agent that can robustly produce plans for long-horizon tasks from multimodal user and environment inputs, and translate them into motor control in Minecraft, a popular yet challenging open-world testbed for generalist agents. To be specific, we chain a multimodal foundation model MineCLIP​ (Fan et al., 2022) and an LLM​ (Brown et al., 2020) together, the resulting multimodal language model (MLM) allows our agent to better understand the task, situations, and environmental feedback. To further enhance the correctness and consistency of planning, especially on long-horizon tasks, we propose to augment the agent with a multimodal memory, which stores both the scenarios and actual plans of the successful planning experiences in the past. By retrieving the relevant memory entries, the planning skill of our MLM-based agent can be strengthened from the agent’s own interactions with the environment in an in-context manner. Finally, JARVIS-1 is able to evolve throughout the gameplay by proposing tasks on its own (i.e. self-instruct) as a means of exploration and saving the obtained experiences in the multimodal memory, therefore facilitating better reasoning and planning. This self-improving ability sparks its potential for a higher level of autonomy.

Our main evaluations are conducted in Minecraft, with more than 200 tasks selected from the Minecraft Universe Benchmark (Lin et al., 2023a), with no demonstration provided. The tasks cover a broad spectrum from the early game (e.g. ObtainCraftingTable) to intermediate and even challenging long-horizon tasks (e.g. ObtainDiamondPickaxe). A glimpse of what JARVIS-1 is able to achieve can be found in Figure 1. JARVIS-1 exhibits strong performances on these tasks, representing an up to 5×\times increase to the previous records. Our ablative analysis then offers a detailed account of how JARVIS-1 approaches this significant progress and becomes the first agent that can robustly obtain the diamond pickaxe with up to 12.5% success rate. What is even more surprising is that, without the need for additional training, JARVIS-1 demonstrates a continuous increase in performance as game time increases in long-horizon tasks. Moreover, JARVIS-1 has demonstrated its potential of self-improve in an exploratory life-long learning experiment, where it needs to propose tasks to progressively explore the world, collect experiences, and sharpen its planning skill using these experiences stored in the multimodal memory.

In summary, JARVIS-1 pilots the effort towards a human-like multi-task and autonomous agent in an open-world, embodied environment like Minecraft. We would like to share the key takeaways of what we have learned during its development as follows:

From LLMs to MLMs. The capability of perceiving multimodal sensory input is critical to planning in a dynamic and open-world world. JARVIS-1 enables this by chaining a multimodal foundation model together with an LLM. Compared to LLM “blindly” produces plans, MLM is able to natively understand the current situation and plan accordingly. Further, rich environmental feedback can be obtained through multimodal perception, therefore helping the self-check and self-explain of the planner spot and fix possible bugs in the plans, enabling stronger interactive planning.

Multimodal memory. Early research has suggested the crucial role that memory mechanisms can serve in the functioning of generalist agents. By outfitting JARVIS-1 with a multimodal memory, we effectively allow it to plan with both pretrained knowledge and its actual experiences in the world, therefore bringing significant improvement to planning correctness and consistency. Compared to canonical RL or planning agents with exploration, no additional model update is needed as the MLM in JARVIS-1 makes it possible to leverage these experiences in an in-context manner.

Self-instruct and self-improve. A sign of generalist agents is the capacity to proactively acquire new experiences and continuously improve themselves. We have demonstrated how JARVIS-1 effectively traverses the environment by executing tasks autonomously generated through its self-instruct mechanism. With multimodal memory teaming up with experiences from the explorations, we have observed consistent improvement, especially in accomplishing more complicated tasks. Ultimately, this aspect of autonomous learning in JARVIS-1 signifies an evolutionary step towards generalist agents that can learn, adapt, and improve over time with minimal external intervention.

Challenges for Open-world Agents

Compared to canonical scenarios with relatively small scale, simple dynamics, and limited tasks, open-world environments impose substantial challenges to building agents that can accomplish a diverse set of tasks (Fan et al., 2022; Guss et al., 2019a, 2021; Kanervisto et al., 2022; Cai et al., 2023a; Wang et al., 2023b; Cai et al., 2023b). In this section, we will review three major challenges we’ve identified during the development of JARVIS-1.

In an open world, there could be various possible paths towards an open-world goal. However, not all of them are plausible or equally efficient given a certain situation (location, inventory status, etc.). For example, building a bed can be done through collecting wool from sheeps , haunting spiders for strings , or trading with villagers . Depending on the current location and its proximity to these subjects, some options can be more viable and more efficient than others. Further, the agent’s own situation can also change throughout the episode, e.g. day and night shifts, weather conditions (bringing different types of danger), and tool usage (it can be broken). To this end, the plan needs to be constantly updated based on the current situation. Figure 2 (left) shows that when attempting the "ObtainDiamondPickaxe" task with a GPT-based planner that produces plans only at the beginning without looking at the current situation, the agent failed to complete the task as opposed to human players and JARVIS-1, which perform situation-aware planning from time to time. We’ve observed that many failures coming from this were attributed to the agent’s inability to adapt to the changing situations including entering a new biome, the tool being used becoming broken, etc.

2 Challenge II: Task Complexity

The second challenge comes from the higher task complexity in open-world environments. Due to the richness of terrains, objects, and action space, tasks in open-world domains usually require substantially long planning horizons as well as good accuracy and precision. For example, the task ObtainEnchantingTable includes more than 20 different sub-goals and therefore demands significantly longer reasoning steps. Meanwhile, many of these sub-goals have to be achieved precisely with the exact object name, quantities, and preconditions, e.g., mine 3 obsidian with diamond pickaxe, craft 1 diamond pickaxe from 3 diamonds and 2 sticks; otherwise, the subsequent sub-goals won’t be executed due to unfulfilled preconditions. To tackle this, we may refer to some approaches in LLM reasoning, e.g. self-debugging (Chen et al., 2023) and turning the planning into an interactive fashion. In Figure 2 (Middle), we’ve shown that as the complexity of the task increases, our JARVIS-1, which uses interactive planning (Wang et al., 2023b) to mitigate the aforementioned issues (details can be found in Section 3.2), elicits more significant advantages over the baseline (GPT) planner.

3 Challenge III: Life-long Learning

Finally, being open world often implies offering an infinite number of tasks. Clearly, it is difficult for an agent to master all tasks or generalize to arbitrary tasks without additional learning. To this end, agents in an open world should be able to learn novel tasks while completing existing tasks, i.e. life-long learning. Furthermore, as many open-world agents employ large models (Wang et al., 2023b; Yuan et al., 2023; Wang et al., 2023a; Zhu et al., 2023), canonical gradient-based learning could be extremely inefficient given the number of new tasks and experiences to learn. Our MLM-based JARVIS-1 tackles this by adopting a memory to save all the experiences on past tasks. By retrieving memory entries relevant to the newly-coming task and putting them into the context as a reference, JARVIS-1 is able to accumulate more experiences as the game continues and strengthen its own planning skills without gradient update. As illustrated in Figure 2 (Right), for instance, both ObtainDiamondPickaxe and ObtainDiamondAxe require gathering almost identical materials. Therefore, they can help each other by using the experiences from the other task. Compared to completing these challenging tasks without any prior experiences, memory-based in-context life-long learning in JARVIS-1 can bring significant advantages.

Multi-task Agent with Memory-Augmented MLM

This section details the architecture of the proposed JARVIS-1 agent. We begin with an overview of the modular agent design in Section 3.1. Next, we elaborate on how to implement an interactive planning scheme with a multimodal language model, which helps with more accurate plans, especially on complex and long-horizon tasks in Section 3.2. Finally, we show how to augment this planning framework with a multimodal memory to allow JARVIS-1 to strengthen its planning skill throughout the episode by in-context life-long learning in Section 3.3 and Section 3.4.

We aim to develop an agent capable of solving long-horizon instruction-following tasks using image observations and human-aligned actions. To accomplish this, we propose a multi-modal agent including an interactive planner, a goal-conditioned controller, and a multimodal memory of multimodal experiences. Upon receiving a task and the current observation, JARVIS-1 first utilizes the MLM to generate a multimodal query (query gen) that retrieves relevant planning experiences from the memory. These experiences will then be used along with the planning instruction to prompt the MLM-based planner. Leveraging its own pretrained knowledge as well as the retrieved reference plans, the planner will ultimately produce a series of KK short-horizon goals g1,…,gKg_{1},\ldots,g_{K} to be executed by the controller. Once the plan is successfully executed, it will be stored in the memory along with the task and the agent situation when it was planned. We also empower JARVIS-1 with life-long learning by combining self-instruct, where JARVIS-1 will propose some tasks for itself to complete as a means of exploration; and self-improve, where multiple JARVIS-1 agents will be running in parallel to gather experiences, therefore helping with better planning later. We provide an illustration in Figure 3.

2 Interactive Planning with MLM

As we have mentioned in Section 2.1 and Section 2.2, the primary challenges for planning in Minecraft come from the requirement of being able to plan for long-horizon tasks under dynamic observations. Confirmed by many prior arts (Wang et al., 2023b, a; Yuan et al., 2023), this makes it exceptionally hard to utilize canonical symbolic planners, which can be much less flexible. To this end, we take a multimodal language model (MLM) as zero-shot planner and combine it with an interactive planning framework to tackle these challenges.

Situation-aware planning with MLM. To achieve situation-aware planning, the planner must take the current observation into account, in addition to the task instruction (Huang et al., 2022a; Yao et al., 2022). Specifically, we begin with translating the multimodal observation into text descriptions. As opposed to letting the MLM caption the scene directly, we first extract keywords of Minecraft items (e.g., "acacia tree", "sheep") from Minecraft wiki and utilizing GPT (Brown et al., 2020) to generate sentences that describe these observations. For example, a generated sentence could be "I can see sheep in the acacia plains". Then the MLM will retrieve the condition sentence according to current visual observation during planning. Additional situation details including biome and inventory status are also converted into text using templates. Finally, we prompt the MLM again (the language part only) into a plan given the task instruction and all the aforementioned textual situation descriptions. Compared to end-to-end alternatives (Brohan et al., 2023; Huang et al., 2023), we find our composable usage of MLM provides higher quality situation descriptions and ultimately, plans with much less hallucination.

Planning with self-check. Our first layer of shield to ensure the correctness of plans involves self-check. Similar to self-debugging​ (Chen et al., 2023), given an initial plan, we ask JARVIS-1 to progressively simulate the plan execution, predict the resulting state after each step (primarily the state of inventory), and evaluate them. By verifying if these states satisfy the goal’s precondition, JARVIS-1 can proactively identify potential plan flaws. Compared to the canonical planner where the agent has to encounter the error first before making a remedy, this upfront plan verification could mitigate the need for the agent to recover (re-plan) from more challenging situations due to plan failure. For instance, if an agent starts digging underground without sufficient wood, it would typically have to return to the surface, which substantially lowers the chance of completing the task.

Planning with environment feedback. Next, our interactive planning framework ventures into allowing JARVIS-1 to quickly recover from failure by leveraging environment feedback in a closed-loop fashion. The process is illustrated in Figure 4. During plan execution, we feed the feedback to the MLM of JARVIS-1 in case there is any execution failure (possibly due to a flawed plan) and utilize its self-explain mechanism (Shinn et al., 2023) to explain the error and locate the bugs in the original plan (we term this as error explanation). Finally, the MLM planner of JARVIS-1 will produce an improved plan based on both the outside environment feedback and the inside retrospective. Compared to other agents that rely on human intervention or privileged environment information (Huang et al., 2022b; Zhu et al., 2023), JARVIS-1 has the ability to speculate about the reasons why current goals cannot be achieved, without the need for additional information or design.

3 Planning with Multimodal Memory in the Loop

To address the life-long learning challenge mentioned in Section 2.3, we equip JARVIS-1 with multimodal memory to allow learning from its own past experiences. We will detail the formulation of the retrieval-augmented planning, query generation, and memory layout below.

Retrival-augmented planning. Retrieval-augmented generation (RAG) (Lewis et al., 2020; Mao et al., 2020) enhances the quality of responses generated by LLMs by incorporating external sources of knowledge to complement the model’s internal representation. We also utilize RAG to enhance JARVIS-1’s long-term planning capability. Compared to official RAG methods leveraging the external knowledge library, we take the collected multimodal memory as the knowledge library and retrieve the interactive experiences as the demonstration prompt to augment the planning results. The formulation is as follows:

where xx, yy, and zz denote instruction, plans, and retrieved memory entries respectively, and pηp_{\eta} and pθp_{\theta} are denoted as retrieval and planning models. Such retrieval-augmented planning method helps JARVIS-1 ground the internal knowledge into the open-ended environments efficiently and leverage the historical interaction feedback to solve the hallucination within LLMs and produce more accurate plans.

Multimodal memory. We have demonstrated the layout of our multimodal memory on the right side of Figure 5. From a high level, it is a key-value memory where the keys are multimodal, comprising both the task and the observation (or situation) made when this memory entry was created. The values are the plans that were successfully executed. Note that, since the plans in an open-world environment like Minecraft are situated (see Section 2.1), there could be multiple entries that are with the same task but different observations and plans. As a result, JARVIS-1 needs to produce multimodal queries based on the current task and situations to retrieve the relevant memory entries.

Query generation via reasoning. When presented with an instruction as a task, we employ query generation via LLM reasoning to decompose the instruction into sub-tasks or related tasks, which will then be used as textual queries to retrieve relevant planning experiences as references for solving the current task. For instance, consider the instruction "craft 1 enchanting table with empty inventory" as shown in Figure 5. JARVIS-1 queries the MLMs to identify the tasks that are required for achieving the main task in a backward search fashion, e.g., “obtain book /diamond /obsidian with empty inventory”. The search depth is bounded for efficiency. Further, instead of relying solely on retrieval based on the text query (Wang et al., 2023a; Zhu et al., 2023), we also propose to append the agent’s current visual observation to the textual query, resulting in a multimodal query to take the situation into account during memory retrieval.

Multimodal retrieval. After obtaining the textual and visual query, we compute the alignment between the query and each trajectory in multimodal memory. We first use the text encoder of the CLIP model to compute the embedding of the query and task key of each entry in memory. We select the memory entries with similarity higher than the confidence threshold as the candidate entries. Then we will compute the visual state embedding of query and states in candidate entires. Then we sort the candidate entries with the visual embedding similarities, which can be formed as:

where szs_{z} and sxs_{x} are the visual key of memory entries and visual query, respectively. Finally, we retrieve the plan of top-k candidate entries as reference prompt zz.

4 Self-improving Agents

Learning in Minecraft with memory. The remaining issue now is where the aforementioned multimodal memory comes from. Inspired by the life-long learning scheme in many close-world and open-world reinforcement learning problems (Abel et al., 2018a, b; Wang et al., 2023a), we propose the following learning approach for augmenting the memory in JARVIS-1: 1) First, we generate a set of tasks, which form some curricula for the agents to complete as means of exploration of the world. During this process, JARVIS-1 produces plans, interacts with the environment, embraces the errors, and stores all these experiences in the memory; 2) After this learning stage, we evaluate JARVIS-1 on various tasks. Therefore, JARVIS-1 is able to produce better plans with the memory teaming up with the planning experiences. In our experiments, we use this as the default setting for all tasks.

Exploration using self-instruct. The key issue to the success of learning with memory is how to effectively acquire useful experiences given a limited amount of time. We propose to use self-instruct (Wang et al., 2022) to generate the dynamic curriculum and guide JARVIS-1 to learn from the interactions with environments. In each round, we prompt the MLM to consider how capable JARVIS-1 is at this point and subsequently select tasks from a task pool to explore. We find that the curriculum almost follows the technical tree-growing direction. To accelerate the learning process, we augment the linear self-instruct to distributed learning in distributed environments with shared memory, i.e. speculative execution (Leviathan et al., 2023). Specifically, we generate multiple executable tasks as candidate task batches and provide them to agents with the same memory for verification and execution in various different environments. Meanwhile, experiences are collected into a shared centralized memory. When all exploration tasks have been accomplished, we move to the next round, until the memory reaches a certain capacity.

Life-long learning. We’ve also observed that the aforementioned learning (where the memory is being filled) can be extended throughout the whole gameplay, where the agent gradually acquires more and more skills. As the gameplay continues, more and more experiences are pouring in, therefore JARVIS-1 can find better references for challenging tasks like ObtainDiamondPickaxe, resulting in an improved success rate on these tasks. Further, there is no gradient update in this thanks to the memory-augmented MLM, i.e. we can do in-context life-long learning. In Section 4.3, we offer exploratory experiments to show the potential of such capability of JARVIS-1.

Experiments

In the experiments, our goal is to 1) evaluate the general performances of JARVIS-1 on the challenging Minecraft tasks, especially on its advantages over baselines that do not (fully) address the aforementioned issues in open-world agents; 2) understand the factors that contributes to the general results; 3) explore the potential of JARVIS-1 in terms of life-long learning and its benefits to long-horizon tasks. To this end, we will first briefly introduce the evaluation settings, then cover the main comparative results and ablation studies, and conclude with an exploratory trial on long-horizon tasks.

We evaluate JARVIS-1 in Minecraft, with tasks selected from the recently introduced Minecraft Universe Benchmark (Lin et al., 2023a). For the reader’s convenience, we provide details on the basic setups below.

Environment setting. To ensure realistic gameplay, the agent needs to utilize observation and action spaces that are similar to those used by humans. Instead of manually designing a custom interface for models to interact with the environment, as done in previous methods such as MineDojo(Fan et al., 2022), GITM(Zhu et al., 2023), and Voyager(Wang et al., 2023a), we opt for using the native human interface provided by Minecraft. This applies to both the observation and action space. The model operates at a speed of 20 frames per second and is required to use a mouse and keyboard interface when interacting with human GUIs. For more information on the detailed descriptions of the observation and action spaces, please refer to the Appendix.

Task setting. In Minecraft, players have access to thousands of items, each with specific acquisition requirements or recipes. For example, stone-type items can only be obtained using a pickaxe, and two planks can be crafted into four sticks (these requirements are available on the Minecraft Wiki1). In survival mode, players must obtain each type of item from the environment or craft/smelt the object item from materials. We choose over 200 tasks from the Minecraft Universe Benchmark (Lin et al., 2023a) for evaluation. These tasks are related to items that can be obtained in the Minecraft overworld. For the convenience of statistics, we have classified them into 11 groups according to recommended categories in Minecraft2 (see Table1). Due to the varying complexity of these tasks, we adopt different maximum gameplay durations (Max. Steps) for each task. The limit is determined by the average time the human players need to accomplish the corresponding task. Other details about each task, such as language instruction, maximum steps, evaluation times, biome, and initial inventory when the agent is born into the world can be found in Appendix Table 5-14.

Evaluation metrics. By default, the agent always starts in survival mode, with an empty inventory. A task is considered a success when the target object is obtained within a specified time. Due to the open-world nature of Minecraft, the world and initial position that the agent is spawned at could vary a lot. Therefore, we conducted at least 30 tests for each task using different seeds and reported the average success rate to ensure a thorough assessment. Further, since we categorize the tasks into groups, we also report mean and variance values for each group for ease of presentation.

2 Main Results

We compare JARVIS-1 with other multi-task instruction-following agents based on LLM, including Instruct GPT​ (Huang et al., 2022a; Ouyang et al., 2022), ReAct​ (Yao et al., 2022), Inner Monologue​ (Huang et al., 2022b), DEPS​ (Wang et al., 2023b). Since some methods are not originally experimented in Minecraft, we reproduce them to conform to the Minecraft specification based on prompt and feedback template design. All LLM-based methods access the LLM model through OpenAI API. And all hyper-parameters of LLM including temperature are kept as default.

The average success rates for every task group are listed in Table 2. JARVIS-1 achieves the best performance with all meta tasks. It is important to note that in Minecraft, the technology tree can be formed by Group Wood, Stone, Iron, Gold, and Diamond. The tasks become increasingly difficult as you progress through the tree. For more difficult tasks such as obtaining a gold ingot or a diamond, the agents typically need to perform more actions and longer goal sequences in order to complete the task. As a result, the success rate of all agents decreases as the difficulty level increases. It is evident that reasoning methods (ReAct​ (Yao et al., 2022) vs. GPT​ (Ouyang et al., 2022; Huang et al., 2022a)) and interactive re-planning with feedback (Inner Monologue​ (Huang et al., 2022b) vs. GPT) effectively enhance the agent’s task performance in an open world. However, these approaches still face challenges when dealing with long-horizon tasks, specifically in the Iron and Diamond group. DEPS​ (Wang et al., 2023b), on the other hand, enables agents to accomplish diamond-related tasks through interactive long-horizon planning accompanied by descriptions and explanations. Nevertheless, its reliability remains very low at approximately 2.5%.

In comparison to DEPS​ (Wang et al., 2023b) without memory, JARVIS-1 demonstrates superior performance even in challenging tasks due to its extensive experience. In diamond-related tasks specifically, the success rate has increased by nearly 3 times (8.99% vs 2.42%). And JARVIS-1 usually only requires 2-3 rounds of re-planning to generate the correct executable plan, whereas DEPS requires more than 6 rounds. This means that JARVIS-1 saves a significant amount of LLM tokens and thinking time, enabling more efficient plan execution and providing additional steps and tokens for handling uncertainty in the environment.

Based on our observations, we have found that the bottleneck for JARVIS-1 in tasks involving diamonds often lies with the Controller’s inability to perfectly execute short-horizon text instructions generated by LLM. Therefore, it is worth exploring methods for generating plans that are easier for the controller to execute or improving the controller’s ability to follow instructions.

We conducted ablation experiments on various Language Models, including OpenAI’s ChatGPT Ouyang et al. (2022) and GPT-4 OpenAI (2023). Among these models, GPT-4 has more parameters and has been proven to outperform ChatGPT in extensive research Wang et al. (2023a). We also select the open-source pre-trained LLaMA2 70B model Touvron et al. (2023). Additionally, we gathered a substantial amount of Minecraft-related text from the internet as training data and further fine-tuned LLaMA2 13B. The experiments were conducted on a subset of Minecraft tasks using different language models. Each JARVIS-1 learns for 4 epochs of interaction with all task sets and evaluates on task subset across at least 20 seeds. The experimental results are presented in Fig. 6.

Table 6 demonstrates that ChatGPT, despite having fewer parameters, achieves nearly identical success rates as GPT-4. This suggests that language models equipped with memory can significantly enhance planning abilities. In Minecraft-related tasks, the open-source pre-trained LLaMA2 70B exhibits a notable performance gap compared to OpenAI models, particularly in long-horizon tasks. However, by finetuning LLaMA2 with fewer parameters, its performance on Minecraft tasks improves substantially. This indicates that the open-source model lacks knowledge specific to Minecraft and requires further finetuning for the successful completion of such tasks.

2.2 Ablation on Memory

We also conduct ablation experiments on the multimodality memory and retrieval methods. We set JARVIS-1 w/o memory module as the baseline agent. We first evaluate JARVIS-1’s performance with different memory sizes (representing different learning stages) as shown in Fig. 7, which demonstrates the effectiveness of self-improving within JARVIS-1. We further conduct the experiments on a subset of Minecraft tasks using three different retrieval methods: retrieval with textual instruction embedding only (Text Memory), synergizing reasoning and retrieval with text embedding (Text Memory+Reasoning), and synergizing reasoning and retrieval with multimodality embedding (Multimodal Memory+Reasoning). Except for the memory and retrieval methods, all others are kept the same. The results are listed in Fig. 8.

The experiments show that reasoning before retrieval can effectively improve retrieval accuracy. Retrieval based on a multimodal state including vision observation and symbolic information (e.g., inventory, location, etc) is better than only considering the text embedding.

3 Long-Horizon Challenges

Most concurrent multi-task agents in Minecraft can only handle short-term tasks and struggle with long-horizon tasks like CraftingDiamondPickaxe. The VPT foundation model​ (Baker et al., 2022) is capable of accomplishing various tasks in Minecraft but lacks the ability to execute human instructions. To address this limitation, Reinforcement Learning is required to fine-tune the VPT foundation model for specific task completion. However, after fine-tuning, VPT may experience a decline in performance for other tasks while focusing on the specified task. In contrast, Steve-1​ (Lifshitz et al., 2023) has implemented goal-conditioned fine-tuning on VPT, enabling it to follow human text instructions while maintaining multitasking capabilities. However, Steve-1 primarily focuses on low-level tasks like obtaining dirt, collecting flowers, and chopping trees. When it comes to long-horizon tasks such as starting from scratch by obtaining a wooden pickaxe, Steve-1 still encounters difficulties.

DEPS​ (Wang et al., 2023b) also utilizes LLM as a planner, but it lacks the ability to learn from experience in different tasks and apply that knowledge to new ones. Additionally, DEPS is limited in its re-planning rounds due to the LM’s context constraints. The experiments reveal that DEPS has a success rate of less than 50% in generating accurate and executable plans for acquiring diamonds. The probability of DEPS successfully obtaining diamonds in the environment is approximately 0.59%. Consequently, DEPS continues to face challenges when attempting to finish long-horizon tasks within the Minecraft world.

Even human players who have mastered the distribution pattern of diamonds achieve success rates of obtaining diamonds and crafting a diamond pickaxe (which requires at least three diamonds) within 10 minutes at approximately 15% and 12%, respectively. JARVIS-1 performs better in the ObtainDiamondPickaxe challenge. Compared to the state-of-the-art model, which has undergone RL-finetuned VPT, JARVIS-1 has more than doubled the success rate of obtaining a diamond pickaxe (6.22% vs 2.5% within 20 minutes).

To increase the chances of obtaining diamonds, we extended the game-playing time to 60 minutes (72000 game-playing steps, as shown in Figure 9). As a result, JARVIS-1’s success rate in acquiring a diamond pickaxe improved from 6.2% to 12.5%. The graph on the right side of Figure 7 illustrates how the success rate of intermediate milestone items changes over time, indicating that JARVIS-1 tends to improve with longer game-playing time. We also conduct two variants of JARVIS-1 with different self-improving curricula: human-written and random-generated. All three JARVIS-1 have collected experiences into memory with the curriculum for 4 epochs before evaluation in 60 minutes. The results show that JARVIS-1 with a GPT-generated curriculum can finish the task within the shortest game-playing steps and achieve the best performance in 60 minutes.

In contrast, VPT’s success rate barely changed when we increased the time from 20 minutes to 60 minutes (from 2.5% to 3%). This can be attributed to Minecraft’s durability system where prolonged underground exploration often leads to pickaxe damage. When JARVIS-1’s pickaxe breaks, it dynamically re-plans based on its current inventory and crafts a new one. However, VPT-RL exhibits perplexing behaviors at this stage by using inappropriate tools for mining stones or crafting unnecessary items. This comparison demonstrates that JARVIS-1 possesses superior generalization and planning abilities for long-horizon tasks.

Note that our method is designed to be multi-task in its nature and not finetuned through imitation learning on specific datasets or reinforcement learning.

Related Works

There have been some methods leveraging the large language model to generate action plans for high-level tasks in embodied environments (Zeng et al., 2022; Dasgupta et al., 2022; Mai et al., 2023; Liu et al., 2023; Zhang et al., 2023; Zhang and Lu, 2023; Gong et al., 2023b). Huang et al. (2022a) decompose natural language commands into sequences of executable actions by text completion and semantic translation, while SayCan generates feasible plans for robots by jointly decoding an LLM weighted by skill affordances from value functions (Brohan et al., 2022b). Some methods also leverage the LLM to produce the program code as plan for better executation (Singh et al., 2022; Liang et al., 2022; Lin et al., 2023b). However, the above methods assume that the initial plan from the LLM is correct. When there are bugs in the initial plan, it’s difficult for the agent to finish the task successfully. Recent research frequently employs LLM as an interactive planner, harnessing its self-updating capabilities to enhance the plan’s executability over time (Wang et al., 2023b; Shinn et al., 2023; Sun et al., 2023). Inner Monologue (Huang et al., 2022b) pilots the front of interactive planning with LLMs, which introduces the feedback (including success detection and scene description) to the planner. However, we found it could still suffer from accumulative planning errors, especially in long-horizon open-world tasks. ReAct (Yao et al., 2022) will reason about the agent state before acting, which indicates that various reasoning methods (Wei et al., 2022; Yao et al., 2023; Wu et al., 2023) are benefitial for planning. LLM-based planning methods often use the fixed pretrained LLM as the agent, while we focus more on life-long and continual learning for agents in open-world environments (Ke et al., 2022b, a; Wang et al., 2023a). For better leveraging historical interaction between agent and environments, an explicit memory (Park et al., 2023; Zhu et al., 2023) for more historical chatting has been leveraged for bigger storage of agent experiences. However, the above methods usually rely only on a text-based environment and struggle to execute plans in partial-observed visual open-world environments.

2 Minecraft Agents

Developing generally capable agents in Minecraft to solve open-world tasks has gained increasing interests (Ding et al., 2023; Fan et al., 2022; Baker et al., 2022; Cai et al., 2023a, b; Zhang and Lu, 2023; Yuan et al., 2023; Zhu et al., 2023). As an early attempt, Oh et al. (2017) studied task generalization in a simple Minecraft environment variant. It designed a two-stage pipeline, first mastering the prerequisite skills with parameterization trick, and then learning a meta controller to execute the instructions. Moving to solve complex long-horizon tasks in Minecraft, works (Oh et al., 2017; Mao et al., 2022; Lin et al., 2021) explored the hierarchical architecture. In recent years, influenced by the trend of large-scale pre-training paradigms, a group of researchers have emerged, who are utilizing vast amounts of internet knowledge to train intelligent agents. Fan et al. (2022) trained a visual-semantic alignment model, MineCLIP, using the correspondences between subtitles and video snippets available on YouTube, and used it to generate intrinsic rewards to guide policy learning. (Baker et al., 2022) utilizes a pre-trained inverse dynamics model to label actions in YouTube videos which are used to learn a foundation policy VPT through imitation learning. By bridging MineCLIP and VPT, Lifshitz et al. (2023) creates a performant instruction-following policy Steve-1 to solve open-world short-horizon tasks using hindsight relabeling and unCLIP tricks. However, Steve-1 can not solve complicated process-oriented tasks due to the expressive capability of its goal space. Cai et al. (2023b) learns to follow reference videos as the instruction by merely watching gameplay videos, which improves the capacity of goal space and reduces the cost of policy training. All of these methods focus on improving the smoothness and robustness of interaction between policy and environment. Inspired by the powerful language understanding and reasoning capabilities of large language models, researchers have begun to build Minecraft agents based on LLMs. Wang et al. (2023a) used LLM to guide the agent to explore the Minecraft world by acquiring diverse skills, making novel discoveries, and generating goal proposals. Zhu et al. (2023) integrated LLM with text-based knowledge and memory to equip the agent with common sense and past experiences for higher reasoning efficiency. Yuan et al. (2023) used LLM to guide the agent to explore the Minecraft world and interact with the environment with reinforcement learning control policies.

Conclusion

We propose a multi-task agent JARVIS-1 designed for the complex environment of Minecraft, which marks a significant advancement in achieving human-like planning within an open-world setting. By leveraging pre-trained Multi-modal Language Models, JARVIS-1 not only effectively interprets multimodal inputs but also adeptly translates them into actions. Its integration of a multimodal memory, which draws from both ingrained knowledge and real-time game experiences, enhances its decision-making capabilities. The empirical evidence of its prowess is evident in its impressive performance across a wide array of tasks in Minecraft. Notably, its achievement in the long-horizon diamond pickaxe task, where it achieved a completion rate that surpasses VPT by up to five times, underscores its potential and the strides made in this domain. This breakthrough sets the stage for the future of more versatile and adaptable agents in complex virtual environments.

Acknowledgments

This work is funded in part by the National Key R&D Program of China #2022ZD0160301, a grant from CCF-Tencent Rhino-Bird Open Research Fund, NSF grants #IIS-1943641, #IIS-1956441, #CCF-1837129, an SRA from Meta and a research gift from Amazon Alexa AI, and a gift from RelationalAI. The authors sincerely thank Dr. Rita Zhang, Zhixiang Dai at NVIDIA for the valuable technical support of GPU computing.

References

Appendix A Implementation Details

Tasks in Minecraft are usually related to mine and craft goals. The mine goals require the agent to collect raw materials from the environment using the appropriate tools. The craft goals ask the agent to use the recipe to generate new items with existing materials in inventory. The mine goals are achieved through STEVE-1 (Lifshitz et al., 2023) with text condition during implementation. The environment can directly executes the craft and smelt actions (craft/smelt with argument), which are same as MineDojo (Fan et al., 2022) .

A.2 Interactive Planner

JARVIS-1 relies on the Multi-modal Language Model for planning, self-checking, and self-explaining, and can accept three types of inputs: visual images, language, and symbolic information (including inventory, located position, home, current life statistics, etc.). Specifically, this is a hybrid model with language processing capabilities derived from the GPT model (OpenAI, 2023). The visual ability comes from MineCLIP (Fan et al., 2022). We collected approximately 1000 Minecraft text data from the internet and calculated the similarity between the current vision observation and these text data. Text above the similarity threshold will be selected into the GPT model’s prompt. Symbolic information is converted into natural language text through a designed template. All modalities are ultimately captured as language and processed by the GPT model.

Different modules in JARVIS-1 (e.g. self-check and self-explain) are completed through MLM based on different prompts. The specific prompt design are shown below.

A.3 Memory

Our memory records every successful trajectory experience of JARVIS-1, including the task goals that the agent needs to execute, the actual goal sequence (plan) executed by the agent, and the state (visual observation and symbolic information returned from the environment) when the agent completes the task. In specific implementation, memory is a list where each trajectory experience is encoded as a dictionary, including the keys task, state, and plan.

Appendix B Environment Setting

Our Minecraft environment is a hybrid between MineRL​ (Guss et al., 2019b) and the MCP-Reborn (github.com/Hexeption/MCP-Reborn) Minecraft modding package. Unlike the regular Minecraft game, in which the server (or the "world") always runs at 20Hz and the client runs as fast as rendering

The environmental observations consist of two parts. The first part is the raw pixels from the Minecraft game that players would see, including overlays such as the hotbar, health indicators, and animations of a moving hand in response to attack or "use" actions. The field of view, GUI scale, and brightness parameters are consistent with VPT (Baker et al., 2022). The second part includes auxiliary information about the agent’s current environment, such as its location and weather conditions. Human players can obtain this information by pressing F3. The specific observation details we include are shown in Table 3.

Note that no high-level observations like voxels and lidar information in Minedojo​ (Fan et al., 2022) can be accessed by agents. During the actual inference process, the controller only perceives the raw pixels and interacts with the environment, which is the same with VPT​ (Baker et al., 2022) models. The agent will access information from the environment to generate the text condition of the controller.

B.2 Action Space

We design a hybrid action space. Some are directly available to human players, including keypresses, mouse movements, and clicks, which are similar to MineRL v1.0 (Guss et al., 2019b) used by VPT (Baker et al., 2022). The keypresses and clicks are binary functional actions, including forward, jump, use and attack etc. In addition to the binary (on/off) keypress actions, our action space also includes mouse movements. When the in-game GUIs (press "E" to open inventory) are closed, the mouse’s X and Y actions control the agent’s yaw and pitch. However, when the GUI is open, camera actions move the mouse cursor on the screen. In Minecraft, precise mouse movements are needed to interact with the inventory for tasks such as crafting and smelting. On the other hand, mining and navigating the world can be done using broader mouse actions. To be enable to achieve both the same action space, we abstract the craft and smelt action with GUI into functional binary actions, which are same as MineDojo (Fan et al., 2022). The detailed action space are described in Table 4.

B.3 Rules

We choose to conduct the test in survival mode of Minecraft 1.16.5. For each environment reset, we have added the following rules:

/difficulty peaceful: Set the difficulty of the environment to peaceful mode.

/gamerule doDaylightCycle false: Set the environment to daytime forever.

/gamerule keepInventory true: Set agent to not drop items upon death. We have added a time limit for each task, within which if the player dies, they will respawn at the spawn point and retain their previous inventory contents.

/effect give @a night_vision 99999 250 true: In order to facilitate the display of agent behavior, we have added night vision effects to the agent.

Appendix C Results and Details of 200+ tasks in Minecraft Universe Benchmark

We list the evaluation task set belows with details including task name, maximum steps, initial inventory, biome, and language instructions. We also show the evaluation times across different seeds and successful episodes rate. Note that all tasks are evaluated in Minecraft 1.16.5 Survival Mode.