Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld
Yijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao, Lusong Li, Li Shen, Xiaodong He, Jing Jiang, Yuhui Shi
Introduction
Embodied multi-modal agents have been acknowledged as a pivotal stride towards achieving Artificial General Intelligence (AGI), as they encompass the potential for a broadened range of intelligent activities . The rise of foundation models brings a glimmer of hope for constructing such agents , and notable efforts from the community have tried to harness them in many decision-making scenarios, e.g., autonomous driving , daily household robots , and complex manipulation tasks .
LLMs can perform as reflex agents interacting with a text world via textual descriptions (observations) of the world states and textual actions. In addition, LLMs exhibit great potential in planning, reflection, and reward shaping. This is mainly due to their reasoning capability and semantic abstraction of the world. However, LLM-based agents cannot be directly applied in a visual world. On the contrary, it is still challenging to train an agent in a visual world due to its complexity, the lack of semantic structures in the perceived pixels, and the noises of the visual signals, which cannot be fully handled by the current vision models such as ViT . Although SOTA large VLMs align LLMs with vision models, their pretraining only focuses on static alignment between image-text pairs so the resulting agents have not been aligned well with the dynamics of the visual world. As shown in Fig. LABEL:fig:teaser (a), even the SOTA VLM, e.g., GPT-4V(ision) , fails to accomplish tasks in the embodied ALFWorld environment . In such a zero-shot setting, GPT-4V tends to mainly rely on its language prior about the objects (laptop) detected in the current step, rather than the alignment between the visual input and the environment dynamics conditioned on the task instruction. How can we finetune a VLM to be an embodied agent with dynamic alignment to the visual world? Moreover, LLMs’ capability is weakened in VLMs since their (partial) inputs are replaced by noisy and inaccurate representations from the vision modules. Furthermore, various VLMs cannot use the most powerful LLMs as their language modules since they are close-sourced. Can we transfer the skills of API LLM reflex agents in a text world to a VLM agent in a visual world?
In this paper, we study how to train a VLM towards an embodied reflex agent in a visual world by aligning it with the visual world dynamics and distilling the skills of an API LLM agent in a parallel text world. Our overarching goal is to build such an Embodied Multi-Modal Agent (EMMA) that can take a textual task description (e.g., from human users) and pixel observations of the state per step to produce a sequence of actions leading to the efficient completion of the task. This is a challenging problem due to (1) the sparsity of the task reward, (2) the noisy visual representations, (3) the hallucinations of the VLM, and (4) the misalignment of VLM’s static representations to the visual world dynamics. While static distillation and imitation of a powerful LLM agent in a parallel text world can potentially address the former two challenges, the latter two can be mitigated by finetuning the VLM as an interactive agent in the embodied visual environment.
To this end, we finetune a VLM agent in an embodied visual world by imitation learning from an API LLM expert (e.g., built on ChatGPT) launched in a parallel text world on the same tasks. Specifically, in each step of EMMA interacting with the visual world, we convert its visual observation into an equivalent textual description sent to the LLM agent, which produces an action for EMMA to imitate. Our cross-modality interactive imitation learning is based on DAgger , which overcomes the cumulative errors and distribution shifts of behavior cloning (BC). As depicted in Fig. LABEL:fig:teaser (b), an InstructBLIP agent finetuned by BC on 170K expert demonstrations produced by a rule-based expert in the visual world still fails to take correct actions based on visual observations. We further improve the DAgger’s objective to be the direct preference optimization (DPO) , which maximizes the preference of LLM actions over VLM actions. To collect better teaching signals retrospective to the VLM student’s actions, the LLM expert is composed of an LLM actor prompted to output expert actions, and an LLM critic prompted for reflection feedback on the VLM agent’s historical trajectories. We maintain a long-term memory storing the LLM feedback, which is then used to induce the actor to improve actions for imitation in future episodes.
We evaluate EMMA and compare it with vision-only agents, LLM agents, and VLM agents deployed to the ALFWorld benchmark, which includes hundreds of tasks in both visual and textual environments. Extensive experiments highlight that EMMA substantially outperforms SOTA VLM agents in visual-only environments by 20%-70% in terms of success rate. In addition, EMMA is the only VLM agent that can generalize to open-vocabulary and free-form test tasks, shedding novel insights on using LLM feedback to train more versatile and generalizable embodied agents in multi-modality environments.
Embodied Multi-Modal Agent
Fig. 2 illustrates the main idea of “Embodied Multi-Modal Agent (EMMA)”, whose detailed training procedures are given in Alg. 1. EMMA is built upon a modularized VLM that can follow instructions and interact with the environment through pixel observations and textual actions. To overcome the training challenges associated with EMMA, such as sparse rewards and distribution shifts, we explore the construction of an LLM expert based on a parallel TextWorld, which can provide EMMA with step-by-step guidance. Hence, in Sec. 2.3, we further discuss how to harness the LLM expert to train EMMA via cross-modality interactive imitation learning. Code is available at https://github.com/stevenyangyj/Emma-Alfworld
In visual environments, the Embodied Multi-Modal Agent (EMMA), denoted as , is designed to process a task description (e.g., instructions from human users) and pixel observations at each step . Its objective is to generate a sequence of high-level textual actions For brevity, we omit in the rest of this paper. towards efficient completion of the task. To achieve this, we draw inspiration from recent advances of large pretrained VLMs , and modularize EMMA’s architecture into three components: (1) a pretrained ViT as the vision encoder that encodes into visual embeddings; (2) a querying transformer (Q-Former) that is tailored to extract the most relevant visual features via the cross-attention between the visual embeddings and query tokens, and then feeds them into an LLM through a linear projection layer; (3) a pretrained LLM as the language decoder that takes the concatenation of the instruction tokens and the output of the linear projection layer to autoregressively generate the textual action . To reduce computational overhead and prevent catastrophic forgetting , we directly use the pretrained ViT, Q-Former, and LLM from InstructBLIP and keep them frozen during finetuning, while updating only the parameters of the linear projection layer, as illustrated in Fig. 2. Such a modularized architecture enables EMMA to integrate any existing pretrained vision models and LLMs in a flexible and computationally-efficient way.
However, deploying EMMA into complex visual environments remains challenging. One of the main obstacles is that the direct use of any pretrained VLM as EMMA’s backbone, without additional finetuning, is suboptimal. This suboptimality arises because existing pretraining only focuses on static alignment between image-text pairs , so the resulting agent may struggle to reason about the dynamics of the environment. As discussed in Sec. 1, even the SOTA VLM, i.e., GPT-4V, fails to accomplish tasks in the embodied ALFWorld environment . In such a zero-shot setting, GPT-4V tends to rely on the linguistic prior of the currently detected objects, rather than the given task instruction and the underlying dynamics of the environment, to guide the interactions. Moreover, finetuning a pretrained VLM using a pre-collected demonstration dataset is also suboptimal due to the diversity of environments and tasks , the scarcity of large-scale expert annotations as well as the challenges posed by distribution shifts . A seemingly natural solution to align EMMA with the environment dynamics is reinforcement learning from environmental feedback (RLEF) , in which reward signals rely on decomposing a task into a sequence of reasonable sub-goals and checking their completion. However, in real-world scenarios, most sub-goals cannot be defined or described precisely, so the reward is sparse; hence, we do not expect RLEF to be effective.
To this end, we propose to leverage an interactive imitation learning (IL) to align EMMA with dynamics of any environment, which however introduces two critical algorithmic challenges: (1) How to obtain a high-quality, accessible, and scalable expert that EMMA can query during IL (Sec. 2.2)? (2) Designing an effective strategy to train EMMA using this expert in complex, diverse, and potentially open-ended environments (Sec. 2.3).
2 LLM Expert from a Parallel TextWorld
Thanks to a series of prompting techniques in the realm of in-context learning, such as chain-of-thought , tree-of-thought , ReAct , and Reflexion , pretrained LLMs have demonstrated impressive zero-shot performance across many decision-making scenarios . Despite the great potential of serving as high-quality and scalable experts, they are only able to interact with the environment via textual descriptions of the states, rather than using raw pixel observations like EMMA. To bridge this gap, we convert each pixel observation into a textual equivalent by extracting its metadata from the simulator , which encompasses attributes such as Observed Objects, Observed Relations, Inventory, and Location. We then employ the Planning Domain Definition Language (PDDL) to describe this metadata and create an equivalent textual description/state using the TextWorld engine . Additional details are available in Appendix 7 and an example of this process is illustrated in Fig. 1. This methodology enables the utilization of any pretrained LLM agent (e.g., based on ChatGPT) from the parallel TextWorld, generating a sequence of actions that facilitates the training of EMMA through cross-modality IL between the two agents.
3 Training EMMA via Cross-Modality Imitation
Given an LLM expert from the parallel TextWorld, our goal is to train a VLM agent in the visual world to closely imitate the the LLM expert’s behaviors. This is equal to minimizing the following objective under the distribution of states induced by .
in which the choice of loss function is dependent on specific scenarios. For instance, it may be the expected cross-entropy loss for the discrete action space, or the expected MSE loss for the continuous action space. In our case, we select DPO loss due to its proven superior performance to the cross-entropy on aligning models with expert preferences within the discrete language space. Hence, Eq. (1) can be extended to the formulation below.
where is the action given by the LLM expert while is the action given by , is the logistic function, and is a hyperparameter controlling the deviation from , i.e. the reference agent obtained by behavior cloning on a demonstration dataset produced by a rule-based expert in the visual world (see Appendix 9 for complete details). The added regularization is essential, as it prevents the agent from deviating too far from the distribution on which the expert is accurate, as well as maintaining the generation diversity and avoiding premature convergence to some easy tasks. In practice, the VLM agent is also initialized to for stabilizing the training process. Since environmental dynamics is both unknown and complex, we cannot compute the distribution of states visited by and can only sample it by rolling out the agent. Hence, Eq. (2) is a non-i.i.d. supervised learning problem due to the dependence of the state distribution on itself, in which naïve behavior cloning faces issues like cumulative error and distribution shift . To address this, we employ an interactive IL algorithm, DAgger , which provably converges to the optimal agent .
As discussed in Sec. 2.2 and Sec. 2.3, we can harness an API LLM, e.g., GPT-4 , to generate a sequence of actions, serving as in Eq. (2). However, these actions may be suboptimal due to sparse environmental feedback or defective in-context instructions . To collect better teaching signals, we introduce a “retrospective LLM expert” that is composed of two specialized models: an actor (), built upon an API LLM and prompted to generate actions based on the task instruction and textual state observations; and a critic (), also based on the same LLM, but designed to analyze EMMA’s historical trajectories and provide reflective feedback. A long-term memory is maintained to store the feedback generated by , which is then used to prompt for improved actions. The complete procedure is detailed in Line 7-10 of Alg. 1, and all prompts are provided in Sec. 6 of the Appendix.
Experiments
Environments. We base our experiments on the ALFWorld benchmark , a cross-modality simulation platform that encompasses a wide range of embodied household tasks. For each task, ALFWorld integrates a visual environment, rendered by the Ai2Thor simulator , with a corresponding textual environment. This textual aspect employs the Planning Domain Definition Language (PDDL) to translate each pixel observation from the simulator into an equivalent text-based observation, and then constructs an interactive environment using the TextWorld engine . Fig. 1 provides an illustrative example of the tasks featured in ALFWorld. The tasks within the ALFWorld benchmark are categorized into 6 types: Pick & Place, Clean & Place, Heat & Place, Cool & Place, Look in Light, and Pick Two Objects & Place. Each task requires an agent to execute a series of text-based actions, such as “go to safe 1”, “open safe 1”, or “heat egg 1 with microwave 1”, following a predefined instruction. These actions involve navigating and interacting with the environment. To provide a comprehensive understanding, we have visualized an example of each task type in Fig. 8 of the Appendix. A task in this benchmark may involve interactions with over 10 objects and require more than 30 steps for a human expert to solve, thus challenging an agent’s capabilities in long-horizon planning, instruction following, and the utilization of commonsense knowledge. For a fair comparison, we follow the same setting as prior work and evaluate all baselines using 134 out-of-distribution (OOD) tasks.
Baselines. To verify the effectiveness of cross-modality imitation learning, we compare our EMMA with several baselines and state-of-the-art (SOTA) agents using the ALFWorld benchmark with both visual and textual environments. The compared agents can be divided into three categories: vision models, language models, and vision-language models. Concretely, vision models, including ResNet-18 and MCNN-FPN , utilize pretrained vision encoders to extract salient features from each pixel observation. The extracted features then serve as input for a Multi-Layer Perceptron (MLP) policy, which is trained by behavior cloning on a pre-collected demonstration dataset. Unlike vision models performing in the visual environment, language models complete exactly the same tasks but in a parallel, text-based environment. BUTLER employs a transformer seq2seq model enhanced with a pointer softmax mechanism . This architecture aggregates previous observations as input to generate text-based actions token-by-token. GPT-BUTLER , a variant of the GPT-2 model , is initially pretrained on a static demonstration dataset and further finetuned using data collected online. ReAct takes a novel approach by utilizing LLMs to generate reasoning traces and task-specific actions in an interleaved manner. This method aids the agent in developing, tracking, and updating its action plans interactively. Reflexion similarly employs an LLM, but it focuses on reflecting upon environmental feedback. It maintains this reflective text in an episodic memory buffer, enhancing the agent’s ability to improve actions in subsequent trials. Similar to Reflexion, a concurrent work, DEPS , also corrects errors in previous LLM-generated actions by integrating descriptions of the action execution process and providing self-explanations for the feedback. Moreover, beyond the single-agent framework, AutoGen exhibits the potential of accomplishing a broad spectrum of tasks through the cooperation of multiple LLM agents. Finally, we consider a range of vision-language models, such as MiniGPT-4 , BLIP-2 , LLaMA-Adaptor , and InstructBLIP , as agents to interact with the visual environment. Unlike pure vision or language models, VLMs are designed to process and integrate both visual and textual data, offering a more holistic understanding of the environment. To align these agents with the specific requirements of the ALFWorld benchmark, they are finetuned on a pre-collected demonstration dataset. This finetuning process is crucial as it enables the agents to comprehend and adhere to ALFWorld’s unique grammar and to develop a basic “gamesense”.
2 Training Details
The architectural design of EMMA is depicted in Fig. 2. At its core, EMMA employs a Query Transformer (Q-Former) to process visual data. This Q-Former extracts features using a frozen ViT encoder. Its output consists of 32 visual tokens, which are then passed through a linear projection layer before fed to a frozen LLM decoder. Similar to other VLM agents, EMMA is also finetuned on a pre-collected demonstration dataset, aligning its basic ability with the ALFWorld benchmark. In order to align EMMA with the dynamics of ALFWorld, we train it by imitating an LLM expert (see Alg. 1). We choose text-davinci-003 developed by OpenAI as our LLM expert because of its established capabilities in reasoning and planning . In this setup, text-davinci-003 serves dual roles: it serves as an actor, providing EMMA with expert actions, and as a critic, analyzing EMMA’s historical trajectories. This analysis generates retrospective feedback, which is then incorporated into the actor’s long-term memory, leading to improved actions in future trials. Further details about the hyperparameters and prompts used in our training procedure are available in Table 1 of the Appendix.
3 Comparison with State of the Art
EMMA sets new SOTA performance in visual ALFWorld. In this section, we compare EMMA with 12 other representative agents on the ALFWorld benchmark, spanning both visual and textual environments (refer to Table 4). We assess two key metrics: the success rate, which is the percentage of trials completed successfully, and the average number of interaction steps required for task completion, with a lower number indicating higher efficiency. EMMA demonstrates superior performance in both metrics, significantly outperforming all VLM agents in visual environments. This achievement underscores the effectiveness of our cross-modality imitation learning approach, as depicted in the learning curve shown in Fig. 4. Furthermore, EMMA’s performance markedly exceeds that of VM agents, highlighting the crucial role of the prior knowledge embedded in VLMs. Intriguingly, EMMA’s performance is comparable with LLM agents that operate using perfectly semantic descriptions of visual observations. This is largely attributed to EMMA’s training strategy of imitating an expert LLM agent, proving to be more efficient than learning from scratch in a purely visual setting. As a result, EMMA stands out as the only VLM agent that substantially surpasses SOTA LLM agents, such as AutoGen and DEPS , in these environments. And its success also directs a potential way to achieve human-level performance in the visual environments of ALFWorld.
EMMA is more robust to noisy observations than LLM agents. While LLM agents exhibit a higher success rate with fewer interaction steps in textual environments, as indicated in Table 4, we hypothesize that this superior performance largely relies on their precise semantic abstraction of the environment. However, such an abstraction might not be feasible in real-world applications. To verify this assumption, we set up a more practical scenario where observations are deliberately perturbed at a specific noise rate. We then compare the robustness of EMMA and a SOTA LLM agent, Reflexion, under these noisy observations. To generate noisy observations, a random portion of the visual observation is cropped, resized, and then used to replace the original observation. Similarly, in the textual observation, random tokens are substituted with arbitrary ones. As illustrated in Fig. 5, with the noise rate increases, EMMA’s performance remains significantly more robust compared to Reflexion. This could be attributed to the vision encoder in the VLM, which is adept at filtering out visual noises. On the other hand, textual noises are directly processed by the LLM, which can substantially impair the performance of LLM-based agents. This finding highlights the potential advantages of VLM agents like EMMA in practical scenarios, in which data is often imperfect and noisy.
4 Ablation Study
Retrospection improves EMMA over time. To assess the importance of the retrospective LLM expert, we present the average success rate of EMMA after each trial, as shown in Fig. 6 (left), and evaluate a key variation: EMMA w/o Retrospection. This variant of EMMA is trained using the same procedure as the original EMMA but removes the retrospective process. Instead, it relies solely on a plain LLM actor to provide relabeled actions. The results show that EMMA with the retrospective mechanism significantly outperforms its counterpart. This finding is crucial as it indicates that the retrospective process is not just a supplementary feature but a fundamental component of EMMA’s architecture that contributes substantially to its enhanced performance.
EMMA benefits from BC initialization. We evaluate the impact of behavior cloning (BC) initialization, a process described in line 3 of Alg. 1, through an ablation study. Fig. 6 (middle) demonstrates that EMMA, when deprived of BC initialization, experiences a slight reduction in the average success rate across 134 unseen tasks compared to its original setup. Despite this decrease, EMMA without BC initialization still outperforms other VLM agents, as clearly shown when compared with the results in Table 4. Furthermore, Fig. 10 in the Appendix breaks down the result by task type. It reveals a consistent but slight drop in performance across 5 out of the 6 task types. These results reflect that while BC initialization contributes positively to EMMA’s overall performance, it is not critical for achieving the notable results we have reported.
DPO enables more effective imitation learning. To evaluate the effectiveness of using DPO loss in Eq. (1), we conducted an ablation study with an alternative version of EMMA, referred to EMMA w/ Cross Entropy (CE) Loss. In this variant, EMMA is optimized using the token-level CE loss, a common objective for finetuning VLMs. The results, as depicted in Fig. 6 (right), reveal that EMMA w/ CE Loss does not achieve the same high success rate as the original EMMA with DPO loss, suggesting that DPO loss contributes to enhancing the upper performance bound of EMMA. In addition, we noted that EMMA w/ CE Loss exhibits faster convergence in the initial stages of training compared to the original EMMA. This premature convergence leads to the agent paying attention to the expert actions from easier tasks, usually addressed in the first few training epochs, which can suppress EMMA’s exploration and learning on more complex tasks.
5 Generalization to Free-form Task Instructions
The ability of AI agents to accurately follow task instructions given in open-vocabulary and free-form text is crucial for their real-world applicability. To assess this, we conducted an additional experiment focusing on the generalization capabilities of EMMA, an agent trained with templated task instructions. In this experiment, we re-evaluated the performance of EMMA and other baseline agents on 134 unseen tasks, using human-annotated instructions instead of the templated ones. These human-annotated instructions include a large amount of OOD verbs and objects, presenting a more realistic and challenging scenario for the agents. To underscore this challenge, we compared the vocabulary distribution between the templated and human-annotated instructions, as shown in Fig. 7 (right). Moreover, we provide a comprehensive analysis of the vocabulary used across both instruction types in Fig. 11 of the Appendix. In Fig. 7 (left), EMMA demonstrates a slight performance decline, while a significant degradation is observed in other baselines. We also note that Reflexion, a SOTA LLM agent, exhibits exceptional generalization to those OOD instructions. According to these empirical results in Fig. 7, we have the following conclusions: (1) EMMA obtains and benefits from the generalization capabilities inherent in the SOTA LLM agent through cross-modality imitation learning; (2) Our work sheds novel insights on using LLM feedback to train more versatile and generalizable embodied agents.
Related Work
Agents based on Foundation Models. Recent research has increasingly focused on harnessing the capabilities of large pre-trained foundation models to build AI agents . These models (e.g., LLMs), benefiting from their commonsense knowledge inherited from Internet-scale pretraining, are able to reason actions according to descriptions of the external environments. For example, given a set of task instructions, LLMs can be elaborately prompted to perform as agents generating high-level step-by-step plans , and each step can be parsed into a sequence of robotic actions that are executed via pretrained policies or available APIs . Furthermore, by using VLMs , plans can also be conditioned on visual inputs that are transformed into language descriptions or token embeddings aligned with LLMs . However, existing foundation models are usually pretrained on static text or text-image datasets and thus may struggle to align with the dynamics of the world. To bridge this gap, we study how to finetune a VLM to be an embodied agent that is dynamically aligned with the world by distilling the cross-modality knowledge from an LLM expert. The work most closely related to ours is EUREKA , which also explores using source information provided by the simulator as the context of an LLM to aid agent training. Instead of directly mimicking the output of the LLM as we did, EUREKA harnesses the coding LLM to generate a desired reward function for a given task and optimizes a policy against the reward function using RL, leading to a more complex and unstable training procedure .
Imitation Learning. Imitation learning is the study of algorithms that improve performance by mimicking an expert’s decisions and behaviors. We summarize three main categories of existing methods in the following: (1) behavior cloning (BC), (2) inverse reinforcement learning (IRL), and (3) the combination of imitation and reinforcement learning. The naïve BC ignores the change in distribution and simply trains a policy that only performs well under the distribution of states visited by the expert. Following works, such as dataset aggregation or policy aggregation , have been proposed to address the limitations of BC. Another line of work, IRL, is a more complicated algorithm framework that learns the reward function from expert demonstrations and then improves the policy using RL with the learned reward. A representative method in this category is generative adversarial imitation learning (GAIL) , in which a policy and a discriminator compete with each other in order to maximize the likelihood of the policy’s behavior matching the expert. The third category of methods usually leverages an IL policy to initialize RL and continues to boost its performance via online collected data from RL . This simple combination significantly improves RL’s sample efficiency and IL’s upper performance bound constrained by the expert. Nevertheless, all of the above methods assume that the expert and the imitator understand the world in the same modality, and thus overlook the fact that the complementary knowledge from other modalities often boosts the model’s accuracy and generalization dramatically .
Conclusion
We create EMMA, an Embodied Multi-Modal Agent, by finetuning a VLM in an embodied visual world with interactive imitation learning from an LLM expert in a parallel text world, who produces better actions and retrospective feedback to VLM’s trajectories. Such imitation learning exhibits substantial advantages over vision or VLM policies directly finetuned in the visual world or finetuned by behavior cloning of a rule-based expert, and SOTA API VLMs such as GPT-4V(ision). As a result, EMMA achieves a comparable success rate and much better robustness in the noisy visual world than its LLM teacher in the easier text world. Furthermore, EMMA shows powerful generalization to open-vocabulary and free-form task instructions, highlighting its potential in real-world scenarios.
References
Appendix
Full Prompts for LLM Expert
In this section, we provide all LLM prompts for the training procedure (Alg. 1) of EMMA. We adopt the prompting technique developed by ReAct but ignore the reasoning traces, i.e., “think” steps, when executing imitation learning between EMMA and the LLM actor. After each trial , the retrospective feedback generated by the LLM critic will be appended to long-term memory . In practice, we bound by a maximum number of stored feedback (usually set to 1-3) to adhere to the max context length of the LLM.
Parallel TextWorld
While the idea of parallel TextWorld is heavily inspired by previous work , we have enhanced the TextWorld engine to create text-based equivalents of each visual environment for training and evaluating language-based agents. This enhancement involves utilizing a combination of the PDDL and Fast Downward to maintain and update the textual state of the simulated environments. Based on the metadata provided by the simulator, we represent a visual state as a list of attributes. These attributes detail the relationships among various entities in the environment, such as positions, goals, and objects. Note that all these attributes, variables, and rules are defined within the framework of PDDL.
Training Details
We provide hyperparameters used for training EMMA in Table 1. These hyperparameters are largely derived from those proposed for finetuning InstructBLIP model . When training, we only update the parameters of linear projection layer while keeping other components frozen, as done during instruction tuning for many existing work . We use the AdamW optimizer with a linear warmup of the learning rate, followed by a linear decay with a minimum learning rate of 0. Moreover, we remove the instruction input of Q-Former, which is used in InstructBLIP, and find this improves performance cross all experiments. Our implementation is heavily inspired by the LAVIS library so the training and evaluation processes use the standard procedure provided by LAVIS.
Collection of Demonstration Dataset
Fine-tuning pretrained VLMs on a pre-collected demonstration dataset via behavior cloning is a critical step, enabling these models to comprehend and follow the unique grammar of ALFWorld as well as to develop a basic “gamesense”. However, the number of task instructions in the original ALFWorld is too limited to yield sufficient data for fine-tuning these large pretrained VLMs effectively. Hence, we propose an automated pipeline, which leverages text-davinci-003 and a rule-based planner to generate a large amount of new instructions and their resulting expert demonstrations, respectively.
To generate a diverse set of new task instructions, we harness the in-context learning capabilities of LLM. Our procedure begins with extracting detailed descriptions from the ALFWorld benchmark for each environment, providing comprehensive information on the number and functional attributes of all items. Then, based on the types of room in these environments, we design different prompts that aim at inducing the LLM to generate task instructions aligned with the features of the target environment. An example of these prompts is detailed in Table 2. For each generated task instruction, we gather demonstrations using a rule-based planner devised by ALFWorld. It’s important to note that this planner operates with an unfair advantage: it considers the environment as fully observable and has complete information of world dynamics, relying on metadata that is not accessible to the agent during training. In summary, our dataset comprises 15247 expert demonstration episodes, amounting to 178585 image-text pairs.