Vision-Language Foundation Models as Effective Robot Imitators
Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Jie Xu, Hongtao Wu, Chilam Cheang, Ya Jing, Weinan Zhang, Huaping Liu, Hang Li, Tao Kong
Introduction
Recent progress in vision-language foundation models (VLM) has presented their exhilarating ability in modeling and aligning the representation of images and words, and the unlimited potential to resolve a wide range of downstream tasks with multi-modality data, for instance, visual question-answering (Li et al., 2023; Zhou et al., 2022), image captioning (Zeng et al., 2022; Wang et al., 2022; Li et al., 2021), human-agent interactions (Liu et al., 2022b; Oertel et al., 2020; Seaborn et al., 2021). These successes, undeniably, encourage people to imagine a generalist robot equipped with such a vision-language comprehension ability to interact naturally with humans and perform complex manipulation tasks.
Therefore, we aim to explore integrating vision-language foundation models to serve as robot manipulation policies. While there have been some previous studies that incorporated large language models (LLMs) and vision-language models (VLMs) into robot systems as high-level planners (Ahn et al., 2022; Driess et al., 2023), making use of them directly for low-level control still poses challenges. Most VLMs are trained on static image-language pairs, whereas robotics tasks require video comprehension for closed-loop control. Additionally, VLM outputs primarily consist of language tokens, which significantly differ in representation compared to robot actions. A recent work (Brohan et al., 2023), namely Robotics Transformer 2 (RT-2), has demonstrated a possible solution for adapting VLMs to low-level robot control. However, democratizing such an expensive framework for all robotics practitioners proves difficult as it utilizes private models and necessitates co-fine-tuning on extensive vision-language data to fully showcase its effectiveness. Consequently, there is an urgent need for robot communities to have a low-cost alternative solution that effectively enables a robot manipulation policy with VLMs.
To this end, we introduce RoboFlamingo, a novel vision-language manipulation framework that leverages publicly accessible pre-trained VLMs to effectively construct manipulation policies for robotics. Specifically, RoboFlamingo is grounded upon the open-source VLM, OpenFlamingo (Awadalla et al., 2023), and resolves the challenge by decoupling visual-language understanding and decision-making. Unlike previous works, RoboFlamingo takes advantage of pre-trained VLMs mainly for understanding vision observations and language instructions at every decision step, models the historical features with an explicit policy head, and is fine-tuned solely on language-conditioned manipulation datasets using imitation learning. With such a decomposition, we only need to combine a small amount of robotics demonstration to adapt the model to downstream manipulation tasks, and RoboFlamingo also offers flexibility for open-loop control and deployment on low-performance platforms. Moreover, benefiting from the pre-training on extensive vision-language tasks, RoboFlamingo achieves state-of-the-art performance with a large margin over previous works, and generalizes well to zero-shot settings and environments. It is worth noting that RoboFlamingo can be trained or evaluated on a single GPU server. As a result, we believe RoboFlamingo can be a cost-effective yet high-performance solution for robot manipulation, empowering everyone with the ability to fine-tune their own robots with VLMs.
Through extensive experiments, we demonstrate that RoboFlamingo outperforms existing methods by a clear margin. Specifically, we evaluate its performance using the Composing Actions from Language and Vision benchmark (CALVIN) (Mees et al., 2022b), a widely-recognized simulation benchmark for long-horizon language-conditioned tasks. Our findings indicate that RoboFlamingo is an effective and competitive alternative for adapting VLMs to robot control, achieving 2x performance improvements compared with the previous state-of-the-art method. Our comprehensive results also yield valuable insights into the use of pre-trained VLMs for robot manipulation tasks, offering potential directions for further research and development.
Related Work
Language can be the most intuitive and pivotal interface for human-robot interaction, enabling non-expert humans to seamlessly convey their instructions to robots for achieving diverse tasks. Consequently, the realm of language-conditioned multi-task manipulation has garnered substantial attention in recent years. Intuitively, such tasks require robots to have a good understanding of not only the visual captures of the outside world, but also the instructions represented by words. With the strong representation ability of pre-trained vision and language models, a lot of previous works have incorporated pre-trained models into the learning framework. Among them, we roughly classify them into the following three categories, which is also illustratively compared in Fig. 1.
While some early works such as Jang et al. (2022); Lynch & Sermanet (2020) trained a vision encoder and a language encoder to learn representations for the input language and vision data from manipulation tasks, some recent work directly takes pre-trained models to obtain great representations, then trains the policy model beyond them from scratch or fine-tuning the whole model. For instance, Jiang et al. (2023) utilizes a pre-trained T5 (Raffel et al., 2020) model to encode the multi-modal prompts, and learn the actions by fine-tuning the T5 model and additionally training an object encoder and attention layers. HULC (Mees et al., 2022a) utilizes the vision encoder of Lynch & Sermanet (2020) trained on the CALVIN dataset (Mees et al., 2022b) and some pre-trained language encoder models such as sentence transformer (Reimers & Gurevych, 2019), and their HULC++ (Mees et al., 2023) also fine-tunes these encoders. Besides, Brohan et al. (2022) proposed RT-1, i.e., robotics transformers, a 35M vision-language-action model (VLA) which tokenizes the action and aligns the vision, language, and action in the token space and is trained on a large amount of real-world manipulation dataset, using the Universal Sentence Encoder (Cer et al., 2018) to obtain the language embedding and the pre-trained EfficientNet-B3 (Tan & Le, 2019) as the vision tokenizer.
Some approaches have exploited large language models (LLMs) as a powerful zero-shot planner, e.g., SayCan Ahn et al. (2022), to generate step-by-step pre-defined plans with human-interactive prompts on given tasks, subsequently instructing different pre-trained low-level skill policies to execute those plans and finish multiple tasks. Compared to other works, the controlling policies do not require any ability to understand instructions, but rely on the pre-trained frozen LLM to select necessary skills.
Driess et al. (2023) proposed 540B PaLM-E model, showing a different way of utilizing the pre-trained vision and language model. Specifically, they choose different pre-trained models to encode the input scene, and the PaLM (Chowdhery et al., 2022) model as the base model, train the model to generate pre-defined multi-step plans described by language by co-fine-tuning the whole VLM end-to-end using both mobile manipulation question-answering data and auxiliary vision-language training data such as image captioning and visual question answering data collected from the web. Similar to SayCan (Ahn et al., 2022), they require low-level control policies to execute the generated plans. Motivated by PaLM-E, Brohan et al. (2023) further introduced RT-2, which is based on RT-1 but is adapted to use large vision-language backbones like PaLI-X (Chen et al., 2023) and PaLM-E (Driess et al., 2023), training the policy utilizing both robot manipulation data and web data. Their method reveals that VLMs have the potential to be adapted into robot manipulation, yet their key co-fine-tuning training strategy requires a large amount of both web-scale data vision-language data and low-level robot actions. Additionally, the VLMs and the data they use are private, making it hard for every robotics practitioner to play on such a solution for their own.
Although these previous models somehow bridge the gap between vision and language on robot manipulation tasks, they either reply on low-level skill policies, like SayCan and PaLM-E; or train a whole large model, such as RT-1; or require a huge amount of vision-language data and computational resources to ensure the model learns the manipulation policy without forgetting the great alignment of vision and language. Compared with these works, our proposed RoboFlamingo is a simple and intuitive solution to easily adapt existing VLMs (OpenFlamingo (Alayrac et al., 2022; Awadalla et al., 2023) used in this paper), only requiring fine-tuning on a small number of manipulation demonstrations. We hope RoboFlamingo provides a different perspective on fully leveraging the ability of VLMs, while requiring less data collection costs and computing consumption to make it an open and easy-to-use solution for everyone.
Background
In this paper, we mainly consider robot manipulation tasks, where the agent (robot) does not have access to the ground-truth state of the environment, but visual observations from different cameras and its own proprioception states. As for the action space, it often includes the relative target pose and open/closed state of the gripper. For instance, in the testbed of CALVIN (Mees et al., 2022b), the observations consist of simulated camera captures from two different views, and the action is a 7-DoF control of a Franka Emika Panda robot arm with a parallel gripper, and the instructions are reaching goals, i.e., the after-the-fact descriptions.
Imitation learning (Pomerleau, 1988; Zhang et al., 2018; Liu et al., 2020; Jang et al., 2022) allows the agent to mimic the manipulation plans from instruction-labeled expert play data , where is the number of trajectories, is the language instruction, and contains preceding states and actions to reach the goal described by the given instruction. The learning objective can be simply concluded as a maximum likelihood goal-conditioned imitation objective to learn the policy :
RoboFlamingo
RoboFlamingo, a generalized robotics agent, excels in resolving language-conditioned manipulation tasks. The key idea is to draw help from pre-trained vision-language models (VLMs) and adapt them to manipulation policies, acquiring the ability of object grounding, language comprehension, vision-language alignment, and long-horizon planning. Particularly, RoboFlamingo looks into one of the popular VLMs, Flamingo (Alayrac et al., 2022), and takes its open-source model OpenFlamingo (Awadalla et al., 2023) as the backbone. The overview of RoboFlamingo is shown in Fig. 2. To adapt large-scale vision-language models to robotic manipulation, RoboFlamingo simply adds a policy head for end-to-end finetuning. It addresses three main challenges: 1) it adapts vision-language models with static image inputs to video observations; 2) it generates robot control signals instead of text-only outputs; 3) it requires a limited amount of downstream robotic manipulation data to achieve high performance and generality with billions of trainable parameters. We will elaborate on the design of RoboFlamingo in this section.
The problem of language-conditioned robot control can be modeled as a goal-conditioned partially observable Markov decision process (GC-POMDP) (Liu et al., 2022a): , where and are the set of states and observations separately, is the action space, is the environment dynamics function, is the initial state distribution, indicate if the task is successful, and is the observation function. Specifically, for each controlling episode, the robot is given a goal, represented by a length- free-form language instruction at every time step , and the observations are typically two images , from a third-perspective camera and a gripper camera. The controlling policy can be modeled as a goal-conditioned policy and the action is typically the desired relative position and pose of the gripper, along with its open/close status.
In our RoboFlamingo, the policy is parameterized by . It consists of a backbone based on Flamingo and a policy head . The backbone takes visual observations and language-represented goals as the input and provides a latent fused representation at each time step for the policy head: . Then the policy head further predicts the action to fulfill the specified goal for the robot: , where is the hidden state from the last step that encodes the history information for decision-making. We will introduce each module in detail in the following sections.
2 The Flamingo Backbone
We adopt the Flamingo backbone for understanding the vision and language inputs at every decision step. Overall, Flamingo encodes the vision observations to the latent tokens by a vision encoder; and then fuses them with language goals through the feature fusion decoder. We explain these parts in detail below.
The vision encoder consists of a vision transformer (ViT) (Yuan et al., 2021) and a perceiver resampler (Alayrac et al., 2022). At every time step , the two-view camera images , are encoded to , consisting of a visual token sequence, through the ViT module:
where represents the visual token sequence at , represents the token number of the encoded output. After encoding, RoboFlamingo utilizes a perceiver resampler to compress the number of visual tokens from to . In detail, the resampler maintains a set of learnable parameters and utilizes the attention mechanism to reduce the number of token sequences to . Formally, the resampler is formulated as:
2.2 Feature Fusion Decoder
3 Policy Head
The output from the feature fusion decoder is trained as the representation of the vision observation and language instruction, which will be further translated into low-level control signals. To achieve this, we simply adopt an additional policy head to predict the action, e.g., the 7 DoF end-effector pose and gripper status. We test various strategies to model the historical observation sequences and behave as the policy head, e.g., a long short-term memory (LSTM) (Hochreiter & Schmidhuber, 1997) network with an MLP for the final prediction; a decoder-only transformer (Brown et al., 2020) similarly with an MLP; or a single MLP that only models single-step information (see Section 5 for more details). Taking the LSTM version as an example, with the vision-language joint embedding sequence , we obtain an aggregated embedding through a max-pooling operation over the token dimension and predict the action as:
where represents the hidden state at , and are the predicted end-effector pose and gripper status.
4 Training Objective
We utilize maximum likelihood imitation learning objectives to fine-tune the proposed pre-trained backbone and the policy head. Concretely, the desired relative pose is optimized via regression loss (we use mean squared error (MSE) loss) and the gripper status uses classification loss (we use binary cross-entropy (BCE) loss):
where is the demonstration for end effector pose and gripper status at timestep , corresponds to the weight of gripper loss.
In the training procedure, we follow the fine-tuning paradigm of OpenFlamingo by only training the parameters of the resampler, the gated cross-attention module of each decoder layer, and the policy head while freezing all other parameters.
Experiments
We conduct extensive experiments to examine the proposed RoboFlamingo solution, and answer how pre-trained VL models (VLMs) benefit language-conditioned robotic manipulation. In short, we investigate RoboFlamingo from the following perspectives:
Effectiveness. We wonder the imitation learning performance of RoboFlamingo by training it on the given demonstration data.
Zero-shot Generalization. We focus on generalization on unseen tasks. In other words, we study how the model will behave given unseen vision contexts like different objects, even with unseen instructions.
Ablation Studies. We further explore the essential factors that matter in adapting VLMs to robot control policy in the framework of RoboFlamingo.
We choose CALVIN (Mees et al., 2022b), an open-source simulated benchmark to learn long-horizon language-conditioned tasks, as our testbed, and the corresponding datasets as our imitation learning demonstration data. CALVIN encompasses a total of 34 distinct tasks and evaluates 1000 unique instruction chains for sequential tasks. In each experiment, the robot is required to successfully complete sequences of up to five language instructions consecutively. The policy for each consecutive task is dependent on a goal instruction, and the agent advances to the subsequent goal only if it successfully accomplishes the current task. The dataset contains four splits for environments A, B, C, and D. Each consists of 6 hours of human-teleoperated recording data (more than 2 million steps) that might contain sub-optimal behavior, and only 1% of that data is annotated with language instructions (24 thousand steps). See Fig. 4 in Appendix A.1 for a more detailed description and visualized examples of the benchmark.
We compare a set of well-performed baselines in CALVIN: (1) MCIL (Lynch & Sermanet, 2020): a scalable framework combining multitask imitation with free-form text conditioning, which learns language-conditioned visuomotor policies, and is capable of following multiple human instructions over a long horizon in a dynamically accurate 3D tabletop setting. (2) HULC (Mees et al., 2022a): a hierarchical method that combines different observation and action spaces, auxiliary losses, and latent representations, which achieved the SoTA performance on CALVIN. (3) RT-1 (Brohan et al., 2022): robotics transformer, which directly predicts the controlling actions by action tokens, as well as vision and language inputs. RT-2 (Brohan et al., 2023) is not experimentally compared since we have no access to their code, data, and model weights.
2 Imitation Performance
We train RoboFlamingo (with the M-3B-IFT backbone) using demonstrations only with language annotation from all 4 splits (A, B, C, and D), and evaluate the imitation performance on episodes sampled on split D ().The performance comparison is shown in Tab. 1. RoboFlamingo outperforms all baseline methods over all metrics by a large margin, even for those methods that are trained on the full set of data. This demonstrates the effectiveness of RoboFlamingo as the solution for robotics manipulation, enabling VLMs to become effective robot imitators.
In addition, the success rate of the subsequent tasks can be regarded as a notion of the generalizability of the manipulation policies, since the initial state of a subsequent task highly relies on the ending state of its former task. The later a task is arranged in the task sequence, the more diverse its initial state is, which will need more powerful visual-language alignment abilities to successfully complete the task. Among all methods, RoboFlamingo achieves the highest success rate over the latter tasks. This demonstrates that RoboFlamingo is able to utilize the visual-language grounding ability of pre-trained VLMs. In the appendix, we further include the results of RoboFlamingo co-trained with COCO and VQA data (Appendix B.1) and compare with recent robotics representation works (Appendix B.2). Appendix B.1 also reveals how the original VL abilities change after fine-tuning.
3 Zero-Shot Generalization
To assess the zero-shot generalization ability, we evaluate RoboFlamingo in two aspects: vision and language. For vision generalization, we train models on splits A, B, and C and test on split D, which presents a different vision context. Our method significantly outperforms baselines in this vision generalization scenario (), as shown in Tab. 1. Regarding language generalization, we enrich the language setting by generating 50 synonymous instructions for each task using GPT-4 (OpenAI, 2023). We then randomly sample instructions during evaluation. Our method exhibits superior performance compared to all baselines in this language generalization setting.
Note that the success rate of RoboFlamingo on subsequent tasks dropped more than HULC does. This may be due to our approach directly using word tokens as input during training, which can result in larger variations for synonymous sentences compared to HULC using a frozen sentence model for embedding instructions. To address this, we freeze the embedding layer of the feature fusion decoder in our method, leading to improved generalization and reduced performance drop.
4 Ablation Studies
In this section, we conduct ablation studies for RoboFlamingo to answer the following questions:
1) How does RoboFlamingo perform with different policy heads/formulations?
2) Does vision-language (VL) pre-training improve downstream robotic tasks?
3) How do critical factors in VL pre-training affect robotic tasks?
We test RoboFlamingo with different policy heads/formulations. In particular, we compare 4 different implementations: (a) takes only the current observation as input to predict actions, which ignores the observation history. (b) takes the history frames into the vision encoder with position embedding, and encodes the history information through the cross-attention layers in the feature fusion decoder. (c) and (d) both utilize the VLM backbone to process single-frame observations and integrate the history with the policy head. explicitly takes the visual history as input to predict the next action. implicitly maintains a hidden state to encode memory and predict the action. See Appendix C.1 for detailed illustration. We compare their best performance on the setting in Fig. 3 (a). performs the worst, indicating the importance of the history information in the manipulation task. performs better than , but is still much worse than and . We hypothesize that this may stem from the fact that the VLM (OpenFlamingo) has only seen image-text pairs during pre-training and cannot process consequent frames effectively. Further, the performance of and are similar, we choose as the default choice due to its simplicity.
4.2 Does VL pre-training improve downstream robotic tasks?
To verify the necessity of VL pre-training, we train the same model without loading the pre-trained parameters of the cross-attention layers and the resampler trained by OpenFlamingo models (denoted as No VL Pre-train). Besides, we also conduct an ablation study to freeze the pre-trained VLM and only train the policy head (denoted as No VL Finetune). As shown in Fig. 3 (b), we can see that vision-language pre-training crucially improves the downstream robotic manipulation by a large margin. Besides, tuning on the VL model itself on robotic tasks is indispensable due to the limited capacity of the policy head.
4.3 How do critical factors in VL pre-training affect robotic tasks?
A larger model usually results in better VL performance. Yet, with full training data in CALVIN, we find that the smaller model is competitive with the larger model (see the comparison in Tab. 2 and Appendix B.4). To further validate the impact of model size on downstream robotic tasks, we train different variants with 10% of language annotated data in CALVIN, which is only 0.1% of the full data. From Tab. 3 we can observe that with limited training data, the performance of VLMs is highly related to the model size. The larger model achieves much higher performance, indicating that a larger VLM can be more data-efficient.
Instruction fine-tuning. Instruction-Finetuning is a specialized technique that utilizes a further pre-training enhancement on the LLM with the IFT dataset (Conover et al., 2023; Peng et al., 2023), which provides a rich repertoire of instruction-following behaviors that inform its capabilities in language-conditioned tasks. We find that LLMs with such a training stage can improve the performance of the policy in both seen and unseen scenarios, revealed by the performance improvements of M-3B-IFT against M-3B, and G-4B-IFT against G-4B shown in Tab. 2.
5 Flexibility of Deployment
Since our RoboFlamingo adopts a structure that separates the perception and policy module and leaves the main computation on the perception module, we could perform open loop control to accelerate the inference of RoboFlamingo. Instead of taking only the next action to execute and performing VLM inference every time for new observations to predict future actions, open-loop control can be achieved by predicting an action sequence (stacked actions) with only one inference given the current observation, therefore alleviating the delay and the test-time computing requirement. However, as indicated in Fig. 3 (c), directly implementing open loop control without re-training may lead to deteriorated performance, retraining the model with jump step demonstration could alleviate the performance drop.
Conclusion and Future Work
This paper explores the potential of pre-trained vision-language models in advancing language-conditioned robotic manipulation. Our proposed RoboFlamingo, based on the pre-trained OpenFlamingo model, showcases state-of-the-art performance on a benchmark dataset. Moreover, our experimental findings highlight the benefits of pre-trained models in terms of data efficiency and zero-shot generalization ability. This research contributes to the ongoing efforts to develop intelligent robotic systems that can seamlessly understand and respond to human language instructions, paving the way for more intuitive and efficient human-robot collaboration. Due to the lack of real-robot data, this paper does not deploy on real-world robotics. To our delight, recent progress on large-scale real robotics data (Padalkar et al., 2023) has shown the potential of fine-tuning large VLMs for real robots, and the most exciting future work is to see how RoboFlamingo will behave in real-world tasks combined with such amount of data.
Acknowledgements
The Shanghai Jiao Tong University team is partially supported by National Key R&D Program of China (2022ZD0114804), Shanghai Municipal Science and Technology Major Project (2021SHZDZX0102) and National Natural Science Foundation of China (62322603, 62076161). The author Minghuan Liu is also supported by the ByteDance Scholarship and Wu Wen Jun Honorary Doctoral Scholarship.
References
Appendix A Environmental Setups
CALVIN (Mees et al., 2022b) is an open-source simulated benchmark for evaluating long-horizon language-conditioned tasks.
As shown in Fig. 4, CALVIN includes four different environments A, B, C, and D, each of which consists of 6 hours of human-teleoperated recording data (more than 2 million trajectories) that might contain sub-optimal behavior, and only 1% of that data is annotated with language instructions (around 24 thousand trajectories). Each split is settled with different settings of objects and environments, aiming to validate the performance, robustness, and generality of policies trained with different data combinations.
This benchmark requires a 7-DOF Franka Emika Panda robot arm with a parallel gripper, utilizing onboard sensors and images from two camera views to successfully complete sequences of up to five language instructions consecutively. This setup further challenges the robot’s ability to transition between various goals. CALVIN encompasses a total of 34 distinct tasks and evaluates 1000 unique instruction chains for sequences. The robot is reset to a neutral position after each sequence to prevent any policy bias resulting from its initial pose. This neutral initialization eliminates any correlation between the initial state and the task, compelling the agent to rely solely on language cues to comprehend and solve the given task. The policy for each consecutive task is dependent on the instruction of the current goal, and the agent advances to the subsequent goal only if it successfully accomplishes the current task.
A.2 Examples of Enriched Instructions
To validate the performance of the policies over diversified language expressions, we utilize GPT4 to augment the language instruction in CALVIN. We showcase the enriched language instructions in Tab. 4. We can see that the enriched instructions do have the same meaning as the original one, yet they are organized with different words. As shown in Table 1, RoboFlamingo can still achieve better performance compared to HULC.
A.3 Computing Resource
All experiments involved in this paper are conducted on a single GPU server with 8 NVIDIA Tesla A100 GPUs, and the default batch size is 6 on each GPU. The MPT-3B model takes 13 hours of training per epoch and achieves the best performance at the 3rd epoch, while the MPT-9B model also takes 26 hours of training per epoch and achieves the best performance at the 4rd epoch.
Appendix B Extended Experimental Results
From Tab. 1, the Enriched setting, we have noticed some evidence that the model may lose some foundation capabilities as the performance loss, which indicates that there is over-fitting during the fine-tuning. To further understand the phenomenon, we conduct further experiments by testing the fine-tuned RoboFlamingo model (the M-3B-IFT variant) on the COCO image caption and VQAv2, which verify our conjecture (see Tab. 6). To prevent such problems, we choose to co-train RoboFlamingo (the M-3B-IFT variant) with VQA and COCO datasets during fine-tuning on the robotics dataset. We test the co-train model on CALVIN, and the COCO image caption, VQAv2 tasks as well, as shown in Tab. 6 and Tab. 5. This provides a solution for fine-tuning VLMs to robotics models while preserving the ability on vision-language tasks, even though it may slightly deteriorate the performance on robotic tasks. In our implementation, we ensure that the model equally incorporates batches of VL and robot data in each epoch. From Fig. 5 and Fig. 5 we could observe a similar performance curve of the Co-trained and Fine-tune version of our model, while Co-trained model achieves higher performance in the early epochs and Fine-tune model ends up higher in the later epochs. One interesting observation is that under the Enriched setting, the performance of the co-trained model also drops, this may indicate the difference between understanding different sentences and aligning vision-language representations (as the pre-trained tasks do).
B.2 Comparison with Pre-trained Robotics Representation Models
We consider comparing our RoboFlamingo with recent pre-trained robotics representation models, such as R3M (Nair et al., 2022) and Voltron (Karamcheti et al., 2023). We loaded the pre-train weights of R3M and Voltron and fine-tuned them on CALVIN data, only training the policy head while freezing their representation parameters. As for Voltron, we also include a version that fine-tunes the representation layers. The results are shown in Tab. 7, which reveals the clear advantage of fine-tuning pre-trained VLMs compared with these specific robotics representation models.
B.3 Fine-tune the Full Model
In the fine-tuning of RoboFlamingo, we follow the training of Flamingo (Alayrac et al., 2022; Awadalla et al., 2023) that only trains the parameters of the resampler, the gated cross-attention module of each decoder layer, and the policy head while freezing all other parameters. This leads RoboFlamingo to have 1B trainable parameters (as shown in Tab. 2). In this part, we show the results of training the full model (the MPT-3B-IFT variant), which has 3B trainable parameters in Tab. 8, revealing an obvious performance deterioration.
B.4 Performance Curves in Training of Different Backbones
Fig. 8 and Fig. 9 show the performance of RoboFlamingo with different VLMs on both and settings in 5 training epochs. It is noticed that most variants converge in 5-epoch training and achieve the best performance, benefiting from the pre-training on extensive vision-language tasks.
B.5 Qualitative Examples
We visualize the task frames and analyze how RoboFlamingo achieve such a great performance. As the example shown in Fig. 7, where RoboFlamingo successfully finishes the entire task sequence, while HULC stucks at the third one. RoboFlamingo only takes a dozen steps to locate and move to the top of the drawer, and simultaneously releases the gripper to complete the task; while HULC keeps moving above the desktop for hundreds of steps and fails to locate the drawer. Furthermore, although both methods are successful for the first two tasks, RoboFlamingo uses significantly fewer steps. This representative episode vividly illustrates that our method is much more effective and efficient and could better generalize to unseen vision context.
B.6 Detailed Imitation Performances on Each Task
We present the detailed imitation performances by tasks in Tab. 9. All model are reported by their best checkpoint.
B.7 Rollout Examples
We present some rollout examples of RoboFlamingo on the split.
Appendix C Additional Details
We illustrate the details of the four policy heads/formulation mentioned in Section 5.4: (a) takes only the current observation as input to predict actions, which ignores the observation history. (b) takes the history frames into the vision encoder with position embedding, and encodes the history information through the cross-attention layers in the feature fusion decoder. (c) and (d) both utilize the VLM backbone to process single-frame observations and integrate the history with the policy head. explicitly takes the visual history as input to predict the next action. implicitly maintains a hidden state to encode memory and predict the action.