PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Liang Chen, Yichi Zhang, Shuhuai Ren, Haozhe Zhao, Zefan Cai, Yuchi Wang, Peiyi Wang, Xiangdi Meng, Tianyu Liu, Baobao Chang

Introduction

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in tackling complex tasks that necessitate a chain of integrated skills, including visual perception, world knowledge, reasoning, action, and more (OpenAI, 2023; Dai et al., 2023a; Liu et al., 2023b; Li et al., 2023c; Zhao et al., 2023).

However, current MLLM benchmarks often evaluate these capabilities individually (Fu et al., 2023; Liu et al., 2023e), overlooking the significant integrated potential that Large Language Models (LLMs) contribute to multimodal models. While some benchmarks like MMMU (Yue et al., 2023) and MathVista (Lu et al., 2023a) require abilities from both the vision and language part, they lack error localization techniques beyond accuracy assessments. This complicates identifying which part of the MLLM malfunctioned when making mistakes—whether it was the visual or the language component—and determines which aspect requires enhancement to enhance overall performance.

To address the challenges of insufficient integrated benchmarking and error localization problems, we introduce PCA-Bench. It arises with MLLM’s applications in embodied AI and decision making, where models called agents need to first process multimodal observation from different environments, reason with the current situation and goal, and finally make an action from a given action space. The abilities in the complex decision making process can be abstracted to Perception, Cognition and Action according to the Perception-Action loop (Fuster, 2004) in Cognitive Science, a fundamental concept that describes how organisms process sensory information to interact with their environment through actions, offering a comprehensive framework for assessment. Figure 1 shows how MLLMs make decisions in the PCA chain.

The instances in PCA-Bench are from three influential domains in embodied decision-making: autonomous driving, domestic robotics, and open-world gaming. As shown in Figure 2, each instance is annotated by human annotators with a 6-element tuple: <<image, question, action candidates, answer, reason, key concept>>. The last three elements serve as anchors for error localization for Action, Cognition and Perception, correspondingly.

PCA-Eval is an anchor-based evaluation protocol, designed to automatically conduct error localization utilizing the powerful semantic parsing ability of LLMs and the anchor information in data annotation. In the past, such localization was both labor-intensive and time-consuming. PCA-Eval with strong LLMs like GPT4 demonstrates a strong kappa correlation with human assessments, reaching 0.8+ average kappa coefficients for perception, cognition, and action scores. The anchor-based evaluation provides the LLMs with groundtruth answers for each sub-score, preventing the systematic bias of LLM evaluators, such as position bias (Wang et al., 2023b; Zheng et al., 2023) in the pair-wise evaluation and verbosity bias (Zheng et al., 2023) in simple preference evaluation. We also compared open state-of-the-art LLMs in PCA-Eval. Though they lag behind close ones in alignment with human assessments, we see large improvement when the model scales up. We believe that with specific training for error localization and improved general ability of open LLMs in the future, they would be more suitable evaluation tools for the reproducible and transparent characteristics.

Aiming at scaling up PCA-Bench, using LLM to synthesize training examples is an increasingly popular method for enhancing models without additional human involvement. We expand this approach to generate more samples following the PCA guideline. Unlike text-based instruction generation methods like Self-Instruct (Wang et al., 2023c), generating instructions in embodied environments poses distinct challenges. It demands not only the creation of textual instructions but also the generation of corresponding precise observations. To address these challenges, we propose Embodied Instruction Evolution (EIE), which integrates external environments with LLMs, thereby extending the LLMs’ ability to data synthesize across various embodied environments, contributing to 7,510 training data in PCA-Bench.

We conduct comprehensive experiments and analysis on PCA-Bench, our findings are summarized as follows:

1. Visual perception and reasoning with world knowledge are two core abilities for an MLLM to make correct decisions in PCA-Bench. GPT4-Vision shows strong zero-shot cross-modal reasoning ability for embodied decision-making tasks, surpassing open-source MLLMs and even Tool-Using LLM-agent.

2. EIE could generate training samples significantly enhancing the performance of open-source MLLMs (surpassing GPT-4V at some scores), validating the effectiveness of the method.

3. PCA-Eval serves as a good error locator. Above the high average kappa coefficient (0.8+) with human assessments and its ability to pinpoint the error source, it can effectively distinguishes whether a model’s correct decisions are fluky or through genuine understanding. This leads to a better ensemble metric for MLLM evaluation named Genuine PCA Score.

PCA-Bench

Multimodal decision-making problems are commonly formalized with a partially observable Markov decision process. For MLLMs F\mathcal{F} tested in PCA-Bench, we care about given the multi-modal observation o∈Oo\in O, the goal description gg, a subset of candidates actions AC⊆AA_{C}\subseteq A, whether the model could make correct action a∈ACa\in A_{C} and give proper reasoning process rr.

As shown in Figure 2, each instance in the benchmark is a 6-element tuple: <<image, question, action candidates, answer, reason, key concept>>. The image is collected from various embodied environments, including transportation scenes, housekeeper environments, and Minecraft. Questions, action candidates, and answers are derived from real tasks within the corresponding environment. The reasons explain why the answer is the best choice for the current image, while the key concept highlights the most question-related aspect of the image.

Unlike traditional visual question-answering datasets that emphasize visual perception (e.g., VQA (Goyal et al., 2017)) or visual reasoning (e.g., NLVR (Suhr et al., 2017)), PCA-Bench mandates accurate observation perception, complex task decomposition, and understanding the outcomes of various actions simultaneously. Compared to embodied simulation environments such as ALFRED (Shridhar et al., 2020) and Minedojo (Fan et al., 2022), PCA-Bench stands out for its focus on high-level actions, proving to be more effective for evaluating MLLMs. This is because high-level actions, which can be readily translated or programmed into low-level actions within their respective domains, are inherently more accessible to LLMs. The high-level actions are more comprehensible for LLMs than the direct low-level actions like action vectors in the simulation environments because (1) the high-level actions are in the form of natural languages, making it easier for LLMs to understand the meaning and connect with world knowledge. (2) LLMs are not grounded with low-level actions during the pretraining or finetuning stage, making it hard for LLMs to understand the consequences of executing an action.

To answer a question in PCA-Bench, the agent must possess the following abilities: (1) Perception: Accurately identify the concept related to the question within the image; (2) Cognition: Engage in reasoning based on image perception and worldly knowledge; (3) Action: Comprehend the potential actions, selecting the one that best aligns with the outcome of the reasoning process. A deficiency in any of these abilities would possibly result in an incorrect answer, posing a significant challenge to the more integrated capabilities of MLLMs.

2 PCA-Eval

For each instance, we prompt the model to deliver an answer comprising a reasoning process rr, and a final action aa, represented as <r,a><r,a>. By comparing the model prediction with the ground truth answer, we can obtain a fine-grained diagnosis of the decision making process as follows:

Perception Score (P-Score) measures the model’s accuracy in perceiving the observation. It is computed based on whether the agent’s reasoning process rr includes the key concept of the instance. A score of 1 is assigned if at least one question-related key concept is described by the agent; otherwise, it is 0. For the top example in Figure 2, the agent should output “clear road” or “no car visible” or other semantically equivalent concepts in its description of the image to get the perception score.

Parsing the model’s output and determining whether it entails the key concept using shallow features of the sentence is not trivial. We leverage LLM to conduct entailment detection, which turns out to have a high alignment with human judgment.

Cognition Score (C-Score) assesses the model’s ability to reason, comprehend, and make informed decisions based on the perceived input data and world knowledge. The score is 1 if the reasoning process is correct, otherwise the score is 0. For the instance in Figure 2, the agent should link the “clear road” to the action “keep driving” based on transportation commonsense to get the score.

Action Score (A-Score) measures the model’s ability to generate appropriate and effective responses or actions based on the perceived input data and the cognitive understanding of the context. The score is assigned a value of 1 if the agent selects the correct action; otherwise, the score is set to 0.

3 Automatic Evaluation

Recent advancements have seen researchers harnessing powerful LLMs for the evaluation of the output of language models. Studies have revealed that the outcomes from LLMs could exhibit remarkable alignment with human judgments Zheng et al. (2023); Wang et al. (2023b, a). In our investigation, we employed GPT-4 to automatically evaluate perception, cognition, and action scores based on the model’s outputs. Our findings underscore a significant agreement between GPT-4 scoring and human evaluation results. This is substantiated by Cohen-Kappa coefficients of 0.71, 0.82, and 0.94 for perception, cognition, and action evaluations, respectively. Experiments of human evaluation and comparison of open LLMs are in section 4.1. For a detailed description of our evaluation tool, kindly refer to Appendix D.

4 Benchmark Dataset Overview

For the test set, the examples are written by 3 human experts for each domain. There are no overlapped environmental observations between the training and test sets. The details of the human annotation pipeline can be found in Appendix B. We introduce the three domains encompassed by our dataset as follows:

In the autonomous driving domain, instances are derived from real-world transportation scenes, which requires the agent to have particular abilities such as traffic sign recognition, obstacle detection, and decision-making at intersections. The dataset aims to evaluate an agent’s ability to perceive and interpret visual information while making safe and efficient driving decisions. The images are collected from TT100K (Zhu et al., 2016) dataset and annotators are instructed to propose an image-conditioned question that is grounded with real actions of vehicles.

Domestic Robot.

The domestic assistance domain features instances from the ALFRED (Shridhar et al., 2020; Kolve et al., 2017) environment, which simulates a housekeeper robot performing tasks within a household setting. These tasks may include object manipulation, navigation, and interaction with various appliances. The environment assesses an agent’s ability to understand and execute complex instructions while navigating and interacting with a dynamic environment. Annotators are asked to select one image from the randomly generated scenes in the environment, propose a question related to the items on the scene, and annotate the full information of the instance.

Open-World Game.

In the open-world game domain, instances are sourced from the Minecraft environment, where agents are tasked with exploring, crafting, and surviving in a procedurally generated world. This dataset evaluates an agent’s ability to reason and plan actions within a complex, open-ended environment, which often requires long-term strategizing and adaptability. Annotators receive predefined tasks from MineDojo (Fan et al., 2022) as a reference during the task generation phase. For each task, we instruct the annotator to sketch a task topology graph, exemplified in Figure 3. The task should be completed under the topological order of the graph, where the event located in the leaf nodes should be finished first. Each node in the task topology graph can be viewed as a step in the sequential decision. We list the in-domain task distribution in Appendix A.

5 Embodied Instruction Evolution

The PCA-Bench benchmark also includes subset of automatic generated samples by Embodied Instruction Evolution(EIE), which is used as training set in our experiment.

The annotation of PCA-Bench examples is a labor-intensive task. As illustrated in Figure 4, we introduce Embodied Instruction Evolution (EIE), a method for automatically augmenting examples in the PCA-Bench format using Large Language Models, such as ChatGPT. This process involves four key steps:

1) Setup of Programmable Interface: Establish a programmable interface with a corresponding template, ensuring that observations in the embodied environment can be generated based on specific parameters.

2) Generation of Seed Tasks: Create initial seed tasks for each environment. These tasks are representative of the general challenges an agent might encounter. We provide ChatGPT with sample tasks and enable it to generate additional seed tasks.

3) Task Specification and Template Filling: For each seed task, we instruct ChatGPT to break down the task into multiple subtasks, following its event topology graph (as seen in Figure 3). This approach mimics the multi-step decision-making process. After determining the subtask names, we use the LLM to populate the environment parameter templates created in Step 1 for each subtask.

4) Observation Generation and Filtering: Generate observations for the environment and implement an automatic process to filter out invalid instances. The filled templates may contain errors, such as incorrect creature names or impossible items, leading to errors during environment creation. When such errors occur, the affected templates are automatically filtered out. For domains without programmable environments (autonomous driving), step 1 and step 4 are not needed, we collect real traffic images and utilize GPT4-Vision to generate seed task based on the image content.

EIE leverages the capabilities of Large Language Models to reduce manual labor and improve the diversity and scalability of PCA-Bench.

Experiments

The test set of PCA-Bench serves as an effective tool for comparing the embodied decision-making and cross-modal reasoning capabilities of various Multimodal Language Learning Models (MLLMs). In this evaluation, the same images and prompts are provided to each model under test. Additionally, to address the challenge of perceiving certain non-visual information from images, details such as “items in hand” and “items in inventory”, particularly relevant in domestic and gaming domains, are directly included in the question prompts.

In our analysis, we benchmark the performance of the most recently open-sourced models, including LLaVA1.5 and Qwen-VL-Chat, as well as the API-only GPT4-V model. All models are evaluated using their default inference configurations to ensure a fair and standardized comparison.

Finetuning with EIE.

In this track, we extend the capabilities of open-source MLLMs by fine-tuning them with the training set generated through our Embodied Instruction Evolution (EIE) method. After the fine-tuning process, these trained models are subjected to the test set of PCA-Bench. We finetune the LLaVA-7b/13b, MMICL and Qwen-VL-Chat models on the training set for 5 epochs. The training details are in Appendix E.

Zero Shot Modality Conversion.

In this track, we introduce and compare a new baseline, termed HOLMES, which utilizes LLM without multimodal perception capabilities. Instead, HOLMES relies on modality conversion APIs for embodied decision-making processes. Within the HOLMES framework, the LLM must continuously invoke various APIs, retrieving and processing return information about the environment. The HOLMES method is illustrated in Figure 7 from Appendix.

We evaluate two LLMs in this track: ChatGPT-3.5-Turbo and GPT-4-0613, comparing their performances against the advanced GPT-4-Vision. Implementation details of the HOLMES framework and the APIs are provided in Appendix 7.

2 Evaluation and Metrics

We use our PCA-Eval evaluation tool proposed in Section 2.3 to automatically assess the output of different models through three lenses: perception (P-Score), cognition (C-Score), and action (A-Score).

3 Main Results

The results of the zero-shot end-to-end track are shown in Table 1. Among all MLLMs, GPT4-V, outperforms existing open-source models by achieving the highest scores of 0.86, 0.7, and 0.68 in the perception, cognition, and action dimensions respectively. This performance represents a 15% action score improvement over its strongest open-source counterpart, LLaVA1.5-13B. The impressive performance of GPT4-V is primarily attributed to its exceptional ability to perceive visual information across different domains and the world knowledge in the language model, particularly in the challenging game domain.

Impact of Finetuning with EIE.

The results of the fine-tuning track are illustrated in Figure 5. Our EIE method has been found to significantly enhance the general decision-making abilities of various models, encompassing perception, cognition, and action. Notably, it has led to an average increase of 0.24 and 0.19 in action scores for the LLaVA1.5-7b and Qwen-VL-Chat models, respectively. Results for LLaVA1.5-13b and MMICL are illustrated in Figure 13, also showing improved performance when trained with EIE. We note that there exist reasoning or perception errors in some of the generated sample due to the hallucination problem of LLM generated content, however they do not influence the overall performance. In some cases, these sub-scores have matched or even surpassed those of the GPT4-V model, demonstrating the potential of the EIE to scale up and apply to different environments.

Comparison Between End-to-End and Modality Conversion Method

In the zero-shot modality conversion track, we conduct an analysis and comparison of the outputs generated by the End2End method with GPT4-V, as well as the HOLMES method with GPT4 and ChatGPT-3.5 in Table 2.

The results show that the HOLMES system based on GPT4 achieves 0.71 Action Score, which is on par with GPT4-V’s performance (0.74). This indicates that, overall, the HOLMES system is able to accurately understand the task goal, split the larger goal into multiple smaller steps, and correctly invoke the relevant APIs to accomplish each step. Specifically, the HOLMES system based on GPT4 can recognize the key concepts in a task, and perceive the state and environment of these concepts through the results returned by APIs. Consequently, the system achieves an average Perception Score of 0.88, which even outperforms GPT4-V’s 0.84. However, compared to End2End methods, HOLMES relies on multi-step reasoning for the final decision, in which reasoning errors tend to accumulate, and thus achieves a lower Cognition Score in both Domestic and Game domains.

On the other hand, we also find that the End2End method effectively mitigates information loss during the modality conversion process. As illustrated in Figure 8 from Appendix, an image depicts a road with several nearby cars. GPT4-V is capable of discerning that the street is not crowded, thereby suggesting that the driver can continue driving.

Conversely, GPT4-HOLMES, while being aware of the number of cars, lacks information about their spatial relation, leading it to recommend slowing down because of the existence of 14 cars. This suggests that the End2End method is superior in perceiving certain visual features that are not captured by the APIs. Conversely, some specialized APIs, such as traffic sign detection, outperform GPT4-V in tasks like traffic sign detection, as they are specifically trained for this task. This could enable the HOLMES method to gather more accurate information than the End2End model.

Discussion

As shown in Table 3, we compare the scoring kappa coefficients with human assessments for different LLMs. We randomly select 300 model outputs equally from different domains and ask 3 human experts to give perception, cognition, and action scores. The final result is based on the majority of three annotators. The result underscores a significant agreement between GPT-4 scoring and human evaluation results. This is substantiated by Cohen-Kappa coefficients of 0.71, 0.82, and 0.94 for perception, cognition, and action evaluations.

We also compare open models as evaluators. We choose one of the best open LLMs, Qwen1.5https://huggingface.co/collections/Qwen series from 7B, 14B to 72B version. Currently open LLMs tend to give wrongly high judgments in all sub-scores. Although currently trailing behind GPT-4 in performance, we anticipate that with targeted training focused on error identification and enhancements in the overall capabilities of open LLMs, these models will become more effective evaluation tools compared to closed models. This is primarily due to the reproducible and transparent nature of open models, which offer significant advantages in the development of evaluation tools.

2 Genuine PCA Score

PCA-Eval could pinpoint cases where the MLLM gets the correct answer by a fluke where perception or cognition score is 0 but the action score is 1. It explains why for some models, the action score is higher than perception and cognition scores. For instance, a model might opt for a conservative action, such as slowing down, even without accurately recognizing snowy weather in the image, resulting in a fluky correct action. In another scenario, if the model exhibits a preference for a specific choice index, it will attain a high action score provided that the evaluation dataset contains a substantial number of correct choices matching the preferred index, a phenomenon attributable to the positional biases inherent in both the model and the dataset. To overcome the mentioned bias when evaluating the genuine ability of MLLM, we propose a new metric Genuine PCA Score. It is equal to one if the perception, cogntion and action scores are all 1 for one model’s response to a question. We find that for all models, there exists significant gap (>10%) between the action score and genuine PCA score in average, revealing that relying on single metric such as choice accuracy is very problematic when conducting model evaluation. In our online leaderboard, both average action score and average genuine PCA score are considered when ranking the candidate models.

3 Alignment between Agent Decisions and Human Values

We have observed instances where the decisions made by the agent contradict human values. Consider the scenario depicted in Figure 9 from Appendix. The image illustrates a crosswalk without pedestrians. The appropriate response would be slowing down, as caution is paramount when approaching a crosswalk, regardless of the presence or absence of pedestrians. However, upon processing the information that the crosswalk is empty, ChatGPT suggests that maintaining the current speed is the optimal action, arguing that the absence of pedestrians eliminates the need to slow down. The rationale provided by ChatGPT is logical, yet it does not align with human values.

Related Work

In recent times, there have been several benchmarks built for evaluating MLLMs, such as MMBench, MME, Seed-Bench, POPE (Liu et al., 2023e; Fu et al., 2023; Li et al., 2023a, e) that assess MLLMs performance from multiple fine-grained dimensions. Visit-Bench, LVLM-eHub, M3IT (Bitton et al., 2023; Xu et al., 2023; Li et al., 2023c) focus on the general instruction following ability. General VQA tasks like OKVQA, VQAv2, Vizwiz, ScienceQA, VSR and IconQA (Marino et al., 2019; Agrawal et al., 2015; Gurari et al., 2018; Lu et al., 2022; Liu et al., 2023a; Lu et al., 2021) focus on visual understanding. MMMU, MathVista, LLaVA-benchmark and MM-Vet (Yue et al., 2023; Lu et al., 2023a; Liu et al., 2023c; Yu et al., 2023) require abilities from the vision part and specific knowledge in the language part. A lack of error localization techniques beyond accuracy assessments is among current benchmarks. This complicates identifying which part of the MLLM malfunctioned when making mistakes. Unlike prior work, PCA-Bench is more relevant to evaluate MLLMs’ ability to utilize integrated abilities to solve one task and make explainable decisions via error localization.

LLM Agent and Embodied Decision Making.

Using LLMs to empower the AI agents (Xi et al., 2023; Liu et al., 2023d; Park et al., 2023; Wang et al., 2023d) becomes more and more promising. Specifically, we can employ LLMs to enhance the decision making ability of the agents (Nakano et al., 2022; Yao et al., 2022; Li et al., 2023d; Song et al., 2023; Li et al., 2023b), expanding their perception and action space through strategies like tool utilization (Schick et al., 2023; Qin et al., 2023; Lu et al., 2023b). This line of research divides the entire decision-making process into two phases: (1) information seeking, usually involving MLLMs to verbalize the current status of AI agents in the vision-based environment with natural language; (2) reasoning and planning with text-based LLMs to decide what the AI agent should do in the next step with textual clues. Although LLM-based agents demonstrate reasoning and planning abilities through techniques like Chain of Thought or problem decomposition (Wei et al., 2023; Yao et al., 2023; Kojima et al., 2022), they inherently lack visual perception, and are limited to the discrete textual content. Therefore, integrating multimodal information can offer agents a broader context and a more precise understanding, such as PaLM-E (Driess et al., 2023), enhancing their environmental perception. However, there is still large gap deploying MLLM in various embodied environments due to the lack of appropriate benchmark and interface linking those two domains while PCA-Bench is an attempt towards that goal.

Conclusion

In this paper, we introduce PCA-Bench, a multimodal benchmark designed to assess the integrated decision-making capabilities of MLLMs. This benchmark features PCA-EVAL, a novel fine-grained automatic evaluation tool that diagnoses decision making processes from three critical perspectives: perception, cognition, and action. To enhance the decision making ability from data perspective, we propose the Embodied Instruction Evolution method to automatically synthesize instruction examples from different environments, which has been proven effective in our main experiments. We believe that powerful MLLMs pave a new and promising way toward decision making in embodied environments and we hope PCA-Bench could serve as a good benchmark in evaluation and error localization for MLLMs’ development.

Limitations

The current scope of PCA-Bench is confined to merely three domains in static environments. One of our future works aims to broaden this scope to encompass more domains and dynamic embodied environments where MLLMs could keep getting feedback, which is closer to real embodied AI scenarios. We do not apply different inference enhancement methods like In-Context Learning and Reflection in the decision making process of MLLMs. We just use the simplest prompting method and leave the exploration of a better cross-modal Chain-of-Thought method for future studies. Currently, PCA-Eval shows the best consistency with human evaluators when using powerful close LLM GPT4, which would bring additional cost to the user of PCA-Eval. We plan to develop and release an open error locator for error localization in the benchmark in the future.

References

Appendix A Examples of PCA-Bench

The PCA-Bench’s data distribution across various domains is outlined in Figure 6. For the Autonomous Driving domain, instances are grouped by their respective task types. In the Domestic Robot domain, instances are grouped by their locations. In the Open-World Game domain, instances are grouped by the tasks they aim to accomplish.

Appendix B Human Annotation Pipelines

The annotation process consists of two stages: (1) Dataset Annotation, and (2) Dataset Refinement. During the initial stage, three annotators are assigned to each domain, adhering strictly to the respective annotation guidelines. They first pinpoint the source images from each domain that are informative and meaningful so that they can write questions for each image. The annotators have the responsibility to ensure every question has only one correct answer and accurate rationales. In the subsequent stage, annotators are instructed to scrutinize the output actions and rationales presented by ChatGPT and check the annotations. This process aims to address the challenge of multiple correct answers, as ChatGPT can furnish comprehensive explanations for its actions. These explanations assist annotators in assessing the acceptability of ChatGPT’s response, particularly when it deviates from the established ground truth answer. This enables annotators to refine annotations to ensure the presence of a single correct answer.

We list three examples of each domain from PCA-Bench, as shown in Figure 10, 11, and 12.

Appendix C Zero Shot Modality Conversion: HOLMES

To optimize the evaluation process of HOLMESOriginally proposed in an early version of this paper (Chen et al., 2023) method, we pre-execute all relevant APIs for each instance within a selected subset of 300 instances from the PCA-Bench test set, recording the results for individual instances. This method enables immediate access to specific API results, eliminating the need to rerun the model for each evaluation instance.

Below is the API description for the traffic domain.

• detect_traffic_sign(): The detection of road traffic signs model utilize YOLO (Redmon and Farhadi, 2018) which trained on the Tsinghua-Tencent 100K dataset (Zhu et al., 2016). TT100K comprises 100,000 images encompassing 30,000 instances of traffic signs. The end-to-end YOLO enables simultaneous detection and classification of traffic signs.

• object_detection(): Objects demanding attention during vehicle operation primarily encompass cars, pedestrians, and bicycles. A surfeit of vehicles can lead to traffic congestion, while the presence of pedestrians or bicycles ahead necessitates cars to decelerate and proceed cautiously. Hence, the object_detection() API predominantly identifies three key object categories: cars, pedestrians, and bicycles. We utilize PMOP (Ren et al., 2023), a model trained on vision-language models through the prompt pre-training method, which enables the detection and counting of the three mentioned objectives by modifying specific class names.

• ocr(): We employ PaddleOCRhttps://github.com/PaddlePaddle/PaddleOCR/tree/release/2.7 to extract textual information from images, providing crucial road data for real-time navigation.

• image_caption(): To initially streamline the road information within the image, we employ the BLIP2-flan-t5-xl to generate an initial caption for the picture. This caption, derived from basic image data, is then utilized as input for the model to facilitate decision-making.

• weather_detection(): Weather detection leverages a pre-trained ResNet50 modelhttps://github.com/mengxianglong123/weather-recognition, derived from a dataset of more than 70,000 weather records. This model extracts weather information from provided images to inform decision-making.

Domestic Robot Domain.

Below is the API description for the Domestic Robot domain.

Game Domain.

Below is the API description for the Game domain (Minedojo).

Note that within the Domestic Robot Domain and Game Domain, APIs can be directly accessed within the virtual environment, allowing for the perception of the surrounding objects and the current picture context.

Appendix D Automatic Evaluation

We utilize the template as shown in Table 4 to query GPT-4, aiming to evaluate its responses and assign scores for perception, cognition, and action. By feeding both the agent’s output and the ground truth answer to GPT-4, based on this template, we can then extract the three distinct scores from the conclusion of GPT-4’s response.

Appendix E Training Details

Table 5 shows the specific parameters used for fine-tuning in different models. The PCA results on the three domains of PCA bench before and after fine-tuning different models are shown in Figure 13.

Appendix F Does Chain-of-Thought Finetuning Improve Cross-modal Reasoning?

Unlike vanilla finetuning, which solely focuses on delivering direct answers, Chain-of-Thought Finetuning necessitates the model to first articulate its reasoning before presenting the answer. This approach has been demonstrated to be a highly effective instruction tuning paradigm for LLMs (Chung et al., 2022; Kim et al., 2023). We have incorporated this methodology in our previous finetuning experiments.

To further evaluate its impact, we conducted an ablation study where the reasoning process was omitted from the target output during the training of MLLMs. We then assessed the variations in action scores on the test set. As depicted in Figure 14, to our surprise, the figures suggest that Chain-of-Thought finetuning exerts a relatively minor influence when compared to conventional label finetuning. We have noticed that similar phenomena has been identified by Zhang et al. (2023) that standard CoT finetuning does not work for MLLMs in their explorations.

We think there are three potential explanations: 1) Task Variation: Contrary to mathematics datasets like GSM8K, the current task doesn’t require multi-step complex reasoning to arrive at the final answer and the automatic generated CoTs have noise. 2) Modality Discrepancy: The CoT capability, inherent in LLMs, is only moderately adjusted for visual input for current open-source MLLMs. This adaptation process could potentially impair the reasoning ability. 3) Short Cut in Pretraining: We think a deeper reason might lie in the short-cut during pretraining period of current open-source MLLMs, which are pretrained on simple image caption task in a large scale. Those captions are usually short and lose a lot of information about the original image. What’s more important is that the reasoning ability of LLM is not utilized during the pretraining stage, which might hurt the reasoning ability of LLM during the SFT period. We defer to future research how to effectively harness the CoT capabilities of LLMs to enhance embodied decision-making processes.