LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Qiao, Ping Luo

Introduction

Large Language Models (LLMs), such as LLaMA , GPT-3 , and Vicuna , have demonstrated remarkable progress in Natural Language Processing (NLP). These models leverage large-scale pre-training data and huge networks to achieve impressive results in NLP benchmarks. Recently, GPT-4 further expanded the impact to the multimodal community, stimulating the rapid development of Large Vision-Language Models (LVLMs) and revolutionizing the landscape of artificial intelligence.

Large Vision-Language Models (LVLM) have achieved remarkable progress in multimodal vision-language learning for various multimodal tasks such as visual question answering and multimodal conversation. Specifically, LVLMs capitalize on the knowledge from LLMs and effectively align visual features with the textual space. Flamingo , a pioneering LVLM, integrates visual features into LLMs through cross-attention layers. Later studies proposed more efficient vision-text interactions , more efficient training methods , and employing instruction tuning .

However, despite the great success, few efforts have been made to provide systematic evaluations of LVLMs. But evaluation plays a critical role in understanding the strengths and weaknesses of LVLMs, thereby guiding their future development. Recent work presents a systematic investigation of object hallucination of LVLMs by proposing a polling-based object probing evaluation method. Moreover, ImageNetVC studies how well LVLMs can master visual commonsense knowledge. Liu et al. comprehensively evaluate the performance of LVLMs in visual recognition with text recognition, such as optical character recognition. GVT evaluates LVLM’s visual semantic understanding and fine-grained perception capabilities. Nevertheless, these studies only evaluate a portion of LVLMs on specific tasks, lacking an overall understanding of LVLM’s capabilities.

In pursuit of a comprehensive evaluation of LVLMs, we build an LVLM Evaluation hub (LVLM-eHub) consolidating 88 representative LVLMs such as InstrucBLIP and MiniGPT-4 . The details about model configuration and training data are listed in Table 1. Our LVLM-eHub consists of a quantitative capability evaluation and an online arena platform, providing a thorough investigation of the selected LVLMs. Specifically, the quantitative capability evaluation extensively evaluates 6 categories of multimodal capabilities of LVLMs including visual perception, visual knowledge acquisition, visual reasoning, visual commonsense, object hallucination, and embodied intelligence (see Fig. 1 (a)), by collecting 4747 standard text-related visual benchmarks. On the other hand, the online arena platform features anonymous randomized pairwise battles in a crowd-sourced manner, providing a user-level model ranking in the open-world question-answering scenario (see Fig. 1 (b & c)).

Our LVLM-eHub comprehensively evaluates LVLMs, revealing several innovative findings. (1) Instruction-tuned LVLM with massive in-domain data suffers from overfitting problem and generalizes poorly in open-world scenarios , such as InstructBLIP (see Fig. 1 (a)). (2) With moderate instruction-following data, instruction-tuned LVLM may cause object hallucination issues, generating objects that are inconsistent with target images in the descriptions. This leads to incorrect answers or renders current evaluation metrics, such as CIDEr for image captioning, ineffective. (3) We find that a multi-turn reasoning evaluation pipeline can mitigate the issue of object hallucination, indicating that developing an effective pipeline for LVLM evaluation is urgent.

The contributions of our work are summarized follows. (1) We propose LVLM-eHub which is the first comprehensive evaluation benchmark for large vision-language models, to our best knowledge. (2) LVLM-eHub provides extensive evaluation on 6 categories of multimodal capabilities of LVLMs in 4747 text-based visual tasks. (3) LVLM-eHub builds an online arena platform for LVLMs, which features anonymous randomized pairwise user-level comparison in a open-world scenario. (4) Our evaluation results reveal several innovative findings, providing a foundational framework for the assessment of innovative strategies aimed at enhancing zero-shot multimodal techniques.

LVLM Evaluation Hub

In this section, we introduce representative LVLMs, multimodal capabilities of interest, and evaluation methods. The whole LVLM Evaluation Hub is illustrated in Fig. 2. Our LVLM evaluation hub compromises 88 representative models including BLIP2 , LLaVa , LLaMA-Adapter V2 , MiniGPT-4 , mPLUG-Owl , Otter , InstructBLIP , and VPGTrans . All models boost vision-language representation learning by utilizing pre-trained image encoders and large language models (LLM). But they differ in training data scale and model configuration as shown in Table 1, where the information is collected from their papers or provided by the authors. For a fair comparison between LVLMs, we collect their checkpoints with parameter sizes less than 1010B. The detailed descriptions of these models are provided in Appendix Sec. A.

We aim to evaluate LVLMs’ capability comprehensively. In particular, we summarize 66 categories of capabilities and collect corresponding benchmarks for quantitative evaluation (see Fig. 2). Please see our Appendix Sec. D for more statistics and details of the collected benchmarks.

Visual Perception. Visual perception is the ability to recognize the scene or objects in images, the preliminary ability of the human visual system. We evaluate this capability of models through image classification (ImgCLs) using the ImageNet1K , CIFAR10 , Pets37 and Flowers102 benchmarks, multi-class identification (MCI) and object counting (OC) using the GVT benchmark. ImgCLs and MCI measure how well an LVLM grasps high-level semantic information, while OC assesses the recognition ability for fine-grained objects.

Visual Knowledge Acquisition. Visual knowledge acquisition entails understanding images beyond perception to acquire knowledge. This evaluation is conducted through Optical Characters Recognition (OCR) using twelve benchmarks (including IIIT5K , IC13 , IC15 , Total-Text , CUTE80 , SVT , SVTP , COCO-Text , WordArt , CTW , HOST , WOST ), Key Information Extraction (KIE) using the SROIE and FUNSD , and Image Captioning (ImgCap) using two benchmarks (including NoCaps and Flickr30K ). The OCR task measures whether a model can accurately identify and extract text from images or scanned documents. The KIE task further poses challenges in extracting structured information from unstructured or semi-structured text. Finally, ImgCap assesses whether a model can generate a good natural language description of the content of an image.

Visual Reasoning. Visual reasoning requires a comprehensive understanding of images and related texts. To evaluate the visual reasoning ability of LVLMs, we utilize three tasks including visual question answering (VQA) (including DocVQA, TextVQA, STVQA, OCR-VQA, OKVQA, GQA, IconQA, Visual Spatial Reasoning (VSR), and Visual Dialog (Visdial).), knowledge-grounded image description (KGID), and visual entailment. For KGID, we use ScienceQA and VizWiz ). For visual entailment task, we use SNLI-VE. These three tasks are in VQA form. A capable LVLM should be able to understand the objects and scenes in an image and can reason to generate answers that are semantically meaningful to the question asked.

Visual Commonsense. Visual commonsense refers to the general visual knowledge commonly shared across the world, as opposed to the visual information specific to a single image. This evaluation tests the model’s understanding of commonly shared human knowledge about generic visual concepts using ImageNetVC and visual commonsense reasoning (VCR) . Specifically, ImageNetVC is utilized for zero-shot visual commonsense evaluation, such as color and shape, while VCR covers various scenes, such as spatial, casual, and mental commonsense.

Object Hallucination. LVLM suffers from the object hallucination problem, i.e., the generated results are inconsistent with the target images in the descriptions . Evaluating object hallucination for different LVLMs help understand their respective weaknesses. To this end, we evaluate the object hallucination problem of LVLMs on the MSCOCO dataset under POPE pipeline .

Embodied Intelligence. Embodied intelligence aims to create agents, such as robots, which learn to solve challenging tasks requiring environmental interaction. Recently, LLM and LVLM exhibited exceptional effectiveness in guiding the agent to complete a series of tasks. In this evaluation, we utilize high-level tasks in EmbodiedGPT and employ Minecraft , VirtualHome , Meta-World , and Franks Kitchen as benchmarks.

2 Online Evaluation with LVLM Arena

Designing quantitative evaluations for LVLM to satisfy all capabilities is challenging, as evaluating LVLM responses constitutes an open-ended problem. Inspired by FastChat , we introduce the LVLM Arena, an online evaluation framework for LVLMs’ pairwise battle with human judgment.

Figure 2 illustrates the LVLM Arena, comprising three primary components: matchmaking, chat, and voting. Initially, two models are sampled from the model zoo. Users then converse side-by-side with the models, who remain anonymous. Subsequently, users vote for the superior model.

Matchmaking. The matchmaking module samples two models in a tournament style based on their Elo rating. However, due to the currently limited size of the model hub, we employ random sampling.

Chat. Users chat side-by-side with two sampled models (which remain anonymous) using images or text inputs. Different from quantitative evaluation, users can chat about anything. Our existing online platform supports only single-round chats due to high computational and memory demands in multi-round chats. Future updates will address this constraint.

Voting. After the chat session, users vote for their preferred model. Four options are available: Model A, Model B, Tie, and Both are bad. The Elo rating is subsequently updated using voting results.

In contrast to limited quantitative evaluations, the LVLM Arena provides an open-world evaluation framework that enables users to chat with models about anything, emulating real-world conditions. Besides, users serve as the judge for the battle, which brings more convincing evaluation results than traditional evaluation metrics.

3 Zero-shot Evaluation

LVLMs are capable of capturing a wide range of multimodal patterns and relationships. We evaluate the aforementioned 66 categories of capabilities of LVLMs in Sec. 2.1 by investigating their zero-shot performance on various tasks. Zero-shot evaluation allows us to evaluate the LVLMs’ ability to generalize to new tasks without training the model, which is competent for large-scale evaluation. To be specific, we treat the zero-shot evaluation as various forms of prompt engineering for different tasks (see Fig. 3) as presented in the following.

Question Answering. Prompting with visual question answering can be used to solve many downstream tasks, which assess how well an LVLM understands the underlying language and visual features. We design proper prompts to ensure that the LLM can produce meaningful results. For example, text prompts of OCR can be "what is written in the image?". Then, we evaluate the answers generated by the LLM using the corresponding metric such as accuracy.

Prefix-based Score. For multi-choice QA tasks, we can utilize a visual encoder to obtain visual prompts for a given image. Then, the visual prompts are prefixed into the text embeddings, which are fed into the LLM. The likelihood of image-text pair can be generated, which is referred to as a prefix-based score. We can obtain a prefix-based score for each text prompt of the candidate’s answer. The answer with the largest prefix-based score is selected as the final answer.

Multi-turn Reasoning. Following IdealGPT , we use a multi-turn reasoning framework to evaluate complex visual reasoning tasks. Specifically, we utilize an LLM such as ChatGPT to generate sub-questions for a given question, an LVLM to provide corresponding sub-answers, and another LLM to evaluate sub-answers’ quality. Such a pipeline iteratively proceeds until a satisfactory answer is obtained.

User Study. Evaluating the quality of the text generated by an LVLM requires a thorough understanding of the underlying language and context. In embodied artificial intelligence tasks, the LVLM generates a plan for the given instruction, which should be evaluated through various aspects such as recognition accuracy and conciseness in answers. It is hard to implement such an evaluation using an existing metric. Thus, user studies are conducted to assess the quality, relevance, and usefulness of the text generated by the LVLM in a specific context. To maintain evaluation fairness, we randomly shuffle the model’s output order and anonymize outputs during evaluation.

Experiment and Analysis

In this section, we perform a zero-shot evaluation to assess the 66 categories of capabilities of LVLMs. Specifically, visual perception ability, visual knowledge acquisition, visual reasoning, visual commonsense understanding, visual object hallucination, and embodied intelligence are assessed in Sec. 3.1 ∼\sim Sec.3.6, respectively. The LVLM arena evaluation result is presented in Sec.3.7. More evaluation findings, evaluation details, quantitative results, and details about evaluation datasets can be found in Appendix Sec. A, Sec. B, Sec. C and Sec. D, respectively.

Visual perception is an important ability of LVLMs. As presented in Sec. 2.1, we evaluate through image classification (ImgCls), multi-class identification (MCI), and object counting (OC). The evaluation details of tasks are provided in Appendix. B.1. The evaluation results are reported in Table 2. We have three observations. (1) mPLUG-Owl and LLaVA perform best on coarse-fined classification tasks (i.e., ImageNet1K and CIFAR10). The commonality is that they update LLM with 158K instruction-following data. (2) InstructBLIP presents good perception ability in fine-grained ImgCls, OC, and MCI tasks. The main reason is that InstructBLIP is fine-tuned on 1.6M VQA data, making it overfit on these tasks. (3) The performances of LVLMs on ImgCls are significantly inferior to supervised SOTA, indicating plenty of room for LVLM’s perception ability.

2 Results on Visual Knowledge Acquisition

Visual knowledge acquisition involves going beyond image perception to acquire deeper understanding and knowledge. In our study, we evaluate the acquisition of visual knowledge through various tasks, namely Optical Character Recognition (OCR), Key Information Extraction (KIE), and Image Captioning, all performed in a Visual Question Answering (VQA) fashion. The evaluation details of tasks are demonstrated in Appendix.B.2. Table 3 shows the zero-shot performance in visual knowledge acquisition, and we have the following observations. First, BLIP2, InstructBLIP, and VPGTrans achieve dominant performance in all tasks. This may be because these models use a large visual encoder (i.e., ViT-g/14) and Q-Former updated with massive image-text pairs. A stronger visual encoder and adaption module can extract better tokens entailed with the global and local context, leading to remarkable improvement in visual knowledge acquisition. Second, InstructBLIP presents consistently the best results on almost all tasks. The main reason is that InstructBLIP overfits these tasks by fine-tuning with massive VQA data.

3 Results on Visual Reasoning

Visual reasoning encompasses the ability to comprehensively understand images and perform cognitive tasks. In this section, we evaluate the visual reasoning ability of LVLMs on various tasks, including Visual Question Answering (VQA), Knowledge-Grounded Image Description (KGID), and Visual Entailment (VE) tasks. The evaluation details of tasks are provided in Appendix.B.3. Table 4 shows the zero-shot performance in visual reasoning, and we have the following observations. First, compared with BLIP2, InstructBLIP again presents better results overall because it overfits many tasks by fine-tuning massive VQA data. Second, compared with BLIP2, instruction-tuned LVLMs, except for InstructBLIP, generally perform worse than BLIP2. The common words in the instruction data often influence the generated content, which can not be evaluated by the current metrics (see Fig. 4). Third, instruction-tuned LVLMs consistently surpass BLIP2 on SNLI-VE where the final answer is obtained by multi-turn reasoning. It shows that instruction-following fine-tuning can produce promising content once a good evaluation scheme is employed. We provide more evidence in Fig. A.3 in Sec. C of Appendix.

4 Results on Visual Commonsense

The visual commonsense evaluation aims to evaluate the model’s comprehension of commonly shared human knowledge about generic visual concepts. We use two challenging visual commonsense benchmarks in a zero-shot setting, including ImageNetVC and Visual Commonsense Reasoning (VCR). The evaluation details of tasks are presented in Appendix.B.4.

As shown in Table 5, we can find that all those LVLMs can partly solve visual commonsense problems. First, InstructBLIP performs best among those LVLMs on the ImageNetVC and VCR dataset. The main reason is that it is fine-tuned on 1.6M fine-grained VQA data, making it adapt to answer visual common questions. Second, LLaVA also performs well in the visual commonsense task. The reason is that it updates LLM with instruction-following data. Third, instruction-tuned LVLMs again surpass BLIP2 on the two Visual Commonsense tasks. It shows that instruction-tuning can provide more effective clues than BLIP2 for visual commonsense. Note that the final answer of VCR is obtained by multi-turn reasoning. It also shows the significant role of a good evaluation scheme in producing promising content for instruction-tuned models.

5 Results on Object Hallucination

Although LVLMs have made significant progress, they still struggle with the issue of hallucination, which refers to their tendency to produce objects that do not align with the descriptions provided in the target images. In this section, we focus on evaluating such object hallucination problems on MSCOCO captioning dataset. Following POPE evaluation pipeline which is a multi-step QA procedure, we prompt LVLMs with multiple Yes-or-No questions. Each image is prompted with 66 Yes-or-NO questions. For example, ‘Is there a person in the image?’. We use accuracy, precision, recall, F1-Score and the ratio of answering ‘Yes’ as evaluation metrics. The evaluations are three datasets including MSCOCO-Random/Popular/Adversarial . From Random to Adversarial, the questions become more and more difficult. From Table 6, we could come to the following conclusions. InstructBlip performs best in the hallucination problem, followed by BLIP2, whose average accuracy reacheS more than 80%. We find that instruction-tuned models, except for InstructBLIP, perform worse than BLIP2 because they tend to answer ‘Yes’ to the question, which shows that LVLMs are prone to generate objects which do not exist in the image. In sec. C.2, we show that such an object hallucination problem can be alleviated by a multi-turn reasoning pipeline, which can be also seen from the experiments on SNLI-VE and VCR.

6 Results on Embodied Intelligence

In this section, we present the evaluation results focusing on embodied intelligence. To appraise the effectiveness of planning outputs using the given image, we conducted a user study involving 15 participants. The study comprised 6 household scenarios carefully selected from VirtualHome . Specifically, the participants rated the generated plans from different LVLM models using a scoring system similar to . The evaluation comprised five dimensions with scores ranging from 1 to 5. These dimensions included object recognition accuracy, spatial relationship understanding, level of conciseness in the response, reasonability of the planning, and executability of the planning. The resulting average scores for the different models among the participants are presented in Table 7 below. Furthermore, in the Appendix C, we present quantitative evaluation results for Franka Kitchen , Minecraft , and Meta-World . Based on the findings, two deductions can be made. Firstly, the use of image-text pairs is consequential in aligning visual-text features. This is evident from the comparison demonstrated in Table 1. Otter and InstructBLIP lacked the training process on image-text pairs which necessitates the alignment of visual reasoning and text description, resulting in a degraded ability of spatial relationship analysis. Conversely, mPLUG-Owl outperformed Otter partially due to 204M image-text training pairs. Secondly, we observe that visual instruction data is essential for embodied tasks. BLIP2 lacked visual instruction tuning, which greatly affected its capability of producing reasonable and executable plans.

7 Results on Online Arena Evaluation

The arena features anonymous and randomized pairwise battles in a crowd-sourced manner. We have collected 634634 and 14251425 pieces of evaluation data up until June 3 and June 13 in 2023, respectively. The collected data shows almost the same number of battle outcomes for ‘Model A wins’ and ‘Model B wins.’ Moreover, 21.821.8% and 22.022.0% battle outcomes are voted as ‘both are bad’ in two copies of collected data, respectively, implying that the current LVLMs still struggle to generate good answers for open-world visual questions. Furthermore, we rank the selected 88 LVLMs with Elo rating using two copies of the collected data by following Fastchat and . As shown in Fig. 1 (b) and (c), mPLUG-Owl, MiniGPT-4, LLaMA-Adapter V2, Otter, and VPGTrans, which are fine-tuned with amounts of instruction-following data with updating many parameters, are the top-3 best models in the open-world VQA scenario according to two ranking lists, indicating the significance of instruction-following tuning and effective parameter update. Moreover, InstructBLIP perform best on in-domain capability evaluation, while being much worse than many instruction-tuned models, implying severe overfitting issue, as shown in Fig. 1.

Discussion and Conlcusion

New Evaluation Metrics. Our quantitative evaluation mainly uses the CIDEr score and accuracy. The CIDEr score is widely used in image captioning and QA evaluation. It measures the similarities between generated and ground-truth answers. However, LVLMs’ responses are diverse, in different styles with the ground truth. As such, the CIDEr score is unsuitable (see Fig. 4 for failure cases). We also tried model-based evaluation, which uses Sentence Transformer to calculate the feature similarity between generated and the ground-truth answers. It is generally more robust but sometimes suffers due to model limitations. Recent studies use the powerful Chat-GPT or GPT-4 as a judge to evaluate LLMs’ responses. However, in LVLM evaluation, GPT is blind to the image and is inaccurate in some cases. We propose LVLM Arena, an innovative evaluation framework that utilizes a 1v1 LVLM battle with human judgment for open-world evaluation, leading to more accurate and realistic evaluation results. However, it requires significant human effort to produce reliable rating results, especially when numerous models exist. Therefore, developing fast, accurate, and generalized evaluation metrics for LVLMs remains an open problem.

A Platform for LVLM Evaluation. We have developed an evaluation framework aimed at comprehensively assessing the performance of LVLM models across six critical capabilities. Each capability encompasses multiple tasks, with several datasets incorporated into each task. Our user-friendly interface allows users to contribute their own datasets and models, facilitating a collaborative and inclusive environment. With just one click, users can effortlessly access a holistic assessment of their target LVLM model through our evaluation platform. We are dedicated to regularly updating the datasets and expanding our support for a wider range of LVLM models on our platform. Users are encouraged to contribute their LVLM models by utilizing our platform’s model inference interface. Additionally, we offer free online inference services for the LVLM models supported by LVLM Arena. This arena not only allows users to vote for their preferred models but also provides an Elo rating rank system that incorporates valuable human feedback, ensuring continuous improvement and refinement.

Conclusion. This paper proposes a comprehensive evaluation benchmark for large vision-language models called LVLM-eHub that incorporates both quantitative performance evaluation and human feedback evaluation. For the quantitative evaluation, we employ 16 tasks spanning over 40+ text-related visual datasets to assess the six essential capabilities of LVLM models. Additionally, we have established an online LVLM Arena to gather human feedback on LVLM models continually. This arena serves as an invaluable resource, providing an Elo rating rank that offers LVLMs ranking in the open-world scenario. Our evaluation results reveal several important findings, stimulating the future development of LVLMs.

References

Appendix A More details about our LVLM-eHub

Our Findings. We present our observations from extensive evaluation experiments in the following.

Instruction-tuned LVLM with massive in-domain data such as InstructBLIP heavily overfits many existing tasks, generalizing poorly in the open-world scenario. As shown in Fig. 1 and Fig. A.1, InstructBLIP achieves the best results in 5 categories of capabilities while lagging behind other instruction-tuned models such as LLaMA-Adapter V2 and mPLUG-Owl in embodied AI and LVLM arena platform. We see that InstructBLIP is fine-tuned on 16M visual question answering pairs (see Table 1, exhibiting in many in-domain tasks such as perception and reasoning tasks. However, Embodied AI tasks require that the model is capable of generating a step-by-step plan for instruction with a given image. Moreover, the arena platform evaluates LVLMs’ ability of visual question answering in open-world scenarios. InstructBLIP overfits in-domain tasks, generalizing poorly in these two real-world tasks.

Instruction-tuned LVLM with moderate high-quality instruction-following data may result in object hallucination issues. The issue means that LVLMs would generate objects that are inconsistent with target images in the descriptions. It either makes the current evaluation metric such as CIDER for image captioning ineffective or generates wrong answers. For instance, LLaMA-Adapter V2 can generate high-quality image captions which yet present a low CIDEr score as shown in Fig. 4. But the high sentence similarities between the generated answer and ground-truth answers measured by Sentence Transformer and GPT3.5 shows that the generated answer is relatively accurate. Therefore, the instruction-tuned models could generate content that cannot be evaluated by existing metrics. It also indicates that it is urgent to develop an effective metric for LVLM evaluation.

In addition, we also find that instruction-tuned LVLMs with moderate high-quality data are more likely to generate wrong answers. As shown in Table 6, LLaMA-Adapter V2, LLaVA, MiniGPT-4, mPLUG-Owl, Otter, and VPGTrans generally present higher accuracy and recall, and lower precision than BLIP2 and InstructBLIP. These models are tuned with moderate high-quality data such as LLaVA-158K or instruction-following data generated by LLM as shown in Table 1. This implies that instruction-tuned LVLMs with moderate high-quality data are prone to answer ‘Yes’ regardless of the accuracy of the answer to the underlying question.

Employing a multi-turn reasoning evaluation framework can mitigate the issue of object hallucination, shedding light on developing an effective metric for LVLM evaluation. In Table 4 and Table 5, we see that instruction-tuned LVLMs with moderate high-quality data can achieve better performance than BLIP on SNLI-VE and VCR tasks under a multi-turn reasoning evaluation pipeline in Sec. 2.3. We also provide more evidence to demonstrate the effectiveness of such an evaluation technique in mitigating object hallucination in Fig. A.3.

A.2 Model Details in LVLM-eHub

BLIP2 pre-trains a lightweight Q-Former on 129M image-text pairs. It follows a two-stage strategy to bridge the modality gap. The first stage bootstraps vision-language representation learning from a frozen image encoder ViT-g/14 in EVA-CLIP . The second stage bootstraps vision-to-language generative learning from a frozen LLM FlanT5-XL , which enables zero-shot instructed image-to-text generation.

LLaVA connects the visual encoder ViT-L/14 of CLIP with the language decoder LLaMA by a lightweight fully-connected (FC) layer. LLaVA first trains the FC layer with 595K image-text pairs while freezing the visual encoder and LLM and then fine-tunes the FC layer and LLM on 158K instructional vision-language data.

LLaMA-Adapter V2 (LA-V2) is a parameter-efficient visual instruction model. Although the visual encoder (ViT-L/14) and LLM are kept frozen, LLaMA-Adapter V2 distributes the instruction-following ability of the whole LLaMA through bias(B)-tuning. In this way, the scale, bias, norm, and prompt parameters are tuned on 200M image captioning data, 158K visual instruction-following data, and 52K language instruction-following data constructed by GPT-4 .

MiniGPT-4 connects the visual encoder and text encoder by an FC layer. It also first trains the FC layer with 5M image-text pairs and then fine-tunes it on 3.5K instructional vision-language data. Despite the simplicity, MiniGPT-4 needs to load a pretrained vision encoder of BLIP2 and Vicuna LLM .

mPLUG-Owl incorporates a visual abstractor, essentially the same as Perceiver Resampler in Flamingo , to bridge pretrained visual encoder ViT-L/14 and LLM (LLaMA) with a two-stage finetuning procedure. It firstly fully finetunes both the visual encoder and visual abstractor on 204M image-text pairs. Then for the second stage, 158K LLaVA-Instruct data is utilized to parameter-efficiently finetune pretrained LLM via LoRA.

Otter is a multimodal model with in-context instruction tuning based on OpenFlamingo which comprises a LLaMA-7B language encoder and a CLIP ViT-L/14. Although the visual and text encoder are frozen, Otter trains extra 1.3B parameters coming from adaption modules on 158K instruction-following data.

InstructBLIP is initialized from a pre-trained BLIP-2 model consisting of a ViT-g/14 image encoder, a Vicuna LLM and a Q-Former to bridge those two. During vision-language instruction tuning, only Q-Former is fine-tuned on 13 visual question-answering datasets.

VPGTrans is a simple transferring technique that adapts a smaller LLM to a larger LLM. It transfers the VPG of BLIP-2 (i.e. ViT-g/14) from OPT6.7B to Vicuna7B by training Q-Former on 13.8M Image-Text pairs. In addition, the VPG and projector are further tuned on MiniGPT-4’s 3.5K self-instruct data instances.

Appendix B Evaluation Details

For ImgCls, we test LVLMs on two coarse-grained benchmarks (i.e., ImageNet1K and CIFAR10) in top-1 accuracy and two fine-grained benchmarks (i.e., Pets37 and Flowers102) in per-class accuracy. Following KOSMOS-1 , the default prompt ‘The photo of the’ is used for all LVLMs for a fair comparison over coarse-grained benchmarks, while it is too general for fine-grained visual perception. When confronted with intricate fine-grained image classification tasks in a zero-shot manner, contemporary multi-modal models often encounter difficulties in accurately generating precise subclass names. To gain a deeper understanding of their capabilities and enable effective comparisons between them, we have heuristically designed prompts “What is the specific category of the dog or cat in the image?” for the Pets37 dataset, and “What is the specific category of the flower in the image?” for the Flowers102 dataset, respectively. Furthermore, the generated coherent sentence-style responses deviate from the standard image classification benchmark. To accommodate this discrepancy, we considered the prediction as correct if the model output contains the correct class name, which is inspired by MultiModal OCR.

For OC task, we test LVLMs on MSCOCO and VCR1.0 . It involves querying the model about the number of objects belonging to an image’s specific class of interest. To this end, we use the prompt ‘Question: How many [obj] are there in the image? Answer:’. The generated answer is then compared with the ground truth. We report accuracy by treating OC as a classification problem.

For MCI task, we also test LVLMs on MSCOCO and VCR1.0 . We ask the model to determine whether a certain object is present or absent by prompting ‘Question: Does [obj] exist in the image. Answer:’. We also report the accuracy by treating MC as a Yes or No classification problem.

B.2 Details of Visual Knowledge Acquisition

For OCR task, we test the selected LVLMs with twelve representative OCR datasets, which are inclusive of IIIT5K, ICDAR 2013(IC13), ICDAR 2015 (IC15), Total-Text, CUTE80, Street View Text (SVT), SVTP-Perspective (SVTP), COCO-Text, WordArt, SCUT-CTW1500 (CTW), heavily occluded scene text (HOST), weakly occluded scene text (WOST). These benchmarks consist of a diverse range of images containing textual information which can make an adequate comparison between LVLMs. The performance of the LVLMs is compared with top-1 accuracy and the prompt we use is ‘what is written in the image?’.

For KIE task, we employ the SROIE and FUNSD benchmarks to evaluate LVLMs, which encompass diverse documents like receipts and forms that require specific information extraction. The performance of LVLMs is evaluated using entity-level F1 scores. Additionally, we utilize information-specific prompts for each piece of information that the model should extract. For instance, in the SROIE benchmark case, we use the prompt ‘what is the name of the company that issued this invoice?’ to extract company information and ‘where was this invoice issued?’ prompt for address information.

For ImgCap task, we utilize two benchmarks, including NoCaps and Flickr30K. Each benchmark provides a collection of images with corresponding captions. In evaluation, CIDEr scores are used to evaluate these models with the prompt ‘what is described in the image?’.

B.3 Details of Visual Reasoning

For VQA task, we utilize nine benchmarks: DocVQA, TextVQA, STVQA, OCR-VQA, OKVQA, GQA, IconQA, Visual Spatial Reasoning (VSR), and Visual Dialog (Visdial). These benchmarks offer a diverse set of question-image pairs, covering a wide range of topics. The task requires LVLMs to not only understand the visual content but also comprehend and reason about the posed questions. For specific evaluation, we employ the Mean Reciprocal Rank (MRR) metric for Visdial and top-1 accuracy for the remaining datasets. These metrics provide insights into the model’s ability to accurately answer questions across the various VQA benchmarks.

For KGID task, it evaluates the LVLM’s capability to generate informative and accurate descriptions of images by incorporating external knowledge. To assess performance, we employ the ScienceQA and VizWiz benchmarks, which consist of images accompanied by textual descriptions and knowledge-based information. Notably, in the case of ScienceQA, we only utilize the samples that contain images.

For VE task, it evaluates the VLPM’s capability to determine the logical relationship between image pairs. We employ the SNLI-VE benchmark, which provides pairs of images along with corresponding textual premises and hypotheses. For efficient evaluation, we randomly select 500 samples from the dev split of the SNLI-VE dataset. We find that a naive QA pipeline is hard to give meaningful predictions. Therefore, as shown in Fig. 3, we employ multi-turn reasoning to solve SNLI-VE. There are three components in multi-turn reasoning pipeline: a Questioner, an Answerer, and a Reasoner. The Questioner first raises a set of sub-questions based on the main question, then Answerer produces the relative sub-answers, and Reasoner decide whether a confident answer is derived to its main question by analyzing the sub-questions and sub-answers. The "Questioner-Answerer-Reasoner" loop keeps iterating until the Reasoner derives a confident final answer or the number of iterations reaches a predefined maximum. Following You et al , three simple prompts are also applied to generate better sub-answers and sub-questions. We use ChatGPT as the Reasoner and Questioner via "gpt-3.5-turbo" API. The eight studied pre-trained LVLMs are served as the Answerers and produce image captions respectively for comparing their abilities to solve visual-entailment problem.

B.4 Details of Visual Commonsense

For ImageNetVC, we evaluate the zero-shot visual commonsense of LVLMs. It contains high-quality QA pairs for various commonsense, including color, shape, mater, comp, and others. Specifically, as shown in Fig. 3, QA pairs in ImageNetVC are first transformed into prompts like ‘[Question] The answer is [Answer].’, and then each prompt is converted into a sequence of tokens. Secondly, the image and text tokens are transformed into a sequence of visual embeddings and a sequence of text embeddings, respectively. Finally, visual embeddings are prefixed into the text embeddings yielding the final embeddings which are put into a frozen pre-trained LVLM to calculate the score. The probability distribution over all answer candidates using softmax is calculated. We use the prefix-based score to choose the final answer with maximum likelihood. Following Xia et al, five similar prompts are utilized to take average values for final evaluation among eight LVLMs.

For VCR, it expects that the LVLMs can find the correct answer among four answer candidates. For efficient evaluation, we randomly select 500 samples from the val split of the VCR dataset. We find that a naive QA prompt cannot give meaningful output. Similar to the SNLI-VE evaluation (Section B.3), we adopt a multi-turn reasoning evaluation technique to solve the VCR task.

Appendix C More Experiments

Throughout our comprehensive evaluation, we discovered that LVLM models are highly sensitive to the choice of prompts. An illustrative example of this sensitivity is observed in the image captioning task, where altering the prompt employed for the VPGTrans model on the NoCaps dataset leads to a substantial variation in performance, ranging from 19.66 to 48.13. Fig. A.2 showcases some examples of the results generated based on different prompts. This sensitivity underscores the significance of carefully selecting and designing prompts to achieve optimal performance in LVLM-based tasks.

C.2 More Results on Object Hallucination

Multi-turn Reasoning Evaluation Can Alleviate Object Hallucination. We also show the hallucination problem by visualizing some examples in Fig. A.3. It can be seen that LVLMs except for BLIP2 and InstructBLIP are prone to generate objects which are inconsistent with the image. Fortunately, such an issue can be mitigated by a multi-turn reasoning evaluation framework. We believe that the reasoning procedure can encourage instruction-tuned models to re-organize the knowledge they grasp and finally generate the right answers. It is significant to investigate how to evaluate instruction-tuned LVLMs in the right way.

C.3 More Results on Embodied Tasks.

In this section, we provide quantitative evaluation results for embodied tasks in addition to the user study discussed in Section 3.6. We selected some representative scenes from Minecraft, Franka Kitchen, and Meta-World benchmarks as shown in Figures A.4 through A.6, and the results for these tasks are provided in Sections C.3.1 to C.3.3.

In Figure A.4, the models were asked to generate feasible plans for the Minecraft agent to reach the opposite shore with a boat floating on the river. All eight models recognized the presence of the floating boat, but only LLaMA-Adapter V2, InstructBLIP, and MiniGPT-4 generated a plan that utilized the boat to help the agent reach the opposite shore more quickly.

In Figure A.5, the models needed to assist the robotic arm in moving the kettle to the top left burner, and we expected the models to analyze where the goal state was achieved from the image. Except for BLIP, all seven models provided a feasible and reasonable plan. Notably, LLaVA recognized that the goal state had already been achieved from the given image. Meanwhile, mPLUG-Owl generated some steps to deal with the situation when the goal state was achieved before execution.

For the Meta-World examples shown in Figure A.6, some models struggled with visual reasoning partially due to the low-resolution image input. mPLUG-Owl detected the basketball in the robot’s hand mistakenly, while others, e.g. LLaVA, MiniGPT-4, and Otter, mistook the robotic arm for a human.

BLIP: This is a Minecraft screenshot of a boat in the water. InstructBLIP: The image depicts a Minecraft scene where the player is standing on a small island in the middle of a large body of water. The player’s goal is to reach the opposite shore as quickly as possible, which involves navigating through the water and possibly encountering various obstacles along the way. In the scene, there are several objects visible, including a boat, a tree, and a rock. The boat appears to be partially submerged in the water, suggesting that the player may need to navigate around it or use it as a means of transportation to reach the opposite shore. The tree serves as a landmark, providing a reference point for the player to orient themselves and determine their progress towards the opposite shore. The rock is also present, adding to the visual interest of the scene. The spatial relationship between the player, the boat, the tree, and the rock can be inferred from the image. The player is located on the small island in the middle of the water, while the boat is partially submerged nearby. The tree serves as a reference point, indicating the player’s proximity to the opposite shore. The rock adds to the visual appeal of the scene, contributing to the overall sense of. LLaMA-Adapter V2: The image shows a Minecraft scene with a small wooden boat floating on a body of water. The boat is positioned near the shore, and there is a path leading to the opposite shore. The boat appears to be in a relatively calm body of water, and it seems to be the primary mode of transportation for the player. The player’s goal is to reach the opposite shore as fast as possible, so they will likely need to use the boat to navigate the water and reach their destination. LLaVA: The image shows a Minecraft scene with a small boat floating in a body of water, surrounded by a sandy beach. There is also a small island in the background. The boat is positioned near the shore, and the player is standing on the beach, looking at the boat. The player’s objective is to reach the opposite shore as fast as possible. The scene appears to be a part of a game, and the player’s actions are likely to be controlled by the game’s mechanics. MiniGPT-4: The image shows a small wooden boat floating in the water. The boat has a small cabin at the front and a mast with a sail at the rear. There is a rocky shore on one side of the boat and a sandy beach on the other. The water is calm and there are no other objects or people in the scene. The sun is shining from the top right corner of the image, casting a bright light on the water and the boat. The boat is positioned in the middle of the water, with the mast facing towards the shore. The goal is to reach the opposite shore as fast as possible, so the player would need to navigate the boat through the water, possibly using a paddle or oars, and avoid any obstacles or hazards in the way. mPLUG-Owl: In the image, a person, likely the user, is standing in the shallows near the water, looking for a way to cross over to a boat that is on the other side of the water. They appear to be in a desert biome, as evidenced by the surrounding landscape. The boat is located at a distance from the shore, and it seems to be partially underwater, suggesting that it could be a partially submerged boat. In addition to the person standing near the water, there are several rocks and trees scattered across the scene. One of the rocks is located closer to the shore, while the others are scattered farther away from it. The trees are also dispersed throughout, some closer to the shore and others further away. The presence of these objects and their spatial relationships suggest that the scene might be part of a desert-themed map or world in Minecraft. Otter: In this image, a person is standing on a raft in the middle of a body of water, with a small boat nearby on the shore. The scene is set in a game like Minecraft, where the player is trying to reach the other side of the water. The raft and the small boat are the primary objects that the player needs to utilize to cross the water safely. The person on the raft is standing near the boat, likely preparing to use it to continue their journey to the other shore. VPGTrans: The image shows a small wooden boat floating in the middle of a large body of water. There is a small island on the other side of the water, and the boat is headed towards it. The water is calm, and there are no visible obstacles in the way of the boat. The boat appears to be in good condition, and there are no visible signs of damage or wear and tear. The island appears to be relatively close, and it should be possible to reach it in a short amount of time.

C.3.2 Franka Kitchen

BLIP: Move the kettle to the top left burner step. InstructBLIP: Here’s a step plan: 1. Analyze the image to determine whether the action goal of moving the kettle to the top left burner has already been achieved or not. 2. If the action goal is not achieved, generate a plan for the robot to move the kettle to the top left burner. 3. Determine the initial position of the kettle on the stove. 4. Calculate the distance between the current position of the kettle and the desired position on the top left burner. 5. Based on the calculated distance, determine the necessary steps for the robot to move the kettle from its current position to the desired position on the top left burner. 6. Implement the necessary steps using the robot’s movement capabilities. 7. Verify that the kettle has been successfully moved to the top left burner. LLaMA-Adapter V2: To move the kettle to the top left burner, the robot would need to follow these steps: 1. Detect the kettle’s position: The robot should first identify the kettle’s current location on the kitchen counter. This can be done by analyzing the image and recognizing the kettle’s shape, size, and color.2. Plan the path: The robot should then calculate the most efficient path to move the kettle to the top left burner. This could involve moving the kettle to a specific location on the counter, then rotating or tilting the kettle to align it with the burner.3. Execute the plan: The robot should then execute the plan. LLaVA: To move the kettle to the top left burner, the robot should follow these steps: 1. Identify the kettle’s current position on the stove. 2. Determine the appropriate path to reach the top left burner. 3. Approach the kettle from the side, ensuring it doesn’t obstruct the path. Move the kettle towards the top left burner, making sure it doesn’t touch any other burners or appliances. 5. Once the kettle is positioned on the top left burner, confirm that it is stable and secure. 6. If the kettle is not yet on the top left burner, repeat steps 2-5 until the kettle is successfully placed on the burner. In the image, the kettle is already on the top left burner, so there is no need to generate the plans. The robot has successfully moved the kettle to the desired location. MiniGPT-4: The robot is standing in the middle of a kitchen, facing a counter with two burners on top and a sink on the left. A kettle is on the counter in front of the right burner. The robot is holding a ladle in its right arm and is reaching towards the kettle with it. Here’s a plan to move the kettle to the top left burner: 1. The robot will take the ladle from its right arm and place it on the counter next to the kettle. 2. The robot will use its left arm to pick up the kettle and place it on the top left burner. 3. The robot will then take the ladle and use it to pour hot water into the kettle. Overall, the action goal is achieved in three steps. mPLUG-Owl: 1. Obtain the position of the kettle on the countertop from the image. If the position of the kettle is not known, move the robot arm to scan the entire countertop area to find it. 3. After locating the kettle, the robot arm should grasp it using its gripper. 4. Using sensors such as cameras, the robot arm should navigate around other objects on the countertop and reach the desired location, which in this case is the left burner of a kitchen stove. Once the robot has arrived at the desired location, it should securely place the kettle on the top left burner. 6. The robot may need to adjust its grip on the kettle or the stove if necessary to maintain stability. 7. After the kettle is securely placed, the robot arm should release its grip and return to its original position or move back to the countertop to perform any additional tasks. Note: If it is already known that the kettle is on the top left burner, the robot may skip to step 4 or proceed to step 6 directly, based on its predefined task. Otter: In the image, the robot is already in the process of lifting the teapot with its arm and is positioning it on top of a burner on the stove. The plan would likely consist of the following steps: 1. Position the arm with the pot securely and safely on the burner. 2. Ensure that the pot is placed at the correct height and distance from the flame to avoid any accidents or damage to the pot. 3. Utilize the robot’s vision system or other sensors to accurately position the pot at the desired location. 4. Once positioned, the arm should maintain a firm grip on the pot to prevent it from falling or shifting during the cooking process. 5. If the pot needs to be lowered or moved to a different burner, the same process would be repeated, with the robot carefully adjusting the position of the pot and maintaining a secure grip. The generated plan should be clear, step-by-step, and easy to follow, ensuring that the robot successfully places the pot on the desired burner while maintaining safety and efficiency. VPGTrans: To move the kettle to the top left burner, the robot should first identify the location of the kettle and the location of the top left burner. Then, the robot should use its arm to pick up the kettle and move it to the top left burner. Once the kettle is in place, the robot should use its arm to adjust the position of the kettle so that it is centered on the burner. Finally, the robot should turn on the burner to heat the water in the kettle.

C.3.3 Meta World

BLIP: A robot is dunking a basketball in a pool. InstructBLIP: The image depicts a 3D basketball game where a robot is attempting to dunk the ball. The scene features a basketball hoop, a basketball, and a robot. In order to successfully dunk the basketball, the robot needs to perform several actions. Here is a detailed plan for the robot to dunk the basketball: 1. Position the robot near the basketball hoop: The robot should be positioned close to the basketball hoop so that it can easily reach the ball when it falls through the hoop. 2. Wait for the ball to fall through the hoop: Once the ball falls through the hoop, the robot should wait for it to land on the ground. 3. Pick up the ball: The robot should pick up the ball as soon as it lands on the ground. 4. Dribble the ball towards the hoop: The robot should dribble the ball towards the hoop, aiming for the basket. 5. Shoot the ball into the hoop: The robot should shoot the ball into the hoop, attempting to dunk it. 6. Celebrate the successful dunk: If the robot successfully dunks the ball, it should celebrate its achievement. LLaMA-Adapter V2: The image shows a basketball hoop with a basketball inside it, and a robot is positioned near the hoop. The robot is likely programmed to dunk the basketball, so it needs to execute a series of steps to achieve this goal. Here is a possible plan for the robot to dunk the basketball: 1. The robot should first identify the basketball and determine its position relative to the hoop. 2. The robot should then move towards the basketball, adjusting its speed and direction to ensure it reaches the ball at the right time. 3. As the robot approaches the basketball, it should extend its arm and grab the ball with its hand. 4. The robot should then lift its arm and place the basketball into the hoop, aiming for the basket. 5. Once the basketball is inside the hoop, the robot should release it. LLaVA: The image shows a basketball court with a basketball hoop, a basketball, and a person holding a remote control. The scene appears to be a simulation or a video game, as the person is controlling the basketball with the remote. The basketball is positioned near the hoop, and the person is likely trying to dunk it. The remote control allows the person to simulate the dunking motion, which is not possible in real life. MiniGPT-4: In this image, there is a basketball player standing on a wooden floor in front of a hoop. The player is holding a basketball in one hand and is preparing to dunk it through the hoop. The player is wearing a basketball uniform, and the hoop is hanging from the ceiling. There is a small orange ball lying on the ground next to the hoop, and a second ball is sitting on the top of the hoop. To write a plan for the robot to dunk the basketball, you could use the following steps: 1. Program the robot to move to the center of the room and stand in front of the hoop. 2. Program the robot to pick up the ball on the ground with its arm. 3. Program the robot to hold the ball with one hand while it uses its other hand to dunk the ball through the hoop. 4. Program the robot to release the ball after it has gone through the hoop. 5. Program the robot to move away from the hoop after it has finished dunking the ball. Overall, this plan would involve the robot using its arm to pick up the ball on the ground, its other hand to dunk the ball through the hoop, and then releasing the ball after it has gone through. mPLUG-Owl: The image features a robot standing on a wooden floor and holding a basketball in its hand. It seems to be attempting to perform a dunk on a basketball hoop. To achieve this, the robot needs to first position itself properly and aim at the hoop. First, it should move towards the basketball hoop by walking on its wheels. It should maintain a straight path and avoid colliding with the walls or any obstacles in its way. Once it reaches the hoop, the robot should adjust its position and align its body with the rim. It should also maintain a stable posture while maintaining the basketball in its hand. Next, the robot should jump up and extend its arm and hand towards the hoop, while simultaneously bending its knees. This will generate enough momentum to launch the basketball toward the rim. As the basketball is in its grasp, the robot should release it with a quick flick of its wrist. The ball’s trajectory should follow the path of the robot’s arm, and it should enter the basket with a satisfying dunk. Otter: The image shows a small toy basketball on a basketball court, with a large toy head of a man, which appears to be a statue or a decoration, nearby. The scene suggests that the toy robot is designed to perform a dunk on the basketball using a ramp. To accomplish this, the robot would need to follow a plan that involves moving to the appropriate position on the court, using the ramp to gain elevation, and timing its jump to make contact with the basketball at the right moment. The plan would likely involve adjusting the robot’s speed, angle, and trajectory to achieve a successful dunk while maintaining balance and coordination throughout the process. The specific details of the plan would vary depending on the design and programming of the robot, but the general idea would be to mimic the motion of a human dunking the ball. VPGTrans: The image shows a basketball hoop in the middle of a basketball court. The hoop is made of metal and has a net attached to it. There is a basketball on the ground next to the hoop. The hoop is at a height of about 10 feet off the ground. The basketball is at a height of about 5 feet off the ground. There is a robot on the left side of the hoop. The robot is facing the hoop and has a net attached to its arm. The robot is at a height of about 3 feet off the ground. The robot’s arm is extended and ready to shoot the basketball into the hoop.

Appendix D Summary of Evaluation Datasets