VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Haodong Duan, Xinyu Fang, Junming Yang, Xiangyu Zhao, Zerun Ma, Yuxuan Qiao, Mo Li, Tianhao Liang, Lin Zhu, Amit Agarwal, Xiaozhe Li, Shengyuan Ding, Jiazi Bu, Ziyu Liu, Zhangyang Qi, Yifei Li, Yuhang Zang, Zhe Chen, Lin Chen, Yuan Liu, Yubo Ma, Hailong Sun, Yifan Zhang, Shiyin Lu, Tack Hwa Wong, Weiyun Wang, Peiheng Zhou, Chaoyou Fu, Junbo Cui, Jixuan Chen, Enxin Song, Song Mao, Junming Lin, Xilin Wei, Jinsong Li, Zeyi Sun, Zhaowei Wang, Zicheng Zhang, Xiaoyi Dong, Junjun He, Pan Zhang, Jiaqi Wang, Dahua Lin, Kai Chen
Introduction
With the rapid development of Large Language Models (LLMs) , Large Multi-Modality Models (LMMs) have also experienced significant advancements. LMMs typically take two or more modalities as input. Most of the research has focused on LMMs for image and text , but research has also been extended to other modalities, such as audiotext , video , or point clouds . Furthermore, there exist LMMs that can simultaneously take more than two modalities as inputs, including proprietary APIs and open-source models . Compared to previous multi-modality models, LMMs, empowered by large language models, exhibit enhanced generalization capability and engage with humans in a variety of conversational styles. These models have not only demonstrated remarkable capabilities in multi-modal perception and reasoning tasks but have also spurred a range of innovative applications.
Quantitative evaluation is crucial in the development of LMMs. As general-purpose models, LMMs must undergo rigorous evaluation in a diverse range of tasks and domains. Comprehensive evaluations not only help users discern the strengths and weaknesses of an LMM, but also offer valuable feedback to developers for ongoing refinement. Unlike ‘pre-GPT’ models, evaluating LMMs on diversified quantitative benchmarks has become a common practice, both in academic works and commercial APIs .
Despite the importance of evaluating LMMs, conducting assessments across dozens of benchmarks can be a daunting task, particularly for small research teams. One must prepare data based on numerous repositories and manage potential environmental conflicts. Moreover, the authors of benchmarks may not provide evaluation results for all LMMs in which users are interested, thus requiring significant effort to compile uncompleted results. To alleviate this challenge, we developed VLMEvalKit, an open-source toolkit designed to facilitate the evaluation of LMMs.
VLMEvalKit aims to provide a comprehensive, user-friendly evaluation framework for researchers and developers to assess existing Large Multi-Modality Models. Currently, the codebase supports more than 70 different large multi-modality models, spaning proprietary APIs and open-source models (see Fig. 1), and over 20 multi-modal benchmarks covering a wide range of tasks and scenarios (see Tab. 1). The codebase’s straightforward design simplifies the integration of new benchmarks or LMMs. Typically, a developer only needs to prepare a single data file or implement a single interface to support new benchmarks or LMMs. Beyond its extensive collection and simplified design, VLMEvalKit significantly eases the work of comprehensive evaluation. Users can launch evaluations across multiple supported LMMs and benchmarks with a single command, generating well-structured evaluation results. The entire process eliminates the need for manual data preparation or post-processing, ensuring a seamless and efficient evaluation experience.
To align with real use cases, VLMEvalKit employs generation-based evaluation across all LMMs and benchmarks. Given their general-purpose nature, LMMs may struggle with adhering to specific formats for solving multiple-choice questions. For a fair comparison, VLMEvalKit utilizes large language models as choice extractors when exact matching fails, thereby mitigating the impact of response styles and enhancing the reliability of evaluation, Leveraging the reliable and reproducible evaluation results provided by VLMEvalKit, we maintain a comprehensive leaderboard to monitor the advancement of LMM development. The leaderboard disseminates valuable insights to the community and has garnered widespread recognition from both academia and industry.
VLMEvalKit is publicly available at https://github.com/open-compass/VLMEvalKit under the Apache 2.0 License. The repository includes the complete source codes along with detailed instructions for installation, evaluation, and further development. In subsequent sections, we will delve into the design and features of VLMEvalKit, present and analyze the evaluation results obtained, and discuss potential future developments.
VLMEvalKit: Design & Features
Fig. 2 displays the major components in VLMEvalKit. In this section, we will describe each single component in detail.
Benchmarks. VLMEvalKit encompasses dozens of multi-modal benchmarks that span a broad spectrum of tasks and scenarios, as detailed in Tab. 1. The benchmarks are first processed into .tsv files, where evaluation samples are stored as individual lines. Each evaluation sample typically includes the index, question, answer, image (one or more images encoded in base64), and choices for multi-choice questions. Each dataset class supports a .build_prompt() interface, which the evaluation script utilizes to construct evaluation samples into multi-modal messages. A multi-modal message comprises an interleaved sequence of contents across various modalities. For illustration, consider the following example:
[ dict(type=’image’, value=’path or url of the image’), dict(type=’image’, value=’path or url of the image’), dict(type=’text’, value=’Please list all the objects that appear in the above images.’) ]
Currently, a multi-modal message primarily incorporates image and text modalities. The format, however, is extensible and can accommodate additional modalities such as audio or point clouds.
LMMs. VLMEvalKit supports over 70 LMMs, including both commercial APIs and open-source models. A unified .generate() interface has been implemented for all LMMs, which accepts a multi-modal message as input and returns the response string. For LMMs that are limited to processing a single image-text pair, the concatenated text messages and the first image are adopted as the default input. To provide flexibility for LMM developers, the .generate() interface also includes dataset_name as an additional argument. An LMM can optionally implement its own .build_prompt() interface to construct custom multi-modal messages, and determine whether to utilize custom multi-modal messages based on the dataset_name flag. The unified interface eases the process of comprehensive evaluation. Each newly supported LMM / benchmark can be directly evaluated with all existing benchmarks / LMMs.
Multi-modal Inference. To expedite the multi-modal inference, we support parallelized inference for both commercial APIs, leveraging Python’s multiprocessing, and open-source models, which are distributed in parallel across multiple GPUs. Additionally, we have implemented a robust inference process so that an interrupted inference process can be resumed with minimal costs of repeated calculation or API calls.
Multi-modal Evaluation. Predictions from the LMMs will be evaluated based on the specific question format to derive the final metrics. Benchmarks supported in VLMEvalKit can be categorized into three primary types: 1. Multi-choice questions (MCQ), where the model must select from given options and respond with the corresponding label (e.g., A, B); 2. Yes-or-No questions (Y/N), requiring a straightforward ‘Yes’ or ‘No’ answer; 3. Open-ended questions, which necessitate a free-form response. Notably, a significant number of LMMs struggle to adhere to the instructions precisely, often producing responses that are not well-formatted for MCQ and Y/N benchmarks. To improve the precision of evaluations and counteract the influence of varied response styles, VLMEvalKit offers the option to integrate LLM-augmented answer extraction specifically for MCQ and Y/N benchmarks. We first adopt exact matching to match the response with the option labels or contents. Should this step fail, the toolkit then prompts an LLM (such as ChatGPT) to match the response with the option that most closely aligns with semantic meaning. This strategy helps us better understand the real performance of LMMs, particularly for commercial APIs .
Another notable challenge in assessing MCQ benchmarks is the inherent variance. When employing random guessing, an LMM may correctly answer questions for -option multi-choice benchmarks. Besides, we find that the outcome of an LMM can be significantly influenced by the order of options. These factors make the evaluation results highly variable and the performance gap between LMMs less discernible. To address this, VLMEvalKit offers an option to evaluate all MCQ benchmarks in Circular mode: options of a MCQ will be shifted in circular times to formulate new questions. The results count only if an LMM accurately answers all circular-shifted MCQs. The CircularEval strategy can more effectively assess the real comprehension of an LMM on MCQ benchmarks, allowing users to identify more pronounced performance disparities between models.
For open-ended benchmarks, VLMEvalKit follows the original practice for conducting the evaluation. Subjective benchmarks like MMVet or LLaVABench adopts GPT-4 for marking, based on the semantic similarity between LMM responses and the reference answer. For VQA benchmarks , we follow the standard practices to calculate accuracies based on heuristic matching.
Evaluation Results
Utilizing VLMEvalKit, one can conduct comprehensive evaluations of an LMM across numerous benchmarks to gain a thorough understanding of its strengths and weaknesses. We publish all evaluation results on OpenVLM Leaderboard. Our core leaderboard is based on evaluations from eight distinct benchmarks: 1. MMBench v1.1 [test] We report the average score of English and Chinese test splits. (all-round capability); 2. MMStar (data contamination); 3. MMMU [val] (multi-modal examination); 4. MathVista [mini-test] (multi-modal math); 5. HallusionBench (hallucination & illusion); 6. AI2D [test] (diagram understanding); 7. OCRBench (text understanding); 8. MMVet (subjective evaluation). The selection encompasses a diverse array of tasks, and the average score across these benchmarks serves as a reliable indicator of the general capabilities of LMMs. In Tab. 2, we present the performance of top-10 commercial APIs and the top-10 open-source models evaluated on this suite of benchmarks.
Commercial APIs. Generally, commercial APIs exhibit a significant performance advantage over open-source LMMs The three low-performing APIs: Qwen-VL-Max, GPT-4v-20231106, and Gemini-1.0-Pro are early versions released in 2023. . Notably, the top-5 performing LMMs are all proprietary APIs, with GPT-4o surpassing the highest-ranked open-source model, InternVL-Chat-V1.5, by over 8 points in average score. API models particularly excel on two benchmarks: MMMU and MMVet . On MMMU, four API models achieve above 60% accuracy, with an additional two surpassing 50%, whereas all open-source LMMs fall below the 50% threshold. Similarly, on MMVet, seven of the top-10 API models score above 60%, outperforming all open-source counterparts. Although the exact reasons for this disparity cannot be definitively confirmed, it is plausible that proprietary API LMMs leverage larger language encoders (at the 100B scale, etc.), which may enhance their knowledge level and better align with human preferences. We also note that the true capabilities of proprietary LMMs might still be underestimated due to content moderation policies employed, which can lead to the rejection of certain question types. Quantitative analysis indicates that the policies result in at least 2% and 4% reduction in average scores for GPT-4o and Gemini-1.5-Pro, respectively.
Open-source LMMs. The top-10 open-source LMMs predominantly utilize small to medium-scale language encoders, with parameter sizes ranging from 4B to 34B. InternVL-Chat-V1.5 stands out as the highest performer among all open-source LMMs, surpassing the -best model by 2 points in both average score and on the comprehensive MMBench benchmark. Leveraging the largest language encoder, LLaVA-NeXT-Yi-34B achieves the top score on MMMU, a benchmark that demands extensive knowledge. InternLM-XComposer2 impresses with its capabilities in multi-modal math and diagram understanding, while two LMMs based on GLM notably excel on OCRBench and MMVet. Additionally, the lightweight Mini-InternVL, featuring a 4B parameter size, demonstrates performance on par with the early versions of GPT-4v and GeminiPro, highlighting its efficiency and effectiveness.
Discussion
We have released VLMEvalKit, an open-source toolkit designed for the evaluation of large multi-modality models. VLMEvalKit encompasses a comprehensive collection of over 70 LMMs and more than 20 multi-modal benchmarks. The codebase is structured with simplicity in mind, facilitating the easy integration of new LMMs or benchmarks. Our framework is thoughtfully designed to extend its capabilities beyond the image modality, we have recently incorporated the evaluation of a video understanding benchmark, MMBench-Video , into VLMEvalKit. Moving forward, the development of VLMEvalKit will focus on expanding the repertoire of LMMs and benchmarks for video and other modalities. We are optimistic that this repository, along with all released resources, will contribute to advancing research in multi-modal learning.