MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Fei Wang, Xingyu Fu, James Y. Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, Tianyi Lorena Yan, Wenjie Jacky Mo, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang, Dan Roth, Sheng Zhang, Hoifung Poon, Muhao Chen
Introduction
The proverb “a picture is worth a thousand words” is often cited to emphasize the richness of visual information hidden in one image . However, an image is only a single projection of the real world captured from a specific angle at a specific moment in time . In contrast, humans naturally observe multiple images – multiple pieces of such projections from discrete moments under various scenes – to perceive and understand the world as a holistic part. Humans excel at synthesizing information from multiple image sources, whether it involves telling stories from a series of cartoon images , drawing comparisons among multiple charts and diagrams to infer holistic new insights , learning from diverse visual experiences such as online lesson slides to adopt new skills , predicting future event actions from past screenshots , or conducting temporal reasoning based on nuanced differences between photographs . Moreover, multi-image input has the advantage of conveying visuospatial ideas directly – combining multiple images of the same scene can reveal spatial relations or other more abstract relations in the world . Multi-image input also overcomes the limitations of resolution that single images face, allowing for better visual perception and understanding .
As multimodal large language models (LLMs) have begun to show superior performance across various single-image tasks, we now expect them to solve hard tasks that require an holistic understanding of multiple images. This work aims at highlighting crucial aspects of multi-image understanding that have been overlooked when evaluating multimodal LLMs, and providing a comprehensive benchmark for robust multi-image reasoning. As shown in Figure 2, current evaluations generally focus on single-image understanding, thereby neglecting the richer, more complex tasks of integrating and reasoning across multiple images. While many of these benchmarks have been popularized as the de facto evaluation measures for influential models like GPT-4-Turbo and Gemini-Pro , this oversight limits the potential of these models to conduct advanced-level multimodal comprehension. Though some recent benchmarks start to include multi-image questions in evaluation (e.g., Mantis-Eval and BLINK ), they are far from being comprehensive in multi-image evaluation that involve multi-persectives, multi-relations and robustness concerns.
In this paper, we introduce MuirBench (Multi-image understanding Benchmark), a comprehensive benchmark designed to rigorously assess and evaluate multi-image understanding by multimodal LLMs. MuirBench encompasses 11,264 images and 2,600 multiple-choice questions spanning across 12 distinctive multi-image understanding tasks, e.g. visual retrieval, cartoon understanding, and attribute similarity, etc. As illustrated in Figure 1, there can be multiple images interleaved in the contexts or questions, or presented as choices in our benchmark. Instances in MuirBench also contain diverse kinds of multi-image relations, e.g. temporal, ordered-pages, or narrative relations, etc. as shown in Figure 4. The questions and choices are either derived from the datasets, or manually written by experts. Additionally, MuirBench adopts a pairwise design approach, where each question-answering instance is paired with a expert-annotated unanswerable counterpart featuring minimal differences following Figure 5. This design ensures a reliable assessment of multimodal LLMs, mitigating the risk of achieving correct answers through vision or language shortcuts. We also include various fine-grained expert annotated labels such as image positions and image types in MuirBench, to facilitate detailed model analysis.
We conduct a comprehensive evaluation on MuirBench using 20 multimodal LLMs of various sizes, including models that accept multi-image inputs and those originally designed for single-image inputs. Experimental results underscore the current limitations of even the most influential multimodal LLMs, e.g. GPT-4o and Gemini Pro, in handling multi-image scenarios. For instance, GPT-4o and Gemini Pro achieve mere 68.0% and 49.3% of accuracy respectively, which are 25.1 % and 43.8% lower than human performance. We also show that multimodal LLMs perform much worse on unanswerable questions than their answerable counterparts, with GPT-4o and Gemini Pro exhibiting accuracy gaps of 26.8% and 21.5%. Furthermore, multimodal LLMs trained solely on single images demonstrate impaired generalization to multi-image contexts. These findings highlight the significance of MuirBench in driving the development of multimodal LLMs in transcending single-image limitations. We believe MuirBench can serve as an effective testbed for holistic multi-image understanding, encouraging the community to cultivate models with a more comprehensive and integrated understanding of the visual world.
Related Work
A number of recent benchmarks have been developed to comprehensively assess the multimodal understanding and reasoning capabilities of multimodal language models (LLMs) . However, most of these benchmarks primarily focus on single-image scenarios. While some benchmarks include multi-image examples , they typically require limited aspects of capacities (e.g., image comparison for MathVista) and do not provide a comprehensive assessment of multimodal LLMs in multi-image scenarios. While some benchmarks feature video understanding or in-context learning , the assessed capabilities are fundamentally different from multi-image understanding. Video understanding focuses on continuous streams of frames capturing dynamic changes over time, while in-context learning focuses on task adaptation using few-shot examples. In contrast, multi-image understanding challenges models to integrate and analyze spatial and contextual cues from varied perspectives, settings, and moments, thereby simulating the way humans process information from multiple visual sources. Recently, there have been dedicated efforts to assess multimodal LLMs in multi-image scenarios. For example, MANTIS-Eval is a human-annotated benchmark comprising 207 examples for multi-image reasoning, such as size perceptions and weight comparisons. DEMON evaluates whether multimodal LLMs can follow zero-shot demonstrative instructions. However, these benchmarks still focus on limited multi-image relations or reasoning processes and lack of robust evaluation. In contrast, MuirBench provides a comprehensive assessment of multimodal LLMs, covering a broader range of multi-image capacities.
2 Multimodal Large Language Models
Inspired by the remarkable achievements in recent LLMs , a series of studies have begun exploring multimodal LLMs that can concurrently interpret visual and linguistic information. However, most of early multimodal LLMs are trained on single-image datasets and overlook the complicated tasks of multi-image understanding . Recent work starts training multimodal LLMs on interleaved image-text corpus such as MMC4 and OBELICS for pretraining as well as Mantis-Instruct for instruction tuning, which enables models to generate texts given multiple images. While some of these models, like Flamingo , Idefics , Emu , and VILA , have demonstrated in-context learning capabilities, there is still a lack of evidence regarding their capabilities in understanding multiple images within independent instances. Although instruction tuned models such as Mantis and GPT-4-Turbo have shown to possess counting and comparison skills over multi-image inputs, their ability in understanding and reasoning over multiple images with different relations across diverse tasks, though critical, remain unexplored. Therefore, we propose MuirBench to conduct comprehensive evaluation and provide insights to further improve their capabilities in handling realistic multi-image tasks.
MuirBench
Our benchmark is meticulously curated for comprehensively assessing multimodal LLMs’ capabilities in holistic multi-image understanding. We introduce the overall design and key features of MuirBench in Section 3.1, and delve deep into the data curation process in Section 3.2.
Focusing on multi-image understanding, MuirBench consists of 11,264 images and 2,600 multiple-choice questions, with an average of 4.3 images per instance. In general, MuirBench adheres to two key design principles. First, it seeks to provide a comprehensive and holistic evaluation on multimodal LLMs’ multi-image understanding capabilities, by containing 12 diverse multi-image tasks covering 10 distinctive multi-image relation categories. Additional fine-grained labels such as input image positions and image types are also included to support comprehensive analysis of models. Second, it seeks to provide a robust evaluation, following a pairwise design where each answerable instance is paired with an unanswerable counterpart featuring minimal differences.
Comprehensive Multi-Image Evaluation. MuirBench provides an comprehensive assessment through 12 distinctive multi-image understanding tasks, with selected examples of each task shown in Figure 6. As illustrated in Figure 4, each task represents 2.5% to 17.8% of the whole benchmark.
[Action Understanding] aims to evaluate the ability of models to understand continuous images in chronological order and match it with an action. [Attribute Similarity] aims to evaluate the ability of models to identify a specific given attribute among multiple images. [Cartoon Understanding] aims to evaluate the ability of models to understand stories conveyed in cartoon images. [Counting] aims to evaluate the ability of models to count the number of specific objects across multiple images. [Diagram Understanding] aims to evaluate the ability of models to understand information conveyed in diagram images. [Difference Spotting] aims to evaluate the ability of models to identify differences across multiple images. [Geographic Understanding] aims to evaluate the ability of models to understand maps and reason upon geographic features. [Image-text Matching] aims to evaluate the ability of models to understand the meaning of a text snippet and match it with the corresponding visual content or vice versa. [Ordering] aims to evaluate the ability of models to order a series of images based on the textual description. [Scene Understanding] aims to evaluate the ability of models to understand a scene comprised of multiple views from multiple surveillance images. [Visual Grounding] aims to evaluate the ability of models to ground a specific object and seek information about it within multiple images. [Visual Retrieval] aims to evaluate the ability of models to retrieval images that contain the same building.
Additionally, MuirBench includes images covering 10 various categories of multi-image relations, such as narrative images conveying stories or ideas, ordered pages of documents and slides providing collective insights, images forming a temporal sequence presenting events, and multiple views of objects or 3D scenes offering a complete vision, with the complete distribution shown in Figure 4. In terms of image presentation, the number of images in each instance ranges from two to nine, while the input positions of images can be the beginning of question, middle of question, end of question, options, and a mix of these positions. MuirBench also exhibits various image types, including but not limited to slides, maps, medical images, drone/satellite images, animations, memes, graphics, and 3D views. The data diversity from the aforementioned perspectives enhances the comprehensiveness of our benchmark. More details can be found in Appendix A.
Robust Evaluation. Existing datasets primarily assess models’ capabilities in solving answerable questions but overlook their ability to recognize what they do not know . In real-world scenarios, there is no guarantee that user queries are answerable. A reliable multimodal LLM should directly indicate when a query is unanswerable rather than providing an answer that is most likely to be correct. In light of this, we pair each answerable instance with an unanswerable counterpart, featuring minimal differences, to provide a more robust evaluation, simulating real-world scenarios. We adopt multiple strategies to manually design the unanswerable instances, with major strategies of image replacing or reordering, question modification, and option modification introduced in Figure 5. More details can be found in Appendix A.
2 Data Collection
Answerable Data Collection. We invest our efforts in collecting multi-image multiple-choice question answering (MCQA) data covering various tasks and multi-image relations. Diverse data attributes enable fine-grained and diagnostic evaluation, while the multiple-choice format ensures deterministic results. To achieve this goal, we consider three sources of data, including existing datasets, dataset derivations, as well as newly collected data. Existing data (40.8%) come from GeneCIS , SeedBench , and IconQA . Derived data (21.7%) reformat data into MCQA format, using multiple strategies including question generation, option rewriting, and single-image QA combination, etc. upon instances from NLVR2 , HallusionBench , ISVQA , and MMBench . New data (37.5%) address certain tasks (e.g. geographic understanding and visual retrieval) that are underrepresented in the aforementioned collection to fulfill a more comprehensive evaluation. We manually create the question and choices for these data based on images from the National Geologic Map Databasehttps://ngmdb.usgs.gov/ngmdb/ngmdb_home.html, University-1652 , PubMed papershttps://pubmed.ncbi.nlm.nih.gov/, and SciDuet slides . Details about curation process and data sources for each task can be found in Appendix A.
Unanswerable Data Collection. As shown in Figure 5, we consider three strategies for modifying an answerable instance to its unanswerable counterpart with minimal changes. We first replace or reorder some images to disrupt the question-image and image-image relations (24.2%). We also modify the question to make it incompatible with the images and options (35.3%). In addition, we replace options to create a scenario with no correct answer (40.5%). For each answerable instance, we apply one of these three strategies. More details can be found in Appendix A.
Quality Control. We employ two types of quality control throughout the annotation process: automatic check with predefined rules, and a manual examination of each instance to filter out any low-quality data. The automatic check verifies valid instance format, answers, metadata values, and the coreference between image placeholders and images (ensuring no redundant image), as well as the accessibility of images. The manual examination is conducted by four experts working in this field, and filters out ambiguous queries, unclear images, and confusing instances.
Experiments
In this section, we first describe the experimental setup and the baselines (Section 4.1). Then we present a comprehensive evaluation of 20 recent multimodal LLMs (Section 4.2). We demonstrate that while humans can answer the questions with high accuracy, MuirBench is challenging for existing models. Finally, we conduct various analyses on multiple experiment settings, including sensitivity to various resolution and error analysis (Section 4.3).
Multimodal LLMs: We evaluate MuirBench on 20 recent multimodal LLMs, including models designed for considering multi-image inputs and those originally designed for single-image inputs. For multi-image input multimodal LLMs, we evaluate on GPT-4o, GPT-4-Turbo , Gemini Pro , Mantis (Idefics2, clip-llama3, and siglip-llama3 versions; 8B) , VILA (v1.5-13B) , Idefics (9B-Instruct and v2-8B) , Emu2 (Chat) and OpenFlamingo (v2-9B) . For single-image input multimodal LLMs, we evaluate on LLaVA (v1.5, NeXT, internLM, and xtuner versions, model size 7B, 13B, and 34B) , Yi-VL-6BMore details are at the official website at https://www.01.ai/, MiniGPT-4-v2 , and CogVLM . We refer the readers to Appendix B for more details.
Evaluation setup: We follow the standard setup as it is in VLMEvalKit , where the temperature is set to 0 and retry is set to 10. For the models that do not support multiple images as input, we concatenate the images to constitute one input. We extract the choice from the models’ output with a set of pre-defined rules. We refer the readers to Appendix C for more details on multi-image concatenation, visual prompting, answer extraction, and the human evaluation protocol.
2 Main Results
Overall performance: As shown in Table 1, the average accuracies of the most advanced multimodal LLMs on MuirBench are no better than 68%, which are still far from enabling satisfactory utility. The mean accuracies of open-source multimodal LLMs that have considered multi-images hover between 23.73% and 44.50%, which fall behind from advanced proprietary LLMs. Notably, there is no obvious correlation between model sizes and performances, indicating the importance of training data and training processes in developing multimodal LLMs with multi-image understanding capabilities. For certain models and tasks, some results are only on par or even below random guessing. We provide more in-depth model analyses in the following and in Appendix D.
In which multi-image tasks do multimodal LLMs show relative strengths and weaknesses? Figure 8 visualizes the accuracies of the best-performing models on MuirBench. We observe that multimodal LLMs perform relatively better on image-text matching, visual retrieval, and diagram understanding. In contrast, multi-image ordering and visual grounding appear to be more challenging for these models, because these tasks require understanding the whole multi-image context and conducting more complicated reasoning processes across images and modalities afterwards.
Can models designed for single-image inputs perform multi-image tasks? In general, models accepting multi-image inputs(e.g., Mantis-8B), even with fewer parameters, perform better than single-image input multimodal LLMs (e.g., LLaVA-NeXT-34B). This observation shows that generalizing from single-image training to multi-image inference is non-trivial. Reasonably, models benefit from multi-image training data and learning processes to develop multi-image understanding capabilities.
3 Analysis
Do multimodal LLMs perform worse on the unanswerable set? Figure 8 compares performances on answerable and unanswerable sets for some best-performing models. All the studied models have severe performance drop when changing answerable instances to unanswerable counterparts. A closer look of the error cases reveals that models often avoid abstention when facing unanswerable questions. These observations not only highlight the importance of assessing model behavior under a more realistic setting, but also show that the pairwise design improves the reliability of MuirBench.
Error analysis of GPT-4o: We randomly sampled 100 error instances made by GPT-4o on MuirBench and meticulously examined them. The most common error category (26% of error cases) is the failure of capturing details in images. The rest 20% of errors are due to inaccurate object counting or reasoning, followed by errors in logical reasoning (18%), identification of the same object in different scenes (14%), and inferring the intents implied by image sequences (12%).
Qualitative Results: Figure 6 presents some qualitative results, one per task. A notable phenomenon is that multimodal LLMs may hallucinate by attempting to find an erroneous option that appears to be likely correct for an unanswerable question rather than abstaining (see examples for cartoon understanding, diagram understanding, visual grounding, and visual retrieval). This illustrates the obvious performance gap between answerable and unanswerable instances in Figure 8. We refer the readers to Appendix D for more in-depth discussions.
Conclusion
In this work, we introduced MuirBench, a comprehensive benchmark designed to provide a robust evaluation on the multi-image understanding capabilities of multimodal LLMs. Experimental results of 20 multimodal LLMs, including the prominent models like GPT-4 and Gemini Pro, revealed substantial limitations in their ability to handle multi-image scenarios. These models showed significant performance deficits compared to human accuracy and struggled more with unanswerable questions in MuirBench. Our findings underscore the need for multimodal LLMs to transcend single-image limitations and achieve more holistic visual comprehension. MuirBench provides a rigorous framework for such assessments, encouraging the community to develop models that can effectively synthesize and reason across multiple visual sources.
Acknowledgments
Special thanks to BLINK authors, especially Wei-Chiu Ma, for providing the figure templates used in this paper.
References
Appendices
Figure 10 presents the overall statistics of MuirBench. Figure 10 shows the data distribution by the type of images. MuirBench covers a wide range of image types, ranging from common types like photography to specific areas such as medical images, slides, and drone and satellite imagery. Figure 12 demonstrates the data distribution by the number of images. MuirBench contains instances ranging from two images to nine images. Figure 12 presents the data distribution by the position of images, including the beginning/middle/end of a question, options, and a mix of these positions.
A.2 Dataset Curation Details
Answerable Data Collection. We invest our efforts in collecting multi-image multiple-choice question answering (MCQA) data covering various tasks and multi-image relations. Diverse data attributes enable fine-grained and diagnostic evaluation, while the multiple-choice format ensures deterministic results. To achieve this goal, we consider three sources of data, including existing datasets, dataset derivations, as well as newly collected data. Existing data come from datasets that focus on a single aspect of multi-image reasoning, such as GeneCIS ; and from datasets not specifically designed for the multi-image setting but containing a portion of multi-image data, such as SeedBench and IconQA . For a fair representation of each task, we sample up to 200 test examples from each dataset. This part contributes 40.8% of the data in the final benchmark. Derived data reformat binary QA, such as NLVR2 and HallusionBench , into MCQA by modifying questions and options; or rewriting open QA, such as ISVQA , into MCQA by adding options; and reconstructing single-image MCQA, such as MMBench , into multi-image MCQA by replacing text options with corresponding images. Similar to those from the existing datasets, we sample up to 200 test examples from each dataset. This part contributes 21.7% of the data in the final benchmark.
New data address certain tasks (e.g. geographic understanding), image relations (e.g. multiview), and types (e.g. medical images) remaining absent or underrepresented in the aforementioned collection to fulfil a more comprehensive evaluation. We present four new datasets: HistoricalMap, UnivBuilding, PubMedMQA, and SciSlides. HistoricalMap requires identifying map patches covering the same regions collected from the National Geologic Map Database.https://ngmdb.usgs.gov/ngmdb/ngmdb_home.html UnivBuilding requires identifying different views of the same building, or buildings from the same universities. The image data are from University-1652 . PubMedMQA contains questions regarding the subfigures from medical papers on PubMed.https://pubmed.ncbi.nlm.nih.gov/ SciSlides consists of questions regarding the slides for paper presentation collected from SciDuet . This part contributes 37.5% of the data in the final benchmark.
Unanswerable Data Collection. As shown in Figure 5, we consider three strategies for modifying an answerable instance to its unanswerable counterpart with minimal changes. We first replace or reorder some images to disrupt the question-image and image-image relations. We also modify the question to make it incompatible with the images and options. In addition, we replace options to create a scenario with no correct answer. For each answerable instance, we apply one of these three strategies. Among all the instances, 24.2% of the unanswerable instances are created by replacing or reordering the images in their answerable counterparts, 35.3% by modifying the questions, and 40.5% by changing the options. This step doubles the size of data, leading to a balanced distribution of answerable and unanswerable instances.
Metadata Annotation. Fine-grained metadata enable a diagnostic analysis of multimodal LLMs’ weaknesses across various aspects. We annotate image relations, tasks, image types, number of images, and image positions for all instances. Among all of these attributes, image relations are a crucial factor that influences the model’s capability for multi-image reasoning, yet they are rarely annotated in existing data. Therefore, we manually annotate them. Tasks and image types are partially annotated in existing data. We match the existing categories with our taxonomy and manually fill in any missing ones. Number of images and image positions are automatically detectable, so we conduct automatic annotation. The annotation interface is shown in Figure 16.
Quality Control. We employ two types of quality control throughout the annotation process: automatic check with predefined rules, and a manual examination of each instance to filter out any low-quality data. The automatic check verifies valid instance format, answers, metadata values, and the coreference between image placeholders and images (ensuring no redundant image), as well as the accessibility of images. The manual examination at last filters out ambiguous queries, unclear images, and instances with other errors, resulting in the retention of 86.3% of instances.
A.3 Multi-image Relations
MuirBench consists of 10 multi-image relations:
Temporal Relation: Images are related by time, showing progression or change over a period. Examples include time-lapse photography or sequential frames from a video.
Ordered Pages: Images are part of a sequence, such as pages in a book or slides in a presentation, where the order conveys meaning.
Complementary Relation: Images that, when viewed together, provide additional information or context that enhances the understanding of the subject. They complement each other by filling in gaps or providing different perspectives.
Cropped/Zoomed Images: One image is a zoomed-in or cropped version of another, focusing on a specific part of the original image to highlight details.
Narrative: A series of images that together tell a story or convey a sequence of events, much like a comic strip or a storyboard.
Scene-Multiview: Multiple images of the same scene taken from different angles or perspectives, providing a more comprehensive view of the scene.
Object-Multiview: Images of the same object captured from various angles or perspectives, useful for understanding the object’s three-dimensional shape.
Overall Similarity: Images that are generally similar in content, style, or subject matter, but not necessarily identical. They might share common themes or visual elements.
Partial Similarity: Images that share some, but not all, elements. They might have overlapping features or subjects but also contain distinct differences.
Independent Images: Images that do not have a clear relation to each other. They are not connected by time, sequence, context, or content.
A.4 Human Evaluation Protocol
Two experts in domain conduct the human evaluation. Each answerable instance and its unanswerable counterparts are randomly assigned to different experts ensuring a fair evaluation. The interface for human evaluation is shown in Figure 13.
Appendix B Baseline Models
We evaluate MuirBench on 20 recent multimodal LLMs, including models designed for considering multi-image inputs and those originally designed for single-image inputs. For most model families, we use the latest and best-performing available checkpoint to date. The list of baseline models are as follows:
(i-ii) GPT-4 is known to be one of the best multimodal models to date. We test with two most up-to-date checkpoints: gpt-4-turbo and gpt-4o. Notice that the GPT-4 performance would change if this specific checkpoint gets updated. (iii) Gemini Pro is one of the most powerful multimodal models, and we use the Gemini 1.0 Pro Vision version of it. (iv-vi) Mantis (Idefics2, clip-llama3, and siglip-llama3 versions; 8B) is a recent strong model specifically finetuned for multi-image related tasks. (vii) VILA (v1.5-13B) , (viii-ix) Idefics (9B-Instruct and v2-8B) , (x) Emu2 (Chat) and (xi) OpenFlamingo (v2-9B) are four recent multimodal models that can take multiple images as input. (xii-xvii) LLaVA (v1.5, NeXT, internLM, and xtuner versions, model size 7B, 13B, and 34B) are included as well. While they’re designed for single-image input, we concatenate all the images in order. (xviii) Yi-VL-6BMore details are at the official website at https://www.01.ai/ has shown great performance recently. (xix) MiniGPT-4-v2 adapts EVA as visual backbone, LLaMA2-chat (7B) as language model backbone, and designs a linear projection layer for visual understanding abilities. (xx) CogVLM adds a trainable visual expert module in the attention and FFN layers to bridge different modalities better. It uses EVA-CLIP as vision encoder and Vicuna as language backbone.
Appendix C Experiment Setting Details
Following ,https://github.com/lupantech/MathVista/blob/9ed0e8b52c0911e31faa75308082af5dcf8e63b2/evaluation/build_query.py#L152 our prompt consists of four parts, the question, options, the hint indicating the answer format, and a prefix of the answer. For images, we insert them into the text to form a coherent prompt. The complete prompt is as follows:
C.2 Evaluation Tool
Following , We use a rule-based automatic toolhttps://github.com/MMMU-Benchmark/MMMU/blob/f3e473e1e7af2c65a56ab66d7b3cf09c5dbaf0b9/eval/utils/eval_utils.py#L10 to extract the exact answer. First, the tool detects if a valid option index appears in the model output. If no direct answer is found, the tool matches the output to the content of each option. If there is still no match, it will randomly select an option as the answer. When more than one valid answer is detected, the tool will use the first one that appears as the final answer.
Appendix D Error Analysis
Do image positions correlate with error rates? We analyze the error rates of varying input positions of images and report the performance of GPT-4o, GeminiProVision, and Mantis-8B-Idefics2. As shown in Figure 15, the highest accuracy is achieved when images are positioned in options, while the highest error rate can be observed when images are in the middle of questions. This consistent trend across different models suggests that the position of images within a question correlates with the error rate. The cause of higher error rates might be that images in the middle or end of a question may interrupt the flow of context processing, increasing complexity and thus reducing model performance. It may also be attributed to the training process. These models may have seen less data with images in the middle during training.
Do unanswerable types correlate with error rates? We further analyze the error rates of varying unanswerable types and report the performance of the same three models in Figure 15. Results show that the error rate also correlates with the type of unanswerable instances. All the three models perform relatively better when we only change the questions to make it incompatible with original images and options. However, all models are confused when the correct option is removed and fail to choose “none of the other options” in this scenario. The performance on unanswerable instances created by reordering or replacing images is divergent. GPT-4o performs much better than the other models in these cases.
Appendix E Limitations
There are several limitations to this work. First, we focus our scope on 2D images, and future research can further extend the idea of work to 3D problems, and include more multi-image tasks and relation categories. We hope our work can guide future efforts in providing robust and faithful evaluation in multimodal benchmarks. Our strategies of creating unanswerable instances, as in Figure 5, do not cover all strategies that can be used to create such instances. Also, we focus our evaluations on multimodal large language models. Future work could include more vision-language foundation models such as Unified-IO 2 and Chameleon .
E.2 Societal impacts
Our work proposes MuirBench, providing a robust evaluation on multi-image tasks using multimodal LLMs. While it includes a comprehensive list of 12 tasks, all of them are in English and could induce bias on multilingual research settings. Also, if misused, the multimodal LLMs may be used to generate harmful vision and text artifacts. Nevertheless, this is not directly related to our research, and the data we curate do not contain personally identifiable information or offensive content. However, more researchers should be encouraged to get involved in research on the safety issues in a multimodal context.
Appendix F License
We release our data under CC-BY 4.0 license. For specific instances we follow their original licenses. The datasets we used and their licenses are as follows:
GeneCIS is released under the CC-BY-NC 4.0 license.https://github.com/facebookresearch/genecis/tree/main?tab=readme-ov-file#license
SEED-Bench is released under the CC-BY-NC 4.0 license.https://huggingface.co/datasets/AILab-CVC/SEED-Bench
IconQA is released under the CC BY-NC-SA license.https://iconqa.github.io/
NLVR2 is released under the CC-BY-4.0 license.https://github.com/lil-lab/nlvr/tree/master?tab=readme-ov-file#licensing
HallusionBench is released under the BSD 3-Clause license.https://github.com/tianyi-lab/HallusionBench?tab=readme-ov-file#license
ISVQA annotation is released under the CC BY-NC-SA 2.0 license.https://github.com/ankanbansal/ISVQA-Dataset/tree/master?tab=License-1-ov-file We only use the images from nuScenes, which is released under the CC BY-NC-SA 4.0 license.https://www.nuscenes.org/terms-of-use
MMBench is released under the Apache-2.0 license.https://github.com/open-compass/MMBench?tab=Apache-2.0-1-ov-file
National Geologic Map Database is free in the public domain.https://www.usgs.gov/faqs/what-are-terms-uselicensing-map-services-and-data-national-map
University-1652 is released under the MIT license.https://github.com/layumi/University1652-Baseline?tab=MIT-1-ov-file#readme
PubMed is a free and public database, with open access articles under a Creative Commons or similar license.https://www.ncbi.nlm.nih.gov/pmc/about/copyright/
SciDuet is released under the Apache 2.0 license with paper slides from ACL, ICML, and NeurIPS.https://github.com/IBM/document2slides?tab=Apache-2.0-1-ov-file
Appendix G Accessibility of MuirBench
The full documentation of MuirBench is on the project page at https://huggingface.co/datasets/MUIRBENCH/MUIRBENCH. For each data entry in MuirBench, it includes metadata of index (idx), task, question, options, answer, image relation, image type, images, and counterpart instance idx.
G.2 Links and Maintenance Plan
MuirBench is hosted on Huggingface/Datasets,https://huggingface.co/datasets/MUIRBENCH/MUIRBENCH where license and metadatahttps://huggingface.co/api/datasets/MUIRBENCH/MUIRBENCH/croissant are also available. We maintain our benchmark on this page and will continually update it. The evaluation code and outputs will be provided to facilitate easy reproduction and analyses of the results in the paper.
G.3 Author Statement
We confirm that we bear all responsibility in case of violation of rights during the collection of data on MuirBench, ensuring accountability and commitment to maintaining ethical standards. We will take appropriate action when needed.
G.4 Intended Uses
The dataset is for academic purposes only and not for commercial usage.