MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, Jianfeng Gao
Introduction
Mathematical reasoning stands as a testament to the intricacies of human intelligence (Kahneman, 2011). It requires rigorous logical thinking, domain-specific knowledge, and the ability to engage in multistep reasoning processes (Lightman et al., 2023). This complexity is observed not only in textual scenarios but also significantly in visual contexts. For instance, when assessing a child’s mathematical and reasoning capabilities, problems are often designed to encompass visual contexts in addition to arithmetic calculations (Stipek & Iver, 1989; Pollitt et al., 2020). At the same time, AI agents with strong mathematical reasoning capabilities in visual contexts have a wide range of real-world applications, such as solving complex problems in educational disciplines (Seo et al., 2015; Wang et al., 2017), helping analysts with logical queries about statistical data (Wu et al., 2023; Yang et al., 2023a), and assisting in theorem proving and scientific discovery in advanced research fields (Taylor et al., 2022; Dong et al., 2023; Trinh et al., 2024).
Numerous datasets have been curated to assess the mathematical reasoning abilities of AI systems, with most presented purely in text form. Some datasets such as ChartQA (Lu et al., 2021a; Dahlgren Lindström & Abraham, 2022; Masry et al., 2022) have explored mathematical reasoning in vision-language settings. However, these datasets tend to either focus on specific tasks, like math word problems, or particular visual contexts, such as geometry problems or bar charts. General-purpose visual question answering (VQA) datasets on natural scenes contain only a small portion of questions necessitating mathematical reasoning, leaving a comprehensive investigation of vision-language reasoning within a mathematical framework largely unexplored.
On the other hand, Large Language Models (LLMs) (OpenAI, 2022; 2023a) and Large Multimodal Models (LMMs) (Google, 2023; OpenAI, 2023b; Team et al., 2023) have exhibited impressive problem-solving skills in many tasks and domains. Recently, some studies have aimed to augment existing LLMs with mathematical and scientific reasoning capabilities using external tools (Lu et al., 2023a; Wang et al., 2023b). However, the ability of these foundation models to perform mathematical reasoning in visual contexts has not been systematically examined. Therefore, it is essential to develop a new benchmark to (1) facilitate the development of mathematical reasoning systems in visually intensive scenarios, and (2) evaluate the research progress of LLMs and LMMs, especially their capabilities in solving rigorous reasoning tasks.
In this paper, we present MathVista, a consolidated Mathematical reasoning benchmark in Visual contexts. We propose a task taxonomy to guide the development of MathVista: (1) we identify seven mathematical reasoning types: algebraic reasoning, arithmetic reasoning, geometry reasoning, logical reasoning, numeric common sense, scientific reasoning, and statistical reasoning; (2) we focus on five primary tasks: figure question answering (FQA), geometry problem solving (GPS), math word problem (MWP), textbook question answering (TQA), and visual question answering (VQA); and (3) we encompass a diverse array of visual contexts, including natural images, geometry diagrams, abstract scenes, synthetic scenes, as well as various figures, charts, and plots. MathVista incorporates 28 existing multimodal datasets, including 9 math-targeted question answering (MathQA) datasets and 19 VQA datasets. In addition, we have created three new datasets (i.e., IQTest, FunctionQA, PaperQA) which are tailored to evaluating logical reasoning on puzzle test figures, algebraic reasoning over functional plots, and scientific reasoning with academic paper figures, respectively. Overall, MathVista consists of 6,141 examples, with 736 of them being newly curated (Table 1). To facilitate fine-grained evaluation, examples are annotated with metadata, including question type, answer type, task category, grade level, visual context, and required reasoning skills. Detailed descriptions of data collection can be found in §2, §C, and §D.
We conduct extensive experiments on MathVista to evaluate the reasoning abilities of 12 foundation models known for their leading performance in mathematical and multimodal reasoning. This ensemble includes three LLMs (i.e, ChatGPT, GPT-4, Claude-2), two proprietary LMMs (i.e., GPT-4V, Bard), and seven open-source LMMs. For LLMs, we examine zero-shot and few-shot settings using two prompting strategies: chain-of-thought (CoT) (Wei et al., 2022b) and program-of-thought (PoT) (Chen et al., 2022b). These LLMs can also be augmented with off-the-shelf visual models for image captioning and OCR. We establish a human performance baseline by engaging qualified human annotators with a high school diploma or higher. We show that MathVista, featuring advanced topics such as college curricula and scientific reasoning, is a very challenging benchmark, with human performance reaching only 60.3% accuracy.
Our results indicate that CoT GPT-4, the best-performing LLM without visual tool augmentations, achieves an overall accuracy of 29.2%. Multimodal Bard, the best-performing LMM, achieves 34.8% (§3.3), which attains only 58% of human performance (34.8% vs 60.3%). When augmented with Bard captions and OCR text, PoT GPT-4 obtains 33.9%, closely matching Multimodal Bard (§3.4). Further analysis indicates that the Multimodal Bard model failures arise from incorrect calculations and hallucinations caused by visual perception and textual reasoning (§3.5).
With MathVista, we report, for the first time, a comprehensive quantitative and qualitative evaluation of GPT-4V (OpenAI, 2023b), the latest multimodal version of GPT-4. Remarkably, GPT-4V achieves a state-of-the-art accuracy of 49.9%, a significant improvement of 15.1% over Multimodal Bard. As illustrated in Figure 1, GPT-4V even surpasses human performance on a set of tasks involving algebraic reasoning and complex visual contexts, which include tables and function plots. Nevertheless, a 10.4% gap in overall accuracy remains when compared to the human baseline, leaving plenty of room for model improvement. Our in-depth analysis (§H) reveals that the superiority of GPT-4V is mainly attributed to its strong capabilities in visual perception and mathematical reasoning. We further highlight its emergent ability for self-verification (§H.5), the use of self-consistency (§H.6), and its ability to drive goal-directed multi-turn human-AI dialogues (§H.7).
The MathVista Dataset
As discussed previously, there is a notable gap in existing benchmarks, which primarily evaluate mathematical reasoning in textual contexts, overlooking the intrinsic visual nature of many mathematical problems. Our dataset, MathVista, is therefore motivated to bridge this gap, offering a robust evaluation benchmark for mathematical reasoning intertwined with visual understanding, thus pushing AI assistants towards general-purpose capabilities. Our benchmark adheres to the following collection guidelines: (1) it covers multiple tasks and topics to mirror real-world applications; (2) it incorporates diverse visual contexts and mathematical skills to foster a well-rounded evaluation; (3) it offers varying levels of challenge to effectively probe and uncover the potential limitations of current models; and (4) it provides robust evaluation settings for deterministic evaluations.
The taxonomy for this work is introduced as follows: We identify seven types of mathematical reasoning: algebraic reasoning, arithmetic reasoning, geometry reasoning, logical reasoning, numeric common sense, scientific reasoning, and statistical reasoning, with detailed definitions provided in §C.1 and examples shown in §C.2. We focus on five primary tasks: figure question answering (FQA), which centers around statistical reasoning over multiple charts and plots; geometry problem solving (GPS), which deals with geometrical topics; math word problem (MWP), which involves arithmetic reasoning in everyday scenarios; textbook question answering (TQA), which usually entails knowledge-intensive reasoning on scientific topics and figures; and visual question answering (VQA). Furthermore, our objective is to account for a diverse array of visual contexts, including natural images, geometry diagrams, abstract scenes, synthetic scenes, multiple charts and plots, scientific figures, tables, function plots, puzzle test figures, and more, with examples shown in §C.3.
2 Data Collection
We collected nine MathQA datasets in multimodal settings, including four for GPS, two for MWP with visual contexts of synthetic scenes, abstract diagrams, and tables, and two for TQA on college curricula (see §C.4). Annotations such as solutions, programs, parsing results, and grounded theorems are also collected, providing demonstration examples for LLMs. Each source dataset is limited to up to 400 examples to ensure a balanced representation of each source in our final compiled benchmark. In total, we collected 2,666 examples.
Many existing VQA datasets feature instances requiring mathematical reasoning abilities, such as arithmetic operations or numeric common sense. Incorporating these datasets enhances problem diversity in terms of tasks, domains, visual contexts, and reasoning skills involved. We reviewed more than 70 datasets, collecting 19 of them that contain math-related instances and are publicly available, as listed in §C.4. Since these datasets are not originally math-targeted, we initially designed heuristic rules to automatically select examples likely to involve mathematical reasoning from a large pool of candidates. Examples with numeric answers or those containing quantity words (as listed in §D.1) in the questions were selected. This automatic filtration yielded 4,949 VQA-format examples, though some false positive examples remained. Therefore, we engaged three expert annotators to manually label these examples to determine if they involve mathematical reasoning (more details in § D.2). Utilizing majority voting and limiting each source dataset to 400 examples, we finalized a collection of 2,739 examples.
While the source datasets we collected encompass multiple visual contexts and mathematical reasoning abilities, certain scenarios remain unaddressed: logical reasoning on puzzle test diagrams, statistical reasoning on functional plots, and scientific reasoning on academic figures. To address these gaps, we introduced three new datasets: IQTest, FunctionQA, and PaperQA, with examples illustrated in Figure 2. IQTest comprises 228 examples requiring inductive reasoning, abstract thinking, pattern prediction, and calculations, sourced from puzzle test figures on online learning platforms. FunctionQA, with 400 examples, emphasizes subtle visual perceptions of functional plots and algebraic reasoning concerning variables, expressions, equations, and functions. PaperQA is a novel dataset featuring questions derived from informative academic illustrations, including tables, figures, and charts from online education resources, with 107 examples sourced from papers released in August 2023 on Huggingfacehttps://huggingface.co/papers.
To ensure data quality, all questions were manually annotated by graduate students in STEM fields and further refined through a rigorous review process. To ensure consistency in annotation, we employed a two-step process. Initially, each dataset was independently annotated by three reviewers, resulting in a high inter-annotation consistency rate of 99.2%. Specifically, among the newly collected 736 questions, only 6 exhibited disagreements in the annotated answers. Then, these discrepancies were resolved through discussion among the entire review team, ensuring a consensus was reached on each example. The GUI of the annotation tool is shown in Figure 23 in §D.3.
3 Metadata Annotation
Fine-grained metadata facilitates a comprehensive analysis of models’ reasoning capabilities across various aspects. To this end, we annotate the examples in MathVista with information including question type, answer type, language, source, category, task, grade level, and visual context, which can be accurately obtained from the details provided in the source datasets. MathVista features seven different types of mathematical reasoning abilities, as categorized in Table 3 (§C.1). Coarse labels of mathematical reasoning can be automatically obtained from the details of the source datasets. To verify the quality of automatic annotation, expert annotators manually label the mathematical reasoning categories from seven candidates for 1,000 examples, using the annotation tool illustrated in §D.4. The results show that 94.1% of the examples from automatic and human annotations have the exact same set of reasoning types, while 98.79% of the individual labels are identical, indicating that the automatic annotation for the labeling of mathematical reasoning is highly accurate.
4 Data Preparation and Release
MathVista consists of 6,141 examples, divided into two subsets: testmini and test. testmini contains 1,000 examples, intended for model development validation or for those with limited computing resources. The test set features the remaining 5,141 examples for standard evaluation. Notably, the answer labels for test will not be publicly released to prevent data contamination, and we will maintain an online evaluation platform. To ensure that each source dataset is well represented in testmini and to maintain a distribution in testmini closely resembling the whole set, we adopted this sampling strategy: (1) first, randomly sample questions with a threshold number of 4 for each source dataset; (2) then, randomly sample the remaining questions for each source dataset on its proportion in the entire set. The KL Divergence and Total Variation (TV) distance between the testmini set and the entire set are 0.008 and 0.035, respectively, suggesting that testmini is close to the distribution of the whole set. We also conducted several quality checks to address any unidentified errors.
5 Data Analysis
The main statistics of MathVista are presented in Table 1. There are two types of questions: multiple-choice and free-form. Answers to free-form questions are categorized as integers, floating numbers, or lists. The large unique number of images, questions, and answers ensures pattern diversity in MathVista. MathVista is derived from 31 source datasets, including three newly annotated datasets to address the missing types of mathematical reasoning over specific visual contexts. Dataset examples in Table 4 (§C.2 ) highlight the richness of mathematical reasoning involved. Examples in §C.3 demonstrate the diverse visual contexts present in MathVista. Further details on data analysis are available in §E.
Experiments
Prior work (Yang et al., 2023b) has studied the reasoning abilities of foundation models in visual settings from a qualitative perspective. In contrast, our goal is to conduct both qualitative and quantitative studies to provide a systematic evaluation of existing foundation models for mathematical reasoning capabilities in visual contexts using MathVista. We introduce a novel benchmarking strategy for MathVista tailored for foundational models (§3.1). The models we have chosen are detailed in §3.2. Quantitative results can be found in §3.3 and §3.4, while the qualitative analysis is provided in §3.5. Given the significant advancements of GPT-4V over other models, we undertake an in-depth comparative study with its peers in various aspects and highlight potential avenues for future research in §H.
Recent LLMs and LMMs have been instructed to generate long responses in conventional settings instead of short text. Therefore, we propose a new strategy for benchmarking MathVista, unlike using human-designed or template matching rules (Lu et al., 2022). The evaluation process consists of three stages: response generation, answer extraction, and score calculation. Initially, the baselines generate responses given the input query, which incorporates the task description, the question, the choices, and the metadata, using the template defined in Table 9 (§F.3). Next, the short answer text is extracted from the detailed response. We propose an answer extractor (§F.2) based on LLMs such as GPT-4, inspired by its remarkable ability for text processing (Wei et al., 2022b). A preliminary study of 200 examples shows that GPT-4 can extract the answer text with more than 99.5% accuracy. Finally, the extracted answer is normalized to a required answer format (e.g., an option letter or an integer), and the target metric scores are computed. Taking advantage of the fact that the instances in MathVista are either multiple-choice questions for textual answers or free-form questions for numerical answers, accuracy scores are used as metrics for deterministic evaluation.
2 Experimental Setup
We evaluate the models on MathVista under three setups: (a) Text-Only LLMs including ChatGPT (OpenAI, 2022), GPT-4 (OpenAI, 2023a), and Claude-2 (Anthropic, 2023) in zero-shot and two-shot settings with Chain-of-Thought (CoT) (Wei et al., 2022b) and Program-of-Thought (PoT) (Chen et al., 2022b), (b) Augmented-LLMs where the LLMs are provided with additional visual information including the generated image captions from Multimodal Bard (Google, 2023) and the detected OCR text from EasyOCR (JaidedAI, 2020), (c) LMMs that include open-source models such as IDEFICS-9B (Laurençon et al., 2023), mPLUG-OWL-LLaMA-7B (Ye et al., 2023), miniGPT-4-LLaMA-2-7B (Zhu et al., 2023a), LLaMA-Adapter-V2-7B (Gao et al., 2023), InstructBLIP-Vicuna-7B (Dai et al., 2023), LLaVA-LLaMA-2-13B (Liu et al., 2023a), LLaVAR Zhang et al. (2023d), and proprietary models such as Bard and GPT-4V. Since GPT-4V does not offer API access, we resorted to manually evaluating it using the playground chatbot. We provide the prompts for LLMs and the hyperparameters used for LMMs in §F.
3 Experimental Results
We compare the performance of several models, including Text-only LLMs, Augmented LLMs, and LMMs on MathVista in Table 2. We include random chance (i.e., one of the options in multiple-choice questions, and empty in the free-form questions) and frequency guess (§F.1) as naive baselines. Additionally, we established a human performance baseline using Amazon Mechanical Turk. Eligible human annotators must have a satisfactory annotating history, successfully pass qualification examples, and possess a high school degree or higher. We asked each annotator to complete five questions within 20 minutes. Further details can be found in §F.6.
Among text-only LLMs, all models outperform the random baselines, with the 2-shot GPT-4 using chain-of-thought (CoT) prompting achieving 29.2%. The limited performance of text-only LLMs suggests that our dataset requires models to reason within visual contexts for optimal results. When equipped with image captions and detected OCR text, augmented LLMs exhibit superior performance compared to their text-only counterparts on MathVista. Specifically, the best-performing augmented LLM is the 2-shot GPT-4 employing program-of-thought (PoT) prompting, which scores 33.9%. This model generates Python programs for execution, thereby promoting rigorous reasoning.
On the LMM side, Multimodal Bard scores a 34.8% accuracy, which is only 58% of human performance at 60.3%. Notably, the best-performing GPT-4V model achieves 49.9%, marking a substantial 15.1% improvement over Bard; however, it still falls 10.4% short of human performance. These gaps highlight that there is a significant scope for further improvements on our benchmark. The open-source models (IDEFICS to LLaVA) achieve underwhelming performance on MathVista. This can be attributed to their lack of math reasoning capabilities, text recognition (useful for math word problems), shape detection (useful for geometrical problems), and chart understanding. Notably, these models utilize different model architectures for processing the vision (e.g., OpenCLIP, CLIP, Vit-G) and language (e.g., LLaMA-1, LLaMA-2), different alignment strategies (e.g., MLP projection in LLaVA, Q-former in InstructBLIP, visual abstractor in mPLUGOwl), and instruction tuning data (e.g., 150K instruction-response pairs from LLaVA data, 3,500 instruction-response pairs from miniGPT-4). While fine-tuned with instruction-following data from text-rich images, LLaVAR does not perform well, indicating that strong text recognition abilities do not guarantee high performance on MathVista, which requires comprehensive visual perception and mathematical reasoning. This underscores that there are immense possibilities for innovations in model, data, or training objectives to improve the zero-shot performance of LMMs on MathVista.
4 Fine-grained Results
We also report fine-grained scores for a comprehensive study of the capabilities of existing models across different tasks (Table 2), mathematical reasoning abilities (Table 2, Figures 1, 33), visual context types (Figures 1, 34), and grade levels (Figure 35). Remarkably, GPT-4V surpasses most other baselines in various categories, with exceptions in problems related to logical reasoning and numeric commonsense reasoning. Notably, GPT-4V surpasses human performance not only in tasks like geometry problem solving (GPS), textbook question answering (TQA), and mathematical reasoning skills such as algebraic reasoning but also in visual contexts including function plots, geometry diagrams, scatter plots, and tables. Please refer to §G.2, §G.3, and §G.4 for more detailed analysis.
We perform an ablation study on the augmented LLMs and present the results in Table 36 (see §G.5). The gap in the performance of the Augmented LLMs can be attributed to poor image captions, which may not adequately describe the math in visual contexts, the inability of the OCR to detect shapes useful for geometrical reasoning, and the lack of mathematical reasoning capabilities. An in-depth study of GPT-4V can be found in §H.
5 Qualitative Analysis
In §3.3, we observe that Multimodal Bard achieves the highest average accuracy on MathVista. Here, we analyze its predictions through human evaluation to understand its mode of success and failure. To do so, we ask the human workers, from Amazon Mechanical Turk (AMT), to study Bard’s predictions given the math question, its associated image, and the ground truth from MathVista dataset for 250 instances. Specifically, workers were instructed to decide whether the predictions contained the correct answer with the correct explanation. If the workers find that the model’s explanation is incorrect, they had to choose whether the wrong explanation was due to various failure modes such as incorrect reasoning with hallucination or wrong calculations. In our setup, we define hallucination as an introduction of incorrect facts, in the model explanation, that is not mentioned in the context of the image or question (e.g., in Figure 39 and Figure 40). More details can be found in §F.7.
We present the distribution of the quality of Bard’s predictions, judged by the human annotators, in Figure 4 (a). We find that 44.6% of the Bard’s predictions had incorrect answers with incorrect explanations. Interestingly, we observe that Bard responds with partial (6.8%) or completely (8.1%) incorrect explanations despite giving the correct answer to the input image and question, highlighting its failure to reach the correct answer for the wrong reasons. In Figure 4 (b), we present the distribution over possible reasons when Bard provides incorrect explanations. Notably, we find that 49.6% of its responses contain hallucinations. Our analysis highlights that hallucination is a major source of errors in the generative foundation models (Lu et al., 2023c; Ji et al., 2023). We also observe that the model responds with correct reasoning but either hallucinates (18.6%) or performs wrong calculations (19.5%) leaving an overall impression of being a wrong explanation.
We also present a few qualitative examples of Bard’s predictions. In Figure 5 (a), we find that Bard generates the correct answer with the correct explanation, including detecting the correct function (i.e., ) and analyzing its properties (i.e., injective) to answer the question. However, in Figure 5 (b), we observe that the model provides the correct answer (i.e., 12) but with an incorrect explanation (i.e., using the law of cosines when the question requires an understanding of the properties of isosceles triangles). We present more examples in §G.9. Overall, our analysis of Bard highlights its modes of failure in detail, which could guide future foundation model design to address these issues.
Augmented with external visual models, CoT GPT-4 and PoT GPT-4 are able to achieve comparable performance with Multimodal Bard. As shown in Figure 6 (a), provided with the accurate OCR text detected in the image, PoT GPT-4 accurately understands the structural information of the image and generates a code snippet to perform precise statistical reasoning. In Figure 6 (b), the caption provides some accurate descriptions of the image (e.g., ) along with hallucination (e.g., , the line passes through ) caused by the external Bard model. Although CoT GPT-4 predicts the correct answer given the partially correct information, the qualities of visual information augmented by external models have an impact on the accurate visual perception and thus the final mathematical reasoning performance. Examples in §G.10 show failure cases due to hallucination caused by external visual models.
Related Work
Several benchmarks (Amini et al., 2019; Cobbe et al., 2021; Mishra et al., 2022; Frieder et al., 2023) have emerged to assess the mathematical reasoning capabilities of LLMs, but most focus solely on text-based tasks. Current benchmarks, such as GSM-8K (Cobbe et al., 2021), exhibit performance saturation. Given the rise of LMMs Li et al. (2023a), there is a need for robust multimodal benchmarks in scientific domains. To address this gap, we introduce a math reasoning dataset that incorporates visual contexts.
VQA datasets (Antol et al., 2015; Gurari et al., 2018; Mobasher et al., 2022) gauge the visual reasoning abilities of LMMs. Recent studies explore assessing LMMs beyond natural images, including abstract scenes, geometry diagrams, figures, charts, documents, and synthetic images (Lu et al., 2021a; Kahou et al., 2017; Masry et al., 2022). In this work, we introduce new datasets (IQTest, FunctionQA, PaperQA) to create a holistic benchmark for evaluating mathematical reasoning.
Generative foundation models like GPT-3, ChatGPT, GPT-4, Claude, and LLaMA have enabled diverse task solutions without fine-tuning. Specialized pretraining methods like PixStruct (Lee et al., 2023), MatCha (Liu et al., 2022), and UniChart (Masry et al., 2023) enhance chart reasoning in visual contexts. Models like LLaVA, miniGPT4, InstructBLIP, and Bard leverage large-scale image-text data, while specialized versions, such as LLaVAR (Zhang et al., 2023d; Ye et al., 2023), emphasize document understanding and math comprehension. Recent works (Bitton et al., 2023; Yu et al., 2023) evaluate instruction-following and reasoning capabilities, underscoring the growing importance of generative foundation models in practical applications. We introduce MathVista as a benchmark to evaluate their math reasoning capabilities in varied visual contexts.
Conclusion
In this work, we introduce MathVista, a benchmark designed to systematically analyze the mathematical reasoning capabilities of state-of-the-art models in visually complex scenarios. Our evaluation of 12 prominent foundation models highlights that significant advancements have been made, especially with the GPT-4V model. However, a substantial gap of 10.4% still exists between GPT-4V, the best-performing model, and human performance. This disparity sets a clear direction for future research, emphasizing the need for models that can seamlessly integrate mathematical reasoning with visual comprehension. Moreover, our exploration of GPT-4V’s self-verification, self-consistency, and chatbot interactions offers valuable insights for future investigations.
References
Appendix A Detailed Related Work
Recently, numerous benchmarks (Amini et al., 2019; Cobbe et al., 2021; Mishra et al., 2022; Frieder et al., 2023) have been proposed to evaluate the mathematical reasoning capabilities of Large Language Models (LLMs). However, most of these are textual only (Lu et al., 2023c), despite a substantial amount of mathematical information and reasoning being encapsulated in visual modalities. Meanwhile, some datasets exhibit performance saturation; for instance, GPT-4 achieves 92.0% accuracy on GSM-8K (Cobbe et al., 2021), a dataset of grade-school mathematics questions. On the other hand, the recent rapid advancement of Large Multimodal Models (LMMs) necessitates the establishment of robust multimodal benchmarks. However, current multimodal reasoning benchmarks provide limited coverage of rigorous and scientific domains (Antol et al., 2015; Kembhavi et al., 2016; Kahou et al., 2017; Mathew et al., 2022), which are key components for creating general-purpose AI assistants. To bridge this gap, it is crucial to develop a robust math reasoning dataset that integrates visual contexts.
High-quality evaluation datasets and benchmarks are a cornerstone for assessing the progress of machine learning models to solve real-world tasks Liao et al. (2021). Prior studies such as VQA (Antol et al., 2015; Goyal et al., 2017), VizWiz (Gurari et al., 2018), and ParsVQA-Caps (Mobasher et al., 2022) assess the general-purpose visual question answering abilities of the LMMs, with or without task-specific training, on open-ended questions about images. In addition, there are several works that focus on evaluating specific skills of the LMMs beyond natural scenes, such as abstract scenes and shapes) (Antol et al., 2015; Lu et al., 2021b; Ji et al., 2022), geometry diagrams (Seo et al., 2015; Lu et al., 2021a; Chen et al., 2022a; Cao & Xiao, 2022), figures and charts (Methani et al., 2020; Masry et al., 2022; Kahou et al., 2017; Chang et al., 2022; Kafle et al., 2018), documents (text in images) (Singh et al., 2019; Mathew et al., 2022; Liu et al., 2023d), or synthetic images (Dahlgren Lindström & Abraham, 2022; Li et al., 2023d; Bitton-Guetta et al., 2023). Besides, there has been significant progress on developing datasets to judge LMMs on skills that require external knowledge (Schwenk et al., 2022; Shah et al., 2019), common sense reasoning (Zellers et al., 2019; Yin et al., 2021), scientific-knowledge (Lu et al., 2022; Kembhavi et al., 2017; 2016), medical understanding (Zhang et al., 2023c; Lau et al., 2018). In this work, we create new datasets (IQTest, FunctionQA, PaperQA) and subsequently design a benchmark for holistic evaluation of the math reasoning capabilities of the LMMs.
Recently, there has been a surge of generative foundation models (Bommasani et al., 2021) that are trained on web-scale data, such as GPT-3, ChatGPT, GPT-4, Claude, LLaMA, LLaMA-Adapter (Brown et al., 2020; OpenAI, 2022; 2023a; Anthropic, 2023; Touvron et al., 2023; Zhang et al., 2023a), with the ability to solve a wide range of downstream tasks (Wei et al., 2022a) without any task-specific finetuning. Prior work has focused on evaluating their abilities to respond to the queries from various disciplines, grounded in text, such as QA, math, medicine, coding and science (Bubeck et al., 2023; Nori et al., 2023; Chen et al., 2021; Fu et al., 2023; Sun et al., 2023; Wang et al., 2023b; Huang et al., 2023; 2022; Liu et al., 2023b; Zhang et al., 2023a). Prior work, such as PixStruct (Lee et al., 2023), MatCha (Liu et al., 2022), and UniChart (Masry et al., 2023), has focused on developing specialized pretraining recipe for improved math and chart reasoning in visual contexts.
On the vision-language side, there are several generative foundation models such as LLaVA, miniGPT4, InstructBLIP, Flamingo, LLaMA-Adapter V2, Multimodal Bard (Liu et al., 2023a; Zhu et al., 2023a; Dai et al., 2023; Alayrac et al., 2022; Awadalla et al., 2023; Gao et al., 2023; Google, 2023) that are trained on vast amount of paired (Schuhmann et al., 2022; Sharma et al., 2018; Lin et al., 2014) and interleaved image-text data (Zhu et al., 2023b). In addition, there has been recent development on specialized versions of these LMMs for document understanding where visual contexts require text recognition, math understanding being one of them (Zhang et al., 2023d; Ye et al., 2023). In recent times, there have been several works, such as Visit-Bench, LVLM-eHub, MMBench (Bitton et al., 2023; Yu et al., 2023; Liu et al., 2023c; Xu et al., 2023; Shao et al., 2023), that assess their instruction-following and reasoning capabilities. As the generative foundation models become more relevant to real-world applications, unlike prior work, we propose MathVista to benchmark their capabilities of math reasoning (logical, arithmetic, statistical) on a diverse set of visual contexts (word problems in images, natural scenes, geometrical shapes, and plots).
We have witnessed the remarkable abilities of large language models (LLMs), and their reasoning capabilities are further enhanced by promoting approaches such as chain-of-thought (CoT) (Wei et al., 2022b), program-of-thought (PoT) (Chen et al., 2022b), and inductive reasoning (Wang et al., 2023a; Tan & Motani, 2023). For example, the feasibility of using LLMs to solve the Abstraction and Reasoning Corpus (ARC) challenge has been verified using zero-shot, few-shot, and context-grounded prompting (Tan & Motani, 2023). In this paper, we evaluate LLMs using zero-shot, few-shot, CoT prompting, PoT prompting, as well as tool-augmented prompting, to explore their potential in solving mathematical reasoning in visual contexts on MathVista. Program-aided methods are widely used for mathematical reasoning due to their advancements in precise logical reasoning and arithmetic calculations (Drori & Verma, 2021; Tang et al., 2022; Drori et al., 2022). In this work, we have developed the LLM baselines with PoT.
Recently, OpenAI released GPT-4V, the multimodal version of GPT-4, which shows promising performance in vision-language reasoning. However, the fine-grained study of its strengths and limitations still remains underexplored. The recent work (Zhang et al., 2023b) contributes pioneering efforts in this field, studying whether large multimodal models (LMMs), like GPT-4V, execute vision and language tasks consistently or independently. As concurrent work, our paper provides, for the first time, a comprehensive quantitative and qualitative study of GPT-4V and other LLMs in mathematical reasoning within visual contexts.
Appendix B Limitations of the Benchmark
Our benchmark, MathVista, makes significant contributions by combining mathematical and visual tasks, a domain where existing models like GPT-4V have shown promise but also face challenges, especially in complex figure understanding and rigorous reasoning. While we have made strides in evaluating model performance, we acknowledge several limitations.
One limitation is the dataset coverage. While MathVista encompasses a broad spectrum of tasks and visual contexts, there may be gaps in the representation of certain types of mathematical problems and visuals. Furthermore, the dataset’s focus on mathematical reasoning within visual contexts, spanning specific domains like science and college-level math, necessitates a more labor-intensive process for collecting high-quality data compared to textual-only or general-purpose datasets. Thus, the scalability and generalizability of our benchmark to other domains remain a concern. Annotations were sourced from original data providers, resulting in only 85.6% of examples (Table 1) having annotations. Due to the heterogeneity of these sources, annotations lack a unified format and structure. For example, the annotations could be logic forms of the problem parsing from Geometry3K (Lu et al., 2021a), natural language solutions from TabMWP (Lu et al., 2023b), and theorems from TheoremQA (Chen et al., 2023). Given the rapid development in foundation models, our study focused exclusively on the most recent and prominent models.
In future iterations, our benchmark will be beneficial to encompass a broader array of problems and visual contexts, while also providing unified and comprehensive annotations. Our benchmark is part of an ongoing research process, and we are committed to maintaining the datasets, such as refining the potential data noise, in response to the community feedback. Also, we are committed to evolving the leaderboard in response to new models.
In conclusion, while there are limitations to our current approach, MathVista represents a significant step forward in the field. We are dedicated to continuously improving our benchmark to better understand and enhance the capabilities of AI in mathematical and visual reasoning.
Appendix C Data Collection Guidelines
Seven mathematical reasoning types are defined in Table 3.
C.2 Mathematical Reasoning Examples
C.3 Visual Context Types
C.4 Source Dataset Summary
The source datasets are summarized in Table 5.
Appendix D Data Collection Details
D.2 Human Labeling of Mathematical Problems
We developed an annotation tool, as illustrated in Figure 22, to enable expert annotators to label problems that involve mathematical reasoning. Annotators were trained using detailed instructions, as shown in Table 7, along with a variety of examples—positive ones that involve mathematical reasoning and negative ones that do not. We provided three labeling options:
Yes - This indicates that the problem involves mathematical reasoning.
No - This indicates that the problem does not involve mathematical reasoning.
Unsure - This option should be selected if it is uncertain whether the problem involves mathematical reasoning. (Annotators are advised to use this option sparingly.)
They may leave comments if they find anything incorrect or offensive for removal at a later stage.
In our study, we employed the Fleiss Kappa score to conduct an inter-annotator agreement analysis among three annotators tasked with labeling examples based on mathematical reasoning. The Fleiss Kappa score is a statistical measure used to evaluate the reliability of agreement between multiple raters, providing a quantifiable metric to assess the consistency across different annotators. A score of 1 indicates perfect agreement, while a score of 0 suggests no agreement beyond what would be expected by chance. Our analysis yielded a Fleiss Kappa score of 0.775, indicating a substantial level of consistency among the annotators. This high degree of agreement underscores the reliability of our annotation process and affirms the quality of the labeled data generated for our study.
D.3 Annotating Three New Datasets
D.4 Human Labeling of Mathematical Reasoning
Appendix E More Dataset Analysis
Apart from English questions, MathVista contains 6.57% non-English questions, including languages such as Chinese and Persian. The multilingual feature necessitates that models be capable of understanding and processing multiple languages to ensure accurate results across the dataset. As illustrated in Table 3, the average number of words in English questions within MathVista is 15.58, while the maximum number of words in a question reaches 213.
Figure 25 further elucidates the distribution of word counts, highlighting the diverse patterns of questions. MathVista features two types of questions: multiple-choice questions and free-form questions. For multiple-choice questions, the average number of choices is 3.4, while the maximum number of choices is 8. In the case of free-form questions, answers can be integers, floating-point numbers, or lists, which can be converted into a standard format. The standard settings in question and answer types facilitate consistent accuracy evaluation for existing models.
Source datasets in MathVista can be categorized into two types: math-targeted VQA datasets, which are originally proposed for assessing mathematical reasoning, and general VQA datasets, which address visual reasoning in everyday scenarios. The distribution proportions of these two categories (55.4% vs. 44.6%, as illustrated in Figure 26) within MathVista enable a balanced examination of mathematical reasoning in both domain-specific and general-purpose applications. The distribution of the five tasks contained within MathVista is visualized in Figure 27. The relatively balanced distribution of these tasks enhances the benchmarking robustness that our dataset provides.
The datasets within MathVista are categorized into four distinct grade levels: elementary school, high school, college, and not applicable, each representing a different level of reasoning complexity and contextual application. The elementary school category aligns with the typical mathematical curriculum of elementary education, introducing basic topics such as arithmetic operations and introductory geometry. High school level questions delve into more complex mathematical concepts such as algebra, geometry, and introductory calculus. The college category encapsulates the highest level of complexity, featuring questions on advanced mathematical and scientific concepts like calculus, linear algebra, and physics. Questions without specific grade levels are categorized as not applicable.
The distribution of questions across these grade levels is visualized in Figure 28. This structured categorization enriches the diversity of MathVista, providing a meaningful framework for evaluating and benchmarking the mathematical and visual reasoning capabilities of various models across different educational contexts, thereby assessing their practical utility and educational relevance.
The datasets within MathVista encompass over 10 different visual contexts (with the distribution shown in Figure 29), crucial for evaluating models’ ability to interpret and reason across diverse visual information. Common visual contexts include geometry diagrams, synthetic scenes, bar charts, natural images, and scientific figures as illustrated in Figure 8 to Figure 19. Less frequent, yet equally important visual contexts such as medical images, word clouds, map charts, radar charts, violin plots, and heatmap charts are depicted in Figure 20 and Figure 21. These visual contexts, ranging from common to specialized representations, challenge the models to decode and reason with varying visual information, contributing to a more robust and comprehensive evaluation. The diversity in visual contexts enriches MathVista, enhancing the benchmarking robustness and providing a solid foundation for understanding the practical utility and domain-specific performance of various models across different domains and applications.
The datasets within MathVista encompass a spectrum of seven distinct mathematical reasoning types, facilitating a thorough evaluation of models’ mathematical reasoning capabilities. Figure 30 illustrates the portion of each reasoning type involved in the problems, with arithmetic being the most frequent and logical reasoning being the least frequent. This distribution reflects the varying degrees of mathematical reasoning required across different problems. Figure 31 further delineates the distribution of reasoning types, showcasing a mean of 1.45. The sparse distribution observed aids in the precise analysis of each type’s performance by the models, providing a nuanced understanding of their strengths and weaknesses across different mathematical reasoning domains. This structured representation of mathematical reasoning types within MathVista not only enriches the dataset but also significantly contributes to a more robust and comprehensive evaluation of models, aiding in the identification of areas for improvement and the development of more proficient mathematical reasoning models.
Appendix F More Details on the Setup
We employ a strategy where the most frequent answers in the testmini set are utilized as predictions for various question and answer types. For multiple-choice questions, the most frequent option is selected based on the number of available options. For instance, option is chosen for questions with two options, aligning with the answer distribution in testmini. Similarly, for questions requiring an answer type of integer, a floating number with one decimal place, a floating number with two decimal places, or a list, we use , , , and $$ respectively, in accordance with the answer distribution observed in testmini.
F.2 Prompt for Answer Extraction
The prompt used to instruct GPT-4 for answer extraction is illustrated in Table 8.
F.3 Prompts for Response Generation
F.4 Prompt for Caption Generation
We instruct Multimodal Bard to generate a detailed description for an input image, aiming to augment current LLMs with visual understanding capabilities. The prompt is shown in Table 10.
F.5 Model Hyperparameters
The hyperparameters for the experiments in §3.2 are set to their default values unless specified otherwise. Table 11 and Table 12 detail specific generation parameters for the various large language models (LLMs) and large multimodal models (LMMs) we evaluated, respectively.
F.6 Human Performance
We conducted a study to evaluate human performance on the testmini subset of the MathVista, utilizing Amazon Mechanical Turk (AMT). Each question from the testmini subset was assigned to five annotators, all of whom have a history of completing more than 5,000 HIT tasks and boast an acceptance score higher than 0.99, to ensure the quality of the results. The study comprised five test questions and two qualification questions, which were to be answered within a 20-minute timeframe. The qualification questions consisted of elementary math word problems requiring basic arithmetic operations (e.g., addition and subtraction). Only annotators who successfully answered the qualification questions were deemed eligible for the study, and their responses were included in the final analysis. Additionally, annotators were requested to provide information regarding their highest level of educational attainment. We retained the results exclusively from annotators who had achieved a high school diploma or higher, as 30.9% of the problems in MathVista are of high-school level difficulty and 10.8% correspond to college-level curricula.
F.7 Multimodal Bard Assessment Task
A screenshot of our AMT worker interface, utilized for the Multimodal Bard assessment task, is provided in Figure 32. The workers were compensated at a rate of $18 per hour.
Appendix G More Experimental Results
Table 13 reports the accuracy scores of two heuristic baselines, two leading augmented LLMs (CoT GPT-4, PoT GPT-4), and one leading LMM (LLaVA-LLaMA-2-13B) on the test subset. The minor differences between scores on the test subset and the testmini subset, as shown in Table 2, suggest that testmini effectively mirrors the test subset, serving as a valuable evaluation subset for model development, especially for those who have limited computing resources.
G.2 Scores for Math Reasoning Types
The accuracy scores across seven mathematical reasoning categories are reported in Table 2, with primary baselines highlighted in Figures 1 and 33. GPT-4V outperforms other baseline models in most mathematical reasoning categories, except for logical reasoning and numeric commonsense reasoning. Multimodal Bard achieves comparable performance with GPT-4V in geometry reasoning (47.8% vs. 51.0%) and algebraic reasoning (46.5% vs. 53.0%), highlighting its enhanced abilities in comprehending geometry diagrams and performing algebraic calculations.
Among open-source LMMs (ranging from IDEFICS to LLaVA), LLaVA achieves the best overall accuracy on MathVista and the highest fine-grained scores for problems in geometry reasoning, logical reasoning, and statistical reasoning. However, these scores still substantially lag behind GPT-4V and Multimodal Bard, indicating a gap in the overall effectiveness of these open-source models compared to more advanced proprietary systems. Despite this, LLaMA-Adapter-V2, tied with LLaVA, outperforms GPT-4V by 2.7% in logical reasoning, and InstructBLIP beats GPT-4V by 0.3% in numeric commonsense, suggesting that specific enhancements in open-source models can lead to superior performance in certain niches. LLaVAR, being on par with Multimodal Bard, which is specifically designed to enhance capabilities in detecting OCR texts and symbols from various forms, including scientific domains, further illustrates the potential of targeted improvements in open-source LMMs to achieve competencies that rival or even exceed those of their proprietary counterparts in specialized areas.
CoT GPT-4, augmented with OCR texts and Bard captions, performs well in scientific reasoning, achieving a gain of 26.2% over random chance, showcasing its superiority in domain-specific knowledge. This performance suggests a significant trend (Shen et al., 2023; Lu et al., 2023a) where the integration of specialized functionalities, such as OCR text recognition and advanced captioning, into LLMs enhances their applicability and accuracy in specific domains. PoT GPT-4 outperforms Multimodal Bard in categories such as arithmetic reasoning, logical reasoning, numeric commonsense reasoning, and statistical reasoning. This superior performance is attributed to its ability to generate high-quality codes for precise mathematical reasoning, illustrating the effectiveness of integrating advanced coding capabilities into language models for complex problem-solving tasks.
G.3 Scores for Various Visual Contexts
Figure 34 illustrates the accuracy scores of leading baselines on MathVista across a diverse range of visual contexts. Remarkably, GPT-4V outperforms human performance in visual contexts of function plots, geometry diagrams, scatter plots, tables, and other types, which aligns with its superiority in terms of related mathematical reasoning types. Other foundation models trail behind humans in visual perception and reasoning across most visual context categories. Multimodal Bard demonstrates comparable performance to humans in questions with a visual context of geometry diagrams, showcasing its promising capabilities in recognizing geometric shapes and relationships. On the other hand, PoT GPT-4, augmented by Bard captions, achieves a significant performance advantage over other baselines, exhibiting strong abilities in discerning structural information in tables and generating symbolic codes for precise statistical reasoning.
G.4 Scores Across Different Grade Levels
Figure 35 displays the average accuracy scores across different grade levels (elementary school, high school, and college) for the leading foundation models, as well as random chance and human performance. Humans exhibit the highest performance on questions at the elementary school level (70.4%), while they fare the worst on college-level questions (52.6%) within MathVista. Foundation model baselines exhibit varying performance behaviors: they achieve better accuracy scores on high school level questions compared to the other two categories.
In addressing elementary school problems, the performance gap between human performance and the best-performing model, GPT-4V, is notably the largest when compared to other grade levels. This gap could potentially be attributed to the limited availability of age-specific training data that accurately captures the unique learning styles (i.e., rich with abstract scenes) of elementary school students. On the other hand, GPT-4V demonstrates an improvement of 20.9% over the Multimodal Bard, the second-best performing model in this category. This improvement suggests that while GPT-4V still lags behind human performance, its ability to tackle elementary-level problems in visually intensive settings has been significantly enhanced.
For high school problems, GPT-4V, with a score of 61.8%, outperforms human performance, which stands at 58.2%. Additionally, the second-best performing model, Multimodal Bard, with a score of 50.3%, is on par with human performance. This disparity might be attributed to the training regimen of the models, which perhaps aligns well with the high school curriculum.
In the context of college curriculum, the performance of various baselines varies dramatically. GPT-4V demonstrates performance comparable to that of humans. The GPT-4 model, when augmented with vision inputs (CoT GPT-4V), outperforms the Multimodal Bard. Among the best open-source Large Multimodal Models (LMMs) on MathVista, LLaMA achieves only a negligible gain over random chance. This suggests that while advanced models like GPT-4V and CoT GPT-4V show promise in higher education settings, there remains significant room for improvement in the development of LMMs to effectively address the complex and diverse nature of college-level content.
G.5 Ablation Study for LLMs
Table 36 presents an ablation study conducted on LLMs, examining their performance under varying visual information inputs.
G.6 LLMs with Different Shots
We explored whether LLMs and Augmented LLMs can benefit from larger numbers of few-shot examples on MathVista, with results reported in Figure 37. In the question-only input setting (a), both Claude-2 and ChatGPT suffer from a performance drop, suggesting that they are more sensitive to the bias in demonstrations, especially in the absence of visual inputs. There is a marginal improvement of 1.4% when the shot number increases from 2 to 4 for GPT-4. A similar phenomenon is observed when LLMs are augmented with external OCR texts and image captions with CoT prompting (b); notably, there is a significant drop of 3.4% when the shot number increases from 2 to 4 for CoT Claude-2. With PoT prompting (c), LLMs like ChatGPT and GPT-4 can obtain gains of 3.4% and 1.4%, respectively, with the shot number increasing from 2 to 4. Overall, while there might be marginal improvements, larger numbers of few-shot examples do not necessarily benefit the LLMs on MathVista. In some settings, LLMs suffer from unstable performance drops. This further indicates that the quality of the augmented information plays a more important role for augmented LLMs.
G.7 LMMs with Different Shots
We conducted an initial study on the few-shot learning ability of the Large Multimodal Model (LMM), specifically IDEFICS (Laurençon et al., 2023), on MathVista. As shown in Figure 38, there is a modest improvement with increased shot numbers, suggesting potential benefits of few-shot learning for LMMs on MathVista.
However, recent studies highlight the instability of LMMs in few-shot settings. For instance, a significant accuracy drop was observed in models like BLIP-2 (Li et al., 2023b) and InstructBLIP (Dai et al., 2023) when applying 4-shot in-context learning in common sense reasoning tasks (Li et al., 2023c). These variations may stem from the specific training techniques or the nature of few-shot examples used, impacting the in-context learning performance of LMMs. Given the rapidly evolving landscape of LMMs, the consistent benefits of few-shot learning remain an open question.
G.8 Hallucinations in Model Explanations
G.9 More Examples for Multimodal Bard
G.10 Comparisons of Different Models
Appendix H A Comparative Study of GPT-4V, Bard, and Other Models
GPT-4 with vision (GPT-4V) is the multimodal version of GPT-4 that is instructed to understand multiple modalities such as texts and images. Due to its remarkable improvements over other AI models (§3.3 and §3.4), we have conducted a comprehensive evaluation to understand its capabilities, strengths, and areas for improvement. Our findings not only validate GPT-4V’s various problem-solving skills but also shed light on developing general-purpose multimodal AI agents.
Given that GPT-4V does not offer API access, we have performed manual evaluations using the playground platformhttps://chat.openai.com/. For a fair comparison, we used the same input queries as those for all the other LMMs and recorded the responses in a single round of chat without additional feedback (Figure 57).
H.2 Leaderboard Scores
The leaderboard in Figure 58 highlights GPT-4V’s substantial advancements over the current LLM and LMM baselines. Notably, there is a 15.1% improvement over the second-best performing Multimodal Bard model. However, a significant gap of 10.4% still exists between GPT-4V and human performance, indicating plenty of room for further improvement by developing new LMMs and tool-augmented LLMs.
H.3 Abilities in Mathematical Reasoning
This section compares the mathematical reasoning ability of GPT-4V with that of other LLMs on MathVista, including LLaMA-Adapter-V2-7B (LLaMA-Adapter-V2 for simplification), LLaVA-LLaMA-2-13B (LLaVA for simplification), and Multimodal Bard.
Algebraic reasoning problems on MathVista require understanding the function plots from figures and inferring their properties. As shown in Figure 1, GPT-4V demonstrates outstanding capabilities in algebraic reasoning, surpassing all competing models and even humans. For instance, GPT-4V accurately identifies the function plot by its equation and subsequently infers its correct properties (Figure 59). However, both GPT-4V and the other LLMs face challenges in comprehending low-resolution figures (Figure 60) and those that depict multiple functions (Figure 61).
H.3.2 Arithmetic Reasoning
Arithmetic reasoning problems in MathVista require accurate fundamental operations in conjunction with understanding diverse visual contexts. As illustrated in Figure 1, GPT-4V exhibits a significant improvement in arithmetic reasoning compared to existing models. For instance, some LLMs struggle with basic arithmetic tasks, such as determining the difference between two values in a bar chart (Figure 62) or computing the probability based on simple statistical data (Figure 63).
H.3.3 Geometry Reasoning
In geometry reasoning, the performance of GPT-4V is comparable to that of humans on MathVista, as demonstrated in Figure 1. Figure 64 and Figure 65, respectively, present two geometry reasoning problems: one at an elementary level and the other at a college level. For both problems, GPT-4V produces the correct answers accompanied by detailed explanations.
H.3.4 Logical Reasoning
Logical reasoning problems represent a different type of question in MathVista. Solving these problems requires abstract thinking to deduce the underlying patterns of numbers or shapes from figures. Current foundation models struggle to effectively tackle logical reasoning problems: GPT-4V achieves only 21.6% accuracy in logical reasoning, which is a modest improvement of 8.1% over random chance, as shown in Table 2. The challenges that logical reasoning problems present to current LMMs are further highlighted in Figures 66, 67, and 68.
H.3.5 Numeric Commonsense Reasoning
Problems involving numeric commonsense reasoning on MathVista require commonsense knowledge about daily objects and celebrities to answer visual questions. However, these problems present significant challenges to existing foundation models, including GPT-4V, as depicted in Figure 1. For instance, Multimodal Bard struggles to understand the optical illusion in an image (Figure 69) and to infer the age gap between two celebrities from another image (Figure 70). Figure 71 poses a question about the maximum volume a beaker can measure. However, GPT-4V lacks commonsense knowledge regarding the use of a beaker, resulting in an incorrect prediction.
H.3.6 Scientific Reasoning
Scientific reasoning represents a distinct mathematical reasoning ability within our MathVista. To tackle problems in this area, a model must not only accurately interpret domain-specific information from figures, but also possess the necessary in-domain knowledge to reason rigorously on scientific topics. Figure 1 shows that GPT-4V substantially outperforms the other foundation models. This superiority is further illustrated by the examples in Figures 72 and 73. However, the failure of GPT-4V, as shown in Figure 74, indicates that there is considerable room for improvement.
H.3.7 Statistical Reasoning
In MathVista, problems encompass a variety of charts, plots, and graphs designed to assess the statistical reasoning capabilities of foundation models. As demonstrated in Figure 1, GPT-4V shows strong statistical reasoning ability. For instance, GPT-4V produces accurate answers for the format-rich table in Figure 75 and the data analysis table in Figure 76.
H.4 Abilities Across Visual Contexts
This section compares the reasoning abilities of GPT-4V with other large multimodal models (LLMs) on MathVista, considering various types of visual contexts. Models used for comparison include LLaMA-Adapter-V2-7B (simplified as LLaMA-Adapter-V2), LLaVA-LLaMA-2-13B (simplified as LLaVA), and Multimodal Bard.
Based on Figure 1, current foundation models lag behind human performance in mathematical reasoning in abstract scenes by a substantial margin. Consider the problems in Figures 77 and 78 that are derived from math word problems found in elementary school curricula. Despite their advanced capabilities, foundation models such as Multimodal Bard and GPT-4V fail to produce the correct responses.
H.4.2 Bar Chart
As shown in Figure 1, foundation models, including GPT-4V, significantly underperform humans in mathematical reasoning when bar charts serve as the visual context. Neither Multimodal Bard nor GPT-4V can solve the problems depicted in Figures 79 and H.4.2, which do not need complex understanding and reasoning.
H.4.3 Function Plot
GPT-4V outperforms other baselines on problems related to function plots and even exceeds human performance. Figures 81 and 82 show questions with digital and hand-drawn function plots, respectively. In both cases, GPT-4V accurately identifies their functions and infers the correct properties.
H.4.4 Geometry Diagram
Geometry diagrams are a distinct type of visual context in MathVista. To answer questions involving these diagrams, a model must comprehend the fine-grained details, including symbols, variables, and relations from the figures. Additionally, it should apply associated theorems before executing calculations to produce final responses. GPT-4V surpasses other models and even humans due to its superior capabilities in geometry recognition and reasoning. In the examples shown in Figures 83 and 84, GPT-4V delivers the correct results through the application of relevant theorems and subsequent calculations.
H.4.5 Line Plot
As evidenced by Figure 1, current models such as GPT-4V do not perform as well as humans in mathematical reasoning involving line plots. We speculate that the low performance is mainly due to the difficulty in detecting OCR text in the figures and accurately grounding the values, as illustrated by the examples in Figures 85 and 86.
H.4.6 Natural Image
MathVista includes questions that require numeric and spatial reasoning based on text and objects in natural images. If models have limited abilities to recognize text (OCR), as shown in Figure 87, or to identify visual objects, as in Figure 88, they are unlikely to generate correct answers to visual questions.
H.4.7 Puzzle Test
Math reasoning with puzzle text figures is challenging for current AI foundation models because interpreting these figures requires discerning underlying patterns from sets of shapes, as illustrated in Figure 89, and numbers, as in Figure 90. There is plenty of room for improvement.
H.4.8 Scatter Plot
A scatter plot is a graphical representation of data points on a two-dimensional axis, where each point represents a value from the dataset. MathVista includes the reasoning task that requires comprehending scatter plots taken from daily life and academic papers, as shown in Figures 91 and 92. Although GPT-4V outperforms other LMMs, such as Multimodal Bard, and even humans in overall accuracy (Figure 1), it often fails in the cases where fine-grained understanding is required, as in Figure 92.
H.4.9 Scientific Scene
Answering questions based on scientific scenes poses a challenge in aligning the scientific concepts present in the question text and those in the accompanying figures. GPT-4V demonstrates its superior ability to reason about scientific scenes compared to other models, as evidenced in Figure 1. In the example of Figure 93, GPT-4V adeptly identifies two organisms in the food web and elucidates their relationship. In another instance, shown in Figures 94 and 95, both Multimodal Bard and GPT-4V are able to use knowledge in the physical domain to effectively ground the symbols and variables depicted in the figure.
H.4.10 Synthetic Scene
Problems involving synthetic scenes require a nuanced understanding of visual objects, such as the numbers, attributes, and positions of these objects, as shown in Figures 96 and 97. Although GPT-4V demonstrates notable advancements over other models, such as Multimodal Bard, it still falls significantly short of human performance, as shown in Figure 1.
H.4.11 Table
Tables serve as a powerful tool to present and summarize large amounts of data in a comprehensible manner. In particular, GPT-4V has shown significant advancements over other foundation models and even surpasses human performance on table-related reasoning tasks, as shown in Figure 1. The example in Figure 98 shows a complex table taken from an academic paper. GPT-4V can accurately pinpoint the target cells among numerous rows and columns. Figure 99 shows a QA task in which the answer needs to be derived from the table regarding the push-up competition. GPT-4V is the only model that can produce the correct answer.
H.4.12 Other Visual Contexts
On the reasoning tasks using other visual contexts, GPT-4V achieves a higher overall accuracy than all the other models, as depicted in Figure 1. For instance, GPT-4V is the only model that is capable of generating the correct answer to the question regarding a violin plot, as shown in Figure 100.
H.5 Self-Verification in GPT-4V
Self-verification is a social psychological theory asserting that people desire others to perceive them as they see themselves. Consequently, individuals will take active measures to ensure that others view them in ways that confirm their stable self-concepts (Talaifar & Swann, 2020).
Interestingly, in our experiments, GPT-4V demonstrates an ability similar to self-verification. The model can inspect its own behaviors during the course of reasoning and can take active actions to correct its mistakes. Note that self-verification we discuss here differs from several recent works on improving the model’s outputs based on external feedback (Peng et al., 2023) or additional generations (Yang et al., 2023b). The examples in Figures 101 and 103 show that GPT-4V, on its own, can inspect a set of candidate answers and identify the one that is valid and meets all the given constraints. The multi-step reasoning example in Figure 102 shows that GPT-4V can verify the validity of (the result of) each reasoning step, and explore alternative approaches if any invalid (intermediate) result is detected (e.g., a negative value for length).
Although self-verification does not guarantee an accurate response even after multiple tries, especially when applying GPT-4V to visual perception or mathematical reasoning in intricate scenarios (see Figure 104), it is instrumental in improving the model performance on MathVista. We also found that GPT-4V’s self-verification is weaker for non-English tasks, such as Mandarin, as shown in Figure 105. It is also worth noting that self-verification does not emerge in other foundation models we studied, or at least it is not as robust as that of GPT-4V. As shown in Figure 106, Multimodal Bard first attempts a natural language solution, followed by a program-assisted one for verification. However, the program-aided solution leads to a different and incorrect prediction.
The emergent ability of self-verification highlights GPT-4V’s potential in solving rigorous reasoning and theorem-proving tasks. One of the most exciting research topics for future work is to develop a mechanism that allows the model to activate self-verification consistently at the right time and to use a set of alternative approaches that maximize the success rate of task completion.
H.6 Self-Consistency for GPT-4V
Self-consistency (Wang et al., 2022) is a decoding strategy for chain-of-thought prompting (Wei et al., 2022b). A diverse set of reasoning paths is sampled, and the most consistent answer is selected as the final prediction. Moving beyond vanilla greedy decoding, this method resorts to the inherent coherence and reliability of multiple reasoning trajectories to produce a more trustworthy conclusion. Self-consistency has been widely employed in LLMs for complex reasoning tasks, such as math word problems and commonsense reasoning.
In our experiments, we validated the effectiveness of using self-consistency for GPT-4V on MathVista. Given a question and context, we ran GPT-4V multiple times to obtain a set of different reasoning paths and then selected the most frequent answer as the final prediction. We found that self-consistency is instrumental in rectifying visual perception errors (Figure 107), correcting calculation mistakes (Figure 108), and mitigating hallucinations (Figure 109). In comparison, self-consistency is less effective when GPT-4V has difficulties in interpreting complex visual contexts (Figures 110, 111) or extracting salient information from images (Figure 112).
H.7 GPT-4V for Multi-Turn Human-AI Interaction
This section investigates the use of GPT-4V for multi-turn human-AI interaction on MathVista, as exemplified in the goal-directed dialog in Figure 113.
We found that GPT-4V is effective in engaging multi-turn goal-directed conversations with users. In particular, GPT-4V can make good use of hints (e.g., user feedback or responses) to guide the conversion to generate desirable results. For instance, it can (1) rectify visual perception errors based on hints (Figure 114), (2) reassess reasoning steps and calculations (Figure 115), (3) correct misinformation using user-provided domain-specific knowledge (Figure 116), and (4) aggregate intricate contexts over multiple turns in a human-AI conversation (Figures 117 and 118).
We also observed failure cases in our evaluation. For instance, GPT-4V struggles to generate correct responses when questions and user hints are ambiguous (Figure 119), or when the model fails to understand abstract shapes and concepts visually (Figure 120). These failures motivate the development of more powerful, conversational foundation models.