G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, Lingpeng Kong

Introduction

Large language models (LLMs) exhibit human-like proficiency in reasoning (Wei et al. 2022; Wang et al. 2022; Zhou et al. 2022) and generation (Ouyang et al. 2022; Touvron et al. 2023), which encourages extensive research on their application in mathematical problem solving (Fu et al. 2023; Gou et al. 2023; Yue et al. 2023b; Luo et al. 2023; Zhao et al. 2023a; Zhao et al. 2023b; Jiang et al. 2023). These problems often require highly sophisticated and symbolic reasoning capabilities, often considered impossible to solve before the era of LLMs.

It is an intuitive approach to use LLMs for mathematical reasoning problems presented in a textual form. Nevertheless, a substantial proportion of mathematical reasoning problems necessitate the comprehension of geometric information. Moreover, even when certain problems do not overtly pertain to geometric information on the surface, the integration of geometrical-based methods often holds significant practical implications (e.g., analytic number theory). With the advent of GPT-4V (OpenAI 2023), Gemini Gemini, a concurrent work, was released one week before our submission. Consequently, our work is primarily benchmarked against GPT4-V and other MLLMs. (Google 2023), and numerous multi-modal large language models (MLLMs) (Zhu et al. 2023; Liu et al. 2023; Dai et al. 2023; Li et al. 2023; Bai et al. 2023; Lai et al. 2023; Gao et al. 2023b; Pi et al. 2023b), recent work has progressively looking into employing MLLMs to tackle geometric reasoning problems in mathematics (Yang et al. 2023; Lu et al. 2023; Yue et al. 2023a).

However, we have observed that even with the most advanced MLLMs, current systems still exhibit limitations in addressing geometric problems due to challenges in accurately comprehending geometric figures. For instance, as demonstrated in Figure 1, GPT4-V often produces inaccurate descriptions for geometric figures. Specifically, the model struggles with understanding the relationships between fundamental elements like points and lines, and in accurately interpreting elements such as the degree of an angle. We presume that the underlying reason for this may be the fact that these MLLMs are typically trained with images and descriptions from the general domain, and the ability to understand such semantics differs significantly from that required for geometric reasoning.

To address this issue, one of the most direct and effective approaches is to enhance current MLLMs by augmenting them with data containing high-quality descriptions of geometric information (Ye et al. 2022a; Meng et al. 2022). However, a significant challenge arises from the limited size of the largest publicly available geometric problem dataset, which contains only a few thousand question-answer pairs. Additionally, the current datasets lack descriptions of geometric images and exhibit a limited range of problem-solving methods, which constrains the model’s ability to understand basic geometric elements and affect its problem-solving capabilities.

In this paper, we propose to synthesize geometric visual-text data leveraging existing datasets via text-only LLMs (e.g., ChatGPT). More specifically, we utilize the geometry characteristic to construct a multi-modal geometry dataset, building upon existing datasets. The data generation process involves incorporating utilizing uniqueness of geometric logic form, geomertic representation uniqueness, geometric scalability, etc (as shown in Figure 2). We term our generated dataset Geo170K, which contains around 60,000 geometric image-caption pairs and more than 110,000 question-answer pairs. This dataset is 28 times larger than GeoQA+, greatly expanding the coverage of geometric problems. With our collected Geo170K, we derive G-LLaVA, a MLLM capable of solving geometric problems, surpassing SOTA MLLMs by a large margin. Specifically, G-LLaVA-13B outperforms LLaVA-13B by 27.4 on GPS minitest split of MathVista (Lu et al. 2023). In addition, with only G-LLaVA-7B, it is able to surpass the powerful GPT4-V on the geometry problem solving questions. Code and data will be available at https://github.com/pipilurj/G-LLaVA.

Related Work

Recent years have witnessed transformative advancements in the development of large language models (LLMs), characterized by a series of pioneering studies (Brown et al. 2020; Scao et al. 2022; Chowdhery et al. 2022; Smith et al. 2022; Hoffmann et al. 2022; Ouyang et al. 2022; Touvron et al. 2023; Bai et al. 2022). These breakthroughs have significantly elevated the capabilities of language understanding and generation, showcasing near-human proficiency across diverse tasks. Concurrently, the success of LLMs has inspired explorations into vision-language interaction, leading to the emergence of multi-modal large language models (MLLMs) (Liu et al. 2023; Li et al. 2023; Dai et al. 2023; Zhu et al. 2023; Dai et al. 2023; OpenAI 2023; Bai et al. 2023; Su et al. 2023; Gao et al. 2023b). These models have exhibited remarkable capabilities in synthesizing detailed descriptions and engaging in dialogue based on visual inputs. However, we observe that even the state-of-the-art MLLMs face challenges in resolving geometric problems using diagrams and figures.

The Geometry problem reasoning is an challenging visual mathematical reasoning problem. Early efforts by Seo et al. 2015; Sachan et al. 2017; Alvin et al. 2017; Sachan and Xing 2017 focused on creating datasets through manual efforts. More recent approaches have introduced enhanced methods and datasets, including Geometry3K (Lu et al. 2021), GeoQA (Chen et al. 2021), GeoQA+ (Cao and Xiao 2022), UniGeo (Chen et al. 2022), UniMath (Liang et al. 2023), and SCA-GPS (Ning et al. 2023), aiming to improve both performance and explainability. However, the scale of current datasets remains limited, and the performance of traditional models in this domain has not achieved the level observed in other areas of mathematical problem solving, particularly when compared to methods that utilize large language models for solving math word problems (Cobbe et al. 2021; Wei et al. 2022; Gou et al. 2023).

Bootstrapping data from pretrained models has long been an active area of research. Ye et al. 2022a; Meng et al. 2022 generates training data using pretrained language models such as GPT-2 for classification tasks. Gao et al. 2023a improves the quality of generated dataset via bi-level approach. Ye et al. 2022b utilizes influence function to select in-context examples to aid data generation. Recently, automatic data generation becomes more ubiquitous with the advent of powerful LLMs such as ChatGPT, a line of recent works utilize ChatGPT-generated data to perform instruction tuning (Wang et al. 2023; Peng et al. 2023; Taori et al. 2023; Liu et al. 2023; Zhu et al. 2023; Bai et al. 2023; Pi et al. 2023a; Su et al. 2023; Yu et al. 2023; Chen et al. 2023; Zhang et al. 2023).

Observation

We observe that most state-of-the-art (SOTA) MLLMs, although being adept at understanding daily visual scenes, have difficulty in comprehending geometric figures, even if they are simple and straightforward for humans. In Figure 1, we demonstrate the descriptions generated by SOTA MLLMs for geometric figure. We observe that severe hallucination exists in all the generated descriptions.

More specifically, we find GPT4-V has difficulty understanding relationships between basic elements like points and lines, and also struggles with precisely interpreting these elements themselves (such as the angle B in Figure 1). Furthermore, smaller MLLMs like LLaVA1.5 and MiniGPT4 demonstrate even greater difficulty in accurately identifying the types of geometric shapes present in a figure.

This inadequacy in interpreting geometric diagrams may be one of the major causes for the failure in solving geometric problems. In contrast, actual geometric diagrams typically exhibit clear and well-defined relationships among their elements. This geometry characteristic can be utilized to develop datasets that help mitigate the above issues and mitigate hallucination.

Geometric Data Generation

While previous efforts have been made to address multi-modal geometry problems (Chen et al. 2021; Chen et al. 2022; Cao and Xiao 2022), the availability of geometry datasets remains limited. The key limitations of existing datasets are threefold: (1) limited data volume (a few thousands for the largest dataset), (2) absence of detailed descriptions for geometric images, and (3) a lack of diversity in problem-solving methodologies and answer pathways. This limitation presents challenges for MLLMs in accurately understanding geometric elements and providing precise geometric solutions.

To address this issue, we utilize the geometry characteristic to construct a multi-modal geometry dataset based upon existing dataset. This dataset includes two parts: an alignment dataset to provide MLLMs with fundamental geometric knowledge and an instruction-tuning dataset to improve the assistant’s ability to understand user instructions and generate accurate geometry solutions.

Image-caption datasets play a significant role in training MLLMs for understanding the context of images, which is essential for aligning image and text modalities. In the field of geometry, there is a lack of such datasets that offer detailed descriptions of geometric diagrams. To address this issue, we propose the generation of image descriptions from labeled question-answer (QA) pairs, as illustrated in Table 1. In particular, we use text-only ChatGPT 3.5 to create image captions based on these human-labeled QA pairs, which can be considered as a type of inverse information recovery. This approach leverages the strong understanding ability of ChatGPT to produce descriptions for geometric diagrams.

1.2 Contrastive QA Pairs for Basic Elements

Our approach also involves generating QA pairs to facilitate the comprehension of geometric diagrams, focusing primarily on their basic elements. The process begins with the interpretation of human-labeled logical forms on Geometry3k (Lu et al. 2021). We employ text-only ChatGPT to convert these logical forms into clear descriptions that cover various geometric elements such as shapes, lines, and points, and their relationships.

After creating these diagram descriptions, the model begins to produce contrastive QA pairs. These pairs are designed to examine different aspects of the diagrams. Questions may explore the presence of certain geometric elements (e.g., "Are there triangular shapes in the diagram?") or check the accuracy of the relationships described (e.g., "Is point D the lies on line BC?"). This method enables the model to comprehend geometric concepts and to analyze and interpret the details in geometric diagrams accurately. The generation example is shown on Table 2.

2 Geometric Instruction Data

After performing alignment leveraging the constructed alignment data, the model is able to better interpret the geometric diagram (Figure 1). However, they are still limited at solving geometric problems. Therefore, we construct an instruction tuning dataset based on existing datasets with the help of powerful LLMs. Specifically, we design a series of strategies to expand the question-answer pairs in existing datasets. The resulting dataset contains more than 110k QA pairs, which is the largest public geometric QA dataset available. We will introduce the proposed strategies in detail below.

As shown in Table 5, we replace the specific values in the original QA pairs with unknown variables and prompt the LLM to construct the solution by solving equation. Such data is helpful for the MLLM to generalize its understanding of the problem, which enables it to apply the similar reasoning and solution steps to different scenarios. The abstraction of the problem by using variables and solving equation helps the LLM focus on the underlying mathematical concepts and relationships, rather than getting caught up in specific numerical values.

2.2 Value Scaling (VS)

As shown in Table 4, we augment the data by scaling the length values in the QA pairs. Note that for the same diagram, the QA pair is still correct if all the lengths in a geometric problem are scaled simultaneously. However, note that it is not the case for quantities such as angles. When different scaling of values are applied, the LLM becomes more flexible in handling different numerical inputs. Involving a range of values that extends beyond the initial training dataset aids in refining the model’s computational and reasoning capabilities, thereby contributing to its generalizability.

2.3 Re-Formulating Condition as Unknown (RCU)

Motivated by(Weng et al. 2023; Yu et al. 2023), we design new multi-modal QA pairs that ask questions backwards, as shown in Table 6. Specifically, we reformulate questions to ask for the values originally present in the condition, and retain the generated data with correct answer only. In this way, the LLM is repeatedly exposed to the relationships between variables, equations, and their solutions. This reinforcement helps the model learn the dependencies and connections between different elements in a mathematical problem.

2.4 Sentence Paraphrase (SP)

We also conduct paraphrasing for both the question and answer pairs, as shown in Table 7. This exposes the LLM to a broader range of phrasing and language variations. This helps the model become more robust in understanding and generating diverse sentence structures. Consequently, it can handle similar questions with different phrasings and provide accurate responses.

Model Architecture and Training

We utilize the LLAVA (Liu et al. 2023) architecture for our model. The model mainly consists of a large language model (LLM) such as LLAMA-2 Touvron et al. 2023, a pretrained vision transformer Radford et al. 2021 (ViT) as image encoder. In addition, a projection layer is required to map the visual features from the image encoder to the same dimension as the LLM.

During inference, given an image and a textual instruction, the image encoder first extracts the visual tokens from the image, which are then mapped to the dimension of LLM’s embedding space via the projection layer. Then, the mapped image features are concatenated with text embeddings to serve as the input to the LLM. Subsequently, the LLM begins to perform next-token-generation.

2 Model Training

We train our G-LLaVA in two phases, namely 1) geometric visual-language alignment, and 2) geometric instruction tuning. In both phases, we leverage the conventional language modeling loss, which can be formulated as follows:

where F\mathcal{F} represents the model. II represents the geometric figure; StarS_{tar} and SinS_{in} represent the target and input sentences, respectively; StartS^{t}_{tar} denotes the ttht^{th} token of target output, and LL stands for length.

Experiments

We generate the alignment data and instruction data utilizing training set of GeoQA+ (Cao and Xiao 2022) and Geometry3K (Lu et al. 2021). More specifically, the contrastive question-answer (QA) pairs in the alignment data are generated using Geometry3K, which features human-labeled logical forms. Note that GeoQA+ covers the training set of GeoQA (Chen et al. 2021), and share the same val/test set as GeoQA (Chen et al. 2021). More details of data split on GeoQA and GeoQA+ is listed in Table 9. Our approach results in 60K alignment data samples, and more than 110K instruction data samples.

We compare our model with other MLLMs on the geometry problems on the minitest split MathVista (Lu et al. 2023), and compare our model with traditional in-domain model on the test split of GeoQA following (Chen et al. 2022; Liang et al. 2023). The geometry problems in MathVista minitest set is collected from four source datasets Geometry3K (Lu et al. 2021), GeoQA+ (Cao and Xiao 2022), GEOS (Seo et al. 2015) and UniGeo (Chen et al. 2022).

We employ ChatGPT (gpt-3.5-turbo-0613) for data generation. A detailed description of our prompts will be provided in the appendix. We use LLaVA (Liu et al. 2023) as our backbone. More specifically, we utilize LLAMA-2 (Touvron et al. 2023) as the language model and employ the visual encoder of a pretrained vision transformer Radford et al. 2021 (ViT). The resolution of the input image is 336 by 336. We conduct experiments with both 7B and 13B LLMs. In the cross-modal alignment process, only the projection linear layer is trainable. During the instruction tuning phase, both the projection linear layer and the language model are trainable.

For training data, as we found the minitest split of MathVista contains some examples of Mix-train.pk of GeoQA+, we remove those samples that also appears in minitest split of MathVista. The learning rate is set to 3e−53e^{-5}. We expand the images into squares during training, where the extended background color is set to white. For image augmentation, we set the maximum translation distance to 0.25 of the length of longer side. If not otherwise specified, the models are trained for 1 epoch for cross-modal alignment and 2 epochs for instruction tuning, respectively. And the batch sizes are set to 6 per GPUs and 32 per GPUs, respectively.

We use accuracy as the metric for evaluation. Note that several prior studies (Chen et al. 2021; Chen et al. 2022; Cao and Xiao 2022) report results using Top-10 accuracy (generating 10 sequences and selecting the first sequence that successfully addresses the problem as the prediction). Our experimental results directly report Top-1 accuracy. During instruction tuning, we enable the model to output the choice in a fixed format. For evaluation, we directly use regular expression to extract the predicted choices from the generated answers. The answer is considered false if the regular expression fails to extract a valid answer.

2 Main Experiment

We compared G-LLaVA with other MLLMs on minitest split of MathVista (Lu et al. 2023) benchmark on Table 8. The results shows that, geometric cross-modal alignment and instructing tuning on our dataset is effective in improve MLLMs’ geometric problem solving ability. Our specific in-domain model G-LLaVA-7B can even surpass the strong GPT4-V on geometric problems.

3 Comparison with Conventional Methods

We additionally compare our method with conventional SOTA methods in geometry problem solving domain. As illustrated in Table 10, our method demonstrates a notable improvement in Top-1 accuracy over the existing SOTA techniques. Moreover, our model’s top-1 accuracy outperforms the baselines’ top-10 accuracy, demonstrating a significant improvement in predictive precision.

4 Performance Across Problem Difficulties

We compare G-LLaVA with the baselines models on problems with different difficulty levels, as shown in Table 11. Specifically, OP represents the number of “operations", or reasoning steps that needs to be taken for solving the problem. The results verify that our G-LLaVA consistently outperforms baseline models by a large margin across various difficulty levels.

5 Performance Across Different Types of Questions

We compare G-LLaVA with the baselines models on problems with different type of questions, as shown in Table 12. The results suggest that G-LLaVA performs better than the baseline models in various geometric problems such as angle, length, and area problems.

6 Effectiveness of Cross-Modal Geometric Alignment

To evaluate the alignment phase’s effectiveness, we conducted the analysis of the model’s performance with and without alignment phase in Table 13. The results suggest that the alignment phase enhances the model’s ability to interpret images, which is also illustrated by the qualitative result in Figure 1.

Conclusion

In this paper, we make the attempt to address the limitations of current MLLMs in solving geometric problems. We propose strategies to enrich the data by leveraging LLMs, resulting in our augmented dataset, Geo170K. With this dataset, our G-LLaVA outperforms GPT-4-V on the geometric split of MathVista benchmark, with as few as 7B parameters. We hope our work provides new insights on improving multimodal LLMs’ ability of solving geometric problems.

References