MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Zhengyang Tang, Xingxing Zhang, Benyou Wang, Furu Wei

Introduction

Large language models (LLMs) have demonstrated remarkable capabilities in problem-solving. However, their proficiency in solving mathematical problems remains inadequate, potentially due to the inherent necessity for multi-step complex reasoning in mathematical problem-solving. Instruction Tuning (Wei et al., 2021) is an effective approach to unlock certain capabilities in LLMs. Unfortunately, this approach is constrained by the limited size of the currently available datasets on mathematical reasoning. For example, the most popular math datasets, GSM8K (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021), each only contains around 7.5K training examples.

An effective method to tackle this challenge is to augment existing high-quality math datasets using frontier LLMs such as GPT-3.5 and GPT-4. For instance, WizardMath (Luo et al., 2023) introduces an array of operations for GPT-3.5 to generate math questions with increased complexity. MetaMath (Yu et al., 2023) bootstraps questions in GSM8K and MATH through answer augmentation, question rephrasing, self-verification and FOBAR questions. The newly generated examples by these methods exhibit substantial similarity to the original examples contained within the training set, which limits their power in generating large scale math datasets.

We therefore propose a conceptually simple and scalable method MathScale, which is less dependent on original training examples. Specifically, we first prompt GPT-3.5 to extract high level concepts (i.e., topics and knowledge points) from existing seed math questions. In this step, we convert concrete math questions to extractions and the dependency to original questions is largely removed. Given these extractions, we then build a concept graph, which is used to estimate the connections between different concepts. Finally, we can instruct GPT-3.5 to generate new math questions based on randomly sampled concepts from the graph. Intuitively, we can generate significantly more examples using different combination of concepts than using augmentation-based methods, since the resulting number of new examples is bounded by the number of augmentation operations. MathScale also bears resemblance to the cognitive mechanisms underlying the process of mathematical learning in humans (Tall, 2013). Tall (2013) argues that the learning process of human involves two distinct steps called concept compression and connection forging. Concept compression mirrors the process of high level concept extraction, while connection forging is similar to our concept graph construction.

Mathematical capability evaluation is another issue arising from the lack of high-quality mathematical datasets. Recently, most LLMs employ GSM8K (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021) for evaluation. However, GSM8K focuses on elementary-level problems, while MATH offers competition-level challenges. There is a clear gap between the two kinds of capabilities measured. Therefore, we introduce MwpBench, a comprehensive and unified benchmark to measure mathematical reasoning capabilities. MwpBench is composed of ten different math word problem datasets (including GSM8K and MATH) and it covers math word problems from elementary school to college level with different difficulty levels. Moreover, MwpBench standardizes evaluations across all datasets with a unified protocol, promoting consistent and fair model comparisons.

MathScale exhibits effective scalability along the size axis of the math dataset that we generate. As a result, we create a mathematical reasoning dataset (MathScaleQA) containing two million math question-answer pairs. We apply MathScaleQA to fine-tune open-source LLMs (e.g., LLaMA-2 and Mistral), resulting in significantly improved capabilities in mathematical reasoning. Evaluated on MwpBench, MathScale-7B achieves 35.0% in micro average accuracy and 37.5% in macro accuracy, outperforming its best peers of equivalent size by 42.9% and 43.7%, respectively.

MwpBench Evaluation Framework

Existing Datasets Our first endeavor is to collate established datasets, including GSM8K (Cobbe et al., 2021), MATH (Hendrycks et al., 2021), TAL-SCQ (TAL, 2023), Math23k (Wang et al., 2017), Ape210k (Zhao et al., 2020), GaokaoBench-Math (Zhang et al., 2023), and AGIEval (Zhong et al., 2023) series (see Table 1). Types of problems of these datasets are different. For example, most datasets contain math word problems, while TAL-SCQ comprises multi-choice questions. Intuitively, multi-choice questions are simpler because LLMs only need to figure out which choice leads to a higher probability. Therefore, we convert all multi-choice questions to math word problems (detailed in Appendix A.1). Secondly, some of the datasets (e.g., Math23k, Ape210k) are not in English and we translate them to English to expand existing math datasets (detailed in Appendix A.2). Note that we translated part of their training sets and full test sets into English.

CollegeMath Existing datasets does not cover college-level mathematics which requires diverse skills such as analytical thinking, logical reasoning, and quantitative analysis. We therefore propose CollegeMath to bridge this gap.

We curated a collection of nine college mathematics textbooks, each addressing a distinct topic (see Table 2 for more details). These textbooks encompass seven critical mathematical disciplines: algebra, pre-calculus, calculus, vector calculus, probability, linear algebra, and differential equations. These textbooks are originally in PDF format and we convert them to text format using the Mathpix APIhttps://docs.mathpix.com/#process-a-pdf, where equations are transformed to LaTeX format. Once converted a textbook to text format, we are ready to extract exercises and their solutions. For each book, we first manually segment the book into chapter and identify pages with exercises and their solutions. Then we extract questions in exercises and their associated short answers (see more details of our prompts in Appendix A.3). In total, this dataset contains 1281 examples for training and 2818 examples for test.

2 Unified Evaluation Protocol

One of the challenges in benchmarking LLMs for mathematical reasoning is the inconsistency across evaluation metrics and protocols used in different work (Touvron et al., 2023; Luo et al., 2023; Yue et al., 2023).

MwpBench aims to evaluate the mathematical reasoning abilities of instruction tuned LLMs using a unified evaluation protocol. We employ zero-shot setting for evaluation and use the accuracy metric. The reason behind that is we believe fine-tuned LLMs should be able to answer questions directly without demonstrations, while in few-shot setting the final results may change with different set of demonstrations. For prompt template, we choose the Alpaca template (Taori et al., 2023) as default, which is the most widely used for instruction tuning (Taori et al., 2023; Luo et al., 2023; Yu et al., 2023). However, we support customized template just in case that LLMs are trained with a different instruction template (e.g., OpenAI ChatGPT template). For decoding, we choose greedy decoding to eliminate randomness in comparisons, selecting the top-1 completion as the solution. To further standardize the evaluation, we carefully implemented the answer extraction and verification processes (with high precision fuzzy match).

We plan to open-source our evaluation framework.

MathScale: Scaling Instruction Tuning for Mathematical Reasoning

We present details of MathScale in this section. MathScale aims to generate large scale Mathematical Reasoning dataset by prompting ChatGPT and it contains four steps.

As shown in Figure 1, MathScale takes seed math questions as input and we use the training set of MwpBench (around 20K math questions). In the first step, we extract high level concepts (i.e., topics and knowledge points) from these seed questions with prompt engineering of GPT-3.5. We aim to extract meta information needed to solve a particular math question. We believe “topics” and “knowledge points” are important meta information for questions. A “topic” refers to the mathematical subject name or the topic name of math book chapter such as “Money and finance” and “Arithmetic operations”. While “knowledge points” refers to more fine grained math concepts (e.g., theorems, skills) in problem solving. Typical examples are “Definition and properties of dot product” or “Converting fractions to whole numbers”. We instruct GPT-3.5 to act as a Math teacher and extract 1 or 2 topics and 1 to 5 knowledge points from a given seed question (see the prompt template in Table 3).

To ensure the diversity of the extracted topics and knowledge points, we use the training set of MwpBench, which includes questions from different sources. We also remove topics and knowlege points that appear only one time to reduce noise. In total, we extracted around 2K topics and 8K knowledge points. The above process mirrors the concept compression described in (Tall, 2013).

2 Concept Graph Construction

Formally, let E={(u,v)∣fco(u,v)>0}E=\{(\mathbf{u},\mathbf{v})|f_{\text{co}}(\mathbf{u},\mathbf{v})>0\} denote edges in CC and fco(u,v)f_{\text{co}}(\mathbf{u},\mathbf{v}) is the edge weight between u\mathbf{u} and v\mathbf{v}. Intuitively, two KPs (or topics) are more likely to be reasonable composition when they have been frequently used to solve the same seed questions. Let wuvw_{\mathbf{uv}} denote the raw co-occurrence count between node u\mathbf{u} and node v\mathbf{v}. The adjusted weight fco(u,v)f_{\text{co}}(\mathbf{u},\mathbf{v}) is defined as follows:

where ε\varepsilon is a small constant introduced to maintain non-zero counts and prevent computational issues.

Concept Composition Given the graph CC, we are ready to sample topics and KPs from it and the sampled topics and KPs are subsequently used to generate new math questions. We use a graph random walk algorithm to create concept compositions.

In the second step, we do a random walk for one to two steps in the topic sub-graph to search for related topics. The probability distribution for the graph random walk is not uniform and defined as follows:

where N(u)\mathcal{N}(\mathbf{u}) denotes the set of nodes adjacent to u\mathbf{u} in the topic sub-graph.

In the third step, we continue to randomly walk in the hybrid topic-KP graph for a single step with the probability distribution calculated as in Equation (2) on the topic-KP graph. So that we now have one sampled KP.

The whole process above is an analogy of the connection forging described in (Tall, 2013).

3 Mathematical Reasoning Data Generation

Furthermore, we apply a decontamination process, where all math questions in the test set of MwpBench are removed.

4 Validation

We observe that sometimes in the newly generated QA pairs, the solution is incorrect. We therefore also tried to add an additional validation process as follows. We first instruction GPT-4 to generate a reference solution for the question and then ask GPT-4 again to validate the GPT-4 solution against the solution generated in the previous step. We assume GPT-4 is more accurate than GPT-3.5. If GPT-4 believe the orignal solution is incorrect, we replace it with the new GPT-4 solution. Small scale experiments (Table 7) show the step does not improve the results. Perhaps because essentially we are trying to distill GPT-3.5 using open source LLMs. Although some solutions are incorrect, they are still help open source LLMs to learn the model distributions of GPT-3.5. Therefore, in our final pipeline, we remove this validation step.

Experiments

In concept extraction (Section 3.1), we use the MwpBench training set, comprising around 20K questions, as the seed questions for our MathScale pipeline and we employ GPT-3.5-Turbo-0613 for the extraction. In total, we obtain 2,018 topics and 8,892 knowledge points. We then construct graphs to establish relationships among these concepts (Section 3.2). The edge weight in the graph is smoothed using Equation (1) and we set ε=1e−5\varepsilon=1e-5. In the concept composition process, treating the iteration through all topic nodes as one epoch, we repeat this process for approximately 1K epochs, resulting 2 million unique concept compositions. Then we instruct GPT-3.5-Turbo-0613 to create 2 million question-answer pairs with these compositions. We also decontaminate the generated datasets by excluding all math questions in the test set of MwpBench. To leverage the precious high quality math reasoning data, we additionally combine the generated data with the training set of MwpBench. We call the resulting dataset MathScaleQA. The validation step (Section 3.4) is excluded from the final pipeline, because we find that the validation step does not improve results (see details in Section 5.3). Example outputs for each step of the pipeline are provided in Appendix A.4.

The questions in MathScaleQA are formatted using the Alpaca prompt (Taori et al., 2023) as follows.

Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: {question} ### Response:

Our training pipeline is adapted from the open-instruct (Wang et al., 2023) toolkit. We utilize the LLaMA-2 7B and 13B models (Touvron et al., 2023) as well as the Mistral 7B model (Jiang et al., 2023) as our backbone models. We use a batch size of 128 and train on the MathScaleQA dataset for 3 epochs using a learning rate of 2e-5. We call the resulting models MathScale-7B, MathScale-13B and MathScale-Mistral-7B. We leave exploration of the LLaMA-2 70B model in future work.

2 Models in Comparison

For a comprehensive evaluation, we select a diverse set of previous LLMs specialized in mathematical reasoning for comparison.

We include the most capable GPT models developed by OpenAI, which are the light-weighted GPT-3.5-Turbo-0613 and the powerful GPT-4-0314. These models are known to be good at mathematical reasoning and serves as the upper bounds.

Open-Source Models: We also compare our model against open-source math models. Specially, we compare with WizardMath (Luo et al., 2023), GAIR-Abel (Chern et al., 2023), MetaMath (Yu et al., 2023), and MAmmoTH (Yue et al., 2023). WizardMath (Luo et al., 2023) is based on evol-instruct (Xu et al., 2023) and reinforcement learning. MetaMath (Yu et al., 2023) is trained on a dataset by augmenting GSM8K (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021) using answer or question side paraphrasing. The dataset used to train MAmmoTH (Yue et al., 2023) comprises a collection of 13 existing math datasets with GPT-4 CoT (Wei et al., 2022) and/or PoT (Gao et al., 2023; Chen et al., 2022) annotations. We evaluate all models using CoT natural language style math solutions. We noticed that some of the models (e.g., GPT-4 and MAmmoTH) can produce code solution of math problems in addition to natural language solutions. For fair comparison, we refrain from comparing using code-interpreter style solutions, because all models above can produce code-interpreter style solutions if the solutions in their training data are replace by GPT annotated code solutions. Also note that WizardMath v1.1 is a Mistral based math model and we do not know how its training data are constructed (the authors did not release any detail of the training data of WizardMath v1.1). We evaluate all models on MwpBench, which contains 10 datasets on mathematical reasoning. We report accuracies of the 10 datasets as well as their micro-average and macro-average. We prompt all models using the Alpaca template (see Section 4.1). (Luo et al., 2023) recommended an improved prompt for during inference (i.e., adding Let’s think step by step after the standard Alpaca template). However, we observe mixed results on MwpBench for some models in comparison. For example, we observe improved results on GSM8K, but decreased results on MATH. We therefore do not use this optimization for all models in comparison.

3 Main Results

As shown in Table 5, MathScale obtains best micro average and macro average scores on MwpBench compared to other models based on LLaMA-2 7B, LLaMA-2 13B or Mistral 7B. Specifically, On average, MathScale-7B achieves a 35.0% (micro) and 37.5% (macro) accuracy across MwpBench, surpassing its best counterparts of equivalent size by 42.9% and 43.7%, respectively. The trends are similar for MathScale-13B and MathScale-Mistral. This also confirms the effectiveness of our MathScaleQA dataset regardless of the backbone model. Note that in GaokaoBench-Math, AGIEval-Gaokao-MATH, and AGIEval-SAT-MATH, there is no training set. Even on these out-of-domain test sets, MathScale-7B wildly outperforms other open-source models in comparison. When compared to frontier LLMs, MathScale-Mistral demonstrates performance parity in both micro and macro averages relative to GPT-3.5-Turbo (see the first block in Table 5). We have also included subset performances on the MATH and CollegeMath datasets in Appendix A.5 to analyze model capabilities across different topics and disciplines.

Analysis and Discussions

As described in Section 3, given a fixed set of math concepts, iterating over concept graphs allows us to generate different compositions of mathematical concepts, thereby synthesizing large amount of new math data. We use LLaMA-2 7B as our base model to study the scaling property of MathScale. When scaling the size of the MathScaleQA dataset, we observe a nearly logarithmic growth in the performance of the MathScale-7b model across all datasets within MwpBench, as depicted in Figure 3. We draw the scaling curve up to two million examples (size of the full MathScaleQA). We also compare MathScale against WizardMath and MetaMath at their respective training sizes. MathScale outperforms both models across all datasets (except for GSM8K) when using an equivalent amount of training data. Given the scaling curves in Figure 3, we anticipate that the performance of MathScale may continue to improve with even more synthetic training examples. Due to resource constraints, we leave the training set scaling beyond two million examples to future work.

2 Ablation on Concept Extraction

In the concept extraction process (Section 3.1), we use all the 20K seed questions. We attempt to answer the following two questions. 1) Does the number of seed questions matter? 2) Does the number of extracted concepts matter? We control the size of resulting training examples to 25K for fast experimentation. In all experiments, we use the LLaMA-2 7B model as our backbone model.

To assess the influence of seed questions, we firstly randomly remove 50% of the seed questions from the MwpBench training set (i.e., we use only 10K seed questions). The results are shown in Table 6. We observe the macro average on MwpBench drops by 2.9%. Further, when we limite the data source of seed questions exclusively to the training sets of GSM8K and MATH, there is a performance decrease of 3.5%. These results above indicate that incorporating of a larger and more diverse set of seed questions is beneficial.

Additionally, we examine the impact of extracted math concepts. As shown in Table 6, by removing half of the topics or knowledge points, we observe a notable decrease in the macro average on the MwpBench. Particularly, removing knowledge points lead to a greater decrease in performance (i.e., -8.6% with 50% knowledge points v.s. -2.3% with 50% of topics). This highlights the essential role that knowledge points play in enhancing the effectiveness of MathScale.

3 On Validating Generated Data

The generated QA pairs in MathScaleQA might be incorrect. Therefore, we introduce a separate validation step in Section 3.4. In this section, we design controlled experiment on 5K generated data from MathScaleQA and again using LLaMA-2 7B as our base model.

We manually annotate 100 randomly chosen generated data points and generate answers with GPT-3.5-Turbo and GPT-4. GPT-4 demonstrate an impressive accuracy of 87%, significantly outperforming the accuracy of 69% by GPT-3.5-Turbo. Therefore, we used GPT-4 to generate reference solutions and validate our synthetic solutions, replacing any incorrect solutions with the GPT-4 reference solutions.

Within the 5K examples, 26% of the solutions are identified as incorrect by GPT-4 and are replaced. We have another two settings with either all GPT-3.5 solutions and GPT-4 solutions. The results are shown in Table 7 and we observe that using original 3.5-Turbo solutions lead to a similar results as using the validation step.

This observation is counter-intuitive. Maybe because training on synthetic data generated from GPT-3.5 is essential distillation. Even if some solutions are incorrect, they may still help to the open-source LLMs (e.g., LLaMA-2 or Mistral) to mimic the distirubtions of GPT-3.5. We also notice that in neural machine translation distillation, the step of validating incorrect translations is also ignored (Kim & Rush, 2016). Therefore, we opt to omit the validation and correction step from the final MathScale pipeline.

4 Performance on a Fresh Math Dataset

While MathScaleQA generated by GPT-3.5 is rigorously decontaminated to prevent overlap with the MwpBench test set, there may still be small chance that some of the test sets have been leaked to GPT-3.5-Turbo or contained in the training data of LLaMA-2. Because GPT-3.5-Turbo uses human annotated queries submitted by users through their APIshttps://openai.com/research/instruction-following. These queries may include test sets such GSM8K. The training set of LLaMA-2 is not released and we are not sure if some examples in test sets of MwpBench are included or not.

To address this issue, we manually curate a new dataset comprising the latest 30 math questions from latest Gaokao Math exam, held in June for China National Higher Education Entrance Examination. We term this dataset, Fresh-GaokaoMath-2023, which we believe Fresh-GaokaoMath-2023 is not likely to be included in the training data of LLaMA-2 or GPT-3.5-Turbo. Because LLaMA-2 and GPT-3.5-Turbo are released before Fresh-GaokaoMath-2023 is created.

We compare our LLaMA-2 7B based model MathScale-7B against two other LLaMA-2 7B based models (i.e., WizardMath-7B and MetaMath-7B) as well as GPT-3.5-Turbo and GPT-4. Results are in Table 8. MathScale consistently surpasses WizardMath and MetaMath, which aligns with the main results shown in Table 5. It demonstrates the robustness and adaptability of MathScale in handling fresh math questions.

Related Work

ChatGPT-based Instruction Tuning A pivotal aspect driving advancements in math instruction tuning is the use of ChatGPT for data synthesis. For instance, WizardMath (Luo et al., 2023) introduced reinforced evol-instruct which integrates five operations: adding constraints, deepening, concretizing, increasing reasoning steps, and complicating input, thereby facilitating comprehensive evolution. Similarly, MetaMath (Yu et al., 2023) employs a bootstrapping strategy for questions, incorporating answer augmentation, rephrasing, self-verification, and FOBAR. While these methods are effective, the breath space is inherently confined to manually designed operations. Our approach seeks to enable ChatGPT to emulate cognitive processes in human mathematical learning, thus overcoming the limitations faced by previous methodologies.

Tool-Integration Instruction Tuning Recent studies have also explored integrating tools into ChatGPT-based instruction tuning for mathematics. ToRA (Gou et al., 2023) combines natural language reasoning with program-based tool usage to synthesize trajectory data. Each trajectory iteratively concatenates reasoning, programming, and program outputs until the final answer is reached. Our current focus is solely on natural language reasoning. While tool integration within the MathScale pipeline is an intriguing prospect, we reserve its exploration for future research.

Conclusions

We propose MathScale, a simple and scalable method to create high-quality mathematical reasoning data using frontier LLMs. We also construct MwpBench, a comprehensive benchmark of Math Word Problems covering K-12, college, and competition level math problems. Evaluated on MwpBench, MathScale-7B achieves state-of-the-art performance across all datasets, surpassing its best peers of equivalent size by 42.9% in micro average accuracy and 43.7% in macro average accuracy, respectively.

Broader Impact

This paper seeks to advance mathematical reasoning by introducing a scalable method for generating high-quality synthetic data with large language models, along with new evaluation benchmarks to foster consistent and fair model comparisons in academia. While our efforts center on assessing mathematical capabilities, it’s crucial to note that the models may exhibit biases not examined in our study. Addressing these biases and ensuring the models’ alignment with societal values is essential, highlighting the need for comprehensive evaluations that encompass both technical performance and ethical considerations.

References

Appendix A Appendix

For datasets like TAL-SCQ (TAL, 2023), GaokaoBench-Math (Zhang et al., 2023), and AGIEval (Zhong et al., 2023), the problems are presented in a multiple-choice format. To eliminate the influence of the problem type and concentrate on the intrinsic ability of LLMs to address mathematical problems, we converted these non-word problems into word problems.

Initially, we identified and filtered out questions that rely heavily on the multiple-choice format. This filtering was done using specific keywords and phrases that are indicative of multiple-choice questions.

A.1.2 Creating Question-Answer Pairs

After filtering out the aforementioned questions, the remaining questions were paired with their corresponding correct answer choices. This transformation resulted in a format where each problem is presented as a word problem followed by its solution.

A.2 MwpBench: Translation of Non-English Problems to English

For several datasets, namely Math23k (Wang et al., 2017), Ape210k (Zhao et al., 2020), GaokaoBench-Math (Zhang et al., 2023), and AGIEval-Gaokao (Zhong et al., 2023), the problems are originally presented in Chinese. To ensure uniformity and mitigate the effects of multilingual representations, we translated these Chinese problems into English. The translation was facilitated by the GPT-3.5-Turbo API. Due to parsing errors encountered during the post-processing, a few examples were excluded. The prompt template employed for the translation request is provided below:

I want you to act as a Math Translator. Your task is to translate Chinese math questions into English math questions. Make sure to keep the original question numbers. Make sure to keep the math formula in Latex format. The translations should be clear, accurate, and easily understandable for students who are native English speakers. # Chinese Math Questions #: # English Math Questions #:

A.3 CollegeMath: Extraction from textbooks

To construct the CollegeMath dataset, we made use of the GPT-3.5-Turbo API to parse and extract questions and answers from raw, segmented LaTeX exercises and their corresponding solutions.

The primary goal was to convert raw, potentially unstructured questions from math textbooks into well-formulated LaTeX-formatted questions. Below is the prompt template we utilized for this extraction process:

I want you to act as a Math Parser. Your task is to convert raw messy questions from a math textbook into well-structured LaTeX-formatted questions. Please ensure to retain the original question numbers. If needed, prepend the original instructions to the parsed questions to make them more comprehensible. If needed, skip the broken questions. #Raw Questions#: ‘‘‘ ‘‘‘ #Well-structured LaTeX-formatted Questions#:

A.3.2 Extracting Answers from Solutions

Similarly, for answers, our aim was to transform raw, messy answers from textbooks into clear, LaTeX-formatted answers. Here’s the template for this task:

I want you to act as a Math Parser. Your task is to convert raw messy answers from a math textbook into well-structured LaTeX-formatted answers. Please ensure to retain the original answer numbers. If needed, skip the broken answers. #Raw Answers#: ‘‘‘ ‘‘‘ #Well-structured LaTeX-formatted Answers#:

By employing the aforementioned prompt templates, we were able to extract a comprehensive set of questions and answers, thereby forming the foundation of the CollegeMath dataset.

A.4 MathScale: Concrete Examples

A set of 30 topics, randomly chosen, is listed below to illustrate the variety:

"Arithmetic operations" "Word problem solving" "Mathematics" "Money and finance" "Problem-solving strategies" "Arithmetic" "Multiplication" "Proportions" "Basic arithmetic operations" "Conversion of units" "Measurement and weight" "Multiplication and addition" "Budgeting" "Basic arithmetic" "Wages and overtime" "Calculating earnings" "Arithmetic Sequences" "Exponential Growth" "Financial calculations" "Problem solving" "Algebraic expressions" "Economics" "Time" "Business and finance" "Ratio and proportion" "Problem-solving" "Time calculations" "Addition" "Distance" "Speed"

A.4.2 More Extracted Knowledge Points

Similarly, we provide a list of 30 knowledge points, chosen at random, to demonstrate the depth and breadth of content:

"Random selection of marbles" "Definition and properties of dot product" "Manipulation of complex numbers" "Calculation of time required to complete a task" "How to apply the concept of a seven-day cycle" "Distinct numbers" "Expectation of a function of a random variable" "Ability to calculate total time" "Combinations of numbers" "Calculation of weekly income" "Relative motion" "Understanding the relationship between centimeters and kilometers" "Diagonalizing a matrix" "Proportional relationships between two quantities" "Ergodic Markov chain" "Addition of values" "Counting the number of cars" "Converting fractions to whole numbers" "Identifying relationships between different variables" "Ability to set up and solve a proportion equation" "Addition and subtraction of matrices" "Using logarithms to solve exponential equations" "Probability of rolling a specific number on a six-sided die" "Divisibility of polynomials" "Application of multiplication to calculate total revenue" "Identifying the highest and lowest scores" "Ability to calculate percentages." "Geometric interpretation of dot product" "Dividing complex numbers" "Understanding weight units"

A.5 Evaluation on Individual Topics

We examine the subset performances on MATH, as shown in Table 9. It is evident that MathScale consistently delivers exceptional results across diverse topics.

We also detail the subset performances on CollegeMath at the College Level. As shown in Table 10, despite the MwpBench training set’s seed questions only encompassing algebra, precalculus, and calculus, MathScale demonstrates robust performance in OOD’s test sets including vector calculus, probability, and linear algebra. However, an area of challenge is differential equations, where all models show limited success.