MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, Weiyang Liu

Introduction

Recent years have witnessed the rapid development of large language models (LLMs) which emerge as the favored approach for various applications and demonstrate multi-dimensional abilities, including instruction following , coding assistance , and mathematical problem-solving . Among various tasks, solving mathematical problems is more challenging as they often require highly complex and symbolic multi-step reasoning capabilities. Although some close-sourced models, e.g., GPT-3.5-Turbo , GPT-4 and PaLM-2 , have demonstrated promising performance on some mathematical problem-solving benchmarks, it is still a mystery how these models are trained and what data these models use. Therefore, how to equip open-source LLMs (e.g., LLaMA ) with good mathematical problem-solving skills remains an open challenge.

To tackle this challenge, two popular lines of research to improve the mathematical problem-solving abilities of LLMs are: prompt-based methods and finetuning-based methods. Prompt-based methods aim to activate the potential capacities of LLMs by choosing suitable prompting inputs without modifying the model parameters. Finetuning-based methods update the open-source LLMs (e.g., LLaMA) under the guidance of some other powerful closed-source LLMs (e.g., GPT-3.5 , GPT-4 ). While prompt-based methods are model-dependent and sensitive to many factors, finetuning-based methods, despite being simple and model-agnostic, heavily rely on effective training data on downstream mathematical questions. Our work aims to improve finetuning-based methods with a novel method to bootstrap available mathematical questions in the training set. Specifically, we propose to bootstrap the questions in both forward and backward reasoning directions. For the forward direction, we have the original and LLM-rephrased questions. For the backward direction, we have the self-verification question and FOBAR question . To construct backward reasoning questions, we mask a token in a question using an identifier “x” and ask the model to predict the masked token if the answer is provided. Different from that apply backward reasoning for inference verification, we use it as a form of question for language model fine-tuning. For answers, we adopt an answer augmentation method based on rejection sampling , where diverse reasoning paths are generated and only those with correct answers are used. After combining both forward and backward mathematical questions with augmented answers, we construct a new dataset for fine-tuning, called MetaMathQA. By fine-tuning LLaMA-2 on MetaMathQA, we obtain our MetaMath model. Our approach is guided by the insight that a mathematical question represents merely a single view of the underlying meta-knowledge. Therefore, question bootstrapping can be viewed as a form of multi-view augmentation in order to enable the transfer of the meta-knowledge. Leveraging the MetaMathQA dataset, MetaMath demonstrates exceptional performance in mathematical reasoning, positioning it among the top performers on widely recognized evaluation benchmarks.

Another motivation behind question bootstrapping is to enlarge the question diversity such that the question distribution can be rich enough to cover more unseen scenarios. We quantify the question diversity of the original questions and our MetaMathQA dataset in Figure 2. The diversity gain indicates how diverse the question is compared to the existing dataset, and larger diversity gain means the new question is more different from the existing dataset. With question bootstrapping, our MetaMathQA dataset is much more diverse than the original dataset. We also observe that the test accuracy without bootstrapped questions rapidly reaches a state of saturation. In contrast, the test accuracy, when using bootstrapped questions, continues to exhibit a steady increase.

Question bootstrapping also has an intrinsic connection to dataset distillation and machine teaching , where the shared target is to construct a training dataset that best facilitates generalization. Unlike both methods that focus on optimizing the training empirical risk, question bootstrapping uses the reasoning diversity of questions as a heuristic proxy and maximizes this diversity by constructing forward, backward and rephrased questions. MetaMath aims to transfer the underlying meta-knowledge to enable strong generalization . Our contributions are listed below:

We propose a novel question bootstrapping method to augment the training dataset, resulting in MetaMathQA. Question bootstrapping rewrites questions with both forward and backward reasoning paths and also leverages LLMs to rephrase the question text.

Based on the MetaMathQA dataset, MetaMath is finetuned from state-of-the-art open-source LLMs (e.g., LLaMA-2), showing excellent elementary mathematical problem-solving capability.

We identify an important factor when creating the MetaMathQA dataset – question diversity. The diversity is particularly important in reasoning directions, and backward reasoning questions are very helpful for LLMs to understand mathematical knowledge without memorization.

We conduct experiments on two standard mathematical reasoning benchmarks: GSM8K and MATH . MetaMath outperforms existing open-source LLMs by a large margin. MetaMath-7B has achieved 66.5%66.5\% on GSM8K (+11.5%+11.5\% compared to the previous best open-source LLM) on GSM8K and 19.8%19.8\% on MATH (+8.7%+8.7\% compared to the previous best open-source LLM).

Our work studies data augmentation for improving the mathematical problem-solving ability of LLMs. Despite being simple, our method significantly outperforms many intricate methods. Our results highlight the importance of data augmentation and also shed light on other reasoning tasks.

Related Work

Large Language Models (LLMs) have achieved great success in various natural language processing tasks, e.g., topic classification , sentiment classification , translation , by few-shot prompting (or in-context learning) . Recently, Wei et al. , Wang et al. show that LLMs with more than 100B parameters (e.g., GPT-3 with 175B, PaLM with 540B ) can solve complex tasks by generating multiple reasoning steps towards the answer when given a few reasoning examples as demonstration. While both GPT-3.5 and GPT-4 have shown promising reasoning ability for complex mathematical tasks like MATH , the performance of open-source models (e.g., LLaMA-1 , LLaMA-2 ) is far from satisfactory.

Learning Mathematical Reasoning for complex math tasks like GSM8K and MATH is one of the most challenging problem in open-source LLMs. Wei et al. enhances the reasoning ability of LLMs by augmenting the output with a sequence of intermediate steps toward the answer. A few methods are proposed to improve the quality of reasoning paths. For example, Complexity-based CoT selects examples with more steps as in-context demonstrations and shows that prompting with more reasoning steps leads to better performance. Self-Consistency samples multiple reasoning paths and selects the final answer by majority voting. Another category of work is finetuning-based methods, which finetunes open-source models (e.g., LLaMA) with the knowledge from some advanced closed-source LLMs . Magister et al. investigates the transfer of reasoning capabilities via knowledge distillation. Yuan et al. proposes to apply rejection sampling finetuning (RFT) to improve mathematical reasoning performance. WizardMath proposes a reinforced evol-instruct method to enhance reasoning abilities by supervised fine-tuning and PPO training . MAmmoTH combines CoT and Program-of-Thought rationales for teaching LLMs to use external tools (e.g., Python interpreter) for solving mathematical problems. Wang et al. propose a constraint alignment loss to finetune LLMs for calibration.

Knowledge Distillation transfers knowledge from a larger teacher model to a smaller student model, achieving promising performance in many applications , Recently, propose to transfer reasoning abilities from LLMs (e.g., GPT-3.5 , PaLM ) to small language models (e.g., T5 , GPT-2 ). For example, Finetune-CoT samples multiple reasoning paths from LLMs and finetune the student model with correct ones, while Self-Improve chooses the one with the highest confidence. Li et al. further feeds the question and ground-truth label to LLMs for prompting its reasoning path. Shridhar et al. proposes to generate sub-questions and solution pairs for training. Small models finetuned by knowledge distillation can achieve similar performance to LLMs on both common sense reasoning (e.g., CommonSenseQA ) and symbol reasoning (e.g., Coin Flip ). However, for solving challenging mathematical problems (e.g., GSM8K ), there is still a large performance gap .

Method

The overview of our method is illustrated in Figure 1. Given a meta-question (a sample in the original mathematical training set), we can generate a series of variants. Specifically, we perform three types of question bootstrapping. Combined with answer augmentation, we present MetaMathQA, a diverse and high-quality mathematical dataset based on GSM8K and MATH. We then present MetaMath, a family of LLMs finetuned on MetaMathQA focusing on elementary mathematical problem-solving.

Generating more reasoning paths is a simple but effective way to augment the training set. For a question qiq_{i}, we use few-shot chain-of-thought prompting with temperature sampling to generate KAnsAugK_{\text{AnsAug}} more reasoning paths {(ri(j),ai(j)):j=1,…,KAnsAug}\{(r_{i}^{(j)},a_{i}^{(j)}):j=1,\dots,K_{\text{AnsAug}}\}: the question is appended to a few in-context reasoning examples, then fed to the LLM for generating its reasoning path ri(j)r_{i}^{(j)} and answer ai(j)a_{i}^{(j)}. We filter out reasoning paths with correct answers as:

2 Question Bootstrapping by LLM Rephrasing

Generating more answers for mathematical questions with LLMs is straightforward, but creating questions is more challenging. The questions in GSM8K and MATH are written by well-educated teachers. Hence, enlarging the question set through manual creation is time-consuming and labor-intensive. To address this issue, we propose rephrasing prompting to generate more questions through the LLM. Specifically, for a question qiq_{i}, we append it to the prompt, which is then fed to the LLM for generating the rephrased question. Example 3.2 shows a generated rephrased question and the complete prompt is shown in Appendix A.1. We adopt temperature sampling to sample KrephraseK_{\text{rephrase}} rephrased questions for each meta-question. For the rephrased questions, it is time-consuming to manually check the consistency compared with the original questions. We propose a supervised method to evaluate the correctness between the rephrased questions and the meta-questions. For each rephrased question q^i(j)\hat{q}_{i}^{(j)}, we use few-shot Chain-of-Thought prompting to generate its reasoning path r^i(j)\hat{r}_{i}^{(j)} and answer a^i(j)\hat{a}_{i}^{(j)}, which is compared with the ground-truth answer ai⋆a_{i}^{\star}. The accuracy of Complexity-based CoT for answering the rephrased question by GPT-3.5-Turbo is 76.30%76.30\%, which is comparable to that of answering the original training questions (80.74%80.74\%). This suggests that the quality of rephrased questions is preserved high while the question diversity is improved. We collect the rephrased questions with correct answers (i.e., a^i(j)=ai⋆\hat{a}_{i}^{(j)}=a_{i}^{\star}) as the augmented data:

3 Question Bootstrapping by Backward Reasoning

Backward reasoning plays an important role in answering many mathematical questions, i.e., starting with a given condition and thinking backward to determine an unknown variable in the question. One specific example between a question and a backward question is illustrated in Example 3.3. However, existing methods (SFT, RFT, WizardMath) have significantly lower accuracy on backward questions, as shown in Figure 7, motivating us to bootstrap backward questions to improve the reasoning ability.

Example 3.2: Question and Backward Question Question: James buys 5 packs of beef that are 4 pounds each. The price of beef is 5.50perpound.Howmuchdidhepay?Answer:Hebought5∗4=20poundsofbeef.Hepaid20∗5.5=5.50 per pound. How much did he pay? Answer: He bought 5*4=20 pounds of beef. He paid 20*5.5=110. The answer is: 110 ✓ Backward Question: James buys x packs of beef that are 4 pounds each. The price of beef is 5.50perpound.Howmuchdidhepay?Ifweknowtheanswertotheabovequestionis110,whatisthevalueofunknownvariablex?Answer:Thetotalweightofthebeefis4∗xbecause4∗5.5=22.…Theansweris:27✗Toimprovethebackwardreasoningabilityoffinetunedmodels,wegeneratemorequestionswhichcanbesolvedinabackwardmanner:anumberinthequestion5.50 per pound. How much did he pay? If we know the answer to the above question is 110, what is the value of unknown variable x? Answer: The total weight of the beef is 4*x because 4*5.5 = 22. … The answer is: 27 ✗ To improve the backward reasoning ability of finetuned models, we generate more questions which can be solved in a backward manner: a number in the questionq_{i}ismaskedby“x”,whiletheLLMisaskedtopredictthevalueof“x”whenitsansweris masked by “x”, while the LLM is asked to predict the value of “x” when its answera_{i}^{\star}$ is provided. Different from forward reasoning, which generates explicit intermediate steps towards the final answer, backward reasoning starts with the answer and generates multiple reasoning steps to predict the masked number. Representative backward reasoning methods include Self-Verification and FOBAR .

In Self-Verification (SV) , the question with the answer is first rewritten into a declarative statement, e.g., “How much did he pay?” (with the answer 110) is rewritten into “He paid 10”.Then,aquestionforaskingthevalueof10”. Then, a question for asking the value of{\bf x}isappended,e.g.,“Whatisthevalueofunknownvariableis appended, e.g., “What is the value of unknown variable{\bf x}$?”. Example 3.3 gives an augmented example. We collect the new questions and their generated reasoning paths with correct answers as the augmented data:

Example 3.3: Self-Verification Question Question: James buys x packs of beef that are 4 pounds each. The price of beef is 5.50perpound.Hepaid110.Whatisthevalueofunknownvariablex?Answer:Tosolvethisproblem,weneedtodeterminethevalueofx,whichrepresentsthenumberofpacksofbeefthatJamesbought.Eachpackofbeefweighs4poundsandcosts5.50 per pound. He paid 110. What is the value of unknown variable x? Answer: To solve this problem, we need to determine the value of x, which represents the number of packs of beef that James bought. Each pack of beef weighs 4 pounds and costs5.50 per pound. The total amount James paid is 110.Wecansetuptheequationasfollows:Numberofpacksofbeef∗Weightperpack∗Priceperpound=Totalamountpaid;x∗4∗110. We can set up the equation as follows: Number of packs of beef * Weight per pack * Price per pound = Total amount paid; x * 4 *5.50 = 110; … The value of x is 5. Self-Verification needs to rewrite the question with answer into a declarative statement, which is challenging for complex questions. To address this issue, FOBAR proposes to directly append the answer to the question, i.e., “If we know the answer to the above question is {a_{i}^{\star}$} , what is the value of unknown variable x?” Example 3.3 shows an example. We collect the new questions along with their correct answers as our augmented data:

4 Finetuning Objective Functions

We merge all the augmented data, including answer-augmented data and bootstrapped questions (Rephrasing, Self-Verification, FOBAR) as:

We finetune a LLM model (parameterized by θ{\bm{\theta}}) on DMetaMathQA\mathcal{D}_{\text{MetaMathQA}} to obtain the MetaMath model by maximizing the log likelihood of the reasoning path conditioned on the question, i.e.,

Although we only consider LLaMA-2 here, MetaMathQA can also be used to finetune other LLMs.

Experiments and Results

Datasets. We use two popular mathematical reasoning benchmarks: (i) GSM8K is a dataset consisting of high-quality grade school math problems, containing 7,473 training samples and 1,319 testing samples; and (ii) MATH dataset consists of high school math competition problems that span seven subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus. It contains 7,500 and 5,000 samples for training and testing, respectively. Questions in GSM8K take between 2 and 8 steps to reach the answer, while MATH is much more challenging.

Models. We use the current state-of-the-art open-source model LLaMA-2 , including three different parameter sizes: 7B, 13B, and 70B, as the base model for fine-tuning. GPT-3.5-Turbo is used for rephrasing questions as well as generating answers in all four augmentations, where the temperature is set to 0.7 as in . The LLaMA-2-7B and LLaMA-2-13B are trained by fully fine-tuning. LLaMA-2-70B is finetuned by QLoRA for computational efficiency. More experimental details can be seen in Appendix A.2.

Baselines. The proposed methods are compared with (i) closed-source models such as GPT-3.5-Turbo , PaLM ; (ii) open-source models such as LLaMA-1 , LLaMA-2 ; (iii) Supervised Fine-Tuning (SFT), which uses the training set of the original GSM8K or MATH datasets; (iv) Rejection sampling Fine-Tuning (RFT) generates and collects correct reasoning paths as augmented data for fine-tuning; (v) WizardMath which generates samples and trains two reward models using ChatGPT https://openai.com/ to select samples for fine-tuning.

Diversity Gain. We use the diversity gain to measure to what extent a new dataset added to a basic dataset can improve the overall data diversity. For a base dataset Dbase={xi=(qi,ri,ai)}i=1N\mathcal{D}_{base}=\{x_{i}=(q_{i},r_{i},a_{i})\}_{i=1}^{N} with NN samples, and a new dataset Dnew={xi=(qi,ri,ai)}i=1M\mathcal{D}_{new}=\{x_{i}=(q_{i},r_{i},a_{i})\}_{i=1}^{M} with M samples, the diversity gain is defined as: Dnew\mathcal{D}_{new} relative to Dbase\mathcal{D}_{base} as: dgain=1M∑xi∈Dnewmin⁡xj∈Dbase(∥f(xi)−f(xj)∥22)d_{gain}=\frac{1}{M}\sum_{x_{i}\in\mathcal{D}_{new}}\min_{x_{j}\in\mathcal{D}_{base}}(\|f(x_{i})-f(x_{j})\|_{2}^{2}), where ff is the feature extractor and we use the OpenAI Embedding API text-embedding-ada-002 for feature extraction. For Figure 2, we change the data size of base data and select a fixed set of 20K new data points that the model has not encountered to form Dnew\mathcal{D}_{new}.

2 Results on GSM8K and MATH

Table 2 illustrates the detailed description of our MetaMathQA collection and Table 3 shows the testing accuracy on GSM8K and MATH. As can be seen, for open-source models with 1-10B parameters, MetaMath achieves the state-of-the-art performance. Compared to the previous best LLM, MetaMath achieves a large improvement of 11.6% on GSM8K and 9.1% on MATH in testing accuracy, showing that finetuning on our MetaMathQA data is effective.

As for LLMs with 11-50B parameters, the proposed MetaMath performs the best. Particularly, on both GSM8K and MATH, MetaMath achieves higher accuracy than SFT, RFT, and WizardMath by a large margin (+7%), demonstrating the effectiveness of the MetaMath data in improving mathematical reasoning ability. Furthermore, for LLMs with 51-70B parameters, again, MetaMath achieves the highest testing accuracy. Particularly, MetaMath is better than GPT-3.5-Turbo on GSM8K, which is used for generating augmented data for finetuning.

3 Effect of Augmentations

In this section, we conduct experiments to study the effect of augmentations in MetaMath. We first finetune the LLaMA-2-7B model on augmented GSM8K (MetaMath-GSM8K) data, and test the finetuned model on GSM8K and MATH. Table 1 shows the testing accuracy of different combinations of augmentations. As can be seen, on GSM8K, the models trained on answer augmentation (AnsAug) or rephrasing augmentation achieve much higher accuracy than SFT, which is only trained on the training set. Combing answer augmentation and rephrasing augmentation data for fine-tuning leads to a slightly higher accuracy, which is further improved by about 4% through merging the FOBAR and SV augmentation data. As for MATH, MetaMath trained only on MetaMahQA-GSM8K data performs better than SFT, suggesting its effectiveness in generalizing to unseen mathematical tasks.

We also conduct an experiment by fine-tuning LLaMA-2-7B on the MetaMathQA-MATH data then evaluate the model on GSM8K and MATH. Table 1 shows the testing accuracy. Again, MetaMath trained on AnsAug or rephrasing augmentation data performs much better than SFT. Furthermore, merging all augmented data together for fine-tuning is better than merging AnsAug and rephrasing augmentation data, demonstrating the effectiveness of SV and FOBAR augmentation data in improving mathematical reasoning ability. Moreover, for the unseen GSM8K task, MetaMath trained on MetaMathQA-MATH data is significantly better than SFT (+20%).

4 Discussion from a Perplexity Perspective

According to the Superficial Alignment Hypothesis proposed by Zhou et al. , the capability of a model is rooted in pretraining, and data from downstream tasks acts to activate the inherent ability of LLMs that has been learned during pretraining. There are two important questions that arise from such a hypothesis: (i) what kind of data is most effective at activating possible latent knowledge, and (ii) why is one dataset better than another at such activation? Our empirical results suggest that, in the mathematical tasks we consider, our MetaMathQA dataset may serve as a superior activator of mathematical knowledge. Yet, why MetaMath yields superior performance than training on the data of correct answer-only or GSM8K CoT is unclear. We speculate that perhaps it is the simplicity of the data that matters. As shown in Figure 4, we compute the perplexity for the under-finetuned LLaMA-2-7B model, in terms of answer-only data, GSM8K CoT, and the subsections of MetaMathQA data. The perplexity of MetaMathQA is significantly lower than the other two datasets. This highlights its inherently easy-to-learn nature, which may be more conducive to eliciting bolstered problem-solving abilities from an LLM. This is also aligned with the findings with TinyStories , where short and easy story data can help LLMs generate content fluently.

5 Discussion from a Diversity perspective

As shown in Figure 2, naively prompting GPT-3.5-Turbo for answer augmentation leads to a clear accuracy saturation. After accuracy saturation, increasing the AnsAug data only yields a limited performance gain. For instance, using 80K answer augmentation data to train a LLaMA-2 7B model leads to a 59.6% accuracy, adding new 20K AnsAug data would only take 0.1% performance gain. This is due to the homogeneity of the additional samples, contributing to a diversity gain of only 0.05 (shown in Figure 4). In comparison, adding the same amount of data generated by question bootstrapping leads to a significant performance boost, which is due to the noticeable diversity gain brought by question bootstrapping. As shown in Figure 4, adding 20K data from Rephrasing, FOBAR, or SV takes an increasing diversity gain, thus causing a 0.4%, 2.3%, and 2.6% accuracy gain, respectively. This experiment demonstrates a positive correlation (the Pearson coefficient is 0.972) between the diversity brought by the bootstrapping methods and accuracy. This is also aligned with the success of MetaMath, which is trained with the diverse MetaMathQA dataset including 4 kinds of data reflecting both the forward and backward reasoning paths.

6 Evaluating the Reversal Mathematical Capability

The Reversal Curse , where LLMs trained from a sentence “A is B” are not able to generalize to answer “B is A”, also aligns with the observation in this paper that LLMs lack backward mathematical reasoning ability. To evaluate the backward mathematical capability, we propose a GSM8K-Backward test set, including 1270 backward questions by using SV and FOBAR to augment the original GSM8K test set (as shown in Example 3.3 and Example 3.3). Figure 7 shows the accuracy comparison of different 7B mathematical LLMs between the GSM8K and GSM8K-Backward datasets. As can be seen, existing LLMs struggle to solve mathematical problems in backward rationales and our MetaMath has a significant improvement on both datasets. Specifically, the ways where different LLMs solve the backward mathematical problem are illustrated through examples in Appendix A.3.

7 Reasoning Paths with Incorrect Answer Can Also Be Useful

We conduct experiments on GSM8K using LLaMA-2-7B to study whether the answer augmentation samples with incorrect answers are helpful for finetuning the LLM. We randomly choose 7,473 reasoning paths with incorrect answers from the generated answers, and we ensure that the size is the same as that of the original training set. From Table 4, we observe that the model finetuned on the augmented data with incorrect answers is actually better than SFT, which is counter-intuitive. We hypothesize that although the final answer is incorrect, some intermediate reasoning steps are correct (see Example 4.7). These reasoning steps can still be useful supervision signals. Our results are also aligned with , where they discover the importance of intermediate process supervision for reasoning.

8 More Data is not Always Better

There are also previous works that augment mathematical reasoning data for fine-tuning . An interesting question is whether combining existing augmented datasets with our MetaMathQA can improve the overall mathematical problem-solving performance. We select the RFT dataset as the external dataset. Figure 7 shows that merging the RFT data into MetaMathQA actually hurts the performance, indicating that the RFT data may not be beneficial to MetaMath. Such a phenomenon is consistently observed in the MetaMathQA dataset under different sizes (from 20K to 100K), and the added RFT dataset is about 47K. The performance drop implies that more augmented data does not always help the generalization.

9 Error Analysis

We have demonstrated that – across multiple scales – our MetaMath models can achieve stellar problem-solving performance. Yet, it is important to consider the characteristics of problems that induce errors in MetaMath and existing open-source mathematical models. In particular, we consider the relationship between question length and model performance. To investigate, we divide the GSM8K test set into three equally-sized subsets based on the different lengths of questions and calculate the accuracy of the models over each subset. We find in Figure 7 that, MetaMath and related methods struggle under longer questions. However, excitingly, MetaMath always obtains superior performance. We see the study of improving model performance with longer question lengths – for instance, by further augmenting the MetaMathQA dataset – as ripe grounds for future work.

Concluding Remarks

In this paper, we focus on improving the mathematical problem-solving abilities of open-source LLMs. By bootstrapping mathematical questions on GSM8K and MATH, we present a high-quality and diverse dataset MetaMathQA, involving forward reasoning and backward reasoning samples. Our family of LLMs finetuned on MetaMathQA, called MetaMath, have achieved state-of-the-art on mathematical benchmarks among all open-source LLMs. Remarkably, MetaMath-7B reaches 66.5%66.5\% on GSM8K and 19.8%19.8\% on MATH, surpassing previous open-source LLMs by a significant margin. Our work further emphasizes the importance of the characteristics of the training data on boosting LLM problem-solving capabilities.

Acknowledgement

The authors would like to sincerely thank Katherine M. Collins from University of Cambridge for her valuable insights and suggestions.

References

Appendix A Prompts

A.2 Experimental Details

Training Details. For the fully fine-tuning setting, we use the AdamW optimizer to train the model with 3 epochs and the batch size is 128. We use 8 NVIDIA A100 GPUs to train the 7B and 13B models, the learning rate is set as 2e-5 with a 3% learning rate warmup. For the 70B model QLoRA fine-tuning, the LoRA rank and alpha are 96 and 16, with a 0.05 dropout between the two matrices. The LoRA matrices are append in both the attention layer and the mlp layer. We use the same AdamW optimizer but with a 1e-4 learning rate and without a learning rate warmup. The Training Prompt A.2 are basically from Alpaca , where the instruction is replaced by the MetaMathQA question.

Prompt 1: Training Prompt Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Response: Prompt 2: Evaluation Prompt Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Response: Let’s think step by step. Evaluation Prompting. Different from the few-shot prompting evaluation for closed-source models, we find that zero-shot prompting is better for finetuned LLMs, which also saves more inference costs. Hence, MetaMath uses the zero-shot Evaluation Prompt A.2 for GSM8K and MATH, where the instruction is replaced by the testing question. We set the temperature as 0 for fine-tuned LLaMA model.

Answer Extraction. Different from the Wei et al. , where they use complex string rules to extract the final answer. In line with WizardMath , MetaMath only extracts the string behind The answer is: as the final answer. To teach the model this extraction method, we append The answer is: {gold answer} to the end of answers in the MetaMathQA dataset, where the gold answer is replaced by the respective question’s answer.

A.3 How do different LLMs solve reversal mathematical problems?