Orca-Math: Unlocking the potential of SLMs in Grade School Math

Arindam Mitra, Hamed Khanpour, Corby Rosset, Ahmed Awadallah

Problem Setup

Frontier Language Models such as GPT-4 have demonstrated capabilities previously unseen in smaller models, most notably the remarkable ability to reason (e.g. mathematical reasoning that requires both language comprehension and mathematical understanding). These capabilities have been largely attributed to the very large scale the model size, the dataset size and ultimately the amount of compute needed for training.

Several recent studies have focused on improved the reasoning abilities of small language models (SLMs). Despite that the extent to which scale is needed for achieving reasoning capabilities is still an open research question.

One of the promising directions of improving the reasoning capabilities of SLMs is using frontier language models, such as GPT-4, to create tailored and high-quality synthetic data that can be used to train the SLM. The high quality of the training data and the ability to elicit richer learning signals (e.g. explanations) have been show to significantly improve SLMs abilities in acquiring skills that had only emerged before at much larger scale.

This paradigm fits under a teacher-student approach where the large model (the teacher) is creating demonstrations for the SLM (the student) to learn from. In this work we further explore this direction with focus on mathematical reasoning on grade school math world problem, using the popular GSM8K benchmark.

Several other studies have demonstrated positive results on GSM8K recently with SLMs, e.g. Phi-GSM , OVM , etc. However, many of them employ ensembling, where outputs of up to 100 model runs are combined to arrive at a more accurate results. Result selection is done using, consensus, majority vote or by using a separate a verifier model to score/verify the outputs and select the best answer. Ensembling provides a substantial boost in accuracy (e.g., Phi-GSM uses top-48 to boost the performance from 68.2 to 81.5, uses top-100 to boost LLAMA-2’s performance from 38.6% to 71.9%). However it comes at a significant increase in cost with multiple calls to the model, generating and verifying a 100 different solutions requires 200 different calls to the models. Additionally, some of them use very larger amounts of data (e.g. 12M for Phi-GSM) or use tools or code to avoid calculation errors.

In this work, we extend the teacher-student paradigm to an iterative learning settings with high-quality synthetic training data as follows:

We create Orca-Math-dataset, a synthetic dataset of 200K math problems, paired with GPT-4-Turbo solutions. The dataset was generated using an agent-based setup, hereby referred as, Agent-Instruct, that not only paraphrases existing problems but aims to expand the problem set both in diversity and difficulty.

We introduce an iterative learning procedure where we: (1) use the dataset for supervised finetuning to train the SLM on demonstrations, (2) allow the SLM to practice generating multiple solutions and (3) use the teacher to provide feedback to the student. The feedback comes in the form of evaluating the solutions generated by the student or providing a teacher solution.

With the supervised finetuning alone, we achieve 81.50% on GSM8k at pass@1 metric. The iterative learning loop further improves the pass@1 to 86.81%. without the need for multiple model calls or the use of verifiers, code execution or any other external tools. The model exceeding much bigger models like LLAMA-2-70B (56.8%) , WizardMath-70B (81.6%), Gemini Pro (86.5% with 32 trials) and GPT-3.5 (77.4%). Most notably it can reach this level with only 200K examples (orders of magnitude less than other datasets).

Dataset Construction: Agent-Instruct

The goal of this step is to create a diverse set of grade school math word problems that contains both easy and hard problems. Towards this goal we create a variety of agents.

We start by collecting sample math word problems from existing open-source datasets, namely NumGLUE , AddSub , ALGES , ASDiv , DRAW , GSM8k , MATHQA , MultiArith , SingeOP , SingleEQ , and SVAMP . We collect a total of 36,21736,217 problems. We utilize the Lila benchmark to collect the datasets. Specifically, we collect problems from the train and validation splits from Lila to construct the seed set. Interested readers, please refer to Lila .

Agent - Ask Me Anything

We expand the seed set by creating multiple word problems from each problem in the seed set. We utilize the subsequent prompt for problem creation.

Your goal is to create multiple word problems from a given word problem and its answer. First convert the question of the word problem into a statement. Then for each number in the converted problem create a new word problem. Here are some examples:

Example 1: Q: Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May?

Replacing question with statement: Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. Natalia sold altogether 72 clips in April and May.

Natalia sold clips to some of her friends in April, and then she sold half as many clips in May. Natalia sold altogether 72 clips in April and May. How many clips did she sell in April?

Natalia sold clips to 48 of her friends in April, and then she sold some clips in May. Natalia sold altogether 72 clips in April and May. What is the ratio of the number clips sold in April to number clips sold in May?

Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May?

Example 2: Q: Weng earns $12 an hour for babysitting. Yesterday, she just did 50 minutes of babysitting. How much did she earn?

Replacing question with statement: Weng earns 12anhourforbabysitting.Yesterday,shejustdid50minutesofbabysitting.Sheearned12 an hour for babysitting. Yesterday, she just did 50 minutes of babysitting. She earned10.

Weng earns a certain amount per hour for babysitting. Yesterday, she just did 50 minutes of babysitting and earned 10. How much does she earn per hour?

Weng earns 12 an hour for babysitting. Yesterday, she just did some babysitting and earned 10. How much time did she spend on babysitting?

Weng earns 12 an hour for babysitting. Yesterday, she just did 50 minutes of babysitting. How much did she earn?

Example 3: Q: Betty is saving money for a new wallet which costs 100. Betty has only half of the money she needs. Her parents decided to give her 15 for that purpose, and her grandparents twice as much as her parents. How much more money does Betty need to buy the wallet?

Replacing question with statement: Betty is saving money for a new wallet which costs 100. Betty has only half of the money she needs. Her parents decided to give her 15 for that purpose, and her grandparents gave her twice as much as her parents. She needs 5 more to buy the wallet.

Betty is saving money for a new wallet. Betty has only half of the money she needs. Her parents decided to give her 15 for that purpose, and her grandparents twice as much as her parents. She needs 5 more to buy the wallet. What is the cost of the wallet?

Betty is saving money for a new wallet which costs 100. She has some money saved, her parents decided to give her 15, and her grandparents gave her twice as much as her parents. Now, Betty needs 5 more to buy the wallet. What is the ratio of the money Betty have saved initially to the cost of wallet?

Betty is saving money for a new wallet which costs 100. She has half of the money she needs, her parents decided to give her some money, and her grandparents gave her twice as much as her parents. Now, Betty needs 5 more to buy the wallet. How much money did her parents give her?

Betty is saving money for a new wallet which costs 100. Betty has only half of the money she needs. Her parents decided to give her 15 for that purpose, and her grandparents also chipped in. Now, Betty needs 5 more to buy the wallet. What is the ratio of the amount given by her grandparents to the amount given by her parents?

Betty is saving money for a new wallet which costs 100. Betty has only half of the money she needs. Her parents decided to give her 15 for that purpose, and her grandparents twice as much as her parents. How much more money does Betty need to buy the wallet?

Example 4: Q: Your teacher is giving a test worth 200 points. There is a total of 30 5-point and 10-point questions. How many 5-point questions are on the test? Answer: 20

Note that, the “Ask Me Anything" agent is generating problems based on the seed in Example 4. Examples 1 to 3 are provided as few-shot demonstrations. This agent creates a total of 120,445120,445 new problems. It is important to note that all the generated problems exhibit a similar narrative to that of the seed word problem. The solutions to these word problems are generated using GPT4-Trubo.

Agent - Suggester & Editor

We further expand the seed set by developing challenging problems. To achieve this, we introduce two new agents, namely Suggester and Editor. This setup involves the establishment of a team of agents working collaboratively to create a dataset geared toward a predefined objective. In this instance, the objective is to modify an existing problem to increase its difficulty.

The Suggester examines a specific problem and proposes several methods for enhancing its complexity without creating the actual problem. Subsequently, the Editor takes the original word problem and the Suggester’s recommendations to generate an updated, more challenging problem. This iterative process can occur over multiple rounds, with each round further increasing the complexity of the previously generated problem. We utilize the AutoGen framework to implement the multi-agent workflow. Here is an example of the iterative process:

Seed Problem: Every hour Joanne has to collect the coins out of the fountain inside the mall. During the first hour, she collected 15 coins. For the next two hours, she collected 35 coins from the fountain. In the fourth hour, she collected 50 coins from the fountain but she gave 15 of them to her coworker so she could buy a soda. How many coins did she have after the fourth hour?

Suggester: Increase the number of hours, add more variables for each hour, and incorporate additional transactions.

Editor: Every hour, Joanne has to collect the coins from two separate fountains inside the mall - Fountain A and Fountain B. During the first hour, she collected 20 coins from Fountain A and 30 coins from Fountain B. In the next hour, she collected 25 coins more than she collected in Fountain A during the previous hour and 10 coins less than she collected in Fountain B during the previous hour. In the third hour, she collected the same number of coins in Fountain A as she collected in the second hour and additional 45 coins this time from Fountain B. During the fourth hour, she collected 90 coins in total, but she gave 20 coins to her coworker to buy a soda and 40 coins to another coworker to buy a meal. After the fifth hour, Joanne decided to donate 15

Suggester: Incorporate more fountains, increase the duration, and add more complex calculations for coins collected and spent or donated.

Editor: Joanne has to collect coins from four separate fountains inside the mall - Fountain A, Fountain B, Fountain C, and Fountain D every hour for 8 hours. In the first hour, she collects 20 coins from Fountain A, 30 coins from Fountain B, 40 coins from Fountain C, and 50 coins from Fountain D. In the subsequent hours, her collections have the following fluctuations:

For Fountain A, she collects 10 more coins each hour than she collected in the previous hour.

For Fountain B, her collections decrease by 5 coins each hour compared to the previous hour.

For Fountain C, she collects double the number of coins each hour compared to the number of coins collected in Fountain A the same hour.

For Fountain D, the number of coins collected remains constant at 50 coins each hour.

At the end of the third hour, she donates 20

During the seventh hour, she has to share coins with her coworkers and splits them as follows: 15 coins for coworker one, 20 coins for coworker two, and she keeps 70

After 8 hours, Joanne decides to count her collected coins. How many coins did she have at the end of the eighth hour?

We allow two rounds of iterations per problem and filter problems where the GPT4-turbo generated answer exceeds 1800 characters. At the end of this process, we collect 37,15737,157 problems.

DMath

Furthermore, we include 6,2166,216 problems sourced from DMath . These problems represent a subset of the 7,9437,943 problems present in the DMath training set, in which the solution computed by GPT4-Turbo aligns with the precise gold-standard answer.

Training

We finetune Mistral-7B on the Orca-Math-200K dataset. We have not used packing. The data is presented in the following instruction format:

The loss is computed only on the answer tokens. We employ a constant learning rate of 1×10−61\times 10^{-6}. The per-device batch size is set to 33. Training is conducted for one epoch on eight A100 nodes, with each node containing eight GPUs.

2 Iterative Learning from both Positive and Negative Signals

To generate additional positive and negative solutions for each problem, we sample four responses from the SFT-tuned model from iteration #1. Specifically, we utilize top_p =0.95=0.95 and temperature =0.7=0.7. This process results in a dataset where each of the 200,000200,000 problems has one GPT4-Turbo generated solution and four student-generated solutions. Subsequently, we employ the prompt defined in GPT4-Based-Exact-Match (See section 4 for details) to assess the alignment between the teacher’s (GPT4-Turbo) answer and the student’s answer. For all solutions where the student-generated answer does not match the teacher’s answer, we label them as negative; otherwise, we label the solution as positive. We then construct the preference dataset as follows:

For each question, qiq_{i} we construct qi+q_{i}^{+}, the set of all positive solutions for qiq_{i}. We treat the teacher solution as positive, thus this set by construction contains at least one element.

For each question, qiq_{i} we also construct qi−q_{i}^{-}, the set of all negative solutions for qiq_{i}. This set can be empty if all the 44 responses are are aligned wrt the teacher’s solution. Infact, this is the case for around 80k80k questions. For such situations, we randomly sample one response from qj−q_{j}^{-} for 44 different qjq_{j} where j≠ij\neq i and ∣qj−∣>0|q_{j}^{-}|>0. Note that, for this special situation ∣qi+∣=4|q_{i}^{+}|=4.

Let, Qi={(qi,ai+,ai−)∣(ai+,ai−)∈qi+×qi−}Q_{i}=\{(q_{i},a_{i}^{+},a_{i}^{-})|(a_{i}^{+},a_{i}^{-})\in q_{i}^{+}\times q_{i}^{-}\} be the preference dataset around qiq_{i}. The final preference dataset is created by taking the union of QiQ_{i} for all qiq_{i} in the training dataset.

Dataset Construction Iteration #3

Let M2 denote the model trained with KTO on the dataset constructed for Iteration #2. We replicate the same procedure for the construction of dataset for Iteration #3; however, we utilize M2 to generate the four responses instead of the SFT-tuned model from iteration #1.

To learn from both positive and negative feedback, we have evaluated the performance of two algorithms: the Direct Preference Optimization (DPO) as described by and the Kahneman-Tversky Optimization (KTO) introduced by . DPO is a simple and popular approach for efficiently fine-tuning language models to align with preferences. Additionally, we have explored the capabilities of KTO, which distinguishes itself by requiring only a binary “yes” or “no” response to assess the quality of an output.

Evaluation

We use exact match as the metric. Given a model generated answer, we prompt GPT4 to extract the final short answer and match it with the gold short answer. We will refer to this metric as, GPT4-based-Exact-Match. The following figure shows the prompt template:

As an expert Math teacher, your role is to evaluate a student’s answer to a word problem. The problem is accompanied by a correct solution provided by the problem setter. It is important to remember that there may be various methods to solve a word problem, so the student’s steps might not always align with those in the problem setter’s solution. However, the final answer, typically a number, should be unique and match the problem setter’s answer. Your task involves analyzing the student’s solution to identify any mistakes and determine whether the answer can be modified to correct the error. If the student’s answer is unfixable, consider creating practice problems to help improve their understanding.

Error Analysis: In one sentence, extract the final answer from the problem setter’s solution and compare it with the student’s answer. Do they match?

Question: Billy is volunteering his time to help people do their taxes. He can help 2 people per hour for 3 hours a day. If he takes 20% of the days between March 1st and April 19th off, and helps people on all the other days. How many people does he help? (Remember there are 31 days in March.)

First find the number of days between March 1st and April 19th: 31 days + 19 days = 50 days

Then subtract the percentage of time Billy takes off from 100% to find the percentage of time he works: 100% - 20% = 80%

Then multiply that number by 80% to find the number of days Billy works: 50 days * 80% = 40 days

Then find how many people Billy can help per day: 2 people/hour * 3 hours/day = 6 people/day

Then multiply the number of people he helps per day by the number of days he works to find the total number of people he helps: 6 people/day * 40 days = 240 people

First, let’s calculate the total number of days between March 1st and April 19th.

March has 31 days, and April has 30 days. So, the total number of days from March 1st to April 19th is:

Now, let’s calculate how many days Billy takes off. He takes 20% of the 50 days off.

So, Billy takes 10 days off. This leaves him with: 50 days - 10 days = 40 days

Billy helps people for 3 hours a day, and he can help 2 people per hour. So, in one day, he can help:

3 hours/day * 2 people/hour = 6 people/day

Now, let’s calculate how many people he helps over the 40 days he works:

Error Analysis: The student’s final answer of helping 240 people matches the problem setter’s solution.

Results

Table 2 captures the performance of several training procedures on the GSM8k test set containing 13191319 word problems. We finetune Mistral-7B for upto three iterations. In first iteration we use supervised finetuning to obtain M1. For the second iteration we compare SFT, DPO and KTO. The KTO trained model performs better in this group. We call this M2 and use M2 to generate the dataset for iteration #3. For third iteration, we compare DPO and KTO where M2 servers as the starting point. We also compare these against three epochs of SFT training on the Orca-Math-200K dataset. For all SFT training we employ a constant learning rate of 1×10−61\times 10^{-6}. The per-device batch size is set to 33 and number-of-epochs is set to 11. For DPO and KTO training jobs, we set beta to 0.30.3, per-device batch size to 33, gradient-accumulation-steps to 1111 and number-of-epochs 11. For DPO and KTO training in iteration #2 we employ a constant learning rate of 1×10−61\times 10^{-6} and for iteration #3 a constant learning rate of 1×10−71\times 10^{-7}.

We study the impact model generated positives by limiting qi+q_{i}^{+} to contain only teacher generated solution. In other words we remove any ai+a_{i}^{+} that is model generated in the creation of the dataset for iteration #2. Table 3 shows the result of training M1 with DPO and KTO on this dataset for one epoch. We reuse the hyperparameters for iteration #2. Irrespective of the training algorithm, we see significant performance drop.

Synthetic Negatives

The preference dataset creation involves synthetic negative creation in the situation where all four responses generated from M1 or M2 are positive. We study the impact of these synthetic negatives by ignoring the questions, qiq_{i}, where all sampled responses are positive (Table 4). This reduces the number of questions for iteration #2 by around 80k80k and for iteration #3 by around 104k104k.

2 Math Benchmarks beyond GSM8k

Table 5 presents the performance of Orca-Math on several other word problem datasets. For ease of evaluation, we selected datasets where the answer to each problem is a single number. The test sets of the benchmarks are obtained from Lila. We employ the GPT4-based exact-match metric, and model responses are generated using greedy decoding.

3 Contamination Check

We never use the test split of GSM8K or any other datasets during training or as seeds for synthetic problem generation. Nevertheless, We take the following approach for detecting any potential text contamination.

We begin by preprocessing the texts, which includes converting all characters to lowercase, removing punctuation, tokenizing the text into individual words, and removing common English stopwords to ensure uniformity in the data.

We then vectorize our text corpus using the Term Frequency-Inverse Document Frequency (TF-IDF) method and determine the cosine similarity between the test and training sets, from which we select the top-k (k=10) most analogous questions for each test query.

Finally, we evaluate the extent of text contamination by counting the number of test questions with the highest n-gram overlap above a preset threshold of 0.5 with their corresponding training set matches. We calculate the overlap of n-grams between pairs of texts using the Jaccard similarity. To conduct a rigorous contamination check, we set n=1. It is important to note that the n-gram overlap, when measured using Jaccard similarity, is a non-increasing function of n.

Upon executing our algorithm, we determined that the count of test questions exhibiting significant n-gram overlap is eight, thus indicating negligible text contamination within our test set according to the defined threshold. When limiting the train set to contain only the seed problems, the count of test questions exhibiting significant n-gram overlap is seven. Note that, for n≥2n\geq 2, the count of test questions exhibiting significant n-gram overlap is zero.

Related Works

The generation of synthetic data through generative artificial intelligence (AI) models has evolved rapidly. Numerous datasets have been proposed for both specialized and generic domains, with math-related datasets being closely related to our work.

Learning from rich signals has also garnered significant attention recently. Several studies , have demonstrated the usefulness of preference learning. In this work, we present a detailed analysis of agent-based synthetic data generation and iterative preference learning in the grade school level math domain. Specifically, we demonstrate the robustness of KTO over DPO and the effectiveness of using model-generated positives to improve model training. We believe this is a preliminary step toward iterative learning and self improvement of small language models in challenging domains.

Conclusions

Our study provides compelling evidence that the mathematical reasoning capabilities of Small Language Models (SLMs) can be substantially enhanced. By employing iterative learning techniques and leveraging both positive and negative signals, we have successfully surpassed the previously perceived 80%80\% barrier on the GSM8k benchmark. Our 7B model, trained with 200K data, achieved an impressive 86.81%86.81\% accuracy. Furthermore, the incorporation of agents in dataset generation has proven to be a valuable approach, enabling the creation of more diverse and interesting datasets. These findings not only highlight the potential for significant improvements in SLM performance but also underscore the importance of innovative learning strategies and dataset generation methods in advancing the creation of powerful SLMs.

References