LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
Zhuoshi Pan, Yu Li, Honglin Lin, Qizhi Pei, Zinan Tang, Wei Wu, Chenlin Ming, H. Vicky Zhao, Conghui He, Lijun Wu
Introduction
Recently, Large Language Models (LLMs) have significantly improved their ability to solve mathematical problems through Supervised Fine-Tuning (SFT). A common strategy involves refining the quality of chain-of-thought (CoT) reasoning data, such as distilling high-quality solutions from advanced models Magister et al. (2023); Yu et al. (2024b). While these methods enhance the model’s capacity to generate step-by-step solutions, they predominantly focus on optimizing correct reasoning trajectories while overlooking the potential of error data. This omission limits the model’s ability to learn from mistakes, thereby constraining its reflective reasoning capability. Reflection—the process of identifying, analyzing, and correcting errors—is a critical component of human problem-solving Stacey et al. (1982). Given the failure to integrate this ability into LLMs, models remain vulnerable to propagating errors during inference without autonomous correction mechanisms.
To address this gap, recent studies have begun exploring methods to cultivate reflection in LLMs by leveraging error data. For instance, some works Xi et al. (2024b); Li et al. (2024f); Qin et al. (2024); Guan et al. (2025) employ external critical models to critique intermediate reasoning steps or use Monte Carlo Tree Search (MCTS) to navigate complex reasoning paths and prune error branches. Others Yan et al. (2024); Han et al. (2024); Zhang et al. (2024a); Yang et al. (2024) propose self-correction frameworks that construct incorrect-correct data for fine-tuning, enabling the model to iteratively revise its outputs. However, these approaches suffer from significant limitations. MCTS-based methods introduce substantial computational overhead and complexity, while self-correction methods rely on naive and inefficient techniques to collect incorrect and correct reasoning trajectories.
In this work, we propose LEMMA to Learn from Errors for MatheMatical Advancement, a novel method to systematically enhance LLMs’ reflective reasoning by constructing and learning from error-corrective trajectories. Our approach begins with a fine-grained categorization of error types in model-generated solutions, ranging from “question misinterpretation” to “calculation error”. Building on this taxonomy, we design an error-type grounded error augmentation strategy that diversifies error data by (1) harvesting mistakes from the target model’s own reasoning traces and (2) guiding advanced models to generate representative errors according to the analyzed error type distributions. For each erroneous solution, we then construct paired reflection data through two complementary mechanisms: Fix & Continue Trajectories, where the mistake is directly corrected within its original context, and Fresh & Restart Trajectories, where a new correct solution is generated from scratch. These trajectories are seamlessly connected via model-aware reflection links—annotations that explain the error’s origin and justify the correction—resulting in coherent training examples.
Experiments across mathematical reasoning benchmarks (e.g., GSM8K, MATH) demonstrate LEMMA’s effectiveness. Models fine-tuned with LEMMA achieve state-of-the-art (SOTA) performance, outperforming both standard SFT baselines and prior error-aware methods (up to 13.3% average accuracy improvement on LLaMA3-8B). LEMMA-trained models also achieve strong generalization ability through evaluation on out-of-distribution (OOD) benchmarks. Further analysis reveals LEMMA can consistently reduce the occurrence of representative error types. In contrast, while fine-tuning on the original training set (SFT) improves overall accuracy, it leads to an increase in certain error types. These results validate that structured learning from errors, guided by systematic analysis, is a powerful yet underutilized lever for advancing mathematical reasoning in LLMs.
Related Work
Due to the scarcity of mathematical reasoning data with detailed, human-annotated reasoning steps Song et al. (2023); Luo et al. (2023), some studies Zelikman et al. (2022); Yuan et al. (2023); Singh et al. (2024); Huang et al. (2023); Tong et al. (2024a) leverage the correct output of the model itself for fine-tuning. This strategy is also known as self-improvement or reject sampling fine-tuning.
Recently, some works Qi et al. (2024); Xi et al. (2024a); Xu et al. (2024); Xi et al. (2024b); Qin et al. (2024); Guan et al. (2025) have begun using Monte Carlo Tree Search (MCTS), Process Reward Models (PRM) Lightman et al. (2024) or critique models to further enhance self-improvement. However, these methods only utilize the correct solutions, neglecting the generated errors. Since models only learn from correct solutions, they struggle to reflect on and self-correct errors they made, leading to error accumulation Zhang et al. (2024a); Han et al. (2024); Yan et al. (2024).
Data augmentation is also a prevalent strategy to enhance model performance on mathematical tasks. Magister et al. (2023) and Yue et al. (2024a) distill reasoning capabilities from LLMs into smaller LMs. Dart-Math Tong et al. (2024b) introduces a difficulty-aware answer augmentation strategy, where more solutions are generated for harder problems. To further boost model performance, several works Tang et al. (2024); Huang et al. (2024); Yue et al. (2024b); Liu et al. (2024); Wang et al. (2024a); Zhou et al. (2024); Ding et al. (2024); Lu et al. (2024a); Li et al. (2024a); Luo et al. (2023); Li et al. (2024c) synthesize more training data by creating new mathematical problems and solutions. For example, MetaMath Yu et al. (2024a) combines answer augmentation and question augmentation, as well as two backward reasoning methods Jiang et al. (2024); Weng et al. (2023), to further augment training data. Our method is orthogonal to these question augmentation methods and can be directly integrated with them.
Reflection and self-correction mechanisms have been proven to be effective in enhancing the performance of large language models (LLMs) across various domains. To encourage models to identify and amend their previous errors, one common approach leverages feedback from an external verifier or critic model Shinn et al. (2024); Renze and Guven (2024); Chen et al. (2024); Li et al. (2024b, 2023); Kim et al. (2023); Wu et al. (2024); Weng et al. (2023); Du et al. (2024); Li et al. (2024d). Alternatively, some research Yang et al. (2024); Zhang et al. (2024a); Han et al. (2024); Yan et al. (2024); Qin et al. (2024); Zhang et al. (2024b); Lu et al. (2024b); Kumar et al. (2024); Singh et al. (2024) focuses on fostering the self-correction capabilities of LLMs during the generation process itself, without external feedback. To collect training data, previous works often use a relatively high temperature during the generation process to introduce errors for later critique and correction Xi et al. (2024b); Yan et al. (2024); Lu et al. (2024b); Zhang et al. (2024a); Han et al. (2024). However, research has indicated that increasing the temperature can lead to nonsensical errors or incoherent text that would not typically occur during standard generation scenarios Lu et al. (2024b); Renze (2024). Moreover, these approaches neglect different correction strategies, which could potentially restrict the model’s ability to reflect and self-correct effectively.
Methodology
To better leverage generated reasoning errors for enhancing the self-reflection and correction capabilities of LLMs, we begin by conducting a systematic analysis of common error types in widely-used models. Building on this analysis, we introduce LEMMA, a novel approach that strategically constructs self-correction data to improve the mathematical reasoning abilities of LLMs. Fig.2 provides an overview of the LEMMA framework.
We begin by defining key components of LEMMA. The generated reasoning trajectory is a sequence of reasoning steps: , where is the predicted answer. A bad trajectory includes both correct steps , incorrect steps and ends with an incorrect answer: . Models equipped with reflection and self-correction capabilities should be able to identify and rectify the incorrect steps , leading to a revised trajectory. The revised trajectory can be viewed as the concatenation of a bad trajectory, a Reflection Phrase (RP), and a correct trajectory , expressed as . Here, RP represents the reflection phrases that pinpoint and correct previous errors while seamlessly transitioning to a correct step. To minimize error accumulation, the model should recognize and correct errors as early as possible. Ideally, the bad trajectory should be a subsequence that ends at the first erroneous step, denoted as . The following paragraphs will detail how we collect the bad sub-trajectory and the good trajectory , ultimately constructing the error-corrective revision trajectory for model training.
2 Error Analysis
To gain a holistic understanding of the mathematical reasoning errors in common LLMs, we conduct a systematic analysis on error types. We use an error taxonomy modified from Li et al. (2024e), as detailed in Tab.1. Fig.3 presents the distribution of error types for different models. Our key findings are: (1) The most common errors include “Question Misinterpretation (QM)”, “Formula Confusion Error (FC)” and “Calculation Error (CA)”. This indicates that the models require improvements in areas such as problem comprehension, formula application, and conceptual understanding. (2) The distribution of error types is relatively consistent across different models. These key insights serve as the foundation for our subsequent error-type grounded error augmentation method.
The above analysis is conducted using greedy decoding generation.The temerature of softmax function is set to 0. However, prior research Xi et al. (2024b); Yan et al. (2024); Lu et al. (2024b); Zhang et al. (2024a) typically uses a relatively high sampling temperature (e.g., in Yan et al. (2024) and in Lu et al. (2024b)) to collect a diverse set of bad trajectories . Hence, we also investigate the effect of temperature on error types. Fig.4(a) depicts how the distribution of error types varies with different softmax function temperatures. As the temperature increases, nonsensical errors, exemplified in Fig.4(b), begin to emerge. In other words, the occurrence of nonsensical errors rises with sampling temperature, whereas this type of error is generally absent in greedy decoding.
3 Erroneous Trajectory Collection
Based on our analysis of the relationship between error types and sampling temperature in Sec.3.2, we opt not to increase the sampling temperature. Instead, we employ a relatively low sampling temperature , which is widely used in mathematical evaluation Xi et al. (2024b); Yan et al. (2024); Zhang et al. (2024b). To mitigate the reduced diversity of error steps at lower temperatures, we propose an error-type grounded mistake augmentation method that systematically generates diverse and meaningful errors for subsequent correction. Specifically, we first determine the error type distribution for each question based on our prior analysis. We then leverage a teacher model (GPT-4o) to intentionally produce erroneous trajectories given an error type, which is sampled from the previously obtained error type distribution for each question. This approach ensures that the introduced errors are both diverse and closely aligned with the error patterns observed in the student model. The prompt is provided in Fig.13 of Appendix A.4. Using this approach, we compile a comprehensive collection of bad trajectories, which consists of: (1) the erroneous trajectories generated by error augmentation and (2) those produced by the student model itself. This strategy mirrors the human learning process, where students not only reflect on and correct their own mistakes but also receive guidance from teachers, who highlight common error-prone steps based on overall performance of all students. Additionally, the teacher model annotates the first error step in each bad trajectory . Starting from this step, the trajectory is truncated to form the bad sub-trajectory , to minimize error accumulation.
4 Revision Trajectory Generation
Upon obtaining the bad trajectory , we proceed with a correction process to generate the final revision trajectory . Inspired by the self-correction process in humans Hoffmann (2018), we explore two correction strategies:
(1) Fix & Continue Revision: In this strategy, the teacher model fixes the student model’s first error step and continues the reasoning process to reach the correct answer. However, as illustrated in Fig.9, there are instances where, despite the initial reasoning steps being correct, they may not be a “smart” way to solve the problem. This can result in a prolonged reasoning trajectory involving complex reasoning and intensive computations, which are more susceptible to errors. To address this limitation, we introduce the “Fresh & Restart” correction strategy as follows.
(2) Fresh & Restart Revision: In this strategy, the teacher model critiques the student model’s errors and then initiates the reasoning process anew, rather than continuing from an erroneous “intermediate” step. We encourage the model to explore alternative solutions using the prompt depicted in Fig.15 in Appendix A.4. This approach emulates human correction processes, where, upon realizing an initial approach is flawed, one may abandon the original reasoning steps and start anew instead of making minor adjustments to the first attempt.
By combining both correction strategies, we generate a diverse set of revision trajectories . Training on the constructed data enables the student model to learn different self-correction strategies. Following Xi et al. (2024b) and Qin et al. (2024), we also employ the teacher model to smooth the entire revision trajectory, adding necessary logical transitions and connections to produce the final training data. Finally, we filter the trajectories based on the correctness of the final answer, retaining only those that lead to a correct answer.
Experiments
We evaluate our method through comprehensive experiments from three key aspects: (1) In-distribution mathematical tasks, (2) Out-of-distribution mathematical tasks, and (3) Reflective mathematical reasoning tasks.
We use the training set of MATH Hendrycks et al. (2021) and GSM8K Cobbe et al. (2021) to generate the error-corrective reasoning trajectory. We utilize LLaMA3-8B to produce the self-generated errors at a temperature of and employ GPT-4o Hurst et al. (2024) as the teacher model to deliberately introduce errors and perform subsequent corrections. Additionally, we employ an open-source model, LLaMA-3.1-Nemotron-70B Wang et al. (2024b), as an alternative teacher model to demonstrate the generalization of our method. For each question, we generate two self-generated errors from the student model and two deliberately introduced errors from the teacher model. For LEMMA (w/ MetaMath), we collect two additional errors. One is from the student model, and the other is generated by the teacher model based on the new questions of MetaMath (Yu et al., 2024a). For each error, we apply both “Fix & Continue” and “Fresh & Restart” correction strategies once. After filtering out the trajectories with incorrect final answers, we obtain error-corrective reasoning trajectories as training data. We fine-tune various base models, including general-purpose models such as LLaMA3-8B and Mistral-7B-v0.1, as well as the math-specialized model DeepSeekMath-7B and Qwen2-Math-7B. Further implementation details are available in Appendix A.3.1.
We use GSM8K Cobbe et al. (2021) and MATH Hendrycks et al. (2021) as the In-Domain evaluation. For Out-of-Domain evaluation, we choose ASDIV Miao et al. (2020), MAWPS Koncel-Kedziorski et al. (2016), Mathematics Davies et al. (2021), SVAMP Patel et al. (2021) and College-Math Tang et al. (2024). Following Zhang et al. (2024b), we also adopt the follow-up QA (FQA) and Error correction (EC) tasks of MathChat Liang et al. (2024), which require the model to reflect on previous generation and perform further reasoning. Unless specified otherwise, we use Pass@1 as the evaluation metric. The performance results using majority voting are detailed in Appendix A.1.3.
We compare LEMMA with four self-correction methods and four data augmentation approaches. For self-correction methods, we consider: (1) Intrinsic Self-Correction (ISC) Han et al. (2024): Teaching small language models to self-correct by training on the constructed self-correction data. (2) S3C-MATH Yan et al. (2024): Employing a step-level sampling approach to generate potentially erroneous steps, followed by reflection and improvement, to construct self-correction data. (3) RefAug Zhang et al. (2024b): Appending a “reflection” part to the original solution, which involves proposing an alternative solution and solving a similar problem. (4) RefAug-90k Zhang et al. (2024b): To eliminate the influence of sample size and the annotation model, we use the official codehttps://github.com/ytyz1307zzh/RefAug of RefAug to generate data with GPT-4o, which aligns with LEMMA in terms of both sample size and annotation model.
For data augmentation approaches, we consider: (1) SFT: Training on the union of GSM8K and MATH training set. (2) Rejection Sampling Fine-tuning (RFT) Yuan et al. (2023): Training on the correct self-generated reasoning trajectories. (3) MetaMath Yu et al. (2024a): Combining answer augmentation, question rephrasing, and two backward reasoning methods Jiang et al. (2024); Weng et al. (2023), to augment training data. (4) GPTAug: Prompting GPT-4o to generate step-by-step solution for each question. Please refer to Appendix A.3.2 for more details regarding baseline implementation.
2 Main Result
Tab.2 lists the performance of different methods. We summarize the key findings as follows.
(1) LEMMA significantly outperforms SOTA baseline methods across most tasks, achieving an average accuracy improvement of at least 2.3% for LLaMA3 and 4.5% for DeepSeekMath. The enhancement is particularly noticeable in challenging tasks such as MATH Hendrycks et al. (2021) and College-Math Tang et al. (2024), where LEMMA surpasses SOTA baselines by at least 4.1% and 2.8% on LLaMA3, respectively. This underscores the efficacy of reflective and self-correction capabilities for solving complex math problems. (2) Interestingly, RFT Yuan et al. (2023) lags behind all reflection and self-correction methods. We attribute this to the inherent limitation of RFT, which solely utilizes the correct self-generated solutions, forgoing the valuable opportunity to learn from failures. (3) Additionally, LEMMA demonstrates strong performance across both in-distribution and out-of-distribution datasets. While some baselines, such as MetaMath, achieve relatively good results on in-distribution datasets, they fall short compared to LEMMA on out-of-distribution datasets. Notably, scaling the data size of RefAug Zhang et al. (2024b) to data (i.e., RefAug-90k) enhances in-distribution performance; however, the improvements on out-of-distribution datasets are limited or even negative.
3 Reflective Math Reasoning Performance
Following Zhang et al. (2024b), we assess the reflective reasoning abilities of LLMs fine-tuned via various methods. Tab.3 presents the results. Notably, LEMMA significantly enhances the reflective reasoning capabilities of models compared to other data augmentation methods, achieving improvements of at least 3.3% and 4.7% in accuracy on MathChat-FQA- and MathChat-EC, respectively. Although some data augmentation approaches, such as MetaMath, have achieved considerable performance gains on multi-turn math question answering (i.e., MathChat-FQA), they fall short in improving error correction ability, with only a 0.5% accuracy increase on MathChat-EC compared to SFT. In comparison to reflection and self-correction methods, such as ISC and RefAug, LEMMA also demonstrates notable superiority. For instance, LEMMA surpasses ISC by 4.1% and 6.2% accuracy points on MathChat-FQA- and MathChat-EC, respectively. Fig.10 and Fig.11 in Appendix A.2 show some output cases where the model fine-tuned with LEMMA performs reflection and self-correction to produce more accurate answers. These results further underscore LEMMA’s advantages in advancing reflective and self-correction capabilities of LLMs.
4 Choice of Teacher Model
We also evaluate the performance of our approach using an open-source teacher model, LLaMA-3.1-Nemotron-70B, instead of GPT-4o, as the teacher model. The results, as shown in Tab.6 of Appendix A.1, demonstrate that LEMMA continues to hold a significant advantage over baseline methods even after the replacement of the teacher model. This suggests that the improvements offered by LEMMA are not attributable to the teacher model itself, but rather to the efficacy of the systematic error introduction and correction strategy.
Analysis
We examine the impact of sample size on the performance of different methods. The results presented in Fig.5, highlight several key observations. (1) LEMMA consistently achieves superior performance across various sample sizes. Notably, as the dataset size increases, the performance gap between LEMMA and other baselines widens, underscoring its scalability potential. (2) LEMMA demonstrates stable performance improvements on both in-distribution (MATH) and out-of-distribution (Mathematics) datasets as the data size grows. In contrast, some baseline methods, such as ISC and RefAug, although showing gains on in-distribution datasets like MATH, tend to plateau or even decline in performance on out-of-distribution datasets. This saturation suggests that these methods might overfit to in-distribution data, lacking the generalization capabilities that LEMMA provides.
2 Analysis on Error Type after Fine-tuning.
We analyze the types of errors generated by the model before and after fine-tuning with LEMMA. We report the error count of the model before fine-tuning (Base), the model fine-tuned on the original dataset (SFT), and the model fine-tuned with our LEMMA approach. The results presented in Fig.6 reveal several insightful trends. Firstly, LEMMA consistently reduces the occurrence of common error types, particularly in categories such as “Question Misinterpretation (QM)” and “Calculation Error (CA)”. Secondly, although fine-tuning with the original training data (SFT) improves overall accuracy, it leads to an increase in certain error types, such as “Confusing Formula Error (FC)”. This can be attributed to limitations in the original training data, which may fail to address specific error patterns and potentially cause overfitting to certain reasoning paths.
3 Ablation study
We conduct ablation studies to assess the contributions of each component within LEMMA. During the erroneous step collection phase, we exclude the error augmentation module, relying solely on errors generated by the model itself, which we denote as “w/o Error Aug.” To ensure that any decline in performance is not merely due to a reduced sample size, we generate the same amount of revision trajectories as LEMMA, which we refer to as “w/o Error Aug. (90k).” We also perform ablation on the error correction strategy by removing the “Fresh & Restart” method from the revision process, labeled as “w/o Fresh & Restart” and “w/o Fresh & Restart (90k)”. We report the accuracy on in-distribution tasks in Tab.4; for out-of-distribution performance, please refer to Tab.9 in Appendix A.1. It is evident that removing the “error augmentation module” results in a significant performance drop. This decline is not due to sample size, as “w/o Error Aug (90k)” also exhibits a 6.2% accuracy decrease in performance on MATH compared to LEMMA. We attribute this decline to the reduced diversity of error steps, as the model relies solely on errors generated by the student model itself. In contrast, the error augmentation module introduces a variety of meaningful errors, enhancing the model’s ability for reflection and self-correction. Furthermore, excluding the “Fresh & Restart” strategy degrades performance. This decline highlights the essential role of the “Fresh & Restart” correction: by enabling the model to reset and reassess problem-solving pathways, it significantly enhances mathematical reasoning capabilities.
Conclusion
In this work, we introduce LEMMA, a novel framework designed to enhance the mathematical reasoning capabilities of LLMs by systematically learning from errors. Based on a comprehensive analysis of error types, LEMMA employs an error-type grounded mistake augmentation strategy and constructs diverse revision pathways using both the Fix & Continue and Fresh & Restart correction strategies. This framework allows models to autonomously detect and correct errors during the generation process, thereby improving their mathematical reasoning abilities. Extensive experiments demonstrate that LEMMA significantly outperforms SOTA baselines.
Limitations
While LEMMA represents a significant advancement in enhancing the mathematical reasoning capabilities of large language models, several limitations persist. Firstly, its focus has been solely on mathematical reasoning tasks, leaving its effectiveness and adaptability in other domains unexplored. Moreover, the synthesized dataset used in LEMMA comprises fewer than examples, which is relatively small compared to data augmentation methods like MetaMath. This raises questions about whether an increase in dataset size could continue enhancing the performance. Investigating the synthesis of additional data to push the performance boundaries of LEMMA could be a future work.
References
Appendix A Appendix
To further validate the robustness of our method across different models, we conduct additional experiments using Mistral-7B-v0.1 and Qwen2-Math-7B as base models. The results, presented in Tab.7, demonstrate that LEMMA consistently outperforms baseline methods on these models. Specifically, LEMMA achieves an average accuracy improvement of at least 2.9% on Mistral-7B-v0.1 and 3.5% on Qwen2-Math-7B. These consistent performance gains across different base models reinforce the robustness of our approach, highlighting LEMMA’s efficacy in enhancing mathematical reasoning capabilities across a diverse range of models.
A.1.2 Experiment using Other Teacher Model
To facilitate the community, we evaluate the performance of our approach using an open-source teacher model, LLaMA-3.1-Nemotron-70B, instead of GPT-4o. The results, presented in Tab.6, indicate that although there is a performance decrease when replacing the teacher model, LEMMA still maintains a significant advantage over baseline methods. This indicates that the improvements achieved by LEMMA do not stem from the teacher model itself but are primarily due to the effectiveness of the systematic error introduction and correction strategy. The consistent improvement further underscores the robustness of LEMMA.
A.1.3 Evaluation using Majority Voting
We present the accuracy results under the Majority@32 setting in Tab.8. The results demonstrate that LEMMA consistently outperforms baseline methods in both Pass@1 and Majority@32 settings across most tasks. Notably, the Majority@32 setting significantly enhances LEMMA’s accuracy compared to Pass@1, particularly on more challenging datasets such as MATH. For instance, LEMMA achieves improvements of 14.8% and 13.0% on MATH for LLaMA3 and DeepSeek-Math, respectively. In contrast, some baseline methods exhibit limited gains under the Majority@32 setting. For example, RefAug-90k shows only 7.6% and 8.1% improvements on MATH for LLaMA3 and DeepSeek-Math, respectively. These findings further underscore LEMMA’s superiority and its compatibility with majority voting.
A.1.4 Ablation Study on Out-of-Distribution Datasets
We report the ablation performance on both in-distribution and out-of-distribution tasks in Tab.9. The results show that removing either the “error augmentation” module or the “Fresh & Restart” correction strategy degrades performance, validating the design of our LEMMA approach.
A.1.5 Error Type Analysis on GSM8K
In this section, we examine the error types on the GSM8K dataset. We observe a similar trend to that on MATH: the distribution of error types is consistent across different models. However, unlike MATH, the primary error types on GSM8K are “Question Misinterpretation (QM)”, “Calculation Error (CA)” and “Confusing Concept Error (CC)”, while “Formula Confusion Error (FC)”, which is common on MATH, is less frequent on GSM8K. This difference stems from the inherent distinctions between the GSM8K and MATH datasets. The MATH dataset is more challenging, often involving complex mathematical formulas, whereas GSM8K is relatively simpler, with many problems requiring only basic arithmetic operations rather than the application of formulas. As a result, formula-related errors are less common on GSM8K.
In Fig.8, we present the changes in error types on GSM8K before and after fine-tuning. The results align with those observed on MATH, demonstrating that LEMMA consistently reduces the frequency of all error types. In contrast, while the overall performance of the model improves after fine-tuning with the original training data (SFT), certain specific type of error increase. This further highlights LEMMA’s ability to systematically address and mitigate a wide range of errors, leading to more robust and reliable mathematical reasoning capabilities.
A.2 Case Study
In Fig.10 and Fig.11, we present examples from GSM8K and MATH, respectively, showcasing the outputs of the LLaMA3 model fine-tuned with our LEMMA data. These examples demonstrate that LEMMA model consciously identify potential errors in its previously generated steps, reflecting upon them and making necessary corrections, or verifying its answers before reaching a final conclusion. This ability explains why our method significantly improves accuracy in mathematical tasks: by enhancing the model’s reflection and correction skills, it can ultimately rectify mistakes and arrive at the correct answer, even if it initially takes a wrong approach or makes careless errors along reasoning path.
A.2.2 Full Error Taxonomy
We present the full error taxonomy in Table 10. Building upon the taxonomy proposed by Li et al. (2024e), we introduce additional error categories to enable a more granular identification of error types. Specifically, we add “Question Misinterpretation Error (QM)”, “Confusing Concept Error (CC)”, and “Nonsensical Output (NO)” to better capture the diverse range of errors that can occur during mathematical reasoning. The expanded taxonomy provides a structured framework for systematically categorizing and addressing the various types of errors encountered in mathematical problem-solving.
A.2.3 Error Type and Corresponding Examples
In this section, we present examples of different error types generated by the model. As shown in Fig.12, we display the problem, the model-generated incorrect answer, the error type label assigned by the model, as well as the model’s explanation for the label of representative error types. It can be observed that the model accurately identifies the first error type. Each error type exhibits distinct characteristics, clearly differentiating them from one another.
A.2.4 Smart Solution v.s. Brute Force Solution
In Fig.9, we illustrate two typical solutions for solving a given problem: a smart solution and a brute force solution. While the brute force method starts with accurate initial steps, it requires complex calculations in the following steps. If the model initially fails to identify the smart solution, simply correcting the first incorrect step in the brute force solution does not easily lead to the correct final answer due to the complexity of subsequent calculations. Consequently, we propose the “Fresh & Restart” correction strategy, which encourages the teacher model to reconsider and generate new solutions. This strategy enables the model to learn a variety of correction techniques, thereby allowing it to rectify errors more flexibly.
A.3 Experiment Setup
We construct the incorrect-correct reasoning trajectories on the training set of MATH Hendrycks et al. (2021) and GSM8K Cobbe et al. (2021). We use GPT-4o Hurst et al. (2024) as the teacher model in our main experiment. Additionally, we employ an open-source model, LLaMA-3.1-Nemotron-70B Wang et al. (2024b), as an alternative teacher model to demonstrate the generalization of our method, which produces similar results, as shown in Tab.6. To collect incorrect-correct reasoning trajectories based on the questions in the MetaMath dataset, we use LLaMA-3.1-Nemotron-70B as the teacher model to reduce computational costs, given that MetaMath is significantly larger than MATH Hendrycks et al. (2021) and GSM8K Cobbe et al. (2021). For trajectory synthesis, we use nucleus sampling with a temperature of and top_p of . Based on our synthesized data, we fine-tune a wide range of base models, including general-purpose models such as LLaMA3-8B and Mistral-7B-v0.1, as well as the math-specialized model DeepSeekMath-7B and Qwen2-Math-7B. We use the LLAMA-Factory package https://github.com/hiyouga/LLaMA-Factory for model training. We adopt a learning rate of 1e-5 with a warmup ratio of . We employ a cosine learning rate scheduler and set the gradient accumulation step to 8 to ensure stable training. All models are trained for 3 epochs. For evaluation, we use official evaluation package in Qwen2.5-Math repository https://github.com/QwenLM/Qwen2.5-Math/tree/main/evaluation. We set the maximum number of generated tokens to and the temperature to 0 for the Pass@1 metric. For the majority voting setting, we set the temperature to and top_p to .
All our experiments were conducted on a server equipped with 8 x A100 GPUs. Training LLaMA3-8B on our synthesized dataset takes approximately 5 hours.
A.3.2 Implementation Details of Baselines
We compare LEMMA with four self-correction methods and four data augmentation approaches. For self-correction methods, we consider: (1) Intrinsic Self-Correction (ISC) Han et al. (2024): Teaching small language models to self-correct by training on the constructed self-correction data. In our reimplementation, we employ GPT-4o instead of GPT-3.5-Turbo to construct the self-correction data, ensuring the improvements are not attributed to model discrepancy. We synthesize data in total, which aligns with LEMMA in quantity, to guarantee a fair comparison. Because the original prompt from their paper, “Please select the correct option from the provided choices and offer a comprehensive problem-solving process” is designed for multi-choice problems. We adapt this to our needs by using the prompt, “Below is an instruction that describes a task. Write a response that appropriately completes the request. ### Instruction: instruction ### Response: Let’s think step by step” for the initial chain-of-thought (COT) generation. We then follow their official prompt “the answer of [Question] is [Ground-Truth]. Please provide a step-by-step explanation for resolving the given problem” to generate the correct solution using GPT-4o. (2) S3C-MATH Yan et al. (2024): Employing a step-level sampling approach to generate potentially erroneous steps, followed by reflection and improvement, to construct self-correction data. Note S3C-MATH synthesizes a total of data based on MetaMath training set. Therefore, it should be compared with LEMMA (w/ MetaMath). (3) RefAug Zhang et al. (2024b): Appending a “reflection” part to the original solution, which involves proposing an alternative solution and solving a similar problem. We use the officially released “reflection” data and augment it with approximately the same amount of synthetic solutions generated by GPT-4-Turbo, as this configuration yields the best results in their paper. Our reimplementation of RefAug on most tasks is slightly better than the original results reported in their paper. (4) RefAug-90k Zhang et al. (2024b): To eliminate the influence of sample size and the annotation model, we employ the official codehttps://github.com/ytyz1307zzh/RefAug of RefAug to generate three correct reflection sections and three correct solutions for each question-answer pair using GPT-4o. This produces data, which aligns with our approach in terms of both sample size and annotation model.
For data augmentation approaches, we consider: (1) SFT: Training on the union of GSM8K and MATH training set. (2) Rejection Sampling Fine-tuning (RFT) Yuan et al. (2023): Training on the correct self-generated reasoning trajectories. We collect a total of data, which aligns with our LEMMA in quantity, isolating the impact of sample size. (3) MetaMath Yu et al. (2024a): Combining answer augmentation, question rephrasing, and two backward reasoning methods, FOBAR Jiang et al. (2024) and Self-Verification Weng et al. , to augment training data. (4) GPTAug: Prompting GPT-4o to generate step-by-step solution for each question. We generate a total of data, consistent in quantity with our LEMMA, to ensure fair comparison.
Please refer to Tab.5 for an overview of data statistics of the different methods.
A.4 Prompt
In Fig.13, Fig.14, and Fig.15, we present the prompts used for error injection, Fix & Continue correction, and Fresh & Restart correction, respectively. The prompts are designed to guide the teacher model in generating erroneous trajectories and correcting them using the two distinct strategies outlined in our methodology.