ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, Weizhu Chen

Introduction

Large language models (LLMs), such as GPT-4 (OpenAI, 2023) and PaLM-2 (Anil et al., 2023), have demonstrated remarkable progress in a wide range of language tasks, particularly in the longstanding challenge of mathematical reasoning (Feigenbaum et al., 1963; Hosseini et al., 2014). However, open-source models, such as LLaMA-2 (Touvron et al., 2023a; b) and Falcon (Penedo et al., 2023), still struggle with advanced mathematical reasoning tasks.

Existing works improve mathematical performance of language models either with step-by-step natural language reasoning (Wei et al., 2022) as illustrated in Fig 2 (a), or by synthesizing and executing programs to obtain the answers (Gao et al., 2022; Chen et al., 2022), as depicted in Fig 2 (b). Both approaches exhibit complementary advantages. Natural language is suitable for semantic analysis, planning, and abstract reasoning (e.g., commonsense reasoning), but struggles with precise computation, symbolic manipulation, and algorithmic processing. Conversely, programs excel in rigorous operations, and can outsource intricate calculations to specialized tools like equation solvers.

To leverage the benefits of both natural language reasoning and program-based tool use, we train open-source models such as LLaMA-2 to reason in a way where natural language reasoning is interleaved with program-based tool use synergistically (as depicted in Fig 2 (c)), thereby largely reducing the gap with closed-source models like GPT-4 in mathematical reasoning. Specifically, we first design the interleaving format of reasoning, curate corresponding interactive tool-use trajectories for mathematical problems from the popular GSM8k (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021) dataset, and then apply imitation learning on the high-quality annotations, leading to a better performance than any existing open-source model. Furthermore, since the curated data is far from exhausting all valid trajectories for a problem, relying solely on imitation learning restricts a model’s output space, hindering the flexibility in exploring plausible trajectories during testing. To improve the diversity of plausible reasoning steps and mitigate improper tool-use behavior, we apply output space shaping which additionally trains the models on both self-sampled valid trajectories and invalid ones that have been corrected by a teacher model (e.g., a 34B model can serve as the teacher for a 7B model). Output space shaping significantly boosts reasoning, allowing open-source models to attain an accuracy exceeding 50% on the competition-level MATH dataset for the first time.

We evaluate the resulting suite of Tool-integrated Reasoning Agents (ToRA) ranging from 7B to 70B on 10 diverse mathematical reasoning datasets. As shown in Fig 1, ToRA series significantly outperform open-source models across all scales. Notably, on the competition-level MATH dataset, ToRA-7B outperforms the previous SoTA WizardMath-70B (Luo et al., 2023) by 22% absolute. ToRA-Code-34B beats GPT-4’s CoT result (Bubeck et al., 2023) by 8.3% absolute (50.8% vs. 42.5%), and is competitive with GPT-4 solving problems with code (GPT-4-Code, 51.8%). In addition, we analyze the benefits and remaining challenges of tool interaction for mathematical reasoning, providing valuable insights for future work.

ToRA: Tool-Integrated Agents for Mathematical Reasoning

ToRA series solve challenging mathematical problems by leveraging both natural language reasoning and program-based tool use. As shown in Fig 2 (c), given a mathematical problem qq, ToRA reasons with natural language, producing r1r_{1}. When reaching a point where program-based tool use is more appropriate for the subsequent task, e.g., equation solving, ToRA generates a program a1a_{1} for tool use following natural language guidance r1r_{1}. The execution output o1o_{1} will be fed to ToRA for subsequent processing including tool use adjustments, sub-tasks solving, or answer finalization. We repeat the process until the model places its answer within “\boxed{}”. The resulting trajectory is denoted as τ=r1a1o1...rn−1an−1on−1rn\tau=r_{1}a_{1}o_{1}...r_{n-1}a_{n-1}o_{n-1}r_{n}, where rnr_{n} contains the answer.

Fig 3 presents the training pipeline of ToRA. We first collect interactive tool-use trajectories on popular mathematical datasets. We then apply imitation learning on the resulting annotations, as well as output space shaping to further refine models’ reasoning behavior.

2 Collecting Interactive Tool-Use Trajectories

Existing mathematical reasoning datasets primarily contain annotations in either natural language or code, posing a challenge for training tool-integrated agents due to the absence of interactive tool-use annotations. To address this, we utilize GPT-4 to synthesize high-quality trajectories on the GSM8k and MATH training sets. We select GSM8k and MATH as they exhibit diverse reasoning patterns, spanning multiple domains and difficulty levels.

We compose instructions along with diverse few-shot examples, utilizing an interleaved format as depicted in Fig 2 (c). These examples showcase interactive tool usage trajectories, incorporating descriptive variable names and combined program outputs. Please refer to Appendix E for the assembled prompts.

We follow Algorithm 1 and feed GPT-4 (G\mathcal{G}) with the composed prompt pp to generate a tool-use trajectory τ\tau for each question qq from the training set. The trajectory is initialized as an empty string τ0\tau_{0}, for each interaction round ii, we first generate a rationale:

where ⊕\oplus means concatenation. If rir_{i} includes an answer within “\boxed{}” (i.e., the stopping condition Stop(rir_{i})), we cease generation, otherwise the model continues to write a program for tool use:

In line with Gou et al. (2023), if the model triggers the code execution stop words like “‘‘‘output”, we supply it with the corresponding execution message and output oio_{i} by calling tools with oi←E(ai)o_{i}\leftarrow\mathcal{E}(a_{i}), facilitating the generation of subsequent steps. Then, we update the trajectory by concatenating it with the newly generated rationale rir_{i}, program aia_{i}, and output oio_{i}:

We repeat the above interaction process until we reach the maximum rounds nn.

We set n=3n=3 and perform inference using GPT-4 with greedy decoding, retaining trajectories that yield correct answers. For questions where GPT-4 fails with greedy decoding, we apply nucleus sampling with a sample size of 10 and keep up to 4 valid trajectories per question. Ultimately, we successfully annotate trajectories for 98.2% of GSM8k questions and 83.1% of MATH questions. After filtering out invalid trajectories with tool-use errors or wrong answers, we obtain 16k annotations which constitute our dataset ToRA-Corpus. Table 1 compares ToRA-Corpus with recently proposed mathematical reasoning datasets, while Table 6 in the Appendix displays MATH annotation accuracy details.

3 Training

We apply imitation learning on ToRA-Corpus by minimizing negative log-likelihood loss on the trajectory τ\tau conditioned on the problem qq:

where M\mathcal{M} is the resulting model. After imitation learning, we can simply apply the same procedure in Algorithm 1 by setting prompt to empty p=""p=\text{""} for inference. Imitation learning leads to state-of-the-art mathematical reasoning performance despite the small scale of ToRA-Corpus.

For each question, ToRA-Corpus mostly demonstrates only one valid interactive tool-use trajectory, which may restrict a model’s output space, rendering it inflexible in exploring plausible trajectories during testing. We therefore propose output space shaping in order to encourage the diversity of plausible reasoning steps and reduce improper tool-use behavior.

In our experiments, we always use CodeLLaMA-34B trained on ToRA-Corpus as the teacher model, and apply sampling with the CodeLLaMA series (ranging from 7B to 34B, with imitation learning). We obtain a total of 233k distinct valid trajectory samples and 69k corrected ones. From this combined dataset, we randomly select up to 4 trajectories per GSM8k and MATH problem, merge them with ToRA-Corpus, and then train all ToRA models on the resulting 69k annotations.

Experiments

We fine-tuned LLaMA-2 (Touvron et al., 2023b) and CodeLLaMA (Rozière et al., 2023) series (ranging from 7B to 70B) using ToRA-Corpus with output space shaping, yielding the ToRA and ToRA-Code series respectively. We used a learning rate of 2e-5 by default except that we used 1e-5 for the 34B and 70B models. We set the global batch size to 128 and used a linear scheduler with a 3% warm-up period for 3 epochs. We trained all models with DeepSpeed ZeRO Stage3 (Rajbhandari et al., 2021) and Flash-Attention 2 (Dao, 2023). We used greedy decoding for all results, with the maximum sequence length set to 2,048 and the maximum number of tool executions set to 3.

2 Evaluation Setup

Datasets We evaluated models on GSM8k (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021), along with 8 out-of-distribution datasets, namely GSM-Hard (Gao et al., 2022), SVAMP (Patel et al., 2021), ASDIV (Miao et al., 2020), TabMWP (Lu et al., 2023), SingleEQ, SingleOP, AddSub, and MultiArith (Koncel-Kedziorski et al., 2016), as illustrated in Table 5 in Appendix. The 10 assorted datasets collectively encompass mathematical problems spanning basic arithmetic to competition level, covering middle and high school curricula and various mathematical domains. The problem formats comprise tabular-based, free-form, and multiple-choice questions, ensuring a thorough assessment of the model’s mathematical reasoning aptitude.

Metrics We report accuracies of predicted answers. Following Lightman et al. (2023), we round numerical values and use sympy https://www.sympy.org for parsing expressions. Since the SingleEQ, SingleOP, AddSub, and MultiArith datasets focus on different aspects of basic arithmetic, we report their average results under the collective term MAWPS (Koncel-Kedziorski et al., 2016) for all methods.

3 Baselines

Proprietary Models We present results from an array of SoTA LLMs, such as OpenAI’s GPT-4, ChatGPT (gpt-3.5-turbo), Google’s PaLM-2, and Anthropic’s Claude-2. By default, we report CoT prompting results, and include PAL (Gao et al., 2022) prompting results for selected models.

Open-Source Models Base models comprise LLaMA-2 and CodeLLaMA with CoT and PAL prompting. Supervised Fine-Tuning (SFT) employs CoT rationales from the original GSM8k and MATH dataset (15k samples) for fine-tuning. Rejection sampling Fine-Tuning (RFT) leverages multiple models to generate diverse reasoning paths for fine-tuning (Yuan et al., 2023). WizardMath augments data using ChatGPT, and conducts SFT and RLHF. Platypus-2, the top model on the LLM Leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, is fine-tuned with Open-Platypus reasoning datasets (Lee et al., 2023). We also compare ToRA with Toolformer (Schick et al., 2023) which is a model trained to utilize calculators.

4 Main Results

Table 2 presents the results of ToRA on 10 mathematical datasets, highlighting the following salient observations: (1) Using interleaved formatting and output space shaping, ToRA consistently surpasses prior state-of-the-art open-source models across all scales, achieving 13% to 19% absolute improvements across 10 tasks. (2) ToRA-70B substantially outperforms ChatGPT with both CoT and PAL prompting on GSM8k (84.3% vs. 80.4%) and MATH (49.7% vs. 38.7%), while ToRA-Code-34B is competitive with GPT-4 solving competition-level MATH dataset with code (50.8% vs. 51.8%). (3) The accuracy of ToRA-Code is about 5% higher than ToRA of the same size, demonstrating that continued training on code data significantly benefits program-based tool use. (4) While rationale-based fine-tuning negatively affects out-of-distribution generalization, ToRA displays superior generalization. For instance, WizardMath-70B underperforms the base model on TabMWP (49.8% vs. 57.5%), while ToRA-70B effectively generalizes to this tabular reasoning task (74.0%). (5) ToRA attains fast zero-shot inference speed, averaging 1.02 tool interaction rounds per problem, while effectively addressing problems that require interactive tool utilization.

5 Ablation Study

To evaluate the efficacy of the reasoning format adopted by ToRA which interleaves rationales with programs, we compared it with Rationale-only and Program-only formats using GPT-4 and LLaMA-2 trained with the same size of data from MATH. As shown in Fig 4, the ToRA method consistently surpasses Rationale-only and Program-only approaches. Remarkably, using LLaMA-2, the ToRA method achieves substantial improvements of 29.0% and 6.7% over Rationale-only and Program-only, respectively. With the closed-source GPT-4, the improvements are 19.1% and 9.8%, respectively. This emphasizes the effectiveness of integrating natural language rationales with programs.

5.2 Effects of Output Space Shaping

We assess the effectiveness of the output space shaping strategies presented in Section 2.3, specifically sampling and correction. As shown in Fig 5 and Table 3: (1) Output space shaping yields a considerable average improvement of 3.4% and 4.0% absolute for GSM8k and MATH, respectively, with greater benefits for smaller models; (2) Applying the sampling strategy results in a 2.7% absolute improvement on average, while additionally incorporating correction offers a more substantial boost of up to 4.5%, without using more training data; (3) Output space shaping benefits even the largest model ToRA-70B, with a notable improvement from 47.3% to 49.7% on MATH. These findings highlight the effectiveness of our shaping strategies across different model sizes and datasets.

6 Analysis

We investigate the benefits, detailed patterns, and remaining challenges of tool interaction for mathematical reasoning on the challenging MATH dataset. Performance breakdowns on all subtopics of MATH are reported in Table 3.

Benefits from Tool-Integration for MATH Sub-topics As shown in Table 3, ToRA outperforms WizardMath by around 45% in Algebra and Number Theory, which is attributed to stimulating and shaping tool-use behavior. Problems from the two sub-topics typically need intricate computation and data manipulation. Algebra mainly focuses on solving equations and application problems, while many Number Theory problems can be tackled using brute-force approaches through code.

Patterns of Library Usage for Problem Solving Fig 6 presents the most frequently used libraries for different sub-topics and the corresponding accuracies of their solutions. Tool-use behavior on different mathematical areas demonstrates distinct patterns. sympy and its internal solvers are primarily employed for algebra-related topics. Precalculus exhibits extensive matrix operations via matrices, resulting in a high accuracy. Number Theory depends on algorithms like gcd and lcm. Geometry mainly uses the rational library for fraction-based computations, while the application of other tools is limited, signifying the potential for improvement.

Detailed Impact of Rationale on Different Topics Table 3 shows that using an interleaved format, in contrast to merely writing the program, leads to significant improvements across all subtopics, especially in Precalculus, Algebra, and Geometry, where notable increases range from 8.6% to 18.8%. Appendix F.1 provides representative examples demonstrating how the rationale aids in planning, multi-round self-correction, and finalizing answers.

Remaining Challenges in Mathematical Reasoning for ToRA To better understand the failure modes and remaining challenges, we manually annotated 100 randomly selected trajectories from the MATH test set, identifying and categorizing their failure modes. The results are shown in Table 4: Primarily, incorrect reasoning steps constitute the primary source of errors for ToRA on complex math reasoning tasks (38%), with some hallucination issues also evident during problem interpretation and answer finalization (5%). Secondly, the misinterpretation of input diagrams contributes significantly to the error rate (21%). This is particularly noticeable in Geometry, Precalculus, and Intermediate Algebra. The diagrams in the MATH dataset are usually detailed in text using the Asymptote language (Hendrycks et al., 2021), thus making it challenging for ToRA to comprehend diagrams purely from textual descriptions. Thirdly, issues with tool usage include Inappropriate Tool Usage (10%), Syntax Error (9%), and Runtime Error (9%). These problems frequently arise when ToRA fails to use tools correctly after several corrections or attempts. There are certain inputs that fail to formalize well as programs (3%), which require abstract reasoning rather than computation. Finally, we also found that there are false negatives when using automatic indicators, i.e., correct predictions that are misjudged as wrong, but the proportion is relatively small (5%).

Conclusion

This paper presents ToRA, a series of novel Tool-integrated Reasoning Agents that synergistically combines natural language rationale with program-based tool-use for mathematical problem solving. Our approach demonstrates the potential of integrating external tools in the reasoning process, enabling language models to effectively tackle complex quantitative tasks. ToRA achieves state-of-the-art performance on 10 diverse mathematical reasoning tasks, substantially outperforming existing rationale-based and program-based approaches. Furthermore, our systematic analysis of the benefits and remaining challenges of tool interaction provides valuable insights for future research, contributing to the development of more advanced and versatile reasoning agents.

Zhibin Gou proposed the interleaved tool-use format of ToRA and curated ToRA-Corpus dataset, implemented the training and evaluation pipeline, conducted experiments and analysis on all datasets, implemented baselines, and was a main contributor to the paper writing. Zhihong Shao proposed the project, conducted preliminary experiments, proposed and implemented the training and evaluation pipelines, proposed and trained all ToRA models with output space shaping as well as ToRA variants in the ablation study, designed and oversaw experimental analysis, and contributed to many parts of the paper writing. Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen provided research mentorship, oversaw project coordination, and advised and contributed to many parts of the writing.

Acknowledgments

Zhibin Gou and Yujiu Yang were supported by the National Natural Science Foundation of China (Grant No. U1903213) and the Shenzhen Science and Technology Program (JSGG20220831110203007). Zhihong Shao and Minlie Huang were supported by the NSFC projects (Key project with No. 61936010 ), and were also supported by the National Science Foundation for Distinguished Young Scholars (with No. 62125604).

References

Appendix A Related Works

Mathematical Reasoning Recent research has greatly improved reasoning in LLMs with step-by-step natural language reasoning (Polu & Sutskever, 2020; Wei et al., 2022; Zhou et al., 2023b; Zhu et al., 2023; Huang et al., 2022; Liang et al., 2023). However, natural language reasoning struggles with complex computations and symbolic manipulations. To overcome the limitations, recent research has exploited tools like calculators (Cobbe et al., 2021; Shao et al., 2022), code interpreters (Mishra et al., 2022), and symbolic solvers (Zhang et al., 2023). Program-based methods (Gao et al., 2022; Chen et al., 2022; Shao et al., 2023a) transform reasoning tasks into program synthesis tasks, thus offering complementary advantages over natural language reasoning, but they face challenges in nuanced reasoning, planning, and error handling (Gou et al., 2023), where natural language reasoning should be more suitable.

Tool-Augmented Language Models Augmenting LLMs with tools can largely alleviate LLMs’ limitations and improve reasoning and generation performance (Parisi et al., 2022; Mialon et al., 2023; Yao et al., 2023). Recent work demonstrates the benefits of integrating retrievers (Borgeaud et al., 2022; Shao et al., 2023b), search engines (Nakano et al., 2021), and multi-tool approaches (Schick et al., 2023; Paranjape et al., 2023; Gou et al., 2023) to improve generation.

Knowledge Distillation Knowledge distillation (KD) transfers knowledge from teacher models to student models (Buciluǎ et al., 2006; Hinton et al., 2015). Using LLM-generated trajectories for fine-tuning is a form of KD (Fu et al., 2023; Taori et al., 2023; Peng et al., 2023; Ho et al., 2023). Our proposed ToRA shows that learning interactive tool-use trajectories is a promising direction to adapt language models to reasoning tasks.

Appendix B Evaluation Datasets

We present statistics and examples of the ten evaluation datasets in Table 5.

Appendix C Additional Experiments and Analysis

Table 6 presents the detailed accuracies of GPT-4 on the MATH dataset. The Tool-integrated Reasoning method used by ToRA significantly outperforms PAL prompting when directly applied to the closed-source GPT-4, further demonstrating the benefits of synergizing natural language reasoning and program-based tool use.

C.2 Effects of # Valid Trajectories for Output Space Shaping

As shown in Fig 7, it is beneficial to increase the number of additional valid trajectories for output space shaping.

C.3 Impact of Output Space Shaping in Relation to Question Difficulty

We compare the effects of output space shaping on MATH problems of different difficulty levels (from level 1 to level 5) in Figure 8, and present the statistics of MATH problems at different levels in Table 7. As can be seen:

Across these different difficulty levels and model sizes, output space shaping generally brings a significant improvement of 4.0% on average across different model sizes.

Output space shaping brings significant improvements for difficult, long problems. E.g., with ToRA-Code-13B, shaping does not significantly improve level 1 to level 2 problems, but it brings a substantial improvement of 5.4% to 5.7% for level 3 to level 5 problems.

After using shaping, ToRA-Code-34B outperforms GPT-4 PAL on problems from Level 1 to Level 4, but there is still a gap at Level 5 (27.3% vs. 30.0%). These problems are usually longer (average about 248.4 characters), require more reasoning steps (>1,000 characters) to solve, and more often include diagram inputs (about 20%). These observations may guide future work to focus more on solving these more difficult problems.

Appendix D Detailed Information of ToRA-Corpus

We provide a more detailed introduction to the data construction process, quality control, and data statistical information, beyond Sec. 2.2.

In our preliminary experiments, we found that the tool-integrated reasoning trajectory format generated by zero-shot prompting was somewhat chaotic. Therefore, we designed a few-shot prompting to control the reasoning format, which effectively improved data quality. On the other hand, we increased the annotation success rate by sampling, ensuring more comprehensive coverage of the training query.

For the data constructed, we filtered out paths that produced incorrect answers by matching them with standard answers. To prevent the model from learning incorrect intermediate reasoning processes, we further filtered out data samples with intermediate program execution errors.

In Table 8, we compared the annotation accuracy (i.e., sample coverage) of the training set on GSM8k, MATH, and MATH subtopics of ToRA-Corpus-Greedy using only the greedy trajectories, and ToRA-Corpus-16k combined with sampled trajectories. Furthermore, in Table 9, we reported the statistical data of ToRA-Corpus-16k, such as the number of samples, average question length, average, minimum, and maximum trajectory length, as shown in the following tables.

As described in Section 2.2, we annotated interactive tool-use trajectories for the training questions from MATH with GPT-4. GPT-4 achieves a success rate below 65% using greedy decoding. As MATH was originally annotated with natural language rationales, to improve the annotation success rate, we tried to provide GPT-4 with the human rationales as hints (Zelikman et al., 2022). However, when using this method, GPT-4 tends to replicate the hints and ignore tool-use outputs especially when the outputs are inconsistent with the hints, thus failing to produce high-quality trajectories. Hence, we deferred the utilization of the already-annotated natural language rationales for future investigations. Instead, we employed nucleus sampling to recall valid trajectories for questions that remained unsolved through greedy decoding. This approach significantly boosted annotation accuracy to 83.1%.

Appendix E Prompts

We present instructions and example few-shot prompts of Tool-integrated Reasoning for querying GPT-4.

Appendix F Examples

F.2 Failure Cases