Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification
Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, Hongsheng Li
Introduction
Large language models (LLMs) (Brown et al., 2020; OpenAI, 2023; Anil et al., 2023) have shown impressive success in various tasks, such as common sense understanding and code generation. However, they still fall short in mathematical reasoning, often producing nonsensical or inaccurate content and struggling with complex calculations. Previous attempts to tackle these challenges include the Chain-of-Thought (CoT) (Wei et al., 2022) framework, which enhances LLMs’ logical reasoning abilities by generating intermediate steps in their reasoning process. Additionally, PAL (Gao et al., 2023) introduces a novel approach by using the Python programming interpreter to improve computational accuracy.
In recent advancements, OpenAI has unveiled an improved version of GPT-4, namely the GPT-4 Code Interpreterhttps://openai.com/blog/chatgpt-plugins##code-interpreterhttps://chat.openai.com/?model=GPT4-Code-interpreter or GPT4-Code, which is proficient at providing logical natural language reasoning, alongside step-by-step Python code. Notably, it can generate and execute code incrementally, and subsequently present the executed code’s output back to the LLM. The addition of code generation and execution to natural language outputs has shown promising results in solving mathematical reasoning problems. Our initial experiments show that GPT4-Code achieved an impressive zero-shot accuracy of 69.7% on the challenging MATH dataset (Hendrycks et al., 2021), marking a significant improvement of 27.5% over GPT-4’s performance (42.2%).
While GPT4-Code has demonstrated proficiency in solving math problems, there has been a notable absence of systematic analysis focusing on understanding and further enhancing its mathematical problem-solving abilities. A key distinction between GPT4-Code and its predecessor, GPT-4, lies in GPT4-Code’s ability to automatically generate and execute code. Therefore, this paper presents pilot experiments that investigate GPT4-Code’s code generation and execution mechanism using specific code-constrained prompts. The analysis reveals that GPT4-Code’s strong performance is not solely due to its code generation and execution abilities, but also its capacity to adjust its problem-solving strategies based on feedback from code execution—a process we term self-debugging (illustrated in Tab. 7 and Tab. 8). Due to the fact that code generation evolves its reasoning step-by-step and performs self-debugging after code execution errors, there is an increased frequency of code usage. Hence, we introduce the concept of Code Usage Frequency to differentiate these unique prompting strategies to quantitatively analyze the impact of code-constrained prompts on GPT4-Code for mathematical problem-solving.
The step-by-step code generation and self-debugging mechanisms highlight the critical role of code in mathematical problem-solving. Nevertheless, the self-debugging mechanism only verifies each step of the generated code while lacks the verification of the reasoning steps and the final answer, which has been demonstrated to be of vital importance to the math problem-solving abilities of LLMs (Cobbe et al., 2021; Lightman et al., 2023; Weng et al., 2023).
We therefore ask the question: can we fully exploit the code generation and self-debugging mechanisms in GPT4-code, so that it can automatically verify and correct its solutions, without extra assistance from other models or users?
To answer this question, we propose a simple yet effective prompting technique termed the explicit code-based self-verification (CSV), which guides GPT4-Code to generate additional code that verifies the answer and adjusts the reasoning steps if there’s a flaw in reasoning. Unlike previous methods that rely on external language models for verification (Lightman et al., 2023; Cobbe et al., 2021), our approach leverages GPT4-Code’s inherent strengths. This approach offers two key benefits: (1) When the verification indicates an answer is False, GPT4-Code can rectify its prior solution and provide an improved alternative. (2) Solutions verified as True tend to be more reliable, akin to human problem-solving. However, even if a solution is self-verified as False, we do not directly abandon it. Instead, we propose a weighted majority voting strategy that incorporates the code-based solution verification results, as opposed to relying exclusively on the frequency of answers. We assign different weights to the solutions according to their verification states, reflecting the solutions’ varying levels of reliability. In alignment with the Code Usage Frequency analysis from our pilot experiments, our explicit code-based self-verification prompt boosts GPT4-Code’s accuracy in mathematical problem-solving with increased code usage.
Empirical study demonstrates the effectiveness of our proposed framework on the MATH, GSM8K, and MMLU-Math datasets using GPT4-Code. Our approach achieves an impressive accuracy of 84.32% on the MATH dataset, greatly outperforming the base GPT4-Code and previous state-of-the-art methods. Additionally, we are making our experimental data on the MMLU-Math and MATH datasets publicly available, enabling result replication and facilitating fine-tuning of open-source LLM models (e.g., LLaMA 2 (Touvron et al., 2023)) to further enhance mathematical problem-solving capabilities with the assistance of code generation.
This paper’s main contributions can be summarized in three key aspects:
This study provides the first systematic analysis of code generation, execution, and self-debugging’s role in mathematical problem-solving. Our findings reveal that GPT4-Code’s impressive mathematical problem-solving proficiency is primarily attributed to its step-by-step code generation and dynamic solution refinement based on code execution outcomes.
We introduce the innovative explicit code-based self-verification (CSV) prompt, which leverages GPT4-Code’s advanced code generation mechanism. This prompt guides the model to verify the answer and then reevaluate its solution with code. CSV not only extends the verification to the logic behind problem-solving but also improves the efficacy of the majority voting method by integrating the verification states.
Additionally, we have contributed to the LLM community by creating two new instruction-following datasets: MATH-code and MMLU-Math-code. These datasets are designed to enhance the mathematical reasoning capabilities of open-source models.
Related work
Chain-of-Thought Reasoning. The Chain-of-Thought (CoT) prompting approach proposed by Wei et al. (2022) is a notable contribution that showcases the multi-step reasoning capabilities of LLMs. By simply adding “Let’s think step by step” before questions, Kojima et al. (2022) implements Zero-shot-CoT, which can serve as a strong zero-shot baseline. Further research extends the reasoning capabilities of CoT by applying majority voting to improve self-consistency (Wang et al., 2023), choosing few-shot examples and output chains with more complex reasoning steps (Fu et al., 2022), breaking down the problem to simpler sub-problems (Zhou et al., 2023), or even expanding Chain-of-thought to Tree-of-Thoughts (Yao et al., 2023). Similar to Zero-shot-CoT, our method apply “step by step”-like prompts to regularize GPT4-Code’s use of code without the careful design of step-by-step few-shot examples. Additionally, We enhance majority voting to verification-guided weighted majority voting, leveraging the results of CSV as voting weights.
Solving Math Problems with Code. Large language models have been found to be less accurate in performing arithmetic calculations, such as addition, subtraction, multiplication, etc (Cobbe et al., 2021; Lewkowycz et al., 2022; Gao et al., 2023; Lu et al., 2022). Consequently, previous works have attempted to solve math problems with the assistance of code. The GSM8K dataset (Cobbe et al., 2021) uses calculation annotations to extract all arithmetic calculations solved by an external calculator: the Python eval function. To further leverage the role of code in LLMs, Program-Aided Language model (PAL) (Gao et al., 2023) as well as Program of Thoughts (PoT) (Chen et al., 2022) interpret the math problems as Python codes and execute the codes with an external Python interpreter to obtain the answer. Although they can get more accurate answers than some non-code methods, many generated codes have execution errors or get wrong answers due to the lack of verification mechanism. Our approach not only utilizes the ability of GPT4-Code to generate multi-step codes and refine codes that fail to run, but also uses CSV to enhance the reliability and accuracy of the answers.
Self-Verification. Human problem solving is not always a one-time success, but rather requires iterative thinking, verification, and refinement. Unlike previous studies that train an additional verifier to verify the correctness of final answers (Cobbe et al., 2021) or intermediate steps (Lightman et al., 2023; Li et al., 2023), Weng et al. (2023) showed the self-verification abilities of LLMs by generating multiple answers and ranking them by self-verification scores. Furthermore, SELF-REFINE proposed by Madaan et al. (2023) iteratively refines its output through self-generated feedback. Unlike these self-verification methods that require LLMs to give verification feedback in natural language, our method applies generated codes to verify the answers and votes on different answers based on the verification results, thus improving the accuracy of the verification and making full use of the information in the verification process.
Method
We first conduct a pilot experiment with GPT4-Code on the challenging MATH dataset (Hendrycks et al., 2021). Remarkably, it achieves an accuracy of 69.7%, significantly surpassing the previous state-of-the-art performance of 53.9% (Zheng et al., 2023). Encouraged by the compelling performance of GPT4-Code, we strive to systematically explore and analyze its underlying code mechanisms. In Sec. 3.1, we illustrate, via our code-constrained prompts design, that GPT4-Code’s robust performance in solving math problems derives not only from its ability to generate accurate step-by-step code, but also from its self-debugging mechanism. In Sec. 3.2, we aim to leverage GPT4-Code’s self-debugging strengths to further improve its mathematical problem-solving ability.
To explore the impact of code usage on GPT4-Code’s math problem-solving capabilities, we adopt a straightforward approach by constraining GPT4-Code’s interaction with code through thoughtfully constructed prompts. Specifically, we introduce two code-constrained prompts and the basic prompt for comparison:
Prompt 1: No code usage is allowed: In response to this prompt, GPT4-Code is prohibited from incorporating code into its solutions. This prompts GPT4-Code to rely solely on Natural Language (NL) reasoning chain, resembling solutions in the CoT framework (Wei et al., 2022). The resulting sequence of reasoning steps is depicted as , with an example given in Fig. 1 (a).
Prompt 2: Code can be used only once: In this prompt setting, GPT4-Code is permitted to employ code exclusively within a single code block to generate the solution, mirroring the PAL approach introduced by Gao et al. (2023). We denote this sequence as , representing a series of Symbolic Language (SL), such as Python. An example is shown in Fig. 1 (b).
Basic Prompt: GPT4-Code is prompted to tackle the problem without any restrictions on code usage. This prompt leads to GPT4-Code’s typical functioning pattern, which can be denoted as , representing a sequential list of reasoning steps, each consisted of both natural language and Python code, with an example shown in Fig. 1 (c).
Apart from the specific example in Fig. 1, we introduce Code Usage Frequency to record the number of the code execution for different prompts. The results of the experiments using these prompts are shown in Fig. 2 (b). This figure illustrates a positive correlation between the better performance of GPT4-Code and the higher Code Usage Frequency. More specifically,
Prompt 1 v.s. Prompt 2: Prompt 1 results in almost negligible code usage, while Prompt 2 results in approximately 1 time’s code usage. Prompt 2 yields an accuracy gain of 6.9 percent over Prompt 1. This suggests that the Python code chains , can improve computational capability more than the natural language chains . This observation is consistent with the findings in previous Python-based prompting methods (Gao et al., 2023; Chen et al., 2022). However, employing code only once comes with an inherent drawback – the model lacks the ability to self-debugging when the code output triggers an error or produces an implausible outcome.
Prompt 2 v.s. Basic Prompt: The Basic Prompt consistently produces solutions that entail multiple instances of code usage, resulting in a large Code Usage Frequency. Additionally, the Basic Prompt exhibits notably enhanced accuracy. These improvements in Code Usage Frequency and accuracy might be attributable to two unique advantages: (1) Generating code in brief and frequent segments, divided among natural language reasoning steps, tends to result in higher accuracy. (2) The model possesses the capability to evaluate the results of code execution and make corrections to solution steps if the outcomes contain bugs or are deemed illogical, as illustrated in Tab. 7 and Tab. 8.
From these observations, it is plausible to enhance and build upon the favorable attributes of GPT4-Code, to further improve its precision in tackling math problems.
2 Explicit Code-based Self-verification Prompting
Inspired by the observations on Code Usage Frequency analysis, we seek to harness the capabilities of GPT4-Code. These capabilities include the model’s aptitude for generating accurate code, evaluating the outcomes of code execution, and automatically adjusting reasoning steps of solutions when needed. However, despite these advantages, GPT4-Code currently falls short in assuring solution correctness. Consequently, our objective is to utilize these strengths to augment solution verification.
To achieve this objective, we propose the technique termed as explicit code-based self-verification (CSV). This method prompts GPT4-Code to explicitly validate its answer through code generation. By implementing this prompt, we introduce an extra verification stage to the solution , referred to as . The verification outcome can be classified as True, False, or Uncertain. An Uncertain classification indicates that GPT4-Code encountered difficulties in identifying an effective method for answer verification, thereby abstaining from delivering a definitive verification result. Leveraging GPT4-Code’s inherent autonomous capabilities, we can formulate the proposed prompting as follows:
An example is presented in Fig. 3 (b). Incorporated with CSV, the model becomes capable of using code to verify answers, then reviewing and adjusting how it arrived at the solution if the verification result is False, aiming at obtaining the correct answer. Upon refining and correcting the initial solution, we anticipate a notable increase in accuracy. It is worth noting that both the verification and rectification stages are code-based. This inevitably results in increased Code Usage Frequency, akin to the aforementioned analysis, which will be further demonstrated in subsequent experiments.
We perform experiments with CSV, and these results can be found in Fig. 2. The experiment here is conducted with GPT4-Code on MATH (Hendrycks et al., 2021). In Fig. 2 (b), the accuracy achieved with our proposed CSV prompt consistently surpasses that of the Basic Prompt across all designated difficulty levelsHuman-perceived easier problems are categorized under Level-1 difficulty as per Hendrycks et al. (2021).. Meanwhile, the Code Usage Frequency receives a clear increase.
Before the advent of GPT4-Code, prior frameworks (Lightman et al., 2023; Cobbe et al., 2021) depended on an external LLM to use natural language for verification and well-designed few-shot example prompts. In contrast, our approach simplifies the process by relying solely on a straightforward prompt for GPT4-Code, all in a zero-shot manner. This enables GPT4-Code to autonomously verify and independently rectify its solutions using the advanced code execution mechanism, thereby eliminating the need for customized few-shot examples.
Given that CSV can effectively verify problem-solving answers, we can naturally integrate the verification states into majority voting, akin to the methodology embraced in self-consistency CoT (Wang et al., 2023). Answers deemed True through verification are generally more trustworthy, reflecting the problem-solving approach seen in human cognition (Newell & Simon, 1972; Wang & Chiew, 2010). This improved reliability can be leveraged in the widely-used majority voting process. To exploit this insight, we introduce verification-guided weighted majority voting, which assigns different weights to the states of the verification process.
In practice, it sometimes occurs that once an answer is confirmed as False, no additional verification is conducted, yielding a False verification state. We allocate corresponding weights these states of True, Uncertain, False: , and , respectively.
Similar to the Self-consistency with CoT (CoT-SC) (Wang et al., 2023) in Fig. 4 (a)(ii), our framework can sample paths. For simplicity, we extract pairs of final answers and their corresponding verification results from solutions, represented as , where and denote the -th final answer and final verification result, respectively.
So the voting score for each candidate answer can be expressed as:
Here, represents a candidate answer, denotes the state of verification, and is an element from the set . Each signifies the degree of confidence associated with its corresponding verification state.
Finally, we select the answer with the highest score from all candidate answers,
where refers to the score of answer according to Eq. 1.
It should be noted that when for all , Eq. 1 becomes equivalent to the naive majority voting employed in Self-Consistency with CoT (CoT-SC) (Wang et al., 2023). Typically, we set , which means that an answer verified true has a greater confidence than the one with an uncertain state of verification, while an answer verified false has the lowest degree of confidence. An example of the calculation process within verification-guided weighted majority voting is illustrated in Fig. 4.
Experiments
The MATH dataset (Hendrycks et al., 2021) is recognized as the most challenging math word problem dataset, as also highlighted by Chen et al. (Chen et al., 2023). Most of our experiments and the corresponding analyses are performed on the MATH benchmark. Tab. 1 compares the performance of the GPT4-Code against other models. GPT4-Code reaches 69.69% on MATH (Hendrycks et al., 2020), largely surpassing the previous state of art result (53.90%), which shows that GPT4-Code exhibits strong abilities in solving math problems and is used as our baseline for ablation study. On top of GPT4-Code, our method further improves its accuracy, raising the result to 73.54% after adding explicit code-based self-verification, and 84.32% after adding both explicit code-based self-verification and verification-guided weighted majority voting (the number of sampled paths is 16). Note that this astonishingly high result is based on the strong abilities of the base model GPT4-Code, and our method amplifies its good qualities of GPT4-Code, with the ability to verify solutions.
Note that although adding Code-based Self-verification can improve the performance of every individual subject, the extent of improvement varies from subject to subject, from 7.6% to only 0.6%. In particular, the Geometry problem only has an increased accuracy of 0.6%, even though the original GPT4-Code accuracy is only 54.0%, which is low among the subjects. This discrepancy may be attributed to the fact that solving geometry problems often requires multi-modality (Chen et al., 2023), a concept beyond the scope of this paper.
2 Performance on other datasets
In addition to the challenging MATH dataset, we have also performed our method on other reasoning datasets such as GSM8K (Cobbe et al., 2021), MMLU-Math, and MMLU-STEM (Hendrycks et al., 2020). The corresponding results can be viewed in Tab. 3 and Tab. 3. When integrated on top of GPT-4-code, our method outperforms other methods in the competition, achieving state-of-the-art results across all datasets. Other subjects in MMLU benchmarks are provided in Fig. 8. A comparative analysis of our results with those of previous state-of-the-art techniques and open-source models is also provided.
Tab. 3 illustrates that verification-guided majority voting is an effective framework to reduce the number of sampled paths, compared to GPT-4 with model selection (Zhao et al., 2023) and PHP (Zheng et al., 2023).
Tab. 3 presents a comparison of our model’s performance with existing models (Hoffmann et al., 2022; Taylor et al., 2022) on the MMLU-Math dataset and with state-of-the-art open-sourced modelshttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard on MMLU-STEM. The open-source models remain significantly outpaced by their closed-source counterparts. To address this gap, we have made the dataset and will make it publicly available in the near future. Our intention is to facilitate the fine-tuning of open-source LLMs. For example, the open-source model LLaMA 2 (Touvron et al., 2023) can potentially utilize this data to further bolster its math reasoning capabilities.
3 Code usage frequency of proposed prompts
Analogous to the approach taken in Sec. 3.1, we gather data to elucidate the correlation between accuracy and Code Usage Frequency across various dimensions - prompts (proposed CSV prompt as well as prompts used in pilot experiments), subjects, and difficulty levels. As shown in Fig. 5,
the model’s behavior is in accordance with our expectations when adding the code-based prompts. Each line in Fig. 5 has an obvious trend of going upwards, proving that the increase of Code Usage Frequency induces a general improvement in accuracy. The performance gain when using more code is more obvious in the higher difficulty levels, while in lower levels, the performance gain is not very prominent, as shown in Fig. 5 (a). Also, the Code Usage Frequency increases steadily with the increase of difficulty levels. This shows that the harder math problems require more frequent code usage, which implies that invoking code multiple times might be an important reason why GPT4-Code have such an advantage in solving difficult math problems. There is a similar trend in Fig. 5 (b).
4 Ablation Study and Discussion
Comparisons between Natural Language and Code-based Self-Verification: to underscore the significance of code in the self-verification stage, we employed a distinct natural language self-verification prompt. In this approach, GPT4-Code is directed to verify the solution through natural language instead of relying on code-based verification, as presented in Tab. 4. The accuracy achieved with this method was slightly lower than that of the Basic Prompt. Moreover, we observed a decline in accuracy for 4 of the 7 subtopics, indicating that relying solely on natural language verification can not only compromise accuracy but also negatively impact performance. In contrast, code-based verification enhances accuracy across all 7 subtopics when compared to the Basic Prompt.
Analysis of Verification-guided Weighted Majority Voting: we initially compiled the confusion matrix (TP/TN/FP/FN), capturing solutions with self-verification that matches the True and False states mentioned in Eq. 1 from five distinct sampled paths. The details of the confusion matrix are presented in Appendix A.1.1. From this data, we computed Precision, Recall, and Accuracy. (Solutions in the True state are seen as positive.) The results are presented in Fig. 6. In comparison to Accuracy, we observed numerical enhancements of 22.3% and 5.6% in the average Precision and Recall, respectively. In particular, the average Precision registered at 95.88%. This implies that the Accuracy has the potential to become much higher, if more solutions reach the verified True state before giving the final answer.
Hyper-parameters ablation in Verification-guided Weighted Majority Voting: we also performed ablation studies on the hyper-parameter in Eq. 1. When the hyper-parameter setting satisfied , the performance of the verification-guided weighted majority voting consistently surpassed that of the naive majority voting methods across all sampled paths. In contrast, when we set the hyper-parameter , the performance under this configuration was worse than the naive majority voting. Therefore, our proposed method, verification-guided weighted majority voting, is easy to tune and robust.
Conclusion and Limitation
In this paper, we begin with pilot experiments on GPT4-Code to explore how its use of code impacts its performance in mathematical reasoning. By analyzing Code Usage Frequency and accuracy, we determine that GPT4-Code’s skill in solving math problems can be largely attributed to its ability to generate and execute code, as well as its effectiveness in adjusting and rectifying solutions when confronted with implausible execution outputs. Expanding on this understanding, we introduce the ideas of explicit code-based self-verification and verification-guided weighted majority voting, with the goal of enhancing GPT4-Code’s mathematical capabilities.
However, there are limitations in our work that we plan to explore further in the future. Firstly, our analysis and improvements are currently focused on GPT4-Code, which is somewhat restrictive. We aim to apply the methods to other LLMs. Secondly, our explicit code-based self-verification and verification-guided weighted majority voting technique could potentially create more accurate datasets. These datasets would include detailed step-by-step code-based solution generation and code-based validation, which could help improve open-source LLMs like LLaMA 2 (Touvron et al., 2023) and enhance their mathematical abilities. Although we haven’t yet investigated this approach, we leave it for future work.
References
Appendix
This Appendix contains two sections. The first section provides the experiment details, including the detailed experiment results on MATH and MMLU datasets. The second section presents some examples of GPT4-Code.
Appendix A Experiment Details
A confusion matrix is a specific table layout that allows visualization of the performance of an algorithm. It’s particularly useful for classification problems, and we utilize it to analyze the performance of our verification process.
The matrix itself is a two-dimensional grid, 2x2, for the binary classification of verification results. Each row of the matrix represents the instances in a predicted class, which is determined by the verification results given by the language model, while each column represents the instances in an actual class, which is determined by the actual correctness of the answer given by the model. Tab. 5 shows how the matrix looks for our verification process:
True Positive (TP): The cases in which the model’s verification result is ‘True’, and the answer is actually correct.
True Negative (TN): The cases in which the model’s verification result is ‘False’, and the answer is actually wrong.
False Positive (FP): The cases in which the model’s verification result is ‘True’, but the answer is actually wrong.
False Negative (FN): The cases in which the model’s verification result is ‘False’, but the answer is actually correct.
This matrix is very helpful for measuring more than just straightforward accuracy, based on which Precision and Recall are two important metrics. They are defined in Eq. 3 and their meanings are as follows:
Precision is the fraction of relevant instances among the retrieved instances. It is a measure of the accuracy of the classifier when it predicts the positive class.
Recall is the fraction of the total amount of relevant instances that were actually retrieved. It is a measure of the ability of a classifier to find all the positive instances.
In other words, precision answers the question “What proportion of verified TRUE answers was actually correct?”, while recall answers “What proportion of actual correct answers was verified TRUE?” Given its meaning, verification-guided voting is bounded to be effective when the precision of verification is high.
A.1.2 Python package usage analysis
Tab. 6 outlines the usage of various Python packages in our experiments. Among them, we found that the sympy package is utilized most frequently, highlighting its central role in the computational tasks performed.
A.2 Detailed experiment result on MMLU dataset
Fig. 8 illustrates that GPT4-Code performs relatively poorly in certain domains, such as engineering and the humanities, with a particularly marked deficiency in virology, where it achieves a score of less than 60%. These observations delineate specific areas that call for further investigation and refinement, thus outlining the direction for future improvements in the model.
Appendix B Examples
In this section, we provide more examples.