A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language Models
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, Mrinmaya Sachan
Introduction
Many natural language understanding situations, such as understanding the financial news, require reasoning with text that includes numbers. However, such mathematical reasoning is challenging for NLP models Cobbe et al. (2021); Mishra et al. (2022b). Mathematical reasoning for text has been an active area of research for a while (Seo et al., 2015; Sachan and Xing, 2017; Sachan et al., 2017, 2018, inter alia), and has also emerged as a key task to track the capabilities of large language models (LLMs) in recent years (Brown et al., 2020; Ouyang et al., 2022; Wei et al., 2022a, inter alia).
However, despite the impressive performance of LLMs on various math reasoning benchmarks (e.g., Ouyang et al., 2022; Chowdhery et al., 2022), it remains unclear whether these models have learned mere artifacts in the data or have truly mastered the mathematical concepts needed to consistently solve all variations of the same problem (Patel et al., 2021; Razeghi et al., 2022; Welleck et al., 2022). In sharp contrast with a large number of papers on improving the performance of LLMs on various types of math-based problems, there has been little effort on behavioral analysis of LLMs for these tasks. Existing methods for understanding the robustness of these models (Patel et al., 2021) rely on manually constructing variations of math problems, and we do not yet have a principled, comprehensive framework for quantifying such robustness.
We apply our framework to study a set of thirteen GPT models with various sizes and training procedures (i.e., instruction-tuned and non-instruction-tuned). We observe that, among non-instruction-tuned language models, the larger ones tend to be more sensitive to changes in the ground-truth result of a math word problem, but not necessarily more robust. However, we observe a different behavior in the instruction-tuned GPT-3 models Ouyang et al. (2022), which show a remarkable improvement in both sensitivity and robustness, although the robustness reduces when problems get more complicated. We additionally investigate the role of size and instruction tuning on the model’s performance with three models of the LLaMA family Touvron et al. (2023) and Stanford Alpaca Taori et al. (2023).
Problem Setup
Template : Mark has trees in his backyard. If he plants more, how many trees will he have? Operands : Operations : {(“”, 1, 2)} Result:
A Causal Framework
We address the research question “Is a model reasoning robustly on MWPs?” by comparing the causal mechanisms of the model’s decisions to a hypothesized human reasoning mechanism. Note that we do not claim to know how humans reason about these problems. We simply propose a reasonable and intuitive way to judge model robustness given a reasonable and intuitive human reasoning mechanism inspired by findings regarding the independence of language and mathematical reasoning in humans Brannon (2005); Monti et al. (2012).
The causal mechanisms of how humans might solve include
1.0.0.2 Model Reasoning Mechanisms.
In contrast, the causal mechanisms of how a model might solve are as follows:
where we are unsure about (1) what part(s) of the model takes into account, and (2) how it operates over the relevant variables.
The model might attend over the question template in two ways: paying attention to the text surface form via the causal path , or text relevant to the math operations via the causal path .
The model might also attend to the operands via a causal path .
If the model learns the correct causal mechanisms as in the human cognitive process, it should capture how the operator and the operands matter to the ground-truth result (via and ) and then the model prediction should be sensitive to any changes in the ground truth, namely . No spurious correlations can directly affect without going through the mediator .
Hence, to answer the question “How robust is the mathematical reasoning of a model on MWPs?” we can answer the following subquestions:
How does change in response to ? By quantifying this, we assess the sensitivity (correct responsiveness) of the model to changes in the problem. In other words, does the model correctly adjust its prediction in response to a change in the correct solution of the problem?
What is the (unwanted) direct causal effect size of , and ? We see the quantities as a measure of the brittleness (i.e., wrong responsiveness) of the model to result-preserving changes in the input. The lower the direct causal effect of and , the more robust the model is.
2 Step 2. Causal Intervention List
After formulating the cognitively-inspired subgraph and defining the undesired causal paths in Figure 2, we list all feasible limited actions that allow us to perform our causal analysis. In the context of MWPs, we use the following interventions:
Direct intervention on all possible ;
Partially controllable interventions on . We can replace the template in two ways:
is affected but is not affected.
3 Step 3. Turning Limited Actions into Causal Effect Sizes
Following the distributional definition of causal effect by Pearl (1995), we quantify the effect of factor in our causal graph using a distance metric between the distributions and . That is,
3.0.0.2 Causal Effects of the Operands.
When intervening on the operands , we can obtain the size of the total causal effect of on , namely
Note that this TCE is not the exact desired quantity, because we want to separate two different paths of how affects : (1) the path , which is the correct decision path that we want the model to pick up (where the model reacts to the change in the ground-truth answer), and (2) the path , which is the spurious correlation that the model might have learned (where the model relies on some spurious correlations with certain numerical values, which could be traced to perhaps their frequencies in the training corpus).
We can quantify the direct causal effect (DCE, i.e., the effect from the directed causal path from a variable to another that does not go through any intermediate variables) Pearl (2001) of on , namely the strength of the direct causal path , by controlling for to be fixed every time we intervene on :
3.0.0.3 Causal Effects of the Text Surface Form.
As for the operands, we can compute both the direct and indirect effects of the surface form representing the math problem. In particular, intervening on without controlling for (intervention 2a in Sec. 3.2), we can compute the total effect, i.e.,
Controlling for the operations (intervention 2b in Sec. 3.2) will instead allow us to obtain the direct causal effect of the surface text:
3.0.0.4 Causal Effects of the Operators.
4 Step 4. Quantifying the Causal Influence
Change in the Prediction. To account for the inability of LMs to capture the continuous property of numbers Jin et al. (2021a), we measure the change in the model’s prediction using an indicator of the “change result” event:
where , and .
Relative Change in Confidence. Inspired by Finlayson et al. (2021), we also highlight the change in terms of the relative difference in the probability assigned to and . We formulate two types of relative change, one quantifying the relative change in the confidence of , and the other quantifying the relative change in the confidence of :
We quantify the overall relative change in confidence (RCC) as the average of the two relative changes above:
Experimental Setup
In this section, we describe the data used to perform the interventions and to measure the causal effects.
For our analyses, we use instances of math word problems from three popular datasets: ASDiv-A Miao et al. (2020), MAWPS Koncel-Kedziorski et al. (2016), and SVAMP Patel et al. (2021). The examples contained in these collections are pairs consisting of a question template with its annotated operations . Each of these pairs can be instantiated multiple times into problems by filling the template with numerical values and computing the ground-truth result (most problems involve two to three operands, i.e., ). We select a set of 437 two-operand and 307 three-operand template-expression pairs that we use to generate pairs of prompts representing an intervention. More details about the prompt generation procedure are in Appendix A. We use to refer to an instantiated template that we use as a prompt.
2 Intervention Data
3 Models to Evaluate
We use our framework to assess the robustness of reasoning in thirteen pre-trained language models. We consider five sizes of the GPT-2 model Radford et al. (2019): distilled Sanh et al. (2019), small, medium, large, and XL. We evaluate four models from EleutherAI that were pre-trained on the Pile Gao et al. (2020): GPT-Neo 1.3B and 2.7B Black et al. (2021), GPT-J-6B Wang and Komatsuzaki (2021), and GPT-NeoX-20B Black et al. (2022). We use HuggingFace Transformers Wolf et al. (2019) to access the models. Additionally, we experiment with a set of instruction-tuned versions of GPT-3 Brown et al. (2020): Instruct Ouyang et al. (2022), Curie, Davinci-002, and Davinci-003.The OpenAI ids for these models are, respectively, davinci-instruct-beta, text-curie-001, text-davinci-002, and text-davinci-003. Experiments with GPT-3 are carried out under the constraints set by the OpenAI APIshttps://openai.com/api/, which prevent us from computing the causal effect using the same procedure as for the other models. We report the details about how the metrics were computed for GPT-3 in Appendix C. In the reported results, we indicate with an asterisk (∗) the metrics that were influenced by this limitation.
Results
In Figure 4, we present a different visualization of the direct causal effect of on the model’s prediction. We report the heatmaps showing the probability assigned by the model to the result of a problem . For Distil-GPT-2 we observe low overall probability assigned to and diagonal patterns indicating consistency in assigning higher probability to specific results (e.g., 10, 20, 30, 40, 50). For the two larger models we notice a higher probability mass assigned to the problem’s result, but less consistency on the prediction of the same result with different sets of operands (this is true for GPT-J in particular). This result is consistent with the observed higher DCE and TCE in larger models: might vary more considerably when intervening on without affecting , but overall the model assigns higher probability weight to the correct result, which correlates with higher sensitivity.
2 Effect of 𝑻𝑻\bm{T} on R𝑅R
3 Overall Insights
Possible explanations for the improved robustness and sensitivity demonstrated by the large GPT-3 models might be the dramatic size increase and extension/enhancement of the training procedure involving instructions. The former idea is aligned with the emergent abilities hypothesis (Wei et al., 2022a), which postulates the existence of skills that are displayed by large-scale models but are not present in smaller-scale models. However, our observations show different performances in versions of GPT-3 Davinci that differ in the training procedure.A high-level description of the training procedures for the models is provided at https://beta.openai.com/docs/model-index-for-researchers. This raises the question of whether the capability of LLMs to reason about math problems benefits from instruction-based tuning. We address this question in the following section.
4 Extending to LLaMA-Based Models
From the results (Figure 6), two notable observations emerge. Firstly, the increased difference between TCE and DCE observed with the increasing size of the LLaMA models suggests that a larger number of parameters can be a significant driver behind robustness/sensitivity improvement. However, this is not necessarily the case across different models: GPT-NeoX-20B shows a smaller TCEcp-DCEcp gap compared to LLaMA 7B (5.2% vs 9.0%). Secondly, the instruction tuning procedure of Alpaca does not seem to help significantly with mathematical computation: the decrease in both TCE and DCE shows that robustness improves at the expense of sensitivity. Nonetheless, overall, when comparing Alpaca compared to its base model, LLaMA 7B, we observe an increase in the gap between TCE and DCE, although this difference is minimal (9.5% vs 9.0%).
The limited improvement of Alpaca might be attributed to its instruction tuning procedure consisting of “a list of user-oriented instructions including email writing, social media, and productivity tools” Taori et al. (2023), which differs from reasoning-intensive tasks. We suggest future work to examine different types of instruction tuning (e.g., focused on reasoning procedures or reinforcement learning from human feedback), which might help the model answer more complex types of questions in a step-by-step manner and more accurately. We hypothesize that the different performances in versions of GPT-3 Davinci might be produced by the specific type of instructions used for training, by the reinforcement learning component Ouyang et al. (2022), or simply by an extension of the language modeling pre-training. It is challenging to pinpoint the exact factor in the training procedure that contributes to this improvement, as specific methodological details are not available.
5 Moving to Three-Operand Problems
Related Work
Causal NLP. Causal inference aims to study the cause and effect from observational and interventional data Pearl (2009); Peters et al. (2017). Traditionally, researchers usually apply causal techniques to phenomena in nature and human society. With the rise of powerful models in NLP, recent research has started to explore the intersection of causal inference and NLP, forming the study of Causal NLP Jin et al. (2022); Feder et al. (2021a).
There are several formulations for Causal NLP: the causality for NLP thread involves using the causal framework for data collection and task formulation Jin et al. (2021c), inspecting the (path-specific) causal effect of certain neurons on predictions Vig et al. (2020); Meng et al. (2022), understanding the causal effect of data and learning paradigm for model performance Ni et al. (2022), and as a way to frame prompts (Lyu et al., 2023); and NLP for causality involves testing the pure causal inference skills of LLMs Jin et al. (2023a, b), and use text as a variable for causal effect estimation Roberts et al. (2020); Veitch et al. (2020); Jin et al. (2021b, 2023c).
The most similar line of research to our work is the application of causal effect estimation on interpreting models’ behavior, such as how models understand syntactic agreement Finlayson et al. (2021), and how interventions in the representations and weights affect the model prediction Feder et al. (2021b). To the best of our knowledge, our work is the first to formulate a causal framework for robustness behavioral tests, and also we are the first to introduce the idea to quantify the differences in the causal mechanisms of human reasoning and model decisions.
Math Reasoning in NLP. A growing body of work tries to improve the math reasoning capability in NLP models Zhang et al. (2020); Geva et al. (2020); Spokoyny et al. (2021), and prompting techniques for LLMs (Cobbe et al., 2021; Shen et al., 2021; Kojima et al., 2022; Wei et al., 2022b; Chowdhery et al., 2022). For analysis, significant attention has been given to models’ ability to understand numerical quantities Wallace et al. (2019); Thawani et al. (2021) and numerical operations Pal and Baral (2021); Berg-Kirkpatrick and Spokoyny (2020); Piękos et al. (2021); Razeghi et al. (2022).
Conclusion
We developed a framework to disentangle and separately measure the effect of different factors influencing the predictions of LLMs for math reasoning. Our results indicate that a drastic increase in both robustness and sensitivity emerges in the GPT-3 Davinci models. Additionally, we study the contribution of size and instruction tuning in the models of the LLaMA family, observing that the Alpaca instruction tuning, while increasing the model’s robustness, does not significantly improve the overall performance. Our framework provides a formalized theory of behavioral testing for math reasoning models and opens new future directions to design behavioral tests of models in a principled way.
Ethical Considerations
As for the ethical practice in this work, the data involved are from existing MWP datasets with no private user information, and available under the MIT license. As for the ethical impact of the use of this work, the study is about providing a metric and analyzing existing models’ robustness, so there is less concern over harmful usage. Rather, it is more about putting checks on existing AI models and helping humans understand them better before use. Potential stakeholders that could benefit from this research include NLP researchers working on math models, practitioners working on various applications involving mathematical reasoning with text, and e-learning design.
Limitations
A key limitation in our work is that LLMs might have seen these math problems. Our work theoretically assumes this is not the case. Another limitation is that for the sake of simplicity, our work makes some assumptions. For example, we assume all numbers in the range of integers 0 to . This would not cover every MWP out there. And future work is needed to generalize our framework to other forms of MWPs. In this work, we are also constrained by the limitations of the OpenAI policy on the GPT-3 API. This limits the number of perturbations we consider in this work as well as the accuracy with which we can estimate our causal distributions. Finally, our work is restricted to English, and extending it to other languages will require us to create an MWP dataset in that language.
Acknowledgments
This material is based in part upon works supported by the German Federal Ministry of Education and Research (BMBF): Tübingen AI Center, FKZ: 01IS18039B; by the Machine Learning Cluster of Excellence, EXC number 2064/1 – Project number 390727645; by the John Templeton Foundation (grant #61156); by a Responsible AI grant by the Haslerstiftung; and an ETH Grant (ETH-19 21-1). Alessandro Stolfo is supported by armasuisse Science and Technology through a CYD Doctoral Fellowship. Zhijing Jin is supported by PhD fellowships from the Future of Life Institute and Open Philanthropy, as well as the travel support from ELISE (GA no 951847) for the ELLIS program. We also thank OpenAI Researcher Access Program for granting our team credits to their API.
References
Appendix A Creation of the Prompts
We consider MWP examples from the union of the three datasets SVAMP, ASDiv-A, and MAWPS. The textual template of a problem consists of a context (describing a real-world state and/or actions) and a question. In order to obtain suitable prompts for the models, we convert the problems’ questions into statements where the result of the problem is expected to be the first token after the prompt. E.g., in the example in section 2, how many trees will he have? is converted into the number of trees that he will have is _. From the MWP templates of the SVAMP/ASDiv-A/MAWPS collection (we consider all splits), we filter out the templates whose questions do not start with How many…, and we use spaCyhttps://spacy.io to identify the subject, the object and the verbs in the sentence. This allows us to convert the last sentence of the template from The number of… is. This way, we obtain 437 statement-based MWP templates for two-operand problems and 307 for three-operand problems. We manually checked a subset of the templates to identify possible mistakes in the conversion procedure.
Appendix B Frequently Asked Questions
In Table 1 we report examples of MWP pairs representing different types of intervention.
B.2 What is the accuracy of the evaluated models on the generated problems?
We report the accuracy of the models considered for evaluation in terms of accuracy at 1 and accuracy at 10. Results are displayed in Figure 8.
B.3 What is the relation between accuracy and the RCC metric?
Moreover, we conduct an additional sanity check as in Patel et al. (2021): removing the question from the MWP templates, we observe a sensitivity-robustness degradation to random guessing (i.e., TCE DCE). This indicates that the measurement of the causal effects within our framework is not affected by patterns in the templates that might have been picked up or memorized by large models.
Appendix C Computation of Causal Effects for GPT-3
The limit for is set by OpenAI to 5. However, for our main set of experiments (i.e., computing the causal effects of , , and ) we were granted an increased limit of to 100. This allowed us to obtain reasonable estimates for the causal effects, as the number of cases in which is not defined is less than of the number of examples that we consider.
In cases when is defined (i.e. when appears in the top token predictions) and is not defined, we compute a lower bound on the relative change using the upper bound on given by the probability of the -th most likely token. This gives us a conservative estimate of . For cases in which is not defined, we cannot say anything about the relative change, and we set . The same applies when swapping and . This procedure is illustrated by Algorithm 1.
C.3 Heatmap Illustration
The heatmap for GPT-3 displayed in Figure 4 was computed by taking the raw probability score produced by the model over the whole vocabulary, as the limit on the available top predicted tokens makes it impossible to normalize it over the set , as done for the other models. The probability was set to 0 when did not appear in the model’s top 5 predictions for the next token after the prompt.
Appendix D Computing Infrastructure & Inference Details
To run our experiments, we used a single NVIDIA TITANRTX with 24GB of memory for all the versions of GPT-2 and GPT-Neo. We used a single NVIDIA A100 with 40GB of memory for GPT-J-6B and a single NVIDIA A100 with 80GB of memory for GPT-NeoX and the LLaMA models (two for the 30B version). We accessed GPT-3 using the OpenAI APIs. The longest run (GPT-J) on the four kinds of experiments corresponding to the four kinds of effects measured took 12 hours, using 500 MWP instances for each of the 437 templates. Due to budget and resource constraints, the experiments on GPT-3, GPT-NeoX, and LLaMA were carried out using 20 examples generated for each template and took 7 hours. Experiment tracking was carried out using Weights & Biaseshttp://wandb.ai/.