MathAttack: Attacking Large Language Models Towards Math Solving Ability
Zihao Zhou, Qiufeng Wang, Mingyu Jin, Jie Yao, Jianan Ye, Wei Liu, Wei Wang, Xiaowei Huang, Kaizhu Huang
Introduction
Solving Math Word Problem (MWP) aims to infer a final answer from the natural language description of a math problem (Wang, Liu, and Shi 2017). With the boom of Large Language Models (LLMs), the research of solving MWP has recently made great progress (Qiao et al. 2022; Uesato et al. 2022; Chang et al. 2023). Most of them work on prompt engineering to improve math solving ability of LLMs (Wei et al. 2022; Zhou et al. 2022; Kojima et al. 2022; Chen et al. 2022; Fu et al. 2022), and LLMs (e.g., ChatGPT) can provide correct reasoning process and the final answer for simple math word problems. Subsequently, they have been progressively applied in the field of intelligence education (Macina et al. 2023). Therefore, it becomes essential to examine the security of LLMs in math solving ability, but this has not attracted much attention so far. To the best of our knowledge, there are only a few works (Zhu et al. 2023; Wang et al. 2023a) to evaluate the robustness of LLMs through attacking prompts (Figure 1(b)). By comparing to prompt-attack, we argue that attacking MWP samples themselves is more direct to reflect the security of LLMs in math solving ability, like Figure 1(c).
On the other hand, general text adversarial attack has made great progress (Li et al. 2018, 2020; Ye et al. 2022). This task aims to generate an adversarial text that is semantically similar to the original text , while victim model can correctly classify but incorrectly classify (Jin et al. 2020; Xu et al. 2020). However, it tends to change mathematical logic by directly applying such techniques of general text adversarial attack. For example, if the word in Figure 1(c) is modified to another number, the mathematical logic will be changed and the original ground-truth will be no longer the correct answer. Therefore, it is essential to preserve the mathematical logic of MWPs, which makes MWP adversarial attack more challenging.
To preserve the mathematical logic of MWPs, we propose MathAttack for attacking the math solving ability of large language models. Figure 2 shows an overview of our MathAttack. We first recognize logical entities, altering these logical entities easily leads to changing the mathematical logic of math word problems. Then we freeze the logical entities, preventing the attacker from modifying logical entities. Finally we attack the LLMs utilizing word-level attacker (Li et al. 2020) while not changing those frozen logical words. With the help of MathAttack and manual check, we propose a new dataset RobustMath, which consists of 300 high-quality MWP adversarial samples and could measure the robustness of LLMs’ math solving ability.
Extensive experiments on our proposed RobustMath dataset and another two math benchmark datasets GSM8K (Cobbe et al. 2021) and MultiAirth (Roy and Roth 2015) show that our MathAttack could effectively attack the math solving ability of LLMs. As far as we know, most works (Zhu et al. 2023; Wang et al. 2023a) focus the robustness of LLMs in general tasks, there are not any comprehensive study on the security of LLMs in math solving ability. To this end, we conduct a a serious of analysis in the experiments and observe the following three points: (1) Transferability of attacking samples. Adversarial samples generated from higher-accuracy LLMs are also effective for attacking LLMs with lower accuracy (e.g., transfer from larger to smaller-size LLMs, or from few-shot to zero-shot prompts); (2) Complex MWPs (such as more solving steps, longer text, more numbers) are more vulnerable to attack; (3) We can improve the robustness of LLMs by using our attacking samples in few-shot prompts.
In summary, our contributions are as follows:
In this paper, we make a first attempt to attack MWP samples to examine the security of LLMs in math solving ability.
We propose a MathAttack for attacking the math solving ability of large language models, including Logical Entity Recognition, Freezing Logical Entity and text Attack.
We propose a new dataset RobustMath by adopting MathAttack and manual check. It consists of 300 high-quality MWP adversarial samples and could measure the robustness of LLMs’ math solving ability.
Extensive experiments show that MathAttack could effectively attack the math solving ability of LLMs. Through the exhaustive analysis, we obtain three findings for the security of LLMs in math solving ability.
Related Work
Recent proposals intend to solve the problem by using sequence or tree generation models. (Wang, Liu, and Shi 2017) presents a sequence-to-sequence (seq2seq) approach to generate the mathematical equation. (Xie and Sun 2019) propose a goal-driven tree-structured (GTS) model to generate the equation tree. This sequence-to-tree approach significantly improves the performance over the traditional seq2seq approaches. (Zhang et al. 2020) adopt a graph-to-tree approach to model the quality relations using graph convolutional networks (GCN). Previous studies (Patel, Bhattamishra, and Goyal 2021) indicate these MWP solvers rely on shallow heuristics to generate equations. With the boom of Large Language Models (LLMs) and the proposal of chain-of-thought (Wei et al. 2022), the math solving ability of the model has recently made great progress. Many research works on prompt engineering to improve math solving ability (Zhou et al. 2022; Kojima et al. 2022; Chen et al. 2022; Fu et al. 2022), they are capable of effortlessly solving simple MWPs, and LLMs are gradually being incorporated in the field of intelligent education (Ji, Han, and Ko 2023; Macina et al. 2023). In this context, examining the security of LLMs in math solving ability becomes essential. In this work, we make a first attempt to examine this security issue by attacking MWP samples.
Large Language Models Attack
Previous proposals have already tried to evaluate the robustness of large language models (Zhuo et al. 2023; Shi et al. 2023). (Wang et al. 2023b) makes the first attempt to systematically evaluate the robustness of LLMs by using robust datasets. Recently, some works propose to address this issue by attacking prompts (Wang et al. 2023a; Zhu et al. 2023). (Wang et al. 2023a) introduces the ICL attack based on TextAttack, which aims to manipulate the prompt only without altering the input. (Zhu et al. 2023) presents PromptBench, a robustness benchmark specifically designed to evaluate the robustness of LLMs against adversarial prompts. Our work differs from theirs in two main aspects: (1) We specifically focus on attacking the MWP sample itself, which provides a more direct approach and fills the gap of non-prompt attacks on LLMs. (2) Their works target general tasks, lacking a comprehensive analysis of the robustness in math solving ability.
MWP Attack
For the MWP solvers, previous works generate some MWP adversarial examples by rule-based methods like reordering the problem description (Kumar, Maheshwary, and Pudi 2021; Patel, Bhattamishra, and Goyal 2021). However, with the development of LLMs, the semantic and logical capabilities of the model have been enhanced, rendering these adversarial examples ineffective. Adversarial MWP sample datasets SVAMP (Patel, Bhattamishra, and Goyal 2021) can be solved well by LLMs like ChatGPT. In this paper, we attack MWP samples of LLMs for the first time and propose a new dataset RobustMath to evaluate the robustness of math solving ability of LLMs. It consists of adversarial examples generated by MathAttack, utilizing simple MWPs from GSM8K and MultiAirth as seed data.
Methodology
Suppose we have a text with words whose ground truth label is . We call an adversarial example when can make the victim model wrong prediction but original correct prediction (), i.e.,
Compared to traditional text attack, math word problem attack need to preserve the mathematical logical of text sample , it is defined as:
The goal of the attack task is to generate an adversarial example among all . Since text data consists of discrete words whose change can be perceived by humans, we always want the optimized adversarial example to be semantically closest to the original text sample . Thus, the objective function of this task can be defined as follows:
where denotes the semantic similarity between and . In this paper, is the large language model and we follow the black-box setting.
The Proposed MathAttack
Figure 2 shows an overview of MathAttack. We firstly recognize logical entities. Altering these logical entities easily leads to changing the logic of math word problems. Then we freeze the logical entities, preventing the attacker from modifying logical entities. Finally we attack the LLMs utilizing word-level attacker while not changing those frozen logical entities.
Logical Entity Recognition
Logical entities are crucial components that constitute logic in math word problems (Kumar, Maheshwary, and Pudi 2022; Li et al. 2022). In order to preserve the logic of a math word problem, it is indispensable to define and identify which entities as logical entities. In this paper, we define the following three types of entities as logical entities. (1) Role Entity: It includes person entity (e.g, Asia in Figure 2). (2) Number Entity: It includes quantity (e.g, $140 in Figure 2), cardinal number and ordinal number. (3) Scenario Entity: It includes time entity and location entity. Altering these environmental factors is easy to change the logic of math word problems too.
Then we employ Named Entity Recognition (NER) model to identify them:
where is a word index set if the word belongs to logical entity type . The symbols , and represent the Role, Number and Scenario Entity respectively. We utilize Spacy https://spacy.io/ as our NER model.
Freezing Logical Entity
It tends to break the original logic of MWP by altering logical entities during the attack process. To this end, we freeze all logical entities in order to prohibit attackers from modifying them:
where denotes the frozen word index set.
Attack
Text attackers are generally classified into three types: char-level, word-level and sentence-level. In MathAttack, We choose word-level attacker because the char-level attacker can distort the semantic meaning of words (like Figure 1(b)) and sentence-level attacker are prone to disrupting the mathematical logic of MWP. The attack process of word-level attacker primarily entails two steps: finding vulnerable words and words replacement.
In order to find vulnerable words, it is necessary to determine which words are significant. Specifically, we first sequentially mask all modifiable words to form new sentences. Afterward, we predict each new sentence to get the drop in the probability of the correct answer. The more it drops, the more important the word is. It is defined as:
where is the important score of and is the function to get the probability of the correct answer. After that, we can get the important scores list . Notice that the length of is not because some words are frozen. Finally, we choose the word which has the max important score as the vulnerable word:
where is the index of the vulnerable word and is the function to pop the word which has the most important score and get its index.
After finding the vulnerable word, we proceed to locate all synonyms of in order to substitute it:
where is the synonyms set of , we sequentially select a word in based on the similarity to then substitute :
where is the sentence by replacing in with . is the function to pop the word in that is most similar to . Notice that if is already empty before popping, we go back to Eqn. (10) and repeat the above process. After obtaining , we perform different actions based on the following situations, if , the final adversarial sample is :
If and , we will keep this word change:
Then go back to Eqn. (12) and repeat the above process. If and , we will abandon this word change then go back to Eqn. (12) and repeat the above process.
In our attacker, we utilize BertAttack (Li et al. 2020) as our backbone, which utilizes [mask] token to mask words and bert embedding to calculate the similarity of words.
Experiments
We choose four mainstream large language models as our victim models.
Flan-T5-large (Chung et al. 2022): Flan-T5-large is a derivative of the Text-to-Text Transfer Transformer (T5) model, developed by Google. It has 760M parameters.
Flan-T5-xl (Chung et al. 2022): Flan-T5-xl is a large version of Flan-T5 than Flan-T5-large, developed by Google. It has 3B parameters.
ChatGLM2 (Du et al. 2022): ChatGLM2 is the second-generation version of the open-source bilingual (Chinese-English) chat model ChatGLM, developed by Tsinghua University. It has 6B parameters.
ChatGPT (OpenAI 2023): Developed by OpenAI, ChatGPT is a large language model trained to generate human-like text. It uses the GPT-3 architecture and has been fine-tuned for more interactive and conversational tasks. In detail, we use the gpt-3.5-turbo API.
We set the temperature = 0 to stabilize the output of LLMs. When attacking victim models, we not only attack them with zero-shot prompt but also few-shot prompt. Specifically, we employ four MWP samples as shots and provide Chain-of-Thought (CoT) (Wei et al. 2022) annotations. This few-shot prompt serves as a method to enhance the math solving ability of LLMs. Similar with other prompts, they are not changed during the attack process.
Datasets
Two math word problems benchmark datasets GSM8K (Cobbe et al. 2021) and MultiArith (Roy and Roth 2015) are adopted in the experiments. However, we only select the subsets for the following considerations by following the previous work (Zhu et al. 2023): (1) we focus on simple MWPs, because hard samples have very lower accuracy not necessary to attack. (2) Owing to the extensive computational requirements of generating single adversarial sample, which necessitates iterating over the entire dataset 100 times on average. Finally, for GSM8K, we firstly remove those hard samples labelled by more than three solving steps, then randomly select half of those remained simple MWPs and obtain 307 MWP samples. For MultiAirth, all MWPs are simple thus we randomly select 150 MWPs similar with the previous work (Zhu et al. 2023).
Metrics
Given a dataset with data instance and label , victim model , an adversarial attack method that generates adversarial examples , we adopt following four metrics:
Similarity: The average semantic similarity between the adversarial sample and the original sample. We use Universal Sentence Encoder (Cer et al. 2018) to measure semantic similarity.
To ensure the correctness in the experiments, we check each adversarial sample manually, and consider adversarial examples which are changed mathematical logic as unsuccessful attacks.
Main results
As shown in Table 1, our approach can effectively attack the math solving ability of large language models. For LLMs with zero-shot, we could get the high ASR, even for ChatGPT, it could achieve an average of 40% on GSM8K (41.15%) and MultiAirth (39.19%). The average Similarity is large than 90%, indicating that we can successfully generate adversarial samples with high similarity and do not alter mathematical logic.
Comparing different LLMs, we can observe that more powerful LLMs (i.e., higher Clean Acc) are more difficult to attack (i.e., lower ASR). For Flan-T5-large and Flan-T5-xl, their robustness in math solving ability is poor, as even a slight disturbance can cause them to predict incorrectly. For ChatGLM2 and ChatGPT, their robustness is noticeably stronger, as our method fails to attack them on some MWP samples.
Furthermore, comparing zero-shot and few-shot, we can see that employing few-shot could enhance the math solving ability of the LLMs and also make them more robust, leading to a lower ASR. For models with stronger in-context ability, the enhancement becomes larger. Like ChatGPT, the Attack Success Rate could decrease from 41.15% to 19.93%. However, we find ChatGLM2 exhibits poor in-context ability which leads math solving ability as well as robustness does not improve with the few-shot prompt.
Fine-grained Analysis
To test the transferability of the generated adversarial samples, we take adversarial samples of model to attack other models . Specifically, we select the samples that can correctly predict as the experimental samples. Subsequently, we provide with adversarial samples generated by attacking on experimental samples. We examine if these adversarial samples can successfully attack model . Here, we propose the metric: Transfer Success Rate (TSR), if an adversarial sample of can successfully attack model then it means transfer success.
In Figure 3, we show the TSR between Y-axis (i.e., model) and X-axis (i.e., model) models, and we can observe that the adversarial samples of larger-size models can attack smaller-size models too. ChatGPT could get 94.44% TSR to Flan-T5-large and 89.47% TSR to Flan-T5-xl. Specifically, we find that the TSR will increase when the math solving ability between models grows wider. As shown in Figure 3, we can see that the adversarial samples of smaller-size models can not attack larger-size models. Flan-T5-large and Flan-T5-xl both show low TSR (6.25% and 6.67%) on ChatGPT.
In order to see the transferability performance between zero-shot and few-shot, we conducted the same experiment on ChatGPT. As shown in Figure 4, the ChatGPT with few-shot can achieve 45.24% TSR to ChatGPT with zero-shot however the reverse is only 20.56%. It indicates that the adversarial samples of LLM with few-shot can attack that with zero-shot. And adversarial samples of LLM with zero-shot can not transfer to that with few-shot.
Analysis on MWPs
To know which MWPs are easier to attack, we investigate the effects of MWPs reasoning steps, problem length and number count on ASR. As shown in Figure 5: (a) with the increase of reasoning steps of ground truth, we can observe that the ASR will increase when the reasoning steps from 2 to 3. Reasoning steps of ground truth can be regarded as a metric to measure the difficulty of MWP. Difficult MWPs are easier to attack. (b) with the increase in problem length, we can observe a gradual increase in the ASR as the length of the math word problems become longer. (c) with the increase in the quantity of numbers in MWPs, we can observe a gradual increase in ASR as the number counts become more. All the above three factors can be used to measure the complexity of an MWP (Fu et al. 2022), therefore, we can draw a conclusion that more complex MWPs are easier to attack.
Using Attacking Samples as Prompts
In the few-shot prompts, we replace normal MWP examples by corresponding adversarial examples generated by our MathAttack but with correct labels, and observe their impact on the math solving ability and robustness of the LLMs. As shown in Table 2, we can see the Clean ACC still maintains a high level of accuracy (87.95% on GSM8K and 98.00% on MultiAirth), because the adversarial examples generated by our MathAttack exhibit high similarity to the original samples. By comparing the Attack Acc and ASR in Table 1, it is surprised to find the use of adversarial examples in the few-shot prompts can enhance the robustness of LLMs (i.e., much lower ASR by comparing to the normal results in Table 1). When we use adversarial examples in the prompt, the LLM could see these examples that are disturbed but still able to predict correctly, therefore they will not affected by some small disturbances when predict. Figure 6 provides a more intuitive visualization, demonstrating that the robustness of LLMs utilizing few-shot prompt can be significantly improved by comparing to zero-shot prompt. When employing adversarial examples as few-shot prompt, it will further strengthen their robustness and the ASR of large language models further decrease. This observation motivates us to enhance the robustness of large language models without compromising their math solving ability by employing the adversarial examples generated by MathAttack as few-shot prompt.
Case Study
Table 3 reports a real case predicted by ChatGPT on the original MWP and its adversarial sample generated by MathAttack. We can find that the adversarial sample generated by MathAttack is similar to the original sample with few changes. For the original sample, ChatGPT can give the correct reasoning process step by step and finally get the correct answer 25. But when MathAttack simply changed class in the original sample to group, ChatGPT can come up with the wrong reasoning process and get the wrong equation (x+10=50), end up with the wrong answer 40. We show more real cases predicted by ChatGPT on adversarial samples in Appendix A. These cases show that the robustness of LLMs in math solving ability still needs to be strengthened.
New MWP Dataset RobustMath
Using the transferability of adversarial samples, we attack ChatGPT by MathAttack to build our RobustMath dataset. Specifically, we first utilize GSM8K and MultiAirth as our seed data then attack ChatGPT to generate adversarial samples. After that, in order to ensure the high quality of RobustMath, we manually check each adversarial sample and filter out samples which change the mathematical logic. Ultimately, our RobustMath has 300 high-quality adversarial samples that can be used to measure the robustness of large language models’ math solving ability. RobustMath is available at this link, and Appendix B shows some samples of RobustMath.
To verify the effectiveness of our RobustMath, we evaluate large language models on RobustMath. In addition to the models mentioned above, we also evaluate large language model that is fine-tuned on MWP datasets. Specifically, we follow (Fu et al. 2023) to finetune Flan-T5-xl with 200k MWP data. In Table 4, we observe that the performance of the LLMs on RobustMath is significantly worse compared to the performance on its original samples. This indicates that our RobustMath can effectively measure the robustness of the model’s math solving ability. When examining the performance of models with zero-shot performance, we can see that as the model’s capability increases, its performance on RobustMath also increases. However, it still does not exceed 37.00%. Moreover, after finetuning, the performance of Flan-T5-xl increases from 6.00% to 12.33%. It indicates that finetuning on specific data could help improve the robustness of models. Observing the performance of models with few-shot prompt, we find that models with a strong in-context ability such as Flan-T5-large and Flan-T5-xl can effectively enhance their performance on RobustMath. In contrast, ChatGLM2 and finetuned Flan-T5-xl which have weaker in-context ability do not exhibit significant improvements on both the original samples set and RobustMath under few-shot prompt.
Conclusion and Future Work
In this paper, we make a first attempt to attack MWP samples to examine the security of LLMs in math solving ability. To preserve the mathematical logic of MWPs, we propose a MathAttack model with a logical entity recognition block. Extensive experiments show that MathAttack could effectively attack the math solving ability. Through the comprehensive experimental analysis, we have three significant findings: (1) Transferability of attacking samples (2) Complex MWPs (such as more solving steps, longer text, more numbers) are more vulnerable to attack, and (3) Attacking samples used in few-shot prompts can improve robustness of LLMs. Furthermore, we propose a new dataset RobustMath by utilizing MathAttack and manual check, which consists of high-quality MWP Adversarial samples and could measure the robustness of LLMs’ math solving ability. We hope our practice and observations can serve as an important attempt to enhance the robustness of LLMs in math solving ability. In the future, we will explore methods such as instruction learning or reinforcement learning to enhance the robustness of the models. As large language models are increasingly being applied in the field of intelligence education, the importance of improving their robustness becomes more significant.
References
Appendix
Table 5 shows three cases predicted by ChatGPT on original samples and their adversarial samples.
B
Table 6 shows eight samples from our proposed RobustMath dataset and their corresponding adversarial samples.