Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, Bo Li

Introduction

Pre-trained language models have achieved state-of-the-art performance over a wide range of Natural Language Understanding (NLU) tasks . However, recent studies reveal that even these large-scale language models are vulnerable to carefully crafted adversarial examples, which can fool the models to output arbitrarily wrong answers by perturbing input sentences in a human-imperceptible way. Real-world systems built upon these vulnerable models can be misled in ways that would have profound security concerns .

To address this challenge, various methods have been proposed to improve the adversarial robustness of language models. However, the adversary setup considered in these methods lacks a unified standard. For example, Jiang et al. 2020, Liu et al. 2020 mainly evaluate their robustness against human-crafted adversarial datasets , while Wang et al. 2021 evaluate the model robustness against automatic adversarial attack algorithms . The absence of a principled adversarial benchmark makes it difficult to compare the robustness across different models and identify the adversarial attacks that most models are vulnerable to. This motivates us to build a unified and principled robustness evaluation benchmark for natural language models and hope to help answer the following questions: what types of language models are more robust when evaluated on the unified adversarial benchmark? Which adversarial attack algorithms against language models are more effective, transferable, or stealthy to human? How likely can human be fooled by different adversarial attacks?

We list out the fundamental principles to create a high-quality robustness evaluation benchmark as follows. First, as also pointed out by , a reliable benchmark should be accurately and unambiguously annotated by humans. This is especially crucial for the robustness evaluation, as some adversarial examples generated by automatic attack algorithms can fool humans as well . Given our analysis in §3.4, among the generated adversarial data, there are only around 10%10\% adversarial examples that receive at least 4-vote consensus among 5 annotators and align with the original label. Thus, additional rounds of human filtering are critical to validate the quality of the generated adversarial attack data. Second, a comprehensive robustness evaluation benchmark should cover enough language phenomena and generate a systematic diagnostic report to understand and analyze the vulnerabilities of language models. Finally, a robustness evaluation benchmark needs to be challenging and unveil the biases shared across different models.

In this paper, we introduce Adversarial GLUE (AdvGLUE), a multi-task benchmark for robustness evaluation of language models. Compared to existing adversarial datasets, there are several contributions that render AdvGLUE a unique and valuable asset to the community.

Comprehensive Coverage. We consider textual adversarial attacks from different perspectives and hierarchies, including word-level transformations, sentence-level manipulations, and human-written adversarial examples, so that AdvGLUE is able to cover as many adversarial linguistic phenomena as possible.

Systematic Annotations. To the best of our knowledge, this is the first work that performs systematic and comprehensive evaluation and annotation over 14 different textual adversarial examples. Concretely, AdvGLUE adopts crowd-sourcing to identify high-quality adversarial data for reliable evaluation.

General Compatibility. To obtain comprehensive understanding of the robustness of language models across different NLU tasks, AdvGLUE covers the widely-used GLUE tasks and creates an adversarial version of the GLUE benchmark to evaluate the robustness of language models.

High Transferability and Effectiveness. AdvGLUE has high adversarial transferability and can effectively attack a wide range of state-of-the-art models. We observe a significant performance drop for models evaluated on AdvGLUE compared with their standard accuracy on GLUE leaderboard. For instance, the average GLUE score of ELECTRA(Large) drops from 93.1693.16 to 41.6941.69.

Our contributions are summarized as follows. (ii) We propose AdvGLUE, a principled and comprehensive benchmark that focuses on robustness evaluation of language models. (iiii) During the data construction, we provide a thorough analysis and a fair comparison of existing strong adversarial attack algorithms. (iiiiii) We present thorough robustness evaluation for existing state-of-the-art language models and defense methods. We hope that AdvGLUE will inspire active research and discussion in the community. More details are available at https://adversarialglue.github.io.

Related Work

Existing robustness evaluation work can be roughly divided into two categories: Evaluation Toolkits and Benchmark Datasets. (ii) Evaluation toolkits, including OpenAttack , TextAttack , TextFlint and Robustness Gym , integrate various ad hoc input transformations for different tasks and provide programmable APIs to dynamically test model performance. However, it is challenging to guarantee the quality of these input transformations. For example, as reported in , the validity of adversarial transformation can be as low as 65.5%65.5\%, which means that more than one third of the adversarial sentences have wrong labels. Such a high percentage of annotation errors could lead to an underestimate of model robustness, making it less qualified to serve as an accurate and reliable benchmark . (iiii) Benchmark datasets for robustness evaluation create challenging testing cases by using human-crafted templates or rules , or adopting a human-and-model-in-the-loop manner to write adversarial examples . While the quality and validity of these adversarial datasets can be well controlled, the scalability and comprehensiveness are limited by the human annotators. For example, template-based methods require linguistic experts to carefully construct reasonable rules for specific tasks, and such templates can be barely transferable to other tasks. Moreover, human annotators tend to complete the writing tasks through minimal efforts and shortcuts , which can limit the coverage of various linguistic phenomena.

Dataset Construction

In this section, we provide an overview of our evaluation tasks, as well as the pipeline of how we construct the benchmark data. During this data construction process, we also compare the effectiveness of different adversarial attack methods, and present several interesting findings.

Tasks. We consider the following five most representative and challenging tasks used in GLUE : Sentiment Analysis (SST-2), Duplicate Question Detection (QQP), and Natural Language Inference (NLI, including MNLI, RTE, QNLI). The detailed explanation for each task can be found in Appendix A.3. Some tasks in GLUE are not included in AdvGLUE, since there are either no well-defined automatic adversarial attacks (e.g., CoLA), or insufficient data (e.g., WNLI) for the attacks.

Dataset Statistics and Evaluation Metrics. AdvGLUE follows the same training data and evaluation metrics as GLUE. In this way, models trained on the GLUE training data can be easily evaluated under IID sampled test sets (GLUE benchmark) or carefully crafted adversarial test sets (AdvGLUE benchmark). Practitioners can understand the model generalization via the GLUE diagnostic test suite and examine the model robustness against different levels of adversarial attacks from the AdvGLUE diagnostic report with only one-time training. Given the same evaluation metrics, model developers can clearly understand the performance gap between models tested in the ideally benign environments and approximately worst-case adversarial scenarios. We present the detailed dataset statistics under various attacks in Table 1. Detailed label distribution and evaluation metrics are in Appendix Table 8.

2 Adversarial Perturbations

In this section, we detail how we optimize different levels of adversarial perturbations to the benign source samples and collect the raw adversarial data with noisy labels, which will then be carefully filtered by human annotators described in the next section. Specifically, we consider the dev sets of GLUE benchmark as our source samples, upon which we perform different adversarial attacks. For relatively large-scale tasks (QQP, QNLI, MNLI-m/mm), we sample 1,000 cases from the dev sets for efficiency purpose. For the remaining tasks, we consider the whole dev sets as source samples.

Existing word-level adversarial attacks perturb the words through different strategies, such as perturbing words with their synonyms or carefully crafted typo words (e.g., “foolish” to “fo01ish”), such that the perturbation does not change the semantic meaning of the sentences but dramatically change the models’ output. To examine the model robustness against different perturbation strategies, we select one representative adversarial attack method for each strategy as follows.

Typo-based Perturbation. We select TextBugger as the representative algorithm for generating typo-based adversarial examples. When performing the attack, TextBugger first identifies the important words and then replaces them with typos.

Embedding-similarity-based Perturbation. We choose TextFooler as the representative adversarial attack that considers embedding similarity as a constraint to generate semantically consistent adversarial examples. Essentially, TextFooler first performs word importance ranking, and then substitutes those important ones to their synonyms extracted according to the cosine similarity of word embeddings.

Context-aware Perturbation. We use BERT-ATTACK to generate context-aware perturbations. The fundamental difference between BERT-ATTACK and TextFooler lies on the word replacement procedure. Specifically, BERT-ATTACK uses the pre-trained BERT to perform masked language prediction to generate contextualized potential word replacements for those crucial words.

Knowledge-guided Perturbation. We consider SememePSO as an example to generate adversarial examples guided by the HowNet knowledge base. SememePSO first finds out substitutions for each word in HowNet based on sememes, and then searches for the optimal combination based on particle swarm optimization.

Compositions of different Perturbations. We also implement a whitebox-based adversarial attack algorithm called CompAttack that integrates the aforementioned perturbations in one algorithm to evaluate model robustness to various adversarial transformations. Moreover, we efficiently search for perturbations via optimization so that CompAttack can achieve the attack goal while perturbing the minimal number of words. The implementation details can be found in Appendix A.4.

We note that the above adversarial attacks require a surrogate model to search for the optimal perturbations. In our experiments, we follow the setup of ANLI and generate adversarial examples against three different types of models (BERT, RoBERTa, and RoBERTa ensemble) trained on the GLUE benchmark. We then perform one round of filtering to retain those examples with high adversarial transferability between these surrogate models. We discuss more implementation details and hyper-parameters of each attack method in Appendix A.4.

2.2 Sentence-level Perturbation

Different from word-level attacks that perturb specific words, sentence-level attacks mainly focus on the syntactic and logical structures of sentences. Most of them achieve the attack goal by either paraphrasing the sentence, manipulating the syntactic structures, or inserting some unrelated sentences to distract the model attention. AdvGLUE considers the following representative perturbations.

Syntactic-based Perturbation. We incorporate three adversarial attack strategies that manipulate the sentence based on the syntactic structures. (ii) Syntax Tree Transformations. SCPN is trained to produce a paraphrase of a given sentence with specified syntactic structures. Following the default setting, we select the most frequent 1010 templates from ParaNMT-50M corpus to guide the generation process. An LSTM-based encoder-decoder model (SCPN) is used to generate parses of target sentences according to the templates. These parses are further fed into another SCPN to generate full sentences. We use the pre-trained SCPNs released by the official codebase. (iiii) Context Vector Transformations. T3 is a whitebox attack algorithm that can add perturbations on different levels of the syntax tree and generate the adversarial sentence. In our setting, we add perturbations to the context vector of the root node given syntax tree, which is iteratively optimized to construct the adversarial sentence. (iiiiii) Entailment Preserving Transformations. We follow the entailment preserving rules proposed by AdvFever , and transform all the sentences satisfying the templates into semantically equivalent ones. More details can be found in Appendix A.4.

Distraction-based Perturbation. We integrate two attack strategies: (ii) StressTest appends three true statements (“and true is true”, “and false is not true”, “and true is true” for five times) to the end of the hypothesis sentence for NLI tasks. (iiii) CheckList adds randomly generated URLs and handles to distract model attention. Since the aforementioned distraction-based perturbations may impact the linguistic acceptability and the understanding of semantic equivalence, we mainly apply these rules to part of the GLUE tasks, including SST-2 and NLI tasks (MNLI, RTE, QNLI), to evaluate whether model can be easily misled by the strong negation words or such lexical similarity.

2.3 Human-crafted Examples

To ensure our benchmark covers more linguistic phenomena in addition to those provided by automatic attack algorithms, we integrate the following high-quality human-crafted adversarial data from crowd-sourcing or expert-annotated templates and transform them to the formats of GLUE tasks.

CheckList We note that both CheckList and StressTest propose both rule-based distraction sentences and manually crafted templates to generate test samples. The former is considered as sentence-level distraction-based perturbations, while the latter is considered as human-crafted examples. is a testing method designed for analysing different capabilities of NLP models using different test types. For each task, CheckList first identifies necessary natural language capabilities a model should have, then designs several test templates to generate test cases at scale. We follow the instructions and collect testing cases for three tasks: SST-2, QQP and QNLI. For each task, we adopt two capability tests: Temporal and Negation, which test if the model understands the order of events and if the model is sensitive to negations.

StressTest proposes carefully crafted rules to construct “stress tests” and evaluate robustness of NLI models to specific linguistic phenomena. We adopt the test cases focusing on Numerical Reasoning into our adversarial MNLI dataset. These premise-hypothesis pairs are able to test whether the model can perform reasoning involving numbers and quantifiers and predict the correct relation between premise and hypothesis.

ANLI is a large-scale NLI dataset collected iteratively in a human-in-the-loop manner. In each iteration, human annotators are asked to design sentences to fool current model. Then the model is further finetuned on a larger dataset incorporating these sentences, which leads to a stronger model. Finally, annotators are asked to write harder examples to detect the weakness of this stronger model. In the end, the sentence pairs generated in each round form a comprehensive dataset that aims at examining the vulnerability of NLI models. We adopt ANLI into our adversarial MNLI dataset. We obtain the permission from the ANLI authors to include the ANLI dataset as part of our leaderboard.

AdvSQuAD is an adversarial dataset targeting at reading comprehension systems. Adversarial examples are generated by appending a distracting sentence to the end of the input paragraph. The distracting sentences are carefully designed to have common words with questions and look like a correct answer to the question. We mainly consider the examples generated by AddSent and AddOneSent strategies, and adopt the distracting sentences and questions in the QNLI format with labels “not answered”. The use of AdvSQuAD in AdvGLUE is authorized by the authors.

We present sampled AdvGLUE examples with the word-level, sentence-level perturbations and human-crafted samples in Table 2. More examples are provided in Appendix A.5.

3 Data Curation

After collecting the raw adversarial dataset, additional rounds of filtering are required to guarantee its quality and validity. We consider two types of filtering: automatic filtering and human evaluation.

Automatic Filtering mainly evaluates the generated adversarial examples along two fronts: transferability and fidelity.

Transferability evaluates whether the adversarial examples generated against one source model (e.g., BERT) can successfully transfer and attack the other two (e.g., RoBERTa and RoBERTa ensemble), given the surrogate models used to generate adversarial examples (BERT, RoBERTa and RoBERTa ensemble). Only adversarial examples that can successfully transfer to the other two models will be kept for the next round of fidelity filtering, so that the selected examples can exploit the biases shared across different models and unveil their fundamental weakness.

Fidelity evaluates how the generated adversarial examples maintain the original semantics. For word-level adversarial examples, we use word modification rate to measure what percentage of words are perturbed. Concretely, word-level adversarial examples with word modification rate larger than 15%15\% are filtered out. For sentence-level adversarial examples, we use BERTScore to evaluate the semantic similarity between the adversarial sentences and their corresponding original ones. For each sentence-level attack, adversarial examples with the highest similarity scores are kept to guarantee their semantic closeness to the benign samples.

Human Evaluation validates whether the adversarial examples preserve the original labels and whether the labels are highly agreed among annotators. Concretely, we recruit annotators from Amazon Mechanical Turk. To make sure the annotators fully understand the GLUE tasks, each worker is required to pass a training step to be qualified to work on the main filtering tasks for the generated adversarial examples. We tune the pay rate for different tasks, as shown in Appendix Table 11. The pay rate of the main filtering phase is twice as much as that of the training phase.

Human Training Phase is designed to ensure that the annotators understand the tasks. The annotation instructions for each task follows , and we provide at least two examples for each class to help annotators understand the tasks. Instructions can be found at https://adversarialglue.github.io/instructions. Each annotator is required to work on a batch of 20 examples randomly sampled from the GLUE dev set. After annotators answer each example, a ground-truth answer will be provided to help them understand whether the answer is correct. Workers who get at least 85%85\% of the examples correct during training are qualified to work on the main filtering task. A total of 100 crowd workers participated in each task, and the number of qualified workers are shown in Appendix Table 11. We also test the human accuracy of qualified annotators for each task on 100 randomly sampled examples from the dev set excluding the training samples. The details and results can be found in Appendix Table 11.

Human Filtering Phase verifies the quality of the generated adversarial examples and only maintains high-quality ones to construct the benchmark dataset. Specifically, annotators are required to work on a batch of 10 adversarial examples generated from the same attack method. Every adversarial example will be validated by 5 different annotators. Examples are selected following two criteria: (ii) high consensus: each example must have at least 4-vote consensus; (iiii) utility preserving: the majority-voted label must be the same as the original one to make sure the attacks are valid (i.e., cannot fool human) and preserve the semantic content.

The data curation results including inter-annotator agreement rate (Fleiss Kappa) and human accuracy on the curated dataset are shown in Table 3. We will provide more analysis in the next section. Note that even after the data curation step, some grammatical errors and typos can still remain, as some adversarial attacks intentionally inject typos (e.g., TextBugger) or manipulate syntactic trees (e.g., SCPN) which are very stealthy. We will retain these samples as their labels receive high consensus from annotators, which means the typos do not substantially impact humans’ understanding.

4 Benchmark of Adversarial Attack Algorithms

Our data curation phase also serves as a comprehensive benchmark over existing adversarial attack methods, as it provides a fair standard for all adversarial attacks and systematic human annotations to evaluate the quality of the generated samples.

Evaluation Metrics. Specifically, we evaluate these attacks along two fronts: effectiveness and validity. For effectiveness, we consider two evaluation metrics: Attack Success Rate (ASR) and Curated Attack Success Rate (Curated ASR). Formally, given a benign dataset D={(x(i),y(i))}i=1N\mathcal{D}=\{(x^{(i)},y^{(i)})\}^{N}_{i=1} consisting of NN pairs of sample x(i)x^{(i)} and ground truth y(i)y^{(i)}, for an adversarial attack method A\mathcal{A} that generates an adversarial example A(x)\mathcal{A}(x) given an input xx to attack a surrogate model ff, ASR is calculated as

For validity, we consider three evaluation metrics: Filter Rate, Fleiss Kappa, and Human Accuracy. Specifically, Filter Rate is calculated by 1−Curated ASRASR1-\frac{\textrm{Curated ASR}}{\textrm{ASR}} to measure how many examples are rejected in the data curation procedures and can reflect the noisiness of the generated adversarial examples. We report the average ASR, Curated ASR, and Filter Rate over the three surrogate models we consider in Table 3. Fleiss Kappa is a widely used metric in existing datasets (e.g., SNLI, ANLI, and FEVER ) to measure the inter-annotator agreement rate on the collected dataset. Fleiss Kappa between 0.4 and 0.6 is considered as moderate agreement and between 0.6 and 0.8 as substantial agreement. The inter-annotator agreement rates of most high-quality datasets fall into these two intervals. In this paper, we follow the standard protocol and report Fleiss Kappa and Curated Fleiss Kappa to analyze the inter-annotator agreement rate on the collected adversarial dataset before and after curation to reflect the ambiguity of generated examples. We also estimate the human performance on our curated datasets. Specifically, given a sample with 5 annotations, we take one random annotator’s annotation as the prediction and the majority voted label as the ground truth and calculate the human accuracy as shown in Table 3.

Analysis. As shown in Table 3, in terms of attack effectiveness, while most attacks show high ASR, the Curated ASR is always less than 11%11\%, which indicates that most existing adversarial attack algorithms are not effective enough to generate high-quality adversarial examples. In terms of validity, the filter rates for most adversarial attack methods are more than 85%85\%, which suggests that existing strong adversarial attacks are prone to generating invalid adversarial examples that either change the original semantic meanings or generate ambiguous perturbations that hinder the annotators’ unanimity. We provide detailed filter rates for automatic filtering and human evaluation in Appendix Table 12, and the conclusion is that around 60−80%60-80\% of examples are filtered due to the low transferability and high word modification rate. Among the remaining samples, around 30−40%30-40\% examples are filtered due to the low human agreement rates (Human Consensus Filtering), and around 20−30%20-30\% are filtered due to the semantic changes which lead to the label changes (Utility Preserving Filtering). We also note that the data curation procedures are indispensable for the adversarial evaluation, as the Fleiss Kappa before curation is very low, suggesting that a lot of adversarial sentences have unreliable labels and thus tend to underestimate the model robustness against the textual adversarial attacks. After the data curation, our AdvGLUE shows a Curated Fleiss Kappa of near 0.6, comparable with existing high-quality dataset such as SNLI and ANLI. Among all the existing attack methods, we observe that TextBugger is the most effective and valid attack method, as it demonstrates the highest Curated ASR and Curated Fleiss Kappa across different tasks.

5 Finalizing the Dataset

The full pipeline of constructing AdvGLUE is summarized in Figure 1.

Merging. We note that distraction-based adversarial examples and human-crafted adversarial examples are guaranteed to be valid by definition or crowd-sourcing annotations, and thus data curation is not needed on these attacks. When merging them with our curated set, we calculate the average number of samples per attack from our curated set, and sample the same amount of adversarial examples from these attacks following the same label distribution. This way, each attack contributes to similar amount of adversarial data, so that AdvGLUE can evaluate models against different types of attacks with similar weights and provide a comprehensive and unbiased diagnostic report.

Dev-Test Split. After collecting the adversarial examples from the considered attacks, we split the final dataset into a dev set and a test set. In particular, we first randomly split the benign data into 9:19:1, and the adversarial examples generated based on 90%90\% of the benign data serve as the hidden test set, while the others are published as the dev set. For human-crafted adversarial examples, since they are not generated based on the benign GLUE data, we randomly select 90%90\% of the data as the test set, and the remaining 10%10\% as the dev set. The dev set is publicly released to help participants to understand the tasks and the data format. To protect the integrity of our test data, the test set will not be released to the public. Instead, participants are required to upload the model to CodaLab, which automates the evaluation process on the hidden test set and provides a diagnostic report.

Diagnostic Report for Language Models

Benchmark Results. We follow the official implementations and training scripts of pre-trained language models to reproduce results on GLUE and test these models on AdvGLUE. The training details can be found in Appendix A.6. Results are summarized in Table 4. We observe that although state-of-the-art language models have achieved high performance on GLUE, they are vulnerable to various adversarial attacks. For instance, the performance gap can be as large as 55%55\% on the SMART (BERT) model in terms of the average score. DeBERTa (Large) and ALBERT (XXLarge) achieve the highest average AdvGLUE scores among all the tested language models. This result is also aligned with the ANLI leaderboard https://github.com/facebookresearch/anli, which shows that ALBERT (XXLarge) is the most robust to human-crafted adversarial NLI dataset .

We note that although our adversarial examples are generated from surrogate models based on BERT and RoBERTa, these examples have high transferability between models after our data curation. Specifically, the average score of ELECTRA (Large) on AdvGLUE is even lower than RoBERTa (Large), which demonstrates that AdvGLUE can effectively transfer across models of different architectures and unveil the vulnerabilities shared across multiple models. Moreover, we find some models even perform worse than random guess. For example, the performance of BERT on AdvGLUE for all tasks is lower than random-guess accuracy.

We also benchmark advanced robust training methods to evaluate whether these methods can indeed provide robustness improvement on AdvGLUE and to what extent. We observe that SMART and FreeLB are particularly helpful to improve robustness for RoBERTa. Specifically, SMART (RoBERTa) improves RoBERTa (Large) over 3.71%3.71\% on average, and it even improves the benign accuracy as well. Since InfoBERT is not evaluated on GLUE, we run InfoBERT with different hyper-parameters and report the best accuracy on benign GLUE dev set and AdvGLUE test set. However, we find that the benign accuracy of InfoBERT (RoBERTa) is still lower than RoBERTa (Large), and similarly for the robust accuracy. These results suggest that existing robust training methods only have incremental robustness improvement, and there is still a long way to go to develop robust models to achieve satisfactory performance on AdvGLUE.

Diagnostic Report of Model Vulnerabilities. To have a systematic understanding of which adversarial attacks language models are vulnerable to, we provide a detailed diagnostic report in Table 5. We observe that models are most vulnerable to human-crafted examples, where complex linguistic phenomena (e.g., numerical reasoning, negation and coreference resolution) can be found. For sentence-level perturbations, models are more vulnerable to distraction-based perturbations than directly manipulating syntactic structures. In terms of word-level perturbations, models are similarly vulnerable to different word replacement strategies, among which typo-based perturbations and knowledge-guided perturbations are the most effective attacks.

We hope the above findings can help researchers systematically examine their models against different adversarial attacks, thus also devising new methods to defend against them. Comprehensive analysis of the model robustness report is provided in our website and Appendix A.9.

Conclusion

We introduce AdvGLUE, a multi-task benchmark to evaluate and analyze the robustness of state-of-the-art language models and robust training methods. We systematically conduct 14 adversarial attacks on GLUE tasks and adopt crowd-sourcing to guarantee the quality and validity of generated adversarial examples. Modern language models perform poorly on AdvGLUE, suggesting that model vulnerabilities to adversarial attacks still remain unsolved. We hope AdvGLUE can serve as a comprehensive and reliable diagnostic benchmark for researchers to further develop robust models.

Acknowledgments and Disclosure of Funding

We thank the anonymous reviewers for their constructive feedback. We also thank Prof. Sam Bowman, Dr. Adina Williams, Nikita Nangia, Jinfeng Li, and many others for the helpful discussion. We thank Prof. Robin Jia and Yixin Nie for allowing us to incorporate their datasets as part of the evaluation. We thank the SQuAD team for allowing us to use their website template and submission tutorials. This work is partially supported by the NSF grant No.1910100, NSF CNS 20-46726 CAR, the Amazon Research Award.

References

Appendix A Appendix

We present a glossary of adversarial attacks considered in AdvGLUE in Table 6 and 7.

A.2 Additional Related Work

We discuss more related work about textual adversarial attacks and defenses in this subsection.

Recent research has shown deep neural networks (DNNs) are vulnerable to adversarial examples that are carefully crafted to fool machine learning models without disturbing human perception . However, compared with a large amount of adversarial attacks in continuous data domain , there are a few studies focusing on the discrete text domain. Most existing gradient-based attacks on image or audio models are no longer applicable to NLP models, as words are intrinsically discrete tokens. Another challenge for generating adversarial text is to ensure the semantic and syntactic coherence and consistency.

Existing textual adversarial attacks can be roughly divided into three categories: word-level transformations, sentence-level attacks, and human-crafted samples. (ii) Word-level transformations adopt different word replacement strategies during attack. For example, existing work applies character-level perturbation to carefully crafted typo words (e.g., from “foolish” to “fo0lish”), thus making the model ignore or misunderstand the original statistical cues. Others adopt knowledge-based perturbation and utialize knowledge base to constrain the search space. For example, Zang et al. 2020 uses sememe-based knowledge base from HowNet to construct a search space for word substitution. Some use non-contextualized word embedding from GLoVe or Word2Vec to build synonym candidates, by querying the cosine similarity or euclidean distance between the original and candidate word and selecting the closet ones as the replacements. Recent work also leverages BERT to generate contextualized perturbations by masked language modeling. (iiii) Different from the dominant word-level adversarial attacks, sentence-level adversarial attacks perform sentence-level transformation or paraphrasing by perturbing the syntactic structures based on human crafted rules or carefully designed auto-encoders . Sentence-level manipulations are generally more challenging than word-level attacks, because the perturbation space for syntactic structures are limited compared to word-level perturbation spaces that grow exponentially with the sentence length. However, sentence-level attacks tend to have higher linguistic quality than word-level, as both semantic and syntactic coherence are taken into considerations when generating adversarial sentences. (iiiiii) Human-crafted adversarial examples are generally crafted in the human-in-the-loop manner or use manually crafted templates to generate test cases . Our AdvGLUE incorporates all of the above textual adversarial to provide a comprehensive and systematic diagnostic report over existing state-of-the-art large-scale language models.

A.3 Task Descriptions, Statistics and Evaluation Metrics

We present the detailed label distribution statistics and evaluation metrics of GLUE and AdvGLUE benchmark in 8.

The Stanford Sentiment Treebank consists of sentences from movie reviews and human annotations of their sentiment. Given a review sentence, the task is to predict the sentiment of it. Sentiments can be divided into two classes: positive and negative.

The Quora Question Pairs (QQP) dataset is a collection of question pairs from the community question-answering website Quora. The task is to determine whether a pair of questions are semantically equivalent.

The Multi-Genre Natural Language Inference Corpus consists of sentence pairs with textual entailment annotations. Given a premise sentence and a hypothesis sentence, the task is to predict whether the premise entails the hypothesis (entailment), contradicts the hypothesis (contradiction), or neither (neutral)

Question-answering NLI (QNLI) dataset consists of question-sentence pairs modified from The Stanford Question Answering Dataset . The task is to determine whether the context sentence contains the answer to the question.

The Recognizing Textual Entailment (RTE) dataset is a combination of a series of data from annual textual entailment challenges. Examples are constructed based on news and Wikipedia text. The task is to predict the relationship between a pair of sentences. For consistency, the relationship can be classified into two classes: entailment and not entailment, where neutral and contradiction are seen as not entailment.

We also show the detailed per-task model performance on AdvGLUE and GLUE in Table 9.

A.4 Implementation Details of Adversarial Attacks

To ensure the small magnitude of the perturbation, we consider the following five strategies: (ii) randomly inserting a space into a word; (iiii) randomly deleting a character of a word; (iiiiii) randomly replacing a character of a word with its adjacent character in the keyboard; (iviv) randomly replacing a character of a word with its visually similar counterpart (e.g., “0” v.s. “o”, “1” v.s. “l”); and (vv) randomly swapping two characters in a word. The first four strategies guarantee the word edit distance between the typo word and its original word to be 1, and that of the last strategy is limited to 2. Following the default setting, in Strategy (ii), we only insert a space into a word when the word contains less than 66 characters. In Strategy (vv), we swap characters in a word only when the word has more than 44 characters.

Concretely, for the sentiment analysis tasks, we set the cosine similarity threshold to be 0.80.8, which encourages the synonyms to be semantically close to original ones and enhances the quality of adversarial data. For the rest of the tasks, we follow the default hyper-parameter to set the cosine similarity threshold to be 0.70.7. Besides, the number of synonyms for each word is set to 5050 following the default setting.

We follow the hyper-parameters from the official codebase, and set the number of candidate words to 48 and cosine similarity threshold to 0.40.4 in order to filter out antonyms using synonym dictionaries, as BERT masked language model does not distinguish synonyms and antonyms.

We adopt the official hyper-parameters in which maximum and minimum inertia weights are set to 0.80.8 and 0.20.2, respectively. We also set the maximum and minimum movement probabilities of the particles to 0.80.8 and 0.20.2, respectively, following the default setting. Population size is set to 6060 in every task.

We follow the T3 and C&W attack and design the same optimization objective for adversarial perturbation generation in the embedding space as:

where the first term controls the magnitude of perturbation, while g(⋅)g(\cdot) is the attack objective function depending on the attack scenario. cc weighs the attack goal against attack cost. CompAttack constrains the perturbation to be close to pre-defined perturbation space, including typo space (e.g., TextBugger), knowledge space (e.g., WordNet) and contextualized embedding space (e.g., BERT embedding clusters) to make sure the perturbation is valid. We can also see from Table 3 that CompAttack overall has lower filter rate than other state-of-the-art attack methods.

We use the pre-trained SCPN models released by the official codebase. Following the default setting, we select the most frequent 1010 templates from ParaNMT-50M corpus to guide the generation process. We first parse sentences from GLUE dev set using Stanford CoreNLP. We used CoreNLP version 3.7.0 in our experiment, along with the Shift-Reduce Parser models.

We follow the hyper-parameters in the official setting where the scaling const is set to 1e41e4 and the optimizing confidence is set to 00. In each iteration, we optimize the perturbation vector for at most 100100 steps with learning rate 0.10.1.

We follow the entailment preserving rules proposed by the official implementation. We adopt all 2323 templates to transform original sentences into semantically equivalent ones. Many common sentence patterns in everyday life are included in these templates.

A.5 Examples of AdvGLUE benchmark

We show more comprehensive examples in Table 10. Examples are generated with different levels of perturbations and they all can successfully change the predictions of all surrogate models (BERT, RoBERTa and RoBERTa ensemble).

A.6 Fine-tuning Details of Large-Scale Language Models

For all the experiments, we are using a GPU cluster with 8 V100 GPUs and 256GB memory.

For RTE, we train our model for 1010 epochs and for other tasks we train our model for 44 epochs. Batch size for QNLI is set to 512512, and for other tasks it is set to 256256. Learning rates are all set to 2e−52e-5.

We follow the official hyper-parameter setting to set the learning rate to 5e−55e-5 and set batch size to 3232. We train ELECTRA on RTE for 1010 epochs and train for 22 epochs on other tasks. We set the weight decay rate to 0.010.01 for every task.

We train our RoBERTa for 1010 epochs with learning rate 2e−52e-5 on each task. The batch size for QNLI is 3232 and 6464 for other tasks.

We train our T5 for 1010 epochs with learning rate 2e−52e-5 on each task. The batch size for QNLI is 3232 and 6464 for other tasks. We follow the templates in original paper to convert GLUE tasks into generation tasks.

We use the default hyper-parameters to train our ALBERT. For example, max training steps for SST-2, MNLI, QNLI, QQP, RTE, is 2093520935, 1000010000, 3311233112, 1400014000, 800800 respectively. For MNLI and QQP, batch size is set to 3232 and for other tasks batch size is set to 128128.

We use the official hyper-parameters to train our DeBERTa. For example, learning rate is set to 1e−51e-5 across all tasks. For MNLI and QQP, batch size is set to 6464 and for other tasks batch size is set to 32.

For SMART(BERT) and SMART(RoBERTa), we use grid search to search for the best parameters and report the best performance among all trained models.

For FreeLB, we test every parameter combination provided by the official codebase and select the best parameters for our training.

We set the batch size to 3232 and learning rate to 2e−52e-5 for all tasks.

A.7 Human Evaluation Details

We present the pay rate and the number of qualified workers in Table 11. We also test our qualified workers on another non-overlapping 100 samples of the GLUE dev sets for each task. We can see that the human accuracy is comparable to , which means that most our selected annotators understand the GLUE tasks well.

The detailed filtering statistics of each stage is shown in Table 12. We can see that around 60−80%60-80\% of examples are filtered due to the low transferability and high word modification rate. Among the remaining samples, around 30−40%30-40\% examples are filtered due to the low human agreement rates (Human Consensus Filtering), and around 20−30%20-30\% are filtered due to the semantic changes which lead to the label changes (Utility Preserving Filtering).

We show examples of annotation instructions in the training phase and filtering phase on MNLI in Figure 2 and 3. More instructions can be found in https://adversarialglue.github.io/instructions. We also provide a FAQ document in each task description page https://docs.google.com/document/d/1MikHUdyvcsrPqE8x-N-gHaLUNAbA6-Uvy-iA5gkStoc/edit?usp=sharing.

A.8 Discussion of Limitations

Due to the constraints of computational resources, we are unable to conduct a comprehensive evaluation of all existing language models. However, with the release of our leaderboard website, we are expecting researchers to actively submit their models and evaluate against our AdvGLUE benchmark to have a systematic understanding of model robustness. We are also interested in the adversarial robustness of large-scale auto-regressive language models under the few-shot settings, and leave it as a compelling future work.

In this paper, we follow ANLI and generate adversarial examples against surrogate models based on BERT and RoBERTa. However, there are concerns that such adversarial filtering may not be able to fairly benchmark the model robustness, as participants may top the leaderboard by producing different errors from our surrogate models. We note that such concerns can be solved given systematic data curation. As shown in our main benchmark results, we observe we successfully select the adversarial examples with high adversarial transferability that can unveil the vulnerabilities shared across models of different architectures. Specifically, we observe a huge performance gap in ELECTRA (Large) that is pre-trained with different data and shown less robust than one of surrogate model RoBERTa (Large).

Finally, we emphasize that our AdvGLUE benchmark mainly focuses on robustness evaluation. Thus AdvGLUE can also be considered as a supplementary diagnostic test set besides the standard GLUE benchmark. We suggest that participants should evaluate their models against both GLUE benchmark and our AdvGLUE to understand both model generalization and robustness. We hope our work can help researchers to develop models with high generalization and adversarial robustness.

A.9 Website

We present the diagnostic report on our website in Figure 4.

Appendix B Data Sheet

We follow the documentation frameworks provided by Gebru et al. 2018.

While recently a lot of methods (SMART, FreeLB, InfoBERT, ALUM) claim that they can improve the model robustness against adversarial attacks, the adversary setup in these methods (ii) lacks a unified standard and is usually different across different methods; (iiii) fails to cover comprehensive linguistic transformation (typos, synonymous substitution, paraphrasing, etc) to recognize to which levels of adversarial attacks models are still vulnerable. This motivates us to build a unified and principled robustness benchmark dataset and evaluate to which extent the state-of-the-art models have progressed so far in terms of adversarial robustness.

University of Illinois at Urbana-Champaign (UIUC) and Microsoft Corporation.

B.2 Composition/collection process/preprocessing/cleaning/labeling and uses:

The answers are described in our paper as well as website https://adversarialglue.github.io.

B.3 Distribution

The dev set is released to the public. The test set is hidden and can only be evaluated by an automatic submission API hosted on CodaLab.

The dev set is released on our website https://adversarialglue.github.io. The test set is hidden and hosted on CodaLab.

Our dataset will be distributed under the CC BY-SA 4.0 license.

B.4 Maintenance

Boxin Wang (boxinw2@illinois.edu) and Chejian Xu (xuchejian@zju.edu.cn) will be responsible for maintenance.

Yes. If we include more tasks or find any errors, we will correct the dataset and update the leaderboard accordingly. It will be updated on our website.

They can contact us via email for the contribution.