Contextualized Perturbation for Textual Adversarial Attack

Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan

Introduction

Adversarial example generation for natural language processing (NLP) tasks aims to perturb input text to trigger errors in machine learning models, while keeping the output close to the original. Besides exposing system vulnerabilities and helping improve their robustness and security (Zhao et al., 2018; Wallace et al., 2019; Cheng et al., 2019; Jia et al., 2019, inter alia), adversarial examples are also used to analyze and interpret the models’ decisions (Jia and Liang, 2017; Ribeiro et al., 2018).

Generating adversarial examples for NLP tasks can be challenging, in part due to the discrete nature of natural language text. Most recent efforts have explored heuristic rules, such as replacing tokens with their synonyms (Samanta and Mehta, 2017; Liang et al., 2019; Alzantot et al., 2018; Ren et al., 2019; Jin et al., 2020, inter alia). Despite some empirical success, rule-based methods are agnostic to context, limiting their ability to produce natural, fluent, and grammatical outputs (Wang et al., 2019b; Kurita et al., 2020, inter alia).

This work presents CLARE, a ContextuaLized AdversaRial Example generation model for text. CLARE perturbs the input with a mask-then-infill procedure: it first detects the vulnerabilities of a model and deploys masks to the inputs to indicate missing text, then plugs in an alternative using a pretrained masked language model (e.g., RoBERTa; Liu et al., 2019). CLARE features three contextualized perturbations: Replace, Insert and Merge, which respectively replace a token, insert a new one, and merge a bigram (Figure 1). As a result, it can generate outputs of varied lengths, in contrast to token replacement based methods that are limited to outputs of the same lengths as the inputs (Alzantot et al., 2018; Ren et al., 2019; Jin et al., 2020). Further, CLARE searches over a wider range of attack strategies, and is thus able to attack the victim model more effectively with fewer edits. Building on a masked language model, CLARE maximally preserves textual similarity, fluency, and grammaticality of the outputs.

We evaluate CLARE on text classification, natural language inference, and sentence paraphrase tasks, by attacking finetuned BERT models (Devlin et al., 2019). Extensive experiments and human evaluation results show that CLARE outperforms baselines in terms of attack success rate, textual similarity, fluency, and grammaticality, and strikes a better balance between attack success rate and preserving input-output similarity. Our analysis further suggests that the CLARE can be used to improve the robustness of the downstream models, and improve their accuracy when the available training data is limited. We release our code and models at https://github.com/cookielee77/CLARE.

CLARE

At a high level, CLARE applies a sequence of contextualized perturbation actions to the input. Each can be seen as a local mask-then-infill procedure: it first applies a mask to the input around a given position, and then fills it in using a pretrained masked language model (§2.1). To produce the output, CLARE scores and descendingly ranks the actions, which are then iteratively applied to the input (§2.2). We begin with a brief background review and laying out of necessary notation.

Adversarial example generation centers around a victim model ff, which we assume is a text classifier. We focus on the black-box setting, allowing access to ff’s outputs but not its configurations such as parameters. Given an input sequence x=x1x2…xn{\mathbf{x}}=x_{1}x_{2}\dots x_{n} and its label yy (assume f(x)=yf({\mathbf{x}})=y), an adversarial example x′{\mathbf{x}}^{\prime} is supposed to modify x{\mathbf{x}} to trigger an error in the victim model: f(x′)≠f(x)f({\mathbf{x}}^{\prime})\neq f({\mathbf{x}}). At the same time, textual modifications should be minimal, such that x′{\mathbf{x}}^{\prime} is close to x{\mathbf{x}} and the human predictions on x′{\mathbf{x}}^{\prime} stay the same. In computer vision applications, minor perturbations to continuous pixels can be barely perceptible to humans, thus it can be hard for one to distinguish x{\mathbf{x}} and x′{\mathbf{x}}^{\prime} (Goodfellow et al., 2015). It is not the case for text, however, since changes to the discrete tokens are more likely to be noticed by humans.

1 Masking and Contextualized Infilling

At a given position of the input sequence, CLARE can execute three perturbation actions: Replace, Insert, and Merge, which we introduce in this section. These apply masks at the given position with different strategies, and then fill in the missing text based on the unmasked context.

A Replace action substitutes the token at a given position ii with an alternative (e.g., changing “fantastic” to “amazing” in “The movie is fantastic.”). It first replaces xix_{i} with a mask, and then selects a token zz from a candidate set Z{\mathcal{Z}} to fill in:

For clarity, we denote replace⁡(x,i)\operatorname{replace}\left(\mathbf{x},i\right) by x~z\widetilde{{\mathbf{x}}}_{z}. To produce an adversarial example,

zz should fit into the unmasked context;

x~z\widetilde{{\mathbf{x}}}_{z} should be similar to x{\mathbf{x}};

x~z\widetilde{{\mathbf{x}}}_{z} should trigger an error in ff.

These can be achieved by selecting a zz such that

zz receives a high probability from a masked language model: pMLM(z∣x~)>kp_{\text{MLM}}(z\mid\widetilde{{\mathbf{x}}})>k;

ff predicts low probability for the gold label given x~z\widetilde{{\mathbf{x}}}_{z}, i.e., pf(y∣x~z)p_{f}(y\mid\widetilde{{\mathbf{x}}}_{z}) is small.

The first two requirements can be met by the construction of the candidate set: Z={\mathcal{Z}}=

V{\mathcal{V}} is the vocabulary of the masked language model. To meet the third, we select from Z{\mathcal{Z}} the token that, if filled in, will cause most “confusion” to ff:

The Insert and Merge actions differ from Replace in terms of masking strategies. The alternative token zz is selected analogously to that in a Replace action.

Insert:

This aims to add extra information to the input (e.g., changing “I recommend …” to “I highly recommend …”). It inserts a mask after xix_{i} and then fills it. Slightly overloading the notations,

Merge:

This masks out a bigram xixi+1x_{i}x_{i+1} with a single mask and then fills it, reducing the sequence length by 1:

zz can be the same as one of the masked tokens (e.g., masking out “New York” and then filling in“York”). This can be seen as deleting a token from the input.

For Insert and Merge, zz is chosen in the same manner as replace action. A perturbation will not be considered if its candidate token set is empty.

In sum, at each position ii of an input sequence, CLARE first: (i)(i) replaces xix_{i} with a mask; (ii)(ii) or inserts a mask after xix_{i}; (iii)(iii) or merges xixi+1x_{i}x_{i+1} into a mask. Then a set of candidate tokens is constructed with a masked language model and a textual similarity function; the token minimizing the gold label’s probability is chosen as the alternative token. The combination of these three operations enables conversion between any two sequences.

CLARE first constructs the local actions for all positions in parallel, i.e., the actions at position ii do not affect those at other positions. Then, to produce the adversarial example, CLARE gathers the local actions and selects an order to execute them.

2 Sequentially Applying the Perturbations

Given an input pair (x,y)({\mathbf{x}},y), let nn denote the length of x{\mathbf{x}}. CLARE chooses from 3n3n actions to produce the output: 3 actions for each position, assuming the candidate token sets are not empty. We aim to generate an adversarial example with minimum modifications to the input. To achieve this, we iteratively apply the actions, and first select those minimizing the probability of outputting the gold label yy from ff.

Each action is associated with a score, measuring how likely it can “confuse” ff: denote by a(x)a({\mathbf{x}}) the output of applying action aa to x{\mathbf{x}}. The score is then the negative probability of predicting the gold label from ff, using a(x)a({\mathbf{x}}) as the input:

Only one of the three actions can be applied at each position, and we select the one with the highest score. This constraint aims to avoid multiple modifications around the same position, e.g., merging “New York” into “Seattle” and then replacing it with “Boston”.

Actions are iteratively applied to the input, until an adversarial example is found or a limit of actions TT is reached. Each step selects the highest-scoring action from the remaining ones. Algorithm 1 summarizes the above procedure. Insert and Merge actions change the text length. When any of them is applied, we accordingly change the text indices of affected actions remaining in A{\mathcal{A}}.

A key technique of CLARE is the local mask-then-infill perturbation. Compared with existing context-agnostic replacement approaches (Alzantot et al., 2018; Jin et al., 2020; Ren et al., 2019, inter alia), contextualized infilling produces more fluent and grammatical outputs. Generating adversarial examples with masked language models is also explored by concurrent work BERTAttack (Li et al., 2020) and BAE (Garg and Ramakrishnan, 2020). Both Li et al. (2020) and Garg and Ramakrishnan (2020) are published concurrently to an initial report of this work.

BERTAttack only replaces tokens and thus can only produce outputs of the same lengths as the inputs. This is analogous with a CLARE model with the Replace action only. BAE entangles replacing and inserting tokens: it inserts only at positions neighboring a replaced token, limiting its attacking capability. Departing from both, CLARE uses three different perturbations (Replace, Insert and Merge), each allowing efficient attacking against any position of the input, and can produce outputs of varied lengths. As we will show in the experiments (§3.3), CLARE outperforms both these methods.

When selecting the attack positions, neither BERTAttack or BAE takes into account the tokens to be infilled, whereas CLARE does. This results in better adversarial attack performance according to our ablation study (§4.1).

CLARE demonstrates the advantage of using RoBERTa over BERT, which was used in the concurent works (§4.1).

Experiments

We evaluate CLARE on text classification, natural language inference, and sentence paraphrase tasks. We begin by describing the implementation details of CLARE and the baselines (§3.1). §3.2 introduces the experimental datasets and the evaluation metrics; the results are summarized in §3.3.

We experiment with a distilled version of RoBERTa (RoBERTadistill{}_{\text{distill}}; Sanh et al., 2019) as the masked language model for contextualized infilling. We also compare to base sized RoBERTa (RoBERTabase{}_{\text{base}}; Liu et al., 2019) and base sized BERT (BERTbase{}_{\text{base}}; Devlin et al., 2019) in the ablation study (§4.1).

The similarity function builds on the universal sentence encoder (USE; Cer et al., 2018).

The victim model is an MLP classifier on top of BERTbase{}_{\text{base}}. It takes as input the first token’s contextualized representation. We finetune BERT when training the victim model.

We compare CLARE with recent state-of-the-art word-level black-box adversarial attack models, including:

TextFooler: a state-of-the-art model by Jin et al. (2020). This replaces tokens with their synonyms derived from counter-fitting word embeddings (Mrkšić et al., 2016), and uses the same text similarity function as our work.

TextFooler+LM: an improved variant of TextFooler we implemented based on Alzantot et al. (2018) and Cheng et al. (2019). This inherits token replacement from TextFooler, but uses an additional small sized GPT-2 language model (Radford et al., 2019) to filter out those candidate tokens that do not fit in the context with calculated perplexity.

BERTAttack: a mask-then-infill approach by Li et al. (2020). It greedily replaces tokens with the predictions from BERT. BAE is not listed as it has a similar performance as BERTAttack (Garg and Ramakrishnan, 2020).

We use the open source implementation of the above baselines provided by the authors. More details are included in Appendix §A.1.

2 Datasets and Evaluation

We evaluate CLARE with the following datasets:

Yelp Reviews (Zhang et al., 2015): a binary sentiment classification dataset based on restaurant reviews.

AG News (Zhang et al., 2015): a collection of news articles with four categories: World, Sports, Business and Science & Technology.

MNLI (Williams et al., 2018): a natural language inference dataset. Each instance consists of a premise-hypothesis pair, and the model is supposed to determine the relation between them from a label set of entailment, neutral, and contradiction. It covers text from a variety of domains.

QNLI (Wang et al., 2019a): a binary classification dataset converted from the Stanford question answering dataset (Rajpurkar et al., 2016). The task is to determine whether the context contains the answer to a question. It is mainly based on English Wikipedia articles.

Table 1 summarizes some statistics of the datasets. In addition to the above four datasets, we experiment with DBpedia ontology dataset (Zhang et al., 2015), Stanford sentiment treebank (Socher et al., 2013), Microsoft Research Paraphrase Corpus (Dolan and Brockett, 2005), and Quora Question Pairs from the GLUE benchmark. The results on these datasets are summarized in Appendix A.2.

Following previous practice (Alzantot et al., 2018), we fine-tune CLARE on training data, and evaluate with 1,000 randomly sampled test instances of lengths ≤100\leq 100. In the sentence-pair tasks (e.g., MNLI, QNLI), we attack the longer sentence excluding the tokens that appear in both.

Evaluation metrics.

We follow previous works (Jin et al., 2020; Morris et al., 2020a), and evaluate the models with the following automatic metrics:

Attack success rate (A-rate): the percentage of adversarial examples that can successfully attack the victim model.

Modification rate (Mod): the percentage of modified tokens. Each Replace or Insert action accounts for one token modified; a Merge action is considered modifying one token if one of the two merged tokens is kept (e.g., merging bigram abab into aa), and two otherwise (e.g., merging bigram abab into cc).

Perplexity (PPL): a metric used to evaluate the fluency of adversaries (Kann et al., 2018; Zang et al., 2020). The perplexity is calculated using small sized GPT-2 with a 50K-sized vocabulary (Radford et al., 2019).

Grammar error (GErr): the absolute number of increased grammatical errors in the successful adversarial example, compared to the original text. Following (Zang et al., 2020; Morris et al., 2020b), we calculate this by the LanguageTool (Naber et al., 2003).https://www.languagetool.org/

Textual similarity (Sim): the cosine similarity between the input and its adversary. Following (Jin et al., 2020; Morris et al., 2020b), we calculate this using the universal sentence encoder (USE; Cer et al., 2018).

The last four metrics are averaged across those adversarial examples that successfully attack the victim model.

3 Results

Table 2 summarizes the results. Overall CLARE achieves the best performance on all metrics consistently across different datasets. Notably, CLARE outperforms BERTAttack, the strongest baseline, by a more than 5.4% attack success rate with fewer average modifications to the text. We attribute this to CLARE’s flexible attack strategies obtained by combining three different perturbations at any position. Interestingly, using contextualized embeddings does not appear to guarantee better fluency: despite fewer modifications to the text, BERTAttack achieves similar perplexity to language-model-augmented TextFooler on three out of the four datasets, while CLARE consistently outperforms both. In terms of grammatical errors, contextualized models (CLARE and BERTAttack) are substantially better than the others, with CLARE performing the best. In terms of similarity, CLARE outperforms all baselines by more than 0.02, a larger gap than BERTAttack’s improvements over TextFooler variants. We observe similar trends on other datasets in Appendix A.2.

Figure 2 compares trade-off curves between attack success rate and textual similarity. We tune the thresholds for constructing the candidate token sets, and plot textual similarity against the attack success rate. CLARE strikes the best balance, showing a clear advantage in success rate with least similarity drop. We observe similar trends for attack success rate and perplexity trade off.

We further conduct human evaluation on the AG News dataset. We randomly sample 300 instances which both CLARE and TextFooler successfully attack. For each input, we pair the adversarial examples from the two models, and present them to crowd-sourced judges along with the original input and the gold label. We ask them which they prefer with a neutral option in terms of (1) having a meaning that is closer to the original input (similarity), and (2) being more fluent and grammatical (fluency and grammaticality). Additionally, we ask the judges to annotate adversarial examples, and compare their annotations against the gold labels (label consistency). We collect 5 responses for each pair on every evaluated aspect. Further details are in Appendix A.3.

As shown in Table 3, CLARE has a significant advantage over TextFooler: in terms of similarity 56% responses prefer CLARE, while 16% prefer TextFooler. The trend is similar for fluency & grammaticality (42% vs. 9%). This observation is consistent with results from automatic metrics. On label consistency, CLARE slightly underperforms TextFooler at 68% with a 95% condidence interval (CI) (66%,70%)(66\%,70\%), versus 70% with a 95% CI (68%,73%)(68\%,73\%). We attribute this to an inherent overlap of some categories in the AG News dataset, e.g., Science & Technology and Business, as evidenced by a 71% label consistency for original inputs.

Closing this section, Table 3.3 compares the adversarial examples generated by TextFooler and CLARE. More samples are listed in Appendix A.4.

We ablate each component of CLARE to study its effectiveness. We evaluate on the 1,000 randomly selected AG news instances (§3.2). The results are summarized in Table 5.

We first investigate the performance of three perturbations when applied individually. Among three editing strategies, using InsertOnly achieves the best performance, with ReplaceOnly coming a close second. MergeOnly underperforms the other two, partly because the attacks are restricted to bigram noun phrases (§3.1). Combining all three perturbations, CLARE achieves the best performance with the least modifications.

To examine the efficiency of attacking order, we compare ReplaceOnly against BERTAttack. Notably, ReplaceOnly outperforms BERTAttack across the board. This is presumably because BERTAttack does not take into account the tokens to be infilled when selecting the attack positions.

We now turn to the two constraints imposed when constructing the candidate token set. Perhaps not surprisingly, ablating the textual similarity constraint (w/o sim⁡>l\operatorname{sim}>l) decreases textual similarity performance, but increases other aspects. Ablating the masked language model yields a better success rate, but much worse perplexity, grammaticality, and textual similarity.

Finally, we compare CLARE implemented with different masked language models. Table 6 summarizes the results. Overall, distilled RoBERTa achieves the fastest speed without losing performance. Since the victim model is based on BERT, we conjecture that it is less efficient to attack a model using its own information.

2 Perturbations by Part-of-speech Tags

In this section, we break down the adversarial attacks by part-of-speech (POS) tags in AG News dataset. We find that most of the adversarial attacks happen to nouns or noun phrases. Presumably, in many topic classification datasets, the prediction heavily relies on some characteristic noun words/phrases. As shown in Table 7, 64% of the Replace actions are applied to nouns. Insert actions tend to insert tokens into noun phrase bigram: two of the most frequent POS bigrams are noun phrases. In fact, around 48% of the Insert actions are applied to noun phrases. This also justifies our choice of only applying Merge to noun phrases.

3 Adversarial Training

This section explores CLARE’s potential in improving downstream models’ accuracy and robustness. Following Tsipras et al. (2018), we use CLARE to generate adversarial examples for AG news training instances, and include them as additional training data. We consider two settings: training with (1) full training data and full adversarial data and (2) 10% randomly-sampled training data and its adversarial data, to simulate the low-resource scenario. For both settings, we compare a BERT-based MLP classifier and a TextCNN (Kim, 2014) classifier without any pretrained embedding.

Whether adversarial examples, as data augmentation, can help achieve better test accuracy? As shown in Table 8, when the full training data is available, adversarial training slightly decreases the test accuracy by 0.2% and 0.5% respectively. This aligns with previous observations (Jia et al., 2019). Interestingly, under the low-data scenario with adversarial training, BERT-based classifier has no accuracy drop, and TextCNN achieves a 2.0% absolute improvement. This suggests that a model with less capacity can benefit more from silver data.

Does adversarial training help the models defend against adversarial attacks? To evaluate this, we use CLARE to attack the classifiers trained with and without adversarial examples. In preliminary experiments, we found that it is more difficult to use other models to attack a victim model trained with the adversarial examples generated by CLARE, than to use CLARE itself. A higher success rate and fewer modifications indicate a victim classifier is more vulnerable to adversarial attacks. As shown in Table 8, in 3 out of the 4 cases, adversarial training helps to decrease the attack success rate by more than 10.3%, and to increase the number of modifications needed by more than 0.8. The only exception is the TextCNN model trained with 10% data. A possible reason can be that it is trained with few data and thus generalizes less well.

These results suggest that CLARE can be used to improve downstream models’ robustness, with a negligible accuracy drop.

An increasing amount of effort is being devoted to generating better textual adversarial examples with various attack models. Character-based models (Liang et al., 2019; Ebrahimi et al., 2018; Li et al., 2018; Gao et al., 2018, inter alia) use misspellings to attack the victim systems; however, these attacks can often be defended by a spell checker (Pruthi et al., 2019; Zhou et al., 2019b; Jones et al., 2020). Many sentence-level models (Iyyer et al., 2018; Wang et al., 2020; Zou et al., 2020, inter alia) have been developed to introduce more sophisticated token/phrase perturbations. These, however, generally have difficulty maintaining semantic similarity with original inputs (Zhang et al., 2020a). Recent word-level models explore synonym substitution rules to enhance semantic meaning preservation (Alzantot et al., 2018; Jin et al., 2020; Ren et al., 2019; Zhang et al., 2019; Zang et al., 2020, inter alia). Our work differs in that CLARE uses three contextualized perturbations that produces more fluent and grammatical outputs.

Text generation with BERT.

Generation with masked language models has been widely studied in various natural language tasks, ranging from lexical substitution (Wu et al., 2019a; Zhou et al., 2019a; Qiang et al., 2020; Wu et al., 2019b, inter alia) to non-autoregressive generation (Gu et al., 2018; Lee et al., 2018; Ghazvininejad et al., 2019; Wang and Cho, 2019; Ma et al., 2019; Sun et al., 2019; Ren et al., 2020; Zhang et al., 2020b, inter alia).

We have presented CLARE, a contextualized adversarial example generation model for text. It uses contextualized knowledge from pretrained masked language models, and can generate adversarial examples that are natural, fluent and grammatical. With three contextualized perturbation patterns, Replace, Insert and Merge in our arsenal, CLARE can produce outputs of varied lengths and achieves a higher attack success rate than baselines and with fewer edits. Human evaluation shows significant advantages of CLARE in terms of textual similarity, fluency and grammaticality. We release our code and models at https://github.com/cookielee77/CLARE.

We would like to thank the reviewers for their constructive comments. We thank NVIDIA Corporation for the donation of the GPU used for this research. We also thank Tongshuang Wu, Guoyin Wang and Shuhuai Ren for their helpful discussions and feedback.

Appendix A Appendix

All pretrained models and victim models based on RoBERTa and BERTbase{}_{\text{base}} are implemented with Hugging Face transformershttps://github.com/huggingface/transformers (Wolf et al., 2019) based on PyTorch (Paszke et al., 2019). RoBERTadistill{}_{\text{distill}}, RoBERTabase{}_{\text{base}} and uncase BERTbase{}_{\text{base}} models have 82M, 125M and 110M parameters, respectively. We use RoBERTadistill\text{RoBERTa}_{\text{distill}} as our main backbone for fast inference purpose. TextFoolerhttps://github.com/jind11/TextFooler and BERTAttackhttps://github.com/LinyangLee/BERT-Attack are built with their open source implementation provided by the authors. In the implementation of TextFooler+LM, we use small sized GPT-2 language model (Radford et al., 2019) to further select those candidate tokens that have top 20%20\% perplexity in the candidate token set. In the adversarial training (§4.3), the small TextCNN victim model (Kim, 2014) has 128 embedding size and 100100 filters for 3,4,53,4,5 window size with 0.50.5 dropout, resulting in 7M parameters.

During the implementation of w/o pMLM>kp_{\text{MLM}}>k in the ablation study (§4.1), we randomly sample 200 tokens and then apply the similarity constraint to construct candidate set, as exhausting the vocabulary is computationally expensive.

Evaluation Metric.

The similarity function sim⁡\operatorname{sim} builds on the universal sentence encoder (USE; Cer et al., 2018) to measure a local similarity at the perturbation position with window size 15 between the original input and its adversary. All baselines are equipped this sim⁡\operatorname{sim} when constructing the candidate vocabulary. The evaluation metric Sim uses USE to calculate a global similarity between two texts. These procedures are typically following Jin et al. (2020). We mostly rely on human evaluation (§3.3) to conclude the significant advantage of preserving textual similarity on CLARE compared with TextFooler.

Data Processing.

When processing the data, we keep all punctuation in texts for both victim model training and attacking. This differs the pre-processing setting in TextFooler (Jin et al., 2020) as we empirically found that removing punctuation makes the victim model vulnerable. Since GLUE benchmark (Wang et al., 2019a) does not provide the label for test set, we instead use its dev set as the the test set for the included datasets (MNLI, QNLI, QQP, MRPC, SST-2) in the evaluation. For the sentence-pair tasks (e.g., MNLI, QNLI, QQP, MRPC), we attack the longer one excluding the tokens appearing in both sentences. This is because inference tasks usually require entailed data to have the same keywords, e.g., numbers, name entities, etc. All experiments are conducted on one Nvidia GTX 1080Ti GPU.

A.2 Additional Results

We include the results of DBpedia ontology dataset (DBpedia; Zhang et al., 2015, Stanford sentiment treebank (SST-2; Socher et al., 2013), Microsoft Research Paraphrase Corpus (MRPC; Dolan and Brockett, 2005), and Quora Question Pairs (QQP) from the GLUE benchmark in this section. Table 9 summarizes come statistics of these datasets. The results of different models on these datasets are summarized Table 10. Compared with all baselines, CLARE achieves the best performance on attack success rate, perplexity, grammaticality, and similarity. It is consistent with our observation in §3.3.

A.3 Human Evaluation Details

For each human evaluation on AG News dataset, we randomly sampled 300 sentences from the test set combining the corresponding adversarial examples from CLARE and TextFooler (We only consider sentences can be attacked by both models). In order to make the task less abstract, we pair the adversarial examples by the two models, and present them to the participants along with the original input and its gold label. We ask them which one they prefer in terms of (1) having more similar a meaning to the original input (similarity), and (2) being more fluent and grammatical (fluency and grammaticality). We also provide them with a neutral option, when the participants consider the two indistinguishable. Additionally, we ask the participants to annotate the adversarial examples, and compare their annotations against the gold labels (label consistency). Higher label consistency indicates the model is better at causing the victim model to make errors while preserving human predictions.

Each pair of system outputs was randomly presented to 5 crowd-sourced judges, who indicated their preference for similarity, fluency, and grammaticality using the form shown in Figure 3. The labelling task is illustrated in Figure 4. To minimize the impact of spamming, we employed the top-ranked 30% of U.S. workers provided by the crowd-sourcing service. Detailed task descriptions and examples were also provided to guide the judges. We calculate pp-value based on 95% confidence intervals by using 10K paired bootstrap replications, implemented using the R Boot statistical package.

A.4 Qualitative Samples

We include generated adversarial examples by CLARE and TextFooler on AG News, DBpeida, Yelp, MNLI, and QNLI datasets in Table A.4 and Table A.4.