Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT

Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, Caiming Xiong

Introduction

Most neural models, such as Recurrent Neural Network Bahdanau et al. (2015), Attentive Convolution Yin and Schütze (2018), BERT Devlin et al. (2018), etc., are evaluated on clean datasets. When deploying these models in real-world scenarios, the models have to address user-generated noisy text. One of the most common noisy text is typos because of human mistakes in typing words, such as character substitution, additional or missing characters. Even a small typo in Fig. 1 may confuse the most advanced models like BERT Devlin et al. (2018), then a question arises: “how robust is BERT with respect to keyboard typos?”

Existing approaches generate malicious adversarial examples as attacks by finding the minimum perturbation on each sample to study the robustness of a model Liang et al. (2017); Ebrahimi et al. (2017); Gao et al. (2018); Li et al. (2018); Alzantot et al. (2018); Ren et al. (2019). These attacking examples are not that meaningful, therefore can be easily recognized by humans. In order to prevent humans from identifying the adversarial examples, the community starts to explore more natural perturbations to create the attacks. Zhao et al. (2017) generates more natural adversarial examples using GANs. Both Ribeiro et al. (2018) and Sato et al. (2018) generate samples with similar semantics. Belinkov and Bisk (2017) presents the first work that studies adversarial typos on the keyboard.

The character distribution on the keyboard results in a special type of noisy examples, which are natural and unintentional. If conventional attacking examples represent the worse case of inputs, the keyboard-constrained adversarial examples represent the more realistic inputs. In this work, we systematically study the robustness of the current state-of-the-art neural model BERT in dealing with those inadvertently generated adversarial inputs. Intensive experiments on sentiment analysis and question answering benchmarks indicate that:

• BERT has unbalanced attention to the typos in the input. Some typo words have a clear influence on the performance; some, instead, show tiny influence. In addition, different typo generation approaches also show different degrees of damages. Mistype is the severest source of typos;

• Machines and humans have different focuses on the typos. BERT pays more attention to the typos in informative words; in contrast, humans can better recognize the typos in some frequent while less informative words;

• The robustness of a system is dependent on the learning algorithms as well as the tasks. We found that the BERT-based question answering system on SQUAD Rajpurkar et al. (2016) is more brittle than the BERT-based sentiment recognizer.

This is the first work that systematically studies the robustness of BERT, a Transformer-style Vaswani et al. (2017) neural model, in addressing noises that appear naturally. Our observations hopefully can provide new insights to the community for building more trustworthy machine intelligence.

Generation of natural adversarial examples

Our adversarial examples are generated with the following principle. We have a pre-trained state-of-the-art model f:X→Yf:X\rightarrow Y for natural language processing tasks, where XX are the text inputs and YY are the corresponding labels. An adversarial keyboard typo example, denoted as xadvx_{adv}, is generated from the original input x∈Xx\in X and the original prediction y∈Yy\in Y based on gradient information or random modification. For each generated example, it will change the model prediction: f(xadv)=yadv≠yf(x_{adv})=y_{adv}\neq y.

To start, a sentence ss is first tokenized into words (or subwords) s=(w1,w2,…,wN)s=(w_{1},w_{2},\ldots,w_{N}) (line 1 in the Algorithm 1), where wiw_{i} is the ii-th item, and NN is the length after tokenization. Let L(w,y)\mathcal{L}(w,y) denote the loss with respect to ww under the ground truth label yy. Then, we can compute the partial derivative of each item wiw_{i} based on the golden output yy as shown in this Equation:

Based on the gradient information, we can track back to the most informative/uninformative word xx though the component ww of the word (line 1). We are interested in generating typos regarding three kinds of words: (1) informative words which have the largest gradient; (2) uninformative words which have the smallest gradient; and (3) random words.

For each word, we consider the following six sources of typos: (1) Insertion: Insert characters or spaces into the word, such as “oh!” →\rightarrow “ooh!”. (2) Deletion: Delete a random character of the word because of fast typing, such as “school” →\rightarrow “schol”. (3) Swap: Swap random two adjacent characters in a word. (4) Mistype: Mistyping a word though keyboard, such as “oh” →\rightarrow “0h”. (5) Pronounce: Wrongly typing due to the close pronounce of the word, such as “egg” →\rightarrow “agg”. (6) Replace-W: Replace the word by the frequent human behavioral keyboard typo based on the statistics.https://en.wikipedia.org/wiki/Wikipedia:Lists_of_common_misspellings

Note that we do not sample characters from random distribution to implement above modifications, all operations are constrained by the character distribution on the keyboard. The whole generation algorithm is demonstrated in the Algorithm 1. There could be more than one typo in each piece of text through keyboard in real life; this work limits the maximal number of typos in one sentence to be KK.

Experiments

We evaluate the start-of-the-art model BERTWe use “bert-base-uncased” throughout. on two NLP tasks: sentiment analysis and question answering. On each benchmark, we report the system performance when putting typos in the informative words (i.e., maximal gradient), uninformative words (i.e., minimal gradient) and random words.

The built-in tokenizer in BERT first performs simple white-space tokenization, then applies WordPiece tokenization Wu et al. (2016). A word can be split into character ngrams (e.g. adversarial→\rightarrow[ad, ##vers, ##aria, ##l], robustness→\rightarrow[robust, ##ness]). “##” is a special symbol to handle the subwords, and we omit it when injecting our typos. An example typo for “robustness” is “robustnesss”, which is split into subwords [robust, ##ness, ##s].

1 Sentiment analysis

We work on the Stanford Sentiment Treebank (SST) Socher et al. (2013) in the binary prediction setting. The standard split has 6920 training, 872 development and 1821 test sentences. Socher et al. (2013) used the Stanford Parser Klein and Manning (2003) to parse each sentence into subphases. The subphases were then labeled by human annotators in the same way as the sentences were labeled. Labeled phrases that occur as subparts of the training sentences are treated as independent training instances

Table 1 lists the results when we try K={0,1,⋯ ,10}K=\{0,1,\cdots,10\} typos in informative words (i.e., “max-grad”), uninformative words (i.e., “min-grad”) and random words. The first column (“max-grad”) shows that the model is highly sensitive to the typos on words with the largest gradient norms. Injecting only a single typo degrades the accuracy by 22.6%, suggesting that the gradient norm is a strong indicator to find task-specific informative words. Another interesting observation is that the accuracy converges to almost the chance-level accuracy (i.e., around 50%). In contrast, the second column (“min-grad”) indicates that the model is not sensitive to the typos in words with the smallest gradient norms. Even with K=10K=10, the accuracy drops merely by 9%, which indicates that in some lucky cases human keyboard typos may not influence the final predictions.

This comparison discovers the unbalanced attention of BERT to the typos in the input. Only the adversarial attacks based on informative parts really matter. However, after K=7K=7, any typos in the words with the largest gradient norms can’t make the model predictions worse any more, because the current most informative words already contain typos, and then it would not change any words in the sentence. In this case, if an adversary wants to attack BERT intentionally, the best strategy is adaptively mixing up “max-grad” and “random” policy for adversarial sample generation.

Next, we further explore the fine-grained influence of the six kinds of word modifications (i.e., “insertion”, “deletion”, “swap”, “mistype”, “pronounce” and “replace-w”) in Fig. 2. Insertion modification has the minimum influence, because sub-words tokenization of BERT would not change much in some cases, such as “apple” →\rightarrow “applee”. Instead, mistype because of fast typing on the keyboard hurts the performance most. The main reason is mistype can generate some uncommon samples, such as “own” →\rightarrow “0wn” or “9wn”.

Comparison between human and machine.

To investigate how well humans can read our typo-injected text, we conducted human evaluation. As studied in McCusker et al. (1981); Michel et al. (2019), some typos are not easily recognized by humans. We invited 10 persons to read and detect typos in 100 examples used in Table 1, then we count the number of sentences where any typos are detected. The detection would be easy if they spend a long time, so we allocated only three seconds for each sentence. Table 2 shows the results with K=1∼5K=1\sim 5, and the scores are the averaged detected counts. Interestingly, the model is sensitive to the “max-grad” setting, humans, instead, are less sensitive—the “min-grad” word typos are more eye-catching for humans. One presumable reason is that the “min-grad” setting often injects the typos into some highly frequent words, such as “the” and “it”. People are over familiar with them and their high frequencies in the text increase the chance of being recognized by humans. Here is an example:

“ut’s a charming and often affecting journey” where “it” is modified to “ut”. This looks strange to humans, but the model prediction does not change.

Sensitivity of word segmentation.

We have observed that humans and the model are sensitive to the different typos (or words). One explanation of the BERT’s sensitivity is that the subword segmentation is sensitive to the typos. For example, our “max-grad” method modifies the keyword “inspire” to “inspird” in the following sentence:

“a subject like this should inspire reaction in its audience; the pianist does not.”

As a result, “inspird” is split into [ins, ##pi, ##rd], which is completely different from the original string. Therefore, one promising direction is to make word segmentation more robust to character-level modifications. To verify this assumption, we trained two other RNN-based classifiers with the widely-used GloVe embeddings (Pennington et al., 2014) and character n-gram embeddings (Hashimoto et al., 2017). Table 3 shows the results with “max-grad” setting. Our BERT-based typos also degrade the scores of the RNN-based models, and we can see that the character information makes the model more robust.

2 Question answering

We work on the SQuAD v1.1 benchmark Rajpurkar et al. (2016). The dataset is randomly partitioned into a training set (80%), a development set (10%), and a blinded test set (10%). Evaluation on the SQuAD dataset consists of two metrics: the exact match score (EM) and F1 score. For this task, we inject the typos into the questions.

Table 4 also reports the results when we cast typos in informative, uninformative or random words. It is surprising that even a single typo is able to decrease the QA model performance dramatically (therefore, we did not increase the typo size KK anymore). Comparing this QA performance and that in sentiment analysis task in Table 1, we notice that the BERT-based QA system is much more brittle than a BERT-based sentiment classifier. It means that the robustness of a NLP system depends on the learning algorithm as well as the task.

Conclusion

This paper has investigated how the state-of-the-art model, BERT, is robust or brittle to keyboard typos. Our experimental results show the different sensitivities to different types of words, and suggest the necessity of considering the robustness of the neural models. We will release our code to reproduce our results, and in future work, we will consider how to make subword-based models more robust to human typos in NLP tasks.

References