Measure and Improve Robustness in NLP Models: A Survey

Xuezhi Wang, Haohan Wang, Diyi Yang

Introduction

NLP models, especially with the recent advances of large pre-trained language models have achieved great progress and gained wide applications in the real world. Despite the performance gains, NLP models are still fragile and brittle to out-of-domain data Hendrycks et al. (2020a); Wang et al. (2019d), adversarial attacks McCoy et al. (2019); Jia and Liang (2017); Jin et al. (2020), or small perturbation to the input Ebrahimi et al. (2018); Belinkov and Bisk (2018). Those failures could hinder the safe deployment of these models in the real world, and impact NLP models’ trustworthiness to users. As a result, an increasing line of work has been conducted to understand robustness issues in the language technologies communities. Still, diverse sets of research across multiple dimensions and numerous levels of depth exist and are scattered across various communities; for instance, using a variety of definitions on a wide range of very different NLP tasks. In this work, we provide a unifying overview of what is robustness in NLP, how to identify robustness failures and evaluate model’s robustness, and systematic ways to improve robustness, as well as a conceptual schema categorizing ongoing research directions. We identify gaps between the to-date robustness work, the technical opportunities, and discuss possible paths forward.

Definitions of Robustness in NLP

The above definition works for a range of NLP tasks like text classification and sequence labeling where yy is defined over a fixed set of discrete labels. For tasks like text generation, robustness is less well defined and can manifest as positional bias Jung et al. (2019); Kryscinski et al. (2019), or hallucination Maynez et al. (2020); Parikh et al. (2020); Zhou et al. (2021). One major challenge here is a lack of robust metrics in evaluating the quality of the generated text Sellam et al. (2020); Zhang et al. (2020b), i.e., we need a reliable metric to determine the relationship between f(x′)f(x^{\prime}) and y′y^{\prime} when both are open-ended texts.

In one line of research, D′\mathcal{D}^{\prime} is constructed by perturbations around input xx to form x′x^{\prime} (x′x^{\prime} typically being defined within some proximity of xx). This topic has been widely explored in computer vision under the concept of adversarial robustness, which measures models’ performances against carefully crafted noises generated deliberately to deceive the model to predict wrongly, pioneered by Szegedy et al. (2013); Goodfellow et al. (2015), and later extended to NLP, such as Ebrahimi et al. (2018); Alzantot et al. (2018); Li et al. (2019); Feng et al. (2018); Kuleshov et al. (2018); Jia et al. (2019); Zang et al. (2020); Pruthi et al. (2019); Wang et al. (2019e); Garg and Ramakrishnan (2020); Tan et al. (2020a, b); Schwinn et al. (2021); Li et al. (2021); Boucher et al. (2022) and multilingual adversaries Yang et al. (2019); Tan and Joty (2021). The generation of adversarial examples primarily builds upon the observation that we can generate samples that are meaningful to humans (e.g., by perturbing the samples with changes that are imperceptible to humans) while altering the prediction of the models for this sample. In this regard, human’s remarkable ability in understanding a large set of synonyms Li et al. (2020) or interesting characteristics in ignoring the exact order of letters Wang et al. (2020b) are often opportunities to create adversarial examples. A related line of work such as data-poisoning Wallace et al. (2021) and weight-poisoning Kurita et al. (2020) exposes NLP models’ vulnerability against attacks during the training process. One can refer to more comprehensive reviews and broader discussions on this topic in Zhang et al. (2020c) and Morris et al. (2020b).

Assumptions around Label-preserving and Semantic-preserving Most existing work in vision makes a relatively simplified assumption that the gold label of x′x^{\prime} remains unchanged under a bounded perturbation over xx, i.e., y′=yy^{\prime}=y, and a model’s robust behaviour should be f(x′)=yf(x^{\prime})=y Szegedy et al. (2013); Goodfellow et al. (2015). A similar line of work in NLP follows the same label-preserving assumption with small text perturbations like token and character swapping Alzantot et al. (2018); Jin et al. (2020); Ren et al. (2019); Ebrahimi et al. (2018), paraphrasing Iyyer et al. (2018); Gan and Ng (2019), semantically equivalent adversarial rules Ribeiro et al. (2018), and adding distractors Jia and Liang (2017). However, this label-preserving assumption might not always hold, e.g., Wang et al. (2021b) studied several existing text perturbation techniques and found that a significant portion of perturbed examples are not label-preserving (despite their label-preserving assumptions), or the resulting labels have a high disagreement among human raters (i.e., can even fool humans). Morris et al. (2020a) also call for more attention to the validity of perturbed examples for a more accurate robustness evaluation.

Another line of work aims to perturb the input xx to x′x^{\prime} in small but meaningful ways that explicitly change the gold label, i.e., y′≠yy^{\prime}\neq y, under which case the robust behaviour of a model should be f(x′)=y′f(x^{\prime})=y^{\prime} and f(x′)≠yf(x^{\prime})\neq y Gardner et al. (2020); Kaushik et al. (2019); Schlegel et al. (2021). We believe these two lines of work are complementary to each other, and both should be explored in future research to measure models’ robustness more comprehensively.

One alternative notion is whether the perturbation from xx to x′x^{\prime} is “semantic-preseving” Alzantot et al. (2018); Jin et al. (2020); Ren et al. (2019) or “semantic-modifying” Shi and Huang (2020); Jia and Liang (2017). Note this is slightly different from the above label-preserving assumptions, as it is defined over the perturbations on (x,x′)(x,x^{\prime}) rather than making an assumption on (y,y′)(y,y^{\prime}), e.g., semantic-modifying perturbations can be either label-preserving Jia and Liang (2017); Shi and Huang (2020) or label-changing Gardner et al. (2020); Kaushik et al. (2019).

2 Robustness under Distribution Shift

Another line of research focuses on (x′,y′)(x^{\prime},y^{\prime}) drawn from a different distribution that is naturally-occurring Hendrycks et al. (2021), where robustness can be defined around model’s performance under distribution shift. Different from work on domain adaptation Patel et al. (2015); Wilson and Cook (2020) and transfer learning Pan and Yang (2010), existing definitions of robustness are closer to the concept of domain generalization Muandet et al. (2013); Gulrajani and Lopez-Paz (2021), or out-of-distribution generalization to unforeseen distribution shifts Hendrycks et al. (2020a), where the test data (either labeled or unlabeled) is assumed not available during training, i.e., generalization without adaptation. In the context of NLP, robustness to natural distribution shifts can also mean models’ performance should not degrade due to the differences in grammar errors, dialects, speakers, languages Craig and Washington (2002); Blodgett et al. (2016); Demszky et al. (2021), or newly collected datasets for the same task but in different domains Miller et al. (2020). Another closely connected line of research is fairness, which has been studied in various NLP applications, see Sun et al. (2019) for a more in-depth survey in this area. For example, gendered stereotypes or biases have been observed in NLP tasks including co-reference resolution Zhao et al. (2018a); Rudinger et al. (2017), occupation classification De-Arteaga et al. (2019), and neural machine translation Prates et al. (2019); Font and Costa-jussà (2019).

3 Connections and A Common Theme

The above two categories of robustness can be unified under the same framework, i.e., whether D′\mathcal{D}^{\prime} represents a synthetic distribution shift (via adversarial attacks) or a natural distribution shift. Existing work has shown a model’s performance might degrade substantially in both cases, but the transferability of the two categories is relatively under-explored. In the vision domain, Taori et al. (2020) investigate models’ robustness to natural distribution shift, and show that robustness to synthetic distribution shift might offer little to no robustness improvement under natural distribution shift. Some studies show NLP models might not generalize to unseen adversarial patterns Huang et al. (2020); Jha et al. (2020); Joshi and He (2021), but more work is needed to systematically bridge the gap between NLP models’ robustness under natural and synthetic distribution shifts.

To better understand why models exhibit a lack of robustness, some existing work attributed this to the fact that models sometimes utilize spurious correlations between input features and labels, rather than the genuine ones, where spurious features are commonly defined as features that do not causally affect a task’s label Srivastava et al. (2020); Wang and Culotta (2020b): they correlate with task labels but fail to transfer to more challenging test conditions or out-of-distribution data Geirhos et al. (2020). Some other work defined it as “prediction rules that work for the majority examples but do not hold in general” Tu et al. (2020). Such spurious correlations are sometimes referred as dataset bias Clark et al. (2019); He et al. (2019), annotation artifacts Gururangan et al. (2018), or group shift Oren et al. (2019) in the literature. Further, evidence showed that controlling model’s learning in spurious features will improve model’s performances in distribution shifts Wang et al. (2019a, b); also, discussions on the connections between adversarial robustness and learning of spurious features have been raised Ilyas et al. (2019); Wang et al. (2020a). Theoretical discussions connecting these fields have also been offered by crediting a reason of model’s lack of robustness in either distribution shift or adversarial attack to model’s learning of spurious features Wang et al. (2021c).

Further, in certain applications, model “robustness” can also be connected with models’ instability Milani Fard et al. (2016), or models having poorly-calibrated uncertainty estimation Guo et al. (2017), where Bayesian methods Graves (2011); Blundell et al. (2015), dropout-based Gal and Ghahramani (2016); Kingma et al. (2015) and ensemble-based approaches Lakshminarayanan et al. (2017) have been proposed to improve models’ uncertainty estimation. Recently, Ovadia et al. (2019) have shown models’ uncertainty estimation can degrade significantly under distributional shift, and call for more work to ensure a model “knows when it doesn’t know” by giving lower uncertainty estimates over out-of-distribution data. This is another example where models can be less robust under distributional shifts, and again emphasizes the need of building more unified benchmarks to measure a model’s performance (e.g., robust accuracy, calibration, stability) under distribution shifts, in addition to in-distribution accuracy.

Robustness in Vision vs. in NLP

Despite the widely study of robustness in vision, the study of robustness in NLP cannot always directly borrow the ideas. We categorize the main differences with the three following points:

The most obvious characteristic is probably the discrete nature of the space of text. This particularly posed a challenge towards the adversarial attack and defense regime when the study in vision is transferred to NLP Lei et al. (2019); Zhang et al. (2020c), in the sense that simple gradient-based adversarial attacks will not directly translate to meaningful attacks in the discrete text space, and multiple novel attack methods are proposed to fill the gap, as we will discuss in later sections.

Perceptible to Human vs. Not

On a related topic, one of the most impressive property of adversarial attack in vision is that small perturbation of the image data imperceptible to human are sufficient to deceive the model Szegedy et al. (2013), while this can hardly be true for NLP attacks. Instead of being imperceptible, the adversarial attacks in NLP typically are bounded by the fact that the meaning of the sentences are not altered (despite being perceptible). On the other hand, there are ways to generate samples where the changes, although being perceptible, are often ignored by human brain due to some psychological prior on how a human processes the text Anastasopoulos et al. (2019); Wang et al. (2020b).

Support vs. Density Difference of the Data Distributions

Another difference is more likely seen in the discussion of the domain adaptation of vision and NLP study. In vision study, although the images from training distribution and test distribution can be sufficiently different, the train and test distributions mostly share the same support (the pixels are always sampled from a 0-255 integer space), although the density of these distributions can be very different (e.g., photos vs. sketches). On the other hand, domain adaptation of NLP sometimes studies the regime where the supports of the data differ, e.g., the vocabularies can be significantly different in cross-lingual studies Abad et al. (2020); Zhang et al. (2020a).

A Common Theme

Despite the disparities between vision and NLP, the common theme of pushing the model to generalize from D\mathcal{D} to D′\mathcal{D}^{\prime} preserves. The practical difference between D\mathcal{D} and D′\mathcal{D}^{\prime} is more than often defined by the human’s understanding of the data, and can differ in vision and NLP as humans perceive and process images and texts in subtly different ways, which creates both opportunities for learning and barriers for direct transfer. Certain lines of research try to bridge the learning in the vision domain to the embedding space in the NLP domain, while other lines of research create more interpretable attacks in the discrete text space (see Table 1 for these two lines of work). How those two lines of research transfer to each other, or complement each other, is not fully explored and calls for additional research.

Identify Robustness Failures

As robustness gained increasing attention in NLP literature, various lines of work have proposed ways to identify robustness failures in NLP models. Existing works can be roughly categorized by how the failures are identified, among which a large portion of work relies on human priors and error analyses over existing NLP models (Section 4.1), and other lines of work adopt model-based approaches (Section 4.2). The identified robustness failure patterns are usually organized into challenging/adversarial benchmark datasets to more accurately measure an NLP model’s robustness. In Table 1, we organize commonly used perturbation types for identifying models’ robustness failures, and in Table 2 we summarize common robustness benchmarks for each NLP task.

An increasing body of work has been conducted on understanding and measuring robustness in NLP models Tu et al. (2020); Sagawa et al. (2020b); Geirhos et al. (2020) across various NLP tasks, largely relying on human priors and error analyses.

Naik et al. (2018) sampled misclassified examples and analyzed their potential sources of errors, which are then grouped into a typology of common reasons for error. Such error types then served as the bases to construct the stress test set, to further evaluate whether NLI models have the ability to make real inferential decisions, or simply rely on sophisticated pattern matching. Gururangan et al. (2018) found that current NLI models are likely to identify the label by relying only on the hypothesis, and Poliak et al. (2018) provided similar augments that using a hypothesis-only model can outperform a set of strong baselines. Kaushik et al. (2019) asked humans to generate counterfactual NLI examples, to better understand what features are causal and encourage models to learn those features.

Question Answering

Jia and Liang (2017) proposed to generate adversarial QA examples by concatenating an adversarial distracting sentence at the end of a paragraph. Miller et al. (2020) built four new test sets for the Stanford Question Answering Dataset (SQuAD) and found most question-answering systems fail to generalize to this new data, calling for new evaluation metrics towards natural distribution shifts.

Machine Translation

Belinkov and Bisk (2018) found that character-based neural machine translation (NMT) models are brittle under noisy data, where noises (e.g., typos, misspellings, etc) are synthetically generated using possible lexical replacements. Data augmentation with artificially-introduced grammatical errors Anastasopoulos et al. (2019) or with random synthetic noises Vaibhav et al. (2019); Karpukhin et al. (2019) can make the system more robust to such spurious patterns. On the other hand, Wang et al. (2020b) showed another approach by limiting the input space of the characters so that the models will be likely to perceive data typos and misspellings.

Syntactic and Semantic Parsing

Robust parsing has been studied in several existing works Lee et al. (1995); Aït-Mokhtar et al. (2002). More recent work showed that neural semantic parsers are still not robust against lexical and stylistic variations, or meaning-preserving perturbations Marzinotto et al. (2019); Huang et al. (2021), and proposed ways to improve their robustness through data augmentation Huang et al. (2021) and adversarial learning Marzinotto et al. (2019).

Text Generation

Existing work found that text generation models also suffer from robustness issues, e.g., text summarization models suffer from positional bias Jung et al. (2019), layout bias Kryscinski et al. (2019), and a lack of faithfulness and factuality Kryscinski et al. (2019); Maynez et al. (2020); Chen et al. (2021b); data-to-text models sometimes hallucinate texts that are not supported by the data Parikh et al. (2020); Wang et al. (2020d). In addition, Sellam et al. (2020); Zhang et al. (2020b) pointed out the deficiency of existing automatic evaluation metrics and proposed new metrics to better align the generation quality with human judgements.

Connection with Dataset Biases

The robustness failures can sometimes be attributed to dataset biases, i.e., biases introduced during dataset collection Fouhey et al. (2018) or human annotation artifacts Gururangan et al. (2018); Geva et al. (2019); Rudinger et al. (2017), which could affect how well a model trained from this dataset generalizes, and how accurately we estimate a model’s performance. For example, Lewis et al. (2021) show there is a significant test-train data overlap in a set of open-domain question-answering benchmarks, and many QA models perform substantially worse on questions that cannot be memorized from training data. In natural language inference, McCoy et al. (2019) show that commonly used crowdsourced datasets for training NLI models might make certain syntactic heuristics more easily adopted by statistical learners. Further, Bras et al. (2020) propose to use a lightweight adversarial filtering approach to filter dataset biases, which is approximated using each instance’s predictability score.

2 Model-based Identification

In addition to the human-prior and error-analysis driven approaches which are usually specific to each task, other lines of work identify robustness failures that are task-agnostic like white-box text attack methods Ebrahimi et al. (2018); Alzantot et al. (2018); Jin et al. (2020), and even input-agnostic like universal adversarial triggers Wallace et al. (2019a) and natural attack triggers Song et al. (2021).

Another line of work proposes to learn an additional model to capture biases, e.g., in visual question answering, Clark et al. (2019) train a naive model to predict prototypical answers based on the question only irrespective of the context; He et al. (2019); Utama et al. (2020a) propose to learn a biased model that only uses dataset-bias related features. This framework has also been used to capture unknown biases assuming that the lower capacity model learns to capture relatively shallow correlations during training Clark et al. (2020). In addition, Wang and Culotta (2020a) identify model shortcuts by training classifiers to better distinguish “spurious” correlations from “genuine” ones based on human annotated examples.

Some work adopts human-in-the-loop to generate challenging examples, e.g., Counterfacutal-NLI Kaushik et al. (2019) and Natural-Perturbed-QA Khashabi et al. (2020). Other work applies model-in-the-loop to increase the likelihood that the perturbed examples are challenging for state-of-the-art models, but it might also introduce biases towards the particular model used. For example, SWAG Zellers et al. (2018) was introduced that fooled most models at the time of publishing but was soon “solved” after BERT Devlin et al. (2019) was introduced. As a result, Yuan et al. (2021) present a study over the transferability of adversarial examples, and Contrast Sets Gardner et al. (2020) intentionally avoid using model-in-the-loop. Further, more recent work adopts adversarial human-and-model-in-the-loop to create more difficult examples for benchmarking, e.g., Adv-QA Bartolo et al. (2020), Adv-Quizbowl Wallace et al. (2019b), ANLI Nie et al. (2020), and Dynabench Kiela et al. (2021).

Improve Model Robustness

Correspondingly, there are multiple lines of directions that try to improve robustness in NLP models. Depending on where and how the intervention is applied, those approaches can be categorized into the following categories: data-driven (Section 5.1), model-based and training-scheme-based (Section 5.2), inductive-prior-based (Section 5.3) and finally causal intervention (Section 5.4).

Data augmentation recently gained a lot of interest, in improving performance in low-resourced language settings, few-shot learning, mitigating biases, and improving robustness in NLP models Feng et al. (2021); Dhole et al. (2021). Techniques like Mixup Zhang et al. (2018), MixText Chen et al. (2020), CutOut DeVries and Taylor (2017), AugMix Hendrycks et al. (2020b), HiddenCut Chen et al. (2021a), have been shown to substantially improve the robustness and the generalization of models. Such mitigation strategies are operated at the data level, and often hard to be interpreted in terms of how and why mitigation works.

Other lines of work deal with spans or regions associated within data points to prevent models from heavily relying on spurious patterns. To make NLP models more robust on sentiment analysis and NLI tasks, Kaushik et al. (2019) proposed curating counterfactually augmented data via a human-in-the-loop process, and showed that models trained on the combination of this augmented data and original data are less sensitive to spurious patterns. Differently, Wang et al. (2021d) performed strategic data augmentation to perturb the set of “shortcuts” that are automatically identified, and found that mitigating these leads to more robust models in multiple NLP tasks. This line of mitigation strategies closely relates to how spurious correlations can be measured and identified, as many of the challenging or adversarial examples (Table 1) can sometimes be used to augment the original model to improve its robustness, either in the discrete input space as additional training examples Liu et al. (2019); Kaushik et al. (2019); Anastasopoulos et al. (2019); Vaibhav et al. (2019); Khashabi et al. (2020), or in the embedding space Zhu et al. (2020); Zhao et al. (2018b); Miyato et al. (2017); Liu et al. (2020).

2 Model and Training-based Approaches

Recent work has demonstrated pre-training as an effective way to improve NLP models’ out-of-distribution robustness Hendrycks et al. (2020a); Tu et al. (2020), potentially due to its self-supervised objective and the use of large amounts of diverse pre-training data that encourages generalization from a small number of examples that counter the spurious correlations. Tu et al. (2020) showed a few other factors can also contribute to robust accuracy, including larger model size, more fine-tuning data, and longer fine-tuning. A similar observation is made by Taori et al. (2020) in the vision domain, where the authors found training with larger and more diverse datasets offer better robustness consistently in multiple cases, compared to various robustness interventions proposed in the existing literature.

Training with a Better Use of Minority Examples

Further, there are several works that propose to robustify the models via a better use of minority examples, e.g., examples that are under-represented in the training distribution, or examples that are harder to learn. For example, Yaghoobzadeh et al. (2021) proposed to first fine-tune the model on the full data, and then on minority examples only.

In general, the training strategy with an emphasis on a subset of samples that are particularly hard for the model to learn is sometimes also referred to as group DRO Sagawa et al. (2020a), as an extension of vanilla distributional robust optimization (DRO) Ben-Tal et al. (2013); Duchi et al. (2021). Extensions of DRO are mostly discussing the strategies on how to identify the samples considered as minority: Nam et al. (2020) trained two models in parallel, where the “debiased” model focuses on examples not learned by the “biased” model; Lahoti et al. (2020) used an adversary model to identify samples that are challenging to the main model; Liu et al. (2021) proposed to train the model a second time via up-weighting examples that have high training losses during the first time.

When to Use Data-driven or Model-based Approaches?

In many cases both the data and the model can contribute to a model’s lack of robustness, hence data-driven and model-based approaches could be combined to further improve a model’s robustness. One interesting phenomenon observed by Liu et al. (2019) is to attribute models’ robustness failures to blind spots in the training data, or the intrinsic learning ability of the model. The authors found that both patterns are possible: in some cases models can be inoculated via being exposed to a small amount of challenging data, similar to the data augmentation approaches mentioned in Section 5.1; on the other hand, some challenging patterns remain difficult which connects to the larger question around generalizability to unseen adversarial and counterfactual patterns Huang et al. (2020); Jha et al. (2020); Joshi and He (2021), which is relatively under-explored but deserves much attention.

3 Inductive-prior-based Approaches

Another thread is to introduce inductive bias (i.e., to regularize the hypothesis space) to force the model to discard some spurious features. This is closely connected to the human-prior-based identification approaches in Section 4.1 as those human-priors can often be used to re-formulate the training objective with additional regularizers. To achieve this goal, one usually needs to first construct a side component to inform the main model about the misaligned features, and then to regularize the main model according to the side component. The construction of this side component usually relies on prior knowledge of what the misaligned features are. Then, methods can be built accordingly to counter the features such as label-associated keywords He et al. (2019), label-associated text fragments Mahabadi et al. (2020), and general easy-to-learn patterns of data Nam et al. (2020). Similarly, Clark et al. (2019, 2020); Utama et al. (2020a, b) propose to ensemble with a model explicitly capturing bias, where the main model is trained together with this “bias-only” model such that the main model is discouraged from using biases. More recent work Xiong et al. (2021) shows the ensemble-based approaches can be further improved via better calibrating the bias-only model. Furthermore, additional regularizers have been introduced for robust fine-tuning over pre-trained models, e.g., mutual-information-based regularizers Wang et al. (2021a) and smoothness-inducing adversarial regularization Jiang et al. (2020).

In a broader scope, given that one of the main challenges of domain adaptation is to counter the model’s tendency in learning domain-specific spurious features Ganin et al. (2016), some methods contributing to domain adaption may have also progressed along the line of our interest, e.g., domain adversarial neural network Ganin et al. (2016). This line of work also inspires a family of methods forcing the model to learn auxiliary-annotation-invariant representations with a side component Ghifary et al. (2016); Wang et al. (2017); Rozantsev et al. (2018); Motiian et al. (2017); Li et al. (2018); Wang et al. (2019c); Vernikos et al. (2020).

Despite the diverse concrete ideas introduced, the above is mainly training for small empirical loss across different domains or distributions in addition to forcing the model to be invariant to domain-specific spurious features. As an extension along this direction, invariant risk minimization (IRM) Arjovsky et al. (2019) introduces the idea of invariant predictors across multiple environments, which was later followed and discussed by a variety of extensions Choe et al. (2020); Ahmed et al. (2020); Rosenfeld et al. (2021). More recently, Dranker et al. (2021) applied IRM in natural language inference and found that a more naturalistic characterization of the problem setup is needed.

4 Causal Intervention

Casual analyses have also been utilized to examine robustness. Srivastava et al. (2020) leverage humans’ common sense knowledge of causality to augment training examples with a potential unmeasured variable, and propose a DRO-based approach to encourage the model to be robust to distribution shifts over the unmeasured variables. Balashankar et al. (2021) study the effect of secondary attributes, or confounders, and propose context-aware counterfactuals that take into account the impact of secondary attributes to improve models’ robustness. Veitch et al. (2021) propose to learn approximately counterfactual invariant predictors dependent on causal structures of the data, and show it can help mitigate spurious correlations in text classification.

5 Connections between Mitigations

Connecting these methods conceptually, we conjecture three different mainstream approaches: one is to leverage the large amount of data by taking advantages of pre-trained models, another is to learn invariant representations or predictors across domains or environments, while most of the rest build upon the prior on what the spurious patterns are and encourage the models to not rely on those patterns. Then the solutions are invented through countering model’s learning of these patterns by either data augmentation, reweighting (the minorities), ensemble, inductive-prior design, and causal intervention. Interestingly, statistical work has shown that many of these mitigation methods are optimizing the same robust machine learning generalization error bound Wang et al. (2021c).

Open Questions

In addition to the challenges mentioned above, we list below a few open questions that call for additional research going forward.

Existing identification around robustness failures rely heavily on human priors and error analyses, which usually pre-define a small or limited set of patterns that the model could be vulnerable to. This requires extensive amount of expertise and efforts, and might still suffer from human or subjective biases in the end. How to proactively discover and identify models’ unrobust regions automatically and comprehensively remains challenging.

Interpreting and Mitigating Spurious Correlations

Interpretability matters for large NLP models, especially key to the robustness and spurious patterns. How can we develop ways to attribute or interpret these vulnerable portions of NLP models and communicate these robustness failures with designers, practitioners, and users? In addition, recent work Wallace et al. (2019c); Wang et al. (2021d); Zhang et al. (2021) show interpretability methods can be utilized to better understand how a model makes its decision, which in turn can be used to uncover models’ bias, diagnose errors, and discover spurious correlations.

Furthermore, the mitigation of spurious correlations often suffers from the trade-off between removing shortcuts and sacrificing model performance Yang et al. (2020); Zhang et al. (2019a). Additionally, most existing mitigation strategies work in a pipeline fashion where defining and detecting spurious correlations are prerequisites, which might lead to error cascades in this process. How to design end-to-end frameworks for automatic mitigation deserves much attention.

Unified Framework to Evaluate Robustness

With a variety of potential spurious patterns in NLP models, it becomes increasingly challenging for developers and practitioners to quickly evaluate the robustness and quality of their models. This calls for more unified benchmarking efforts such as CheckList Ribeiro et al. (2020), Reliability Testing Tan et al. (2021), Robustness Gym Goel et al. (2021) and Dynabench Kiela et al. (2021), to facilitate fast and easy evaluation of robustness.

User Centered Measures and Mitigation

Instead of passively detecting spurious correlations from a post-processing perspective, how to approach robustness from a user centric perspective needs further investigation. Based on the dual-process models of information processing, humans use two different processing styles Evans (2010). One is a quick and automatic style that relies on well-learned information and heuristic cues. The other is a qualitatively different style that is slower, more deliberative, and requires more reflective reasoning. Would these well-learned information and heuristic rules be leveraged to help design better human priors to measure and mitigate spurious correlations? If users or stakeholders are involved in this process, collecting a set of test cases where a system might perform well for the wrong reasons could help design sanity tests.

Connections between Human-like Linguistic Generalization and NLP Generalization

Linzen (2020) argue NLP models should behave more like humans to achieve better generalization consistently. It is interesting to note that how humans process information in NLP tasks exactly is still under exploration, and to what extent models should leverage human-knowledge is still a debatable topic.http://www.incompleteideas.net/IncIdeas/BitterLesson.html Nonetheless, if we can better understand and utilize the robustness properties in human perception, we can potentially advance models’ robustness in a more meaningful way.

Conclusion

In this paper, we provided a unifying overview over robustness definitions, evaluations and mitigation strategies in the NLP domain. We also highlighted open challenges in this area to motivate future research, encouraging people to think deeply about more comprehensive benchmarks, transferability and validity of adversarial examples, unified framework to evaluate and improve robustness, user-centered measures and mitigation, and finally how to potentially achieve human-like linguistic generalization more meaningfully.

Acknowledgements

The authors would like to thank reviewers for their helpful insights and feedback. This work is funded in part by a grant from Google.

References