Explaining NLP Models via Minimal Contrastive Editing (MiCE)
Alexis Ross, Ana Marasović, Matthew E. Peters
Introduction
Cognitive science and philosophy research has shown that human explanations are contrastive (Miller, 2019): People explain why an observed event happened rather than some counterfactual event called the contrast case. This contrast case plays a key role in modulating what explanations are given. Consider Figure 1. When we seek an explanation of the model’s prediction “by train,” we seek it not in absolute terms, but in contrast to another possible prediction (i.e. “on foot”). Additionally, we tailor our explanation to this contrast case. For instance, we might explain why the prediction is “by train” and not “on foot” by saying that the writer discusses meeting Ann at the train station instead of at Ann’s home on foot; such information is captured by the edit (bolded red) that results in the new model prediction “on foot.” For a different contrast prediction, such as “by car,” we would provide a different explanation. In this work, we propose to give contrastive explanations of model predictions in the form of targeted minimal edits, as shown in Figure 1, that cause the model to change its original prediction to the contrast prediction.
Given the key role that contrastivity plays in human explanations, making model explanations contrastive could make them more user-centered and thus more useful for their intended purposes, such as debugging and exposing dataset biases Ribera and Lapedriza (2019)—purposes which require that humans work with explanations Alvarez-Melis et al. (2019). However, many currently popular instance-based explanation methods produce highlights—segments of input that support a prediction (Zaidan et al., 2007; Lei et al., 2016; Chang et al., 2019; Bastings et al., 2019; Yu et al., 2019; DeYoung et al., 2020; Jain et al., 2020; Belinkov and Glass, 2019) that can be derived through gradients (Simonyan et al., 2014; Smilkov et al., 2017; Sundararajan et al., 2017), approximations with simpler models (Ribeiro et al., 2016), or attention (Wiegreffe and Pinter, 2019; Sun and Marasović, 2021). These methods are not contrastive, as they leave the contrast case undetermined; they do not tell us what would have to be different for a model to have predicted a particular contrast label.Free-text rationales Narang et al. (2020) can be contrastive if human justifications are collected by asking “why… instead of…” which is not the case with current benchmarks Camburu et al. (2018); Rajani et al. (2019); Zellers et al. (2019).
As an alternative approach to NLP model explanation, we introduce Minimal Contrastive Editing (\mice)—a two-stage approach to generating contrastive explanations in the form of targeted minimal edits (as shown in Figure 1). Given an input, a fixed Predictor model, and a contrast prediction, \micegenerates edits to the input that change the Predictor’s output from the original prediction to the contrast prediction. We formally define our edits and describe our approach in §2.
We design \miceto produce edits with properties motivated by human contrastive explanations. First, we desire edits to be minimal, altering only small portions of input, a property which has been argued to make explanations more intelligible (Alvarez-Melis et al., 2019; Miller, 2019). Second, \miceedits should be fluent, resulting in text natural for the domain and ensuring that any changes in model predictions are not driven by inputs falling out of distribution of naturally occurring text. Our experiments (§3) on three English-language datasets, Imdb, Newsgroups, and Race, validate that \miceedits are indeed contrastive, minimal, and fluent.
We also analyze the quality of \miceedits (§4) and show how they may be used for two use cases in NLP system development. First, we show that \miceedits are comparable in size and fluency to human edits on the Imdb dataset. Next, we illustrate how \miceedits can facilitate debugging individual model predictions. Finally, we show how \miceedits can be used to uncover dataset artifacts learned by a powerful Predictor model.Our code and trained Editor models are publicly available at https://github.com/allenai/mice.
\mice: Minimal Contrastive Editing
This section describes our proposed method, Minimal Contrastive Editing, or \mice, for explaining NLP models with contrastive edits.
Contrastive explanations are answers to questions of the form Why p and not q? They explain why the observed event happened instead of another event , called the contrast case.Related work also calls it the foil Miller (2019). A long line of research in the cognitive sciences and philosophy has found that human explanations are contrastive (Van Fraassen, 1980; Lipton, 1990; Miller, 2019). Human contrastive explanations have several hallmark characteristics. First, they cite contrastive features: features that result in the contrast case when they are changed in a particular way (Chin-Parker and Cantelon, 2017). Second, they are minimal in the sense that they rarely cite the entire causal chain of a particular event, but select just a few relevant causes (Hilton, 2017). In this work, we argue that a minimal edit to a model input that causes the model output to change to the contrast case has both these properties and can function as an effective contrastive explanation. We first give an illustration of contrastive explanations humans might give and then show how minimal contrastive edits offer analogous contrastive information.
As an example, suppose we want to explain why the answer to the question “Q: Where can you find a clean pillow case that is not in use?” is “A: the drawer.”Inspired by an example in Talmor et al. (2019): Question: “Where would you store a pillow case that is not in use?” Choices: “drawer, kitchen cupboard, bedding store, england.” If someone asks why the answer is not “C1: on the bed,” we might explain: “E1: Because only the drawer stores pillow cases that are not in use.” However, E1 would not be an explanation of why the answer is not “C2: in the laundry hamper,” since both drawers and laundry hampers store pillow cases that are not in use. For contrast case C2, we might instead explain: “E2: Because only laundry hampers store pillow cases that are not clean.” We cite different parts of the original question depending on the contrast case.
In this work, we propose to offer contrastive explanations in the form of minimal edits that result in the contrast case as model output. Such edits are effective contrastive explanations because, by construction, they highlight contrastive features. For example, a contrastive edit of the original question for contrast case C1 would be: “Where can you find a clean pillow case that is not in use?”; the information provided by this edit—that it is whether or not the pillow case is in use that determines whether the answer is “the drawer” or “on the bed”—is analogous to the information provided by E1. Similarly, a contrastive edit for contrast case C2 that changed the question to “Where can you find a clean dirty pillow case that is not in use?” provides analogous information to E2.
2 Overview of \mice
We define a contrastive edit to be a modification of an input instance that causes a Predictor model (whose behavior is being explained) to change its output from its original prediction for the unedited input to a given target (contrast) prediction. Formally, for textual inputs, given a fixed Predictor , input of tokens, original prediction and contrast prediction , a contrastive edit is a mapping such that .
We propose \mice, a two-stage approach to generating contrastive edits, illustrated in Figure 2. In Stage 1, we prepare a highly-contextualized Editor model to associate edits with given end-task labels (i.e., labels for the task of the Predictor) such that the contrast label is not ignored in \mice’s second stage. Intuitively, we do this by masking the spans of text that are “important” for the given target label (as measured by the Predictor’s gradients) and training our Editor to reconstruct these spans of text given the masked text and target label as input. In Stage 2 of \mice, we generate contrastive edits using the Editor model from Stage 1. Specifically, we generate candidate edits by masking different percentages of and giving masked inputs with prepended contrast label to the Editor; we use binary search to find optimal masking percentages and beam search to keep track of candidate edits that result in the highest probability of the contrast labels given by the Predictor.
3 Stage 1: Fine-tuning the Editor
In Stage 1 of \mice, we fine-tune the Editor to infill masked spans of text in a targeted manner. Specifically, we fine-tune a pretrained model to infill masked spans given masked text and a target end-task label as input. In this work, we use the Text-to-Text Transfer Transformer (T5) model (Raffel et al., 2020) as our pretrained Editor, but any model suitable for span infilling can in principle be the Editor in \mice. The addition of the target label allows the highly-contextualized Editor to condition its predictions on both the masked context and the given target label such that the contrast label is not ignored in Stage 2. What to use as target labels during Stage 1 depends on who the end-users of \miceare. The end-user could be: (1) a model developer who has access to the labeled data used to train the predictor, or (2) lay-users, domain experts, or other developers without access to the labeled data. In the former case, we could use the gold label as targets, and in the latter case, we could use the labels predicted by Predictor. Therefore, during fine-tuning, we experiment with using both gold labels and original predictions of our Predictor model as target labels. To provide target labels, we prepend them to inputs to the Editor. For more information about how these inputs are formatted, see Appendix B. Results in Table 2 show that fine-tuning with target labels results in better edits than fine-tuning without them.
4 Stage 2: Making Edits with the Editor
In the second stage of our approach, we use our fine-tuned Editor to make edits using beam search Reddy (1977). In each round of edits, we mask consecutive spans of of tokens in the original input, prepend the contrast prediction to the masked input, and feed the resulting masked instance to the Editor; the Editor then generates edits. The masking procedure during this stage is gradient-based as in Stage 1.
In one round of edits, we conduct a binary search with levels over values of between values to to efficiently find a value of that is large enough to result in the contrast prediction while also modifying only minimal parts of the input. After each round of edits, we get ’s predictions on the edited inputs, order them by contrast prediction probabilities, and update the beam to store the top edited instances. As soon as an edit is found that results in the contrast prediction, i.e., , we stop the search procedure and return this edit. For generation, we use a combination of top-k (Fan et al., 2018) and top-p (nucleus) sampling (Holtzman et al., 2020).We use this combination because we observed in preliminary experiments that it led to good results.
Evaluation
This section presents empirical findings that \miceproduces minimal and fluent contrastive edits.
We evaluate \miceon three English-language datasets: IMDB, a binary sentiment classification task (Maas et al., 2011), a 6-class version of the 20 Newsgroups topic classification task (Lang, 1995), and Race, a multiple choice question-answering task (Lai et al., 2017).We create this 6-class version by mapping the 20 existing subcategories to their respective larger categories—i.e. “talk.politics.guns” and “talk.religion.misc” “talk.” We do this in order to make the label space smaller. The resulting classes are: alt, comp, misc, rec, sci, and talk.
Predictors
can be used to make contrastive edits for any differentiable Predictor model, i.e., any end-to-end neural model. In this paper, for each task, we train a Predictor model built on RoBERTa-large (Liu et al., 2019), and fix it during evaluation. The test accuracies of our Predictors are 95.9%, 85.3% and 84% for Imdb, Newsgroups, and Race, respectively. For training details, see Appendix A.1.
Editors
Our Editors build on the base version of T5. For fine-tuning our Editors (Stage 1), we use the original training data used to train Predictors. We randomly split the data, 75%/25% for fine-tuning/validation and fine-tune until the validation loss stops decreasing (for a max of 10 epochs) with of tokens masked, where is a randomly chosen value in $b=3s=4n_{2}n_{2}m=15p=0.95k=3050n_{2}$, the generations produced by the T5 Editors sometimes degenerate; see Appendix C for details.
Metrics
We evaluate \miceon the test sets of the three datasets. The Race and Newsgroups test sets contain 4,934 and 7,307 instances, respectively.For the Newsgroups test set, there are 7,307 instances remaining after filtering out empty strings. For Imdb, we randomly sample 5K of the 25K instances in the test set for evaluation because of the computational demands of evaluation. A single contrastive edit is expensive and takes an average of seconds per Imdb instance ( tokens). Calculating the fluency metric adds an additional average of seconds per Imdb instance. For more details, see Section 5.
For each dataset, we measure the following three properties: (1) flip rate: the proportion of instances for which an edit results in the contrast label; (2) minimality: the “size” of the edit as measured by the word-level Levenshtein distance between the original and edited input, which is the minimum number of deletions, insertions, or substitutions required to transform one into the other. We report a normalized version of this metric with a range from 0 to 1—the Levenshtein distance divided by the number of words in the original input; (3) fluency: a measure of how similarly distributed the edited output is to the original data. We evaluate fluency by comparing masked language modeling loss on both the original and edited inputs using a pretrained model. Specifically, given the original -length sequence, we create copies, each with a different token replaced by a mask token, following Salazar et al. (2020). We then take a pretrained t5-base model and compute the average loss across these copies. We compute this loss value for both the original input and edited input and report their ratio—i.e., edited original. We aim for a value of 1.0, which indicates equivalent losses for the original and edited texts. When \micefinds multiple edits, we report metrics for the edit with the smallest value for minimality.
2 Results
Results are shown in Table 1. Our proposed Grad \miceprocedure (upper part of Table 1) achieves a high flip rate across all three tasks. This is the outcome regardless of whether predicted target labels (first row, 91.5–100% flip rate) or gold target labels (second row, 94.5–100% flip rate) are used for fine-tuning in Stage 1. We observe a slight improvement from using the gold labels for the Race Predictor, which may be explained by the fact that it is less accurate (with a training accuracy of ) than the Imdb and Newsgroups classifiers.
achieves a high flip-rate while its edits remain small and result in fluent text. In particular, \miceon average changes 17.3–33.1% of the original tokens when predicted labels are used in Stage 1 and 18.5–33.5% with gold labels. Fluency is close to 1.0 indicating no notable change in mask language modeling loss after the edit—i.e., edits fall in distribution of the original data. We achieve the best results across metrics on the Imdb dataset, as expected since Imdb is a binary classification task with a small label space. These results demonstrate that \micepresents a promising research direction for the generation of contrastive explanations; however, there is still room for improvement, especially for more challenging tasks such as Race.
In the rest of this section, we provide results from several ablation experiments.
We investigate the effect of fine-tuning (Stage 1) with a baseline that skips Stage 1 altogether. For this No-Finetune baseline variant of \mice, we use the vanilla pretrained T5-base as our Editor. As shown in Table 1, the No-Finetune variant underperforms all other (two-stage) variants of \micefor the Imdb and Newsgroups datasets.We leave Race out from our evaluation with the No-Finetune baseline because we observe that the pretrained T5 model does not generate text formatted as span infills; we hypothesize that this model has not been trained to generate infills for masked inputs formatted as multiple choice inputs. Fine-tuning particularly improves the minimality of edits, while leaving the flip rate high. We hypothesize that this effect is due to the effectiveness of Stage 2 of \miceat finding contrastive edits: Because we iteratively generate many candidate edits using beam search, we are likely to find a prediction-flipping edit. Fine-tuning allows us to find such an edit at a lower masking percentage.
Gradient vs. Random Masking
We study the impact of using gradient-based masking in Stage 1 of the \miceprocedure with a Rand variant, which masks spans of randomly chosen tokens. As shown in the middle part of Table 1, gradient-based masking outperforms random masking when using both predicted and gold labels across all three tasks and metrics, suggesting that the gradient-based attribution used to mask text during Stage 1 of \miceis an important part of the procedure. The differences are especially notable for Race, which is the most challenging task according to our metrics.
Targeted vs. Un-targeted Infilling
We investigate the effect of using target labels in both stages of \miceby experimenting with removing target labels during Stage 1 (Editor fine-tuning) and Stage 2 (making edits). As shown in Table 2, we observe that giving target labels to our Editors during both stages of \miceimproves edit quality. Fine-tuning Editors without labels in Stage 1 (“No Label”) leads to worse flip rate, minimality, and fluency than does fine-tuning Editors with labels (“Label”). Minimality is particularly affected, and we hypothesize that using target end-task labels in both stages provides signal that allows the Editor in Stage 2 to generate prediction-flipping edits at lower masking percentages.
In this section, we compare \miceedits with human contrastive edits. Then, we turn to a key motivation for this work: the potential for contrastive explanations to assist in NLP system development. We show how \miceedits can be used to debug incorrect predictions and uncover dataset artifacts.
We ask whether the contrastive edits produced by \miceare minimal and fluent in a meaningful sense. In particular, we compare these two metrics for \miceedits and human contrastive edits. We work with the Imdb contrast set created by Gardner et al. (2020), which consists of original test inputs and human-edited inputs that cause a change in true label. We report metrics on the subset of this contrast set for which the human-edited inputs result in a change in model prediction for our Imdb Predictor; this subset consists of instances. The flip rate of \miceedits on this subset is 100%. The mean minimality values of human and \miceedits are (human) and (\mice), and the mean fluency values are (human) and (\mice). The similarity of these values suggests that \miceedits are comparable to human contrastive edits along these dimensions.
We also ask to what extent human edits overlap with \miceedits. For each input, we compute the overlap between the original tokens changed by humans and the original tokens edited by \mice. The mean number of overlapping tokens, normalized by the number of original tokens edited by humans, is . Thus, while there is some overlap between \miceand human contrastive edits, they generally change different parts of text.\miceedits explain Predictors’ behavior and therefore need not be similar to human edits, which are designed to change gold labels. This analysis suggests that there may exist multiple informative contrastive edits for a single input. Future work can investigate and compare the different kinds of insight that can be obtained through human and model-driven contrastive edits.
2 Use Case 1: Debugging Incorrect Outputs
Here, we illustrate how \miceedits can be used to debug incorrect model outputs. Consider the Race input in Table 3.2, for which the Race Predictor gives an incorrect prediction. In this case, a model developer may want to understand why the model got the answer wrong. This setting naturally brings rise to a contrastive question, i.e., Why did the model predict the wrong choice (“twice”) instead of the correct one (“only once”)?
The \miceedit shown offers insight into this question: Firstly, it highlights which part of the paragraph has an influence on the model prediction—the last few sentences. Secondly, it reveals that a source of confusion is Mark’s joke about having traveled in George’s plane twice, as changing Mark’s dialogue from talking about a “first and…last” trip to a single trip results in a correct model prediction.
edits can also be used to debug model capabilities by offering hypotheses about “bugs” present in models: For instance, the edit in Table 3.2 might prompt a developer to investigate whether this Predictor lacks non-literal language understanding capabilities. In the next section, we show how insight from individual \miceedits can be used to uncover a bug in the form of a dataset-level artifact learned by a model. In Appendix D, we further analyze the debugging utility of \miceedits with a Predictor designed to contain a bug.
3 Use Case 2: Uncovering Dataset Artifacts
Manual inspection of some edits for Imdb suggests that the Imdb Predictor has learned to rely heavily on numerical ratings. For instance, in the Imdb example in Table 3.2, the \miceedit results in a negative prediction from the Predictor even though the edited text is overwhelmingly positive. We test this hypothesis by investigating whether numerical tokens are more likely to be edited by \mice.
We analyze the edits produced by \mice(Gold + Grad) described in §3.1. We limit our analysis to a subset of the 5K instances for which the edit produced by \micehas a minimality value of 0.05, as we are interested in finding simple artifacts driving the predictions of the Imdb Predictor; this subset has 902 instances. We compute three metrics for each unique token, i.e., type :
and report the tokens with the highest values for the ratios and . Intuitively, these tokens are removed/inserted at a higher rate than expected given the frequency with which they appear in the original Imdb inputs. We exclude tokens that occur 10 times from our analysis.
Results from this analysis are shown in Table 4. In line with our hypothesis, we observe a bias towards removing low numerical ratings and inserting high ratings when the contrast prediction is positive, and vice versa when is negative. In other words, in the presence of a numerical score, the Predictor may ignore the content of the review and base its prediction solely on the score (as in the Imdb example in Table 3).
In this section, we reflect on \mice’s shortcomings. Foremost, \miceis computationally expensive. Stage 1 requires fine-tuning a large pretrained generation model as the Editor. More significantly, Stage 2 requires multiple rounds of forward and backward passes to find a minimal edit: Each edit round in Stage 2 requires decoded sequences with the Editor, as well as forward passes and backward passes with the Predictor (with the first edit round), where is the beam width, is the number of search levels in binary search over the masking percentages, and is the number of generations sampled for each masking percentage. Our experiments required 180 forward passes, 180 decoded sequences, and 3 backward passes for edit rounds after the first.
While efficient search for targeted edits is an open challenge in other fields of machine learning (Russell, 2019; Dandl et al., 2020), this problem is even more challenging for language data, as the space of possible perturbations is much larger than for tabular data. An important future direction is to develop more efficient methods of finding edits.
This shortcoming prevents us from finding edits that are minimal in a precise sense. In particular, we may be interested in a constrained notion of minimality that defines an edit as minimal if there exists no subset of that results in the contrast prediction. Future work might consider creating methods to produce edits with this property.
The problem of generating minimal contrastive edits, also called counterfactual explanations Wachter et al. (2017),Formally, methods for producing targeted counterfactual explanations solve the same task as \mice. However, not all contrastive explanations are counterfactual explanations; contrastive explanations can take forms beyond contrastive edits, such as free-text rationales Liang et al. (2020) or highlights Jacovi and Goldberg (2020). In this paper, we choose to refer to \miceedits as “contrastive” rather than “counterfactual” because we seek to argue for the utility of contrastive explanations of model predictions more broadly; we present \miceas one method for producing contrastive explanations of a particular form and hope future work will explore different forms of contrastive explanations. has previously been explored for tabular data Karimi et al. (2020) and images Hendricks et al. (2018); Goyal et al. (2019); Looveren and Klaise (2019) but less for language. Recent work explores the use of minimal edits changing true labels for evaluation (Gardner et al., 2020) and data augmentation (Kaushik et al., 2020; Teney et al., 2020), whereas we focus on minimal edits changing model predictions for explanation.
There exist limited methods for automatically generating contrastive explanations of NLP models. Jacovi and Goldberg (2020) define contrastive highlights, which are determined by the inclusion of contrastive features; in contrast, our contrastive edits specify how to edit (vs. whether to include) features and can insert new text.See Appendix D for a longer discussion about the advantage of inserting new text in explanations, which \miceedits can do but methods that attribute feature importance (i.e. highlights) cannot. Li et al. (2020a) generate counterfactuals using linguistically-informed transformations (LIT), and Yang et al. (2020) generate counterfactuals for binary financial text classification using grammatically plausible single-word edits (REP-SCD). Because both methods rely on manually curated, task-specific rules, they cannot be easily extended to tasks without predefined label spaces, such as Race.LIT relies on hand-crafted transformation for NLI tasks based on linguistic knowledge, and REP-SCD makes antonym-based edits using manually curated, domain-specific lexicons for each label. Most recently, Jacovi et al. (2021) propose a method for producing contrastive explanations in the form of latent representations; in contrast, \miceedits are made at the textual level and are therefore more interpretable.
This work also has ties to the literature on causal explanation Pearl (2009). Recent work within NLP derives causal explanations of models through counterfactual interventions Feder et al. (2021); Vig et al. (2020). The focus of our work is the largely unexplored task of creating targeted interventions for language data; however, the question of how to derive causal relationships from such interventions remains an interesting direction for future work.
Counterfactuals Beyond Explanations
Concurrent work by Madaan et al. (2021) applies controlled text generation methods to generate targeted counterfactuals and explores their use as test cases and augmented examples in the context of classification. Another concurrent work by Wu et al. (2021) presents Polyjuice, a general-purpose, untargeted counterfactual generator. Very recent work by Sha et al. (2021), introduced after the submission of \mice, proposes a method for targeted contrastive editing for Q&A that selects answer-related tokens, masks them, and generates new tokens. Our work differs from these works in our novel framework for efficiently finding minimal edits (\miceStage 2) and our use of edits as explanations.
Connection to Adversarial Examples
Adversarial examples are minimally edited inputs that cause models to incorrectly change their predictions despite no change in true label Jia and Liang (2017); Ebrahimi et al. (2018); Pal and Tople (2020). Recent methods for generating adversarial examples also preserve fluency Zhang et al. (2019); Li et al. (2020b); Song et al. (2020)Song et al. (2020) propose a method to produce fluent semantic collisions, which they call the “inverse” of adversarial examples.; however, adversarial examples are designed to find erroneous change in model outputs; contrastive edits place no such constraint on model correctness. Thus, current approaches to generating adversarial examples, which can exploit semantics-preserving operations (Ribeiro et al., 2018) such as paraphrasing (Iyyer et al., 2018) or word replacement (Alzantot et al., 2018; Ren et al., 2019; Garg and Ramakrishnan, 2020), cannot be used to generate contrastive edits.
Connection to Style Transfer
The goal of style transfer is to generate minimal edits to inputs to result in a target style (sentiment, formality, etc.) (Fu et al., 2018; Li et al., 2018; Goyal et al., 2020). Most existing approaches train an encoder to learn style-agnostic latent representation of inputs and train attribute-specific decoders to generate text reflecting the content of inputs but exhibiting a different target attribute (Fu et al., 2018; Li et al., 2018; Goyal et al., 2020). Recent works by Wu et al. (2019) and Malmi et al. (2020) adopt two-stage approaches that first identify where to make edits and then make them using pretrained language models. Such approaches can only be applied to generate contrastive edits for classification tasks with well-defined “styles,” which exclude more complex tasks such as question answering.
We argue that contrastive edits, which change the output of a Predictor to a given contrast prediction, are effective explanations of neural NLP models. We propose Minimal Contrastive Editing (\mice), a method for generating such edits. We introduce evaluation criteria for contrastive edits that are motivated by human contrastive explanations—minimality and fluency—and show that \miceedits for the Imdb, Newsgroups, and Race datasets are contrastive, fluent, and minimal. Through qualitative analysis of \miceedits, we show that they have utility for robust and reliable NLP system development.
is intended to aid the interpretation of NLP models. As a model-agnostic explanation method, it has the potential to impact NLP system development across a wide range of models and tasks. In particular, \miceedits can benefit NLP model developers in facilitating debugging and exposing dataset artifacts, as discussed in §4. As a consequence, they can also benefit downstream users of NLP models by facilitating access to less biased and more robust systems.
While the focus of our work is on interpreting NLP models, there are potential misuses of \micethat involve other applications. Firstly, malicious actors might employ \miceto generate adversarial examples; for instance, they may aim to generate hate speech that is minimally edited such that it fools a toxic language classifier. Secondly, naively applying \micefor data augmentation could plausibly lead to less robust and more biased models: Because \miceedits are intended to expose issues in models, straightforwardly using them as additional training examples could reinforce existing artifacts and biases present in data. To mitigate this risk, we encourage researchers exploring data augmentation to carefully think about how to select and label edited instances.
We also encourage researchers to develop more efficient methods of generating minimal contrastive edits. As discussed in §5, a limitation of \miceis its computational demand. Therefore, we recommend that future work focus on creating methods that require less compute.
Appendix A Training Details
A.2 Editor Models
Appendix B Data Processing
We remove newline and tab tokens (
, \t, \n) in all datasets, as these are tokenized differently by our Predictors (RoBERTa-large) and Editors (T5). For Newsgroups, we also remove headers, footers, and quotes.
For Imdb and Newsgroups Editors, we simply prepend target labels to the masked original inputs. For Race, we give the question, context, all answer options, and the correct choice as input to the Race Editor. We only mask the context. See Table 5 for examples.
We noticed that generations sometimes degenerate when we decode from T5 with a large masking percentage . For example, sentinel tokens are sometimes generated out of consecutive order. We attribute this to the large difference between masking percentages we use (up to 55%) and masking percentage used during T5 pretraining (15%). Specifically, we observed that generations tend to degenerate after the the 28th sentinel token. Thus, we heuristically reduce the number of sentinel tokens by combining neighboring sentinel tokens that are separated by 1-2 tokens into one sentinel token.
When the output degenerates, we do the following: In-fill the mask tokens with the “good” parts of the generation (i.e. parts with correctly ordered sentinel tokens), and replace the remaining mask tokens with the original text; get the contrast label probabilities from for these intermediate in-filled candidates; of these, take the candidates with the highest probabilities and use as input to generate new candidates. If one of the partially-infilled candidates results in the contrast label, we return this as the edited input.
In §4, we illustrate how \miceedits can be used to debug both individual predictions and natural dataset artifacts learned by a model. Here, we further explore the utility of \miceedits in debugging through Data Staining Sippy et al. (2020): We design a “buggy” Predictor and evaluate whether \miceedits can recover the bug.
We create a buggy Race Predictor by introducing an artifact into the Race train set. This artifact is the presence of the phrase “It is interesting to note that” in front of the correct answer choice. We introduce this artifact as follows: We filter the Race train data to contain instances for which the correct answer choice is contained by some sentenceA sentence “contains” the correct answer choice if the answer has at least a 4-gram overlap with the sentence. and the overlapping sentence does not have a higher degree of n-gram overlap with some other (incorrect) choice. After filtering, 11,188 of 87,866 train instances remain. We then prepend “It is interesting to note that” to the overlapping sentence to design a correlation between the location of this phrase and the correct answer choice; our goal is to encourage a Predictor to learn to predict the multiple choice option closest to this buggy phrase as the correct answer. If there are multiple overlapping sentences, we choose the one with the most overlap with the answer choice. We randomly sample from this filtered subset such that 10% of the train data contains this artifact. Our buggy Race Predictor is trained on this modified data using the same set-up from §A.1, except that we use a batch size of 2 and 32 gradient accumulation steps.
The test accuracies of our original and buggy Race Predictors are both 84%, and so we cannot use this measure to select the better classifier. We ask whether \miceedits can be used for this purpose. One such edit is shown in Table 8. We observe that the signal from the edit, which contains both the manual artifact “It is interesting to note that” and the contrast prediction “three,” is enough to overpower the signal from the explicit assertion that “All” is the correct answer (“And all of them finally went to Pecking University”) such that the Predictor’s prediction changes to “Three.” This edit thus provides evidence that some heuristic may have been learned by the predictor. Considering multiple \miceedits can validate such a hypothesis: We find that of the edits produced by \micereflect this bug (i.e. contain the phrase “interesting to note that”); in other words, they do uncover the manually inserted bug.
Furthermore, \miceedits are able to uncover the artifact because they can insert new text. For instance, in the edit in Table 5, the buggy phrase “It is interesting to note that” is not part of the original input. Applying saliency-based explanation methods, such as gradient attribution, to the buggy Predictor’s prediction would not reveal the Predictor’s reliance on the manual artifact, as the buggy phrase is not already present in the text. This difference highlights a key advantage of \miceover existing instance-based explanation methods that attribute feature importance, which can only cite text already present in original inputs.