AEON: A Method for Automatic Evaluation of NLP Test Cases

Jen-tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He, Yuxin Su, Michael R. Lyu

Introduction

NLP software has become increasingly popular in our daily lives. For example, NLP virtual assistant software, such as Siri and Alexa, receives billions of requests (Maggio, 2018; SafeAtLast, 2021) while Google Translate App translates more than 100 billion words per day (Turovsky, 2016). With the development of Deep Neural Networks (DNNs), the performance of NLP software has been largely boosted. Equipped with the SOTA model (Vaswani et al., 2017), Microsoft question answering robot surpasses humans on conversational question answering task. In addition, the performance of machine comprehension (Gao et al., 2020), text generation (Li et al., 2020) and machine translation (Wang et al., 2022) has been significantly improved. However, NLP software can produce erroneous results, leading to misunderstanding, financial loss, threats to personal safety, and political conflicts (Okrent, 2016; Ong, 2017).

To discover erroneous behaviors in NLP software, researchers have designed various software testing techniques (He et al., 2020; Ribeiro et al., 2020; Li et al., 2019; Zang et al., 2020; Chen et al., 2021b). A test case for NLP software is in the form of a text (e.g., a sentence) and its label, where the label is the expected correct output of the NLP software. In theory, most of these testing techniques modify part(s) of the input text (e.g., a word/character substitution/insertion/deletion) under the assumption that the generated test case preserves an equivalent or similar semantic meaning. Typically, these techniques take labeled texts as inputs and output the mutated texts and the corresponding labels.

However, it is still challenging for current testing techniques to produce practical test cases of high quality. Specifically, tiny modification in a text can change its semantic meaning, which invalidates the common assumption that the semantic meaning of the original text and that of the generated text should remain equivalent or similar, further rendering the possibility of changing the corresponding labels (Michel et al., 2019; Morris et al., 2020a). For example, removing “not” from the text “I do not like the movie” changes its semantic meaning and further changes its label for a sentiment analysis task from “negative” to “positive”, resulting in a test case with an incorrect label and further a false alarm. Moreover, existing testing approaches cannot guarantee the fluency and naturalness of the generated test cases. Many word-level testing approaches introduce grammar errors and punctuation errors, and sometimes they introduce words that do not exist or are rarely used (Michel et al., 2019). Although these test cases may trigger “software errors” (e.g., unexpected software behaviors), it is important to first ensure the quality of the test cases in terms of semantic consistency and naturalness before finding more errors.

According to our user study, many of the NLP test cases generated by existing approaches are of low quality because of the following two issues: Inconsistent issue and Unnatural issue. These issues can lead to false alarms in testing and unnaturalness in language. In this paper, we say an NLP test case is of high quality if it does not have any of these issues. As shown in Table 1, a high-quality test case preserves the semantics of the original text and reads smoothly. The first Inconsistent case changes the semantics to the opposite while the second one changes the subjects. Two Unnatural cases hurt the fluency and naturalness of natural language by introducing either non-existing words or wrong grammar. It is unlikely that these low-quality test cases can contribute to improving NLP software in practice.

Hence, an automatic quality evaluation metric that can help filter out low-quality test cases generated by the existing testing techniques is highly in demand. Nevertheless, designing an automatic quality evaluation metric for NLP test cases is highly challenging. First, existing testing criteria are mainly based on coverage metrics, such as code coverage for traditional software (Chen et al., 2001) and neuron coverage for deep neural networks (Pei et al., 2017), which cannot be directly leveraged to detect false alarms and evaluate the quality of a natural language test case. Second, general semantic similarity evaluation metrics fail to detect Inconsistent issues under this scenario. Specifically, (1) most of the words in the original text and the generated text are the same while existing metrics evaluate the semantic similarity based on all the words in the text and thus, the impact of the mutated word(s) easily vanish; (2) a word may have different meanings in different contexts, making it difficult to compare only the mutated word(s). Third, existing work on naturalness evaluation metric either relies on human evaluation (Morris et al., 2020a) or qualitative analysis (e.g., part-of-speech checking (Garg and Ramakrishnan, 2020)), while we need an automatic and quantitative naturalness evaluation metric.

To address these problems, we introduce AEON, a method for Automatic Evaluation Of NLP test cases. AEON takes a text pair ¡original text, generated text¿ as input and outputs scores regarding semantic similarity and syntactic correctness, aiming for detecting Inconsistent and Unnatural issues, respectively. We use AEON to analyze the quality of NLP test cases generated from four popular testing techniques (Jia et al., 2019; Zang et al., 2020; Garg and Ramakrishnan, 2020; Ribeiro et al., 2020) on five datasets (Pang and Lee, 2005; Zhang et al., 2015; Williams et al., 2018; Bowman et al., 2015; Wang et al., 2019b) which cover three typical NLP tasks, namely natural language inference, sentiment analysis, and semantic equivalence. We conduct a comprehensive human evaluation on the semantic similarity and language naturalness between the original texts and the generated test cases, and we check whether AEON’s score aligns with human evaluation or not. The results show that AEON achieves the Average Precision (AP), Area Under Curve (AUC), and Pearson Correlation Coefficient (PCC) scores of 0.688, 0.742, and 0.922, outperforming the best baseline metric by 10%, 8.1%, and 7.8% respectively. On the evaluation of human judgment of language naturalness, AEON also surpasses all baselines and achieves the average AP, AUC, PCC scores of 0.69, 0.63, 0.82. These results demonstrate the effectiveness AEON on detecting false alarms and evaluating the language naturalness of NLP test cases. We also show that the high-quality test cases selected by AEON can significantly improve the accuracy and robustness of NLP software via model training. Our contributions can be summarized as:

We conduct a comprehensive user study on the test cases generated by existing NLP software testing techniques and find that 85% of them suffer from two issues: Inconsistent and Unnatural, resulting in a false alarm rate of 44%.

We introduce AEON, the first approach to quantitatively evaluate the quality of NLP test cases from semantics and language naturalness, addressing two main quality issues of NLP test cases mentioned above.

AEON is employed to evaluate the test cases generated by four testing techniques on five widely-used datasets, which shows that AEON achieves the best performance in terms of average AP, AUC, and PCC on all datasets.

The implementation of AEON, the raw experimental results, and the human annotation on the test case quality are available on Githubhttps://github.com/CUHK-ARISE/AEON.

Preliminaries

Though many papers have proposed testing techniques for Computer Vision (CV) software (e.g., face recognition system) (Tian et al., 2018; Xie et al., 2019; Pei et al., 2017; Carlini and Wagner, 2017), the characteristics of natural language make NLP software testing distinguished from that in CV software. The most significant difference between NLP test cases and CV test cases is that the input space of textual data is not as continuous as images, making every mutation in the original text perceptible. In addition, in natural language, mutating a single word can cause considerable semantic differences, which further leads to the risk of changing the correct label of the text. Therefore, when NLP testing techniques assign the label of the original text to the generated test cases, lots of false alarms occur.

Current testing techniquesIn this paper, we consider papers on attacking NLP models as a line of research on testing NLP software because the adversarial examples generated by these techniques can be regarded as test cases for NLP software. for NLP software can be roughly divided into four categories: character-level, word-level, sentence-level, and multi-level (Zhang et al., 2020d). Character-level techniques (Li et al., 2019) mutate a few characters that do not affect human reading comprehension. Word-level techniques (Ren et al., 2019; Ribeiro et al., 2020) are based on word substitution, usually using synonyms sets or Pre-trained Language Models (PLMs). Sentence-level techniques (He et al., 2021) change the whole structure of the sentences either by adding a sentence to the original texts or transforming the entire texts into another semantically similar format. Those combining different levels of techniques (Liang et al., 2018) can be categorized into multi-level techniques. In particular, word-level techniques significantly outperform others in terms of efficiency (Zang et al., 2020), applicability, and usefulness in robust training (Zang et al., 2020; Ren et al., 2019). However, this kind of technique suffers more from low-quality test cases (Morris et al., 2020a). Thus, we focus on test cases generated by word-level testing techniques.

From the perspective of combinatorial optimization, generating test cases with word-level techniques can be formulated as a searching problem, where we substitute each word in the original text to other words in our vocabulary. The whole search space is the number of words in original text NN (where we substitute) times the vocabulary size VV (word candidates). In general, these techniques include diverse modules to prune the search space, which can be classified into three components: target word selection, word substitution, and generation constraints (Zhang et al., 2020d; Morris et al., 2020b). Table 2 presents the modules of the four selected testing techniques in terms of the three components. A suitable target word selection method can decrease NN while a proper word substitution method can cut back VV. Constraints are commonly applied to ensure that the synthesized texts preserve semantic meaning and are syntactically correct.

2. Problem Definition

The task of this paper is to design an automatic evaluation metric that can reflect test case quality in terms of semantic consistency and naturalness, which facilitates the detection of Inconsistent (false alarms) and Unnatural test cases.

Approaches and Implementation

This section introduces the details of AEON whose input is a text pair ¡original text, generated text¿ and outputs are a semantic score and a syntactic score. AEON consists of two parts: SemEval (Semantic Evaluator), which captures the semantic difference between input text pair, and SynEval (Syntactic Evaluator), which assesses how likely the generated test case will be used (i.e., written or typed) by real users. These two components aim to address Inconsistent and Unnatural issues, respectively. In the rest of this section, we will introduce the details of the key components of the two evaluators.

SemEval aims to solve the two challenges mentioned above. (1) The influence of the mutated position can easily vanish when taking average since most words in the original text and the generated test case are the same. (2) Metrics comparing words without contexts can neglect their alternative meanings (i.e., polysemy). To this end, we propose to combine Levenshtein distance (Levenshtein, 1966) and sentence embedding model to evaluate the semantic similarity in the NLP testing scenario. The approach is surprisingly effective considering its simplicity, which is shown in Alg. 1.

After tokenizing the input texts (line 1-2), which converts all words and punctuation as individual tokens, SemEval extracts small patches of text where the two inputs differ using Levenshtein distance (line 3). With the help of Levenshtein distance, we can find all mutated positions in linear time. Next, it applies a PLM to obtain the embeddings of all tokens in the two inputs (line 4-5). Current PLMs (Devlin et al., 2019; Liu et al., 2019; Conneau et al., 2017; Gao et al., 2021) can be leveraged to project tokens to the embedding space, so this module can be replaced easily when more powerful PLMs are proposed. Then, for all patches and the whole text, we compute the cosine similarity defined by aTb∥a∥⋅∥b∥\frac{a^{T}b}{\lVert a\rVert\cdot\lVert b\rVert} in line 7-9 and line 12, respectively. Note that we extract totally five tokens as the patch for a mutation happened in position ii by [i−2:i+2][i-2:i+2]. For a mutation at the beginning or the end of a sentence, we extract the first or last three tokens as our patch. Finally, we compute the minimum and the average numbers among all patch similarities (line 13-14) and combine them with the text similarity using two hyper-parameters, λ1\lambda_{1} and λ2\lambda_{2} (line 15). After this convex combination, we obtain the output of SemEval, namely Sim(x, x^)Sim(x,\ \hat{x}).

We tackle challenge (1) by considering the minimum and average patch similarities. For challenge (2), AEON extract the mutated position along with its context, which can improve its ability to understand semantics. Consider an example:

If we only consider the mutated position charged and indicted, the similarity is high since they are synonyms in the meaning of “being accused”. However, charged here means ”to put electricity into an electrical device”. This kind of relationship can be captured by its context, which is modeled in the PLMs.

2. SynEval

Since synthesized test cases may include grammar errors, punctuation errors, or produce rarely used words and phrases, it is vital to use an automatic and quantitative metric to filter out these Unnatural test cases. Note that this kind of sentence rarely appears in real-world natural languages, hence they are treated as noises and ignored during the training process of PLMs (Devlin et al., 2019; Liu et al., 2019; Lan et al., 2020). Intuitively, how natural a sentence is can be reflected by the probability that the sentence has the same distribution as its training data, which can be estimated by PLMs. Therefore, SynEval is designed to measure naturalness through the perplexity of PLMs. Perplexity, in its formal definition, is the exponential form of the cross entropy of the given sentence (Jelinek et al., 1977), having the form of:

where xix_{i} is the ii-th word in the sentence and x1:i−1x_{1:i-1} is the first to i−1i-1-th words in the sentence. perplexity:X→[1, ∞)perplexity:\mathcal{X}\rightarrow[1,\ \infty) measures how confused the PLM is when it sees xix_{i} given x1:i−1x_{1:i-1}, the greater the more confused (i.e., worse).

The recently proposed BERT-like models, including BERT (Devlin et al., 2019), RoBERTa (Liu et al., 2019), and ALBERT (Lan et al., 2020) which trained on billions of sentences, are powerful PLMs for modeling this probability. However, BERT and its variants are bi-directional, taking not only x1:i−1x_{1:i-1} but also xi+1:nx_{i+1:n} as input. Therefore, we need to replace x1:i−1x_{1:i-1} with x\ix_{\backslash i} in Eq. 1, where x\ix_{\backslash i} denotes the input sentence with its ii-th word being [MASK]. Since our semantic evaluator outputs similarity scores in (0, 1](0,\ 1] (the greater, the more similar), we adopt ∏i=1NP(xi∣x\i)N\sqrt[N]{\prod_{i=1}^{N}P(x_{i}|x_{\backslash i})} for SynEval, having the same value range of (0, 1](0,\ 1] (the greater, the better).

Alg. 2 illustrates the implementation of SynEval. First we tokenize the input (line 1). Then, for each token in the input, we replace it with the special token [MASK] (line 4). Feeding the masked text to the PLM, we can obtain the prediction of the masked position, which is a probability distribution over the entire vocabulary (line 5). Next, we find out the probability that the PLM thinks the masked position can be filled with the original token and record it as the perplexity of this token (line 6-7). Finally, we compute the minimum and the average numbers among all perplexities and combine them using a hyper-parameter ϕ\phi. The score after this convex combination is the output of SynEval, namely Nat(x^)Nat(\hat{x}).

Experimental Design and Settings

In this paper, we focus on the following four research questions:

RQ1: What is the quality of the test cases generated by existing testing techniques (Section 5.1)?

RQ2: How effective is AEON (Section 5.2)?

RQ3: How can AEON help in testing NLP software? (Section 5.3)

RQ4: How can AEON help in improving NLP model? (Section 5.4)

To answer the RQs, the first step is generating test cases, i.e., testing NLP software. We choose to test the APIs provided by Hugging Face Inc.https://huggingface.co/, the largest NLP open-source community, on five widely-used datasets across three typical tasks: sentiment analysis, natural language inference, and semantic equivalence.

Datasets. Sentiment analysis aims at classifying the polarity (either positive or negative) of the sentiment of given texts. The inputs of natural language inference tasks are two pieces of texts, namely Premise and Hypothesis, and the target is to predict whether the Hypothesis is a contradiction, entailment, or neutral to the given Premise. If the Premise can infer the Hypothesis, the output is entailment; if the Premise can infer NOT Hypothesis, the output is contradiction; otherwise, the output is neutral. The inputs of semantic equivalence tasks are two pieces of text, namely question 1 and question 2, and the objective is to judge if the meaning of the two given questions is equivalent. We select five datasets, namely MR, Yelp, SNLI, MNLI, and QQP, for our experiments, whose details are shown in Table 3. MR and Yelp are crawled from the internet, so the data contain noises such as HTML tags, HTML encodings, HTML entity names, and hyperlinks, which will make the generated test case hard to read. To eliminate the influence of noisy data in our human evaluation, we convert HTML texts to plain texts and remove hyperlinks using regular expressions.

Testing. To be more specific, we choose five BERT-based APIshttps://huggingface.co/textattack/bert-base-uncased-rotten-tomatoes https://huggingface.co/textattack/bert-base-uncased-yelp-polarity https://huggingface.co/textattack/bert-base-uncased-MNLI https://huggingface.co/textattack/bert-base-uncased-snli https://huggingface.co/textattack/bert-base-uncased-QQP for five different datasets. According to the statistics given by Hugging Face Inc, these APIs are downloaded more than 30k times every month on average. Using the testing techniques described in Table 2 implemented by TextAttack (Morris et al., 2020b) with their default settings, we generate test cases for all datasets (APIs). We select 400 original texts for each dataset using each technique, resulting in 8,000 test cases. After testing the APIs with our test cases, 3,262 test cases (40.8%) are reported as software errors.

2. Human Evaluation

We aim to find out whether the reported cases really trigger the erroneous behaviors of NLP software, in other words, whether they are false alarms. To this end, we design and launch a user study.

Design. Following (Celikyilmaz et al., 2020), we propose a unified framework to measure the quality of generated test cases. The quality is defined from four perspectives, including Naturalness, Consistency, Human Label, and Difficulty:

Consistency: From “1 strongly disagree” to “5 strongly agree”, how much do you think the two sentences have the same meaning? Consistency quantifies the semantic similarity between the original text and the changed text.

Naturalness: From “1 very bad” to “5 very good”, how fluent and natural do you think this sentence is? Naturalness measures the fluency and grammar of the examples, including grammar errors, punctuation errors, and spelling errors (unrecognizable words).

Human label: Ask humans to do the tasks of the given datasets. It is a task-specific question and records the human judgment of classification answers.

Difficulty: From “1 very easy” to “5 very hard”, how difficult for you to make the decision? Difficulty reflects how difficult the task is for humans.

Based on our definition in Sec. 1, high-quality text cases should have high naturalness and consistency scores. Human label and difficulty are used to classify the human evaluation results. We also ask annotators whether these test cases have other problems/issues that we have not identified. The responses show that Inconsistent issue and Unnatural issue can cover all their concerns.

Crowdsourcing. We distribute our questionnaire on Qualtricshttps://www.qualtrics.com/, a platform to design, share, and collect questionnaires. We recruit crowd workers on Prolifichttps://prolific.co/, a platform to post tasks and hire workers. Since our questions require a high level of reading comprehension and inference skills in English, we require Prolific workers to have a bachelor’s degree or above and have English as their first and most fluent language. Since we focus on false alarms, we randomly sample 100 test cases per dataset that are reported as software errors for human evaluation. In total, we choose 500 test cases and generate 2,000 questions. For each question, we ask three workers to give their judgment to reduce the variance. Therefore, we ask 150 workers to complete all questionnaires. It takes each worker 15-25 minutes to answer around 40 questions in a questionnaire, and each worker is paid about 5 pounds per hour. The total cost is 300 pounds.

3. Baselines

We select diverse test case evaluation metrics as baselines from two categories: Neuron Coverage (NC) metrics and NLP-based metrics, which are summarized in Table 4.

NC-based. NC and its variants are commonly-used for evaluating test cases. Different from AEON, NC-based metrics mainly aim at the evaluation of a test set instead of a test case. In our experiments, we consider NC-based metrics in two ways. (1) For basic Neuron Coverage (NC) (Pei et al., 2017) and Neuron Boundary Coverage (NBC) (Ma et al., 2018a), we calculate the NC scores of each generated test case. (2) For Top-kk Neuron Coverage (TKNC) (Ma et al., 2018a) and Bottom-kk Neuron Coverage (BKNC) (Xie et al., 2019), they cannot be adapted to a single test case (e.g., TKNC produces the same coverage for single test cases), thus we compute the number of neurons covered by the generated test case but not by the original text. Intuitively, changes in texts may be reflected in neuron activation. Note that the comparison with NC-based metrics is not apples-to-apples because NC-based metrics mainly evaluate the quality of a test set. We include the comparison here for the completeness of our discussion.

NLP-based. Since the main reason behind false alarms is that the generated test cases cannot keep equivalent or similar semantic meaning with the original text, we include multiple semantic similarity metrics for the baselines of SemEval. Evaluating the semantic similarity of texts has long been a complex problem in NLP research. Previous metrics can be divided into corpus-based, knowledge-based, and DNN-based. DNN-based metrics outperform other methods and have served as a breakthrough in semantic similarity research (Chandrasekaran and Mago, 2021). We consider a corpus-based metric, BLEU, a knowledge-based metric, Meteor, and four DNN-based metrics, InferSent, SBERT, SimCSE, and BERTScore. For embedding models, namely InferSent, SBERT, and SimCSE, we report the semantic similarity based on cosine similarity because it is used by most of the researchers (Garg and Ramakrishnan, 2020; Gao et al., 2021; Zhang et al., 2020c) and Euclidean distance yields similar results in all our experiments.

4. Evaluation Criteria

We compute three criteria: AP, AUC, and PCC, to discover the correlation between human judgment (Sec. 4.2) and the automatic evaluation metrics, including AEON. We treat the scoring systems as binary classification systems, the human judgment as ground truth, and draw their Precision-Recall curve (P-R curve) and Receiver Operating Characteristic curve (ROC curve) to calculate AP and AUC. P-R curve shows the trade-off between recall (i.e., true positive rate) and precision, while the ROC curve depicts the trade-off between true positive rate and false positive rate. AP and AUC represent the area under P-R curve and ROC curve, respectively. An excellent binary classification system tends to have high AP and AUC scores. Then, we check whether our scores are correlated with human judgment using PCC, the covariance of two variables divided by the product of their standard deviations, which can be written in the form of:

PCC is able to show how linearly correlated two variables are. Note that negative PCC value indicates that the two variables are negatively correlated.

Experimental Results

We average the consistency, naturalness, and difficulty scores as the respective final scores. We use the label that most workers agree as the final human label. For each test case, we decide whether it is an Inconsistent case or an Unnatural case based on the consistency and naturalness scores. Finally, if the human label differs from the given label, the test case is considered a false alarm.

To show how severe the problem in NLP test case quality is, we draw the Venn diagram of the generated test cases based on the human annotation results. As is shown in Fig. 1, 44% of them change the label and thus are false alarms. In other words, there are only 1,435 cases triggering software errors in all 8,000 test cases. 57% of them are not natural enough, while 71% fail to preserve the semantic meaning. Only 15% of them have good language naturalness and preserve the semantic meaning, which are counted as high-quality test cases. Besides the statistical information, we have two more observations. First, though the majority of Inconsistent cases are false alarms, a few Inconsistent cases do not change the label. These test cases only account for 11%, and the rest can be categorized into the two issues. Second, bad naturalness can sometimes hurt semantic meaning, resulting in test cases that are both Inconsistent and Unnatural. This is because the unnatural part can eliminate some key information in texts and further change the semantics.

2. RQ2: The Effectiveness of AEON

Since AEON is designed to evaluate the semantic similarity and language naturalness of NLP software test cases, we assess the two modules, SemEval and SynEval, to validate the effectiveness of our approach. We use default settings for all baselines, and we select k=192k=192 (one-fourth of neurons in each layer) for TKNC and BKNC. We set λ1=0.1, λ2=0.2\lambda_{1}=0.1,\ \lambda_{2}=0.2 for SemEval, and ϕ=0.6\phi=0.6 for SynEval.

We draw P-R curves and ROC curves for the semantic scores calculated by SemEval as well as the other baselines mentioned in Sec. 4.3 and consistency from human evaluation. Then we compute AP and AUC scores, which are shown in Table 5. Our method achieves higher AP and AUC values averaged on all datasets and baselines, showing the strong ability of SemEval to filter out Inconsistent cases. The results also validate the effectiveness of SemEval on capturing subtle semantic changes. We calculate PCC between the semantic score and human-annotated consistency for each method. As shown in Table 5, our approach achieves about 0.920.92 PCC on average, which significantly outperforms all the baselines. This shows that the score of the SemEval aligns well with human evaluation.

NC-based metrics achieve decent performance and surpass many NLP-based metrics, especially on MNLI and QQP datasets, indicating that neuron activation patterns can reflect text semantic changes. As for NLP-based metrics, BLEU and Meteor perform the worst since they cannot handle highly overlapped texts. The BLEU and Meteor scores for text pairs are always high since most of the words in the original texts and the generated test cases are the same. DNN-based metrics cannot perform well because of three main reasons. (1) Word embeddings usually lack semantic information. For instance, the embeddings of [reject] and [accept] calculated by BERT (Devlin et al., 2019) have high cosine similarity of 0.8460.846, while such word substitution changes the correct label in sentiment analysis. Another example is that the cosine similarity between embeddings of [Tom] and [Jack] is 0.9780.978, which hurts in natural language inference tasks. (2) Baselines that employ token matching (Zhang et al., 2020c) are prone to mistakenly matching multiple words to a single word. Consider the following example:

The word [not] in the generated text can be matched to the second [not] in the original sentence, resulting in a high similarity score. (3) Models based on contrastive learning fail due to the lack of data with subtle differences in their training set. These models are mainly trained on natural language inference datasets (Conneau et al., 2017; Gao et al., 2021), which can hardly cover the cases where two sentences have few but vital differences.

Considering different datasets, in sentiment analysis tasks, AEON and all other baselines perform better on MR than on Yelp since the texts in MR in shorter and simpler. For natural language inference tasks, though MNLI is more complicated than SNLI, it is surprising to observe that SNLI has lower PCC scores than MNLI. There are negative PCC scores when using BLEU and Meteor, indicating the negative correlation between the baselines and human evaluation. We think the reason behind this is that the complexity of MNLI lies in the diversity of contexts, and changes in contexts typically will not change the corresponding labels.

2.2. SynEval

To evaluate the performance of SynEval, we draw P-R curves and ROC curves and compute AP and AUC, treating SynEval as a binary classifier to recognize Unnatural cases. We also calculate PCC between SynEval and naturalness score from human evaluation, which is included in Table 6. We can observe that though NC-based metrics have an excellent performance on detecting Inconsistent test cases, they fall short of measuring language naturalness. The performance varies significantly on different datasets. The PCC scores of NBC, TKNC, and BKNC show a negative correlation on MNLI, while positively correlated on other datasets. We infer the reason behind this may be that the models make the decision based on the appearance of certain words or phrases, ignoring whether the input texts have good language naturalness. In addition to using BERT (Devlin et al., 2019) in SynEval that is presented in Table 6, we also try other language models including RoBERTa (Liu et al., 2019) and ALBERT (Lan et al., 2020), among which BERT achieves the highest AUC and AP values, averaged on all datasets. Note that traditional grammar checkers are not suitable for this task because they do not provide quantitative results, and they cannot reveal the error-free yet strange sentences that people rarely write.

The impact of the hyperparameters. If we set the proportion of the minimal semantic score from 0 to 1, we can observe that the performance increases at first, then remains stable at the same level, and finally drop when it gets close to 1. We balance this trade-off using a grid search for lambdas and phis. These parameters can be generalized to other datasets and NLP tasks since we adopt the same parameters and consistently achieve good performance for all selected datasets in our experiments. We also test different patch lengths ll for extracting [i−l:i+l][i-l:i+l]. In particular, l=1l=1 does not work well because most NLP models use BPE (Byte Pair Encoding) (Sennrich et al., 2016) for tokenization, which may divide a word into smaller tokens, making it extract only part of a word. Long patches (e.g., l≥5l\geq 5) suffer from the same problem as average scores, i.e., the impact of mutation vanishes after averaging. In our experiments, l=3l=3 and l=4l=4 lead to similar results. To reduce computation cost, we select l=4l=4 for this parameter.

3. RQ3: Test Case Selection Using AEON

This paper aims to propose a metric that facilitates NLP software testing by evaluating the quality of test cases. In this section, we utilize AEON to filter out low-quality test cases. We conduct experiments to verify whether the test cases selected by AEON enjoy better semantic consistency and language naturalness. Specifically, AEON can be utilized to filter out Inconsistency and Unnatural test cases to improve the quality of test cases in average. For SemEval, we set different thresholds for different tasks. We choose multiple thresholds for semantic similarity score because whether the label will change depends on the given task. Consider this pair of texts (original and generated) which is inconsistent in semantics:

If they are in a sentiment analysis dataset, the label remains unchanged (i.e., positive). However, if they appear in a natural language inference dataset as premises, and the hypothesis is “I went out for the movie”, the label changes from contradiction to entailment. Therefore, to best filter out those false alarms, we set thresholds as 0.87, 0.90, and 0.91 for sentiment analysis, natural language inference, and semantic equivalence, respectively. The thresholds are computed with a balance between true positive rate and false positive rate. From the thresholds, we can see that the three tasks need more semantic similarity increasingly to ensure the preservation of labels, which aligns with the characteristics of the datasets. For language naturalness, we set the threshold as 0.21.

We generate 500 test cases that are reported to trigger some software errors, including various datasets and testing techniques mentioned in Table 3 and Table 2 respectively. Then we check whether human evaluation has improved before and after applying AEON to select high-quality test cases. The results are shown in Table 7. The average consistency and naturalness scores of the 500 test cases are 2.627 and 2.916, which are below 3 (considered as Inconsistency and Unnatural test cases) in average. The false alarm rate is 0.44. After selecting test cases whose SemEval and SynEval scores are above the thresholds with the help of AEON, the quality of test cases is significantly enhanced. The scores increase to 3.357 and 3.305, considered high-quality test cases on average. The consistency score improves by 27.8%, and the naturalness score improves by 13.3%. The false alarm rate is 26.2%, showing a significant improvement of 40.6%. The results demonstrate the effectiveness of AEON on high-quality test case selection.

We choose one of the generated test cases as an example to illustrate the performance of our SemEval and SynEval compared to other baselines.

AEON achieves a semantic score of 0.58 and a syntactic score of 0.22. From the semantic side, the sentiment of the original example is positive. However, the sentiment of the generated test case is negative because [terrible] is a negative adjective. This test case is not only Inconsistent but also a false alarm since the label of this example is changed. Therefore, an excellent semantic metric should give this test case a low score to filter it out. Our method, SemEval, gives the text pair a score of 0.58, which is far below the threshold of 0.87, indicating that the test cases cannot preserve the semantics and should be filtered out. AEON works effectively because our design to consider patch similarity identifies that the substitution ([powerful]→\rightarrow[terrible]) dramatically changes the semantic meaning. From the syntactic side, the generated test case reads smoothly without difficulty comprehending its meaning, suggesting that it has good language naturalness. The case obtains a score of 0.22 given by SynEval, which is above the threshold of 0.21 and indicates that the test case is not an Unnatural case. All in all, our method outperforms other baselines both in semantic and syntactic perspectives on this example.

4. RQ4: Improving NLP Software with AEON

Although NLP software testing is a promising research direction, it incurs an important yet unavoidable question: can the test cases be utilized to improve NLP software? To further show how high-quality test cases selected by AEON can help in improving NLP software, we add test cases that the model misclassifies to the training set and conduct model re-training. Accuracy is verified on the test set of the given task, while robustness is evaluated using the success rate of adversarial attacks. In this section, we run experiments to verify whether AEON can further improve the robustness and accuracy in model re-training.

We focus on fine-tuning a pre-trained BERT on MR dataset for sentiment analysis task for simplicity and better reproducibility. We first generate the test cases using the testing techniques mentioned in Table 2 using the entire training set of MR for seeds (original texts). Then we consider two settings: (1) randomly select as many test cases as 5% to 25% cases in the training set to train the model; (2) rank all the test cases with AEON in descending order and select the same size as (1) to train the model. After training, which takes five epochs to reach convergence, we evaluate the models’ accuracy using the MR test set and robustness using an adversarial attack method, PWWS (Ren et al., 2019).

As shown in Fig. 2(a), models trained with ranked test cases outperform the models with randomly selected test cases in terms of accuracy. In addition, Fig. 2(b) shows that models trained with ranked test cases are more robust (the lower the attack success rate is, the more robust the model is). At the beginning of each figure, we observe improvement in accuracy and robustness on both lines. This is because test cases add generalization ability to models. As we use more additional data in training, the noisy data issue surfaces (i.e., false alarms and low-quality test cases) and starts to harm the model accuracy and robustness. We can observe that no matter what ratio of test cases we use, the models trained with AEON ranked data achieve higher accuracy and robustness thanks to the high-quality test cases it selects. In contrast, the accuracy and robustness drop quickly in models trained with randomly selected test cases. In conclusion, adding low-quality test cases can easily hurt both accuracy and robustness, while adding a reasonable amount of high-quality test cases selected by AEON leads to model improvement in terms of both accuracy and robustness.

5. Discussion

The validity of out user study. To alleviate workers’ negligence in the annotation process, we double-check the results in two cases. (1) The workers feel it challenging to select one specific label. Specifically, we check test cases with high difficulty scores (above 3.5), or where annotators returned three different labels (e.g., three workers give the label of contradiction, entailment, and neutral in natural language inference tasks). (2) The human label is different from the original label (i.e., label changes), while the human-annotated consistency score is still high, which is counter-intuitive since if the label is changed, then the semantic meaning must have changed. We find out 21 cases for case (1) and 66 examples for case (2) and hire two annotators with two years of experience in NLP research to re-evaluate them. Afterward, we compute the Kappa score to check inter-rater reliability. The Fleiss’ Kappa of the classification task is 0.76 averaged on all datasets, implying substantial agreement among the annotators. We do not apply Cohen’s Kappa since we have three annotators for each case.

The importance of Unnatural test cases. We evaluate the naturalness of test cases and filter out Unnatural ones because they are rarely seen in real-world scenarios. Though previous work intentionally add spelling and grammar errors to data to improve the robustness of the NLP models (Jones et al., 2020), we noticed annotators mentioned that they could not understand the weird grammar or expression of some generated texts. In addition, unnaturalness can hide important semantic information, which changes the semantics and further renders the possibility of generating false alarms. Consider the following example:

In this example, the unnaturalness changes the semantics of the sentence and further reverses from positive sentiment to negative sentiment while making readers confused. Thus, we think Unnatural issues are important, and they can also degrade the quality of NLP software test cases.

Limitations of our proposed approaches. Our proposed approaches consist of two parts, SemEval for semantic semilarity and SynEval for language naturalness. To apply SemEval, we require both the generated test cases and the original texts at the same time, thus limiting SemEval to testing techniques of metamorphic testing, where the metamorphic relation assumes the transformation will change or will not change the semantics of the original text. Under the assumption that the transformation will not change the semantics, we filter out those with lower SemEval scores, while under the assumption that the transformation will change the semantics, we filter out those with higher SemEval scores. SynEval can be applied to any textual test case.

Related Work

With the improvement of Artificial Intelligence (AI) models, companies tend to deploy AI in real-world applications like autonomous driving and neural machine translation (Pouyanfar et al., 2019). However, AI software inherits the deficiencies of AI models that they are prone to erroneous behavior given particular inputs (Athalye et al., 2018; Carlini et al., 2016; Carlini and Wagner, 2017; Du et al., 2020; Goodfellow et al., 2015; Xiang et al., 2019; Wu et al., 2019, 2021). A line of research has been conducted to test AI software systems to address this problem. Specifically, they test software based on convolutional neural network and feed forward neural network (Henriksson et al., 2019; Gambi et al., 2019; Kim et al., 2019; Ma et al., 2018c; Tian et al., 2018; Zhang et al., 2018; Wu et al., 2020; Zhang et al., 2022), software based on recurrent neuron network (Du et al., 2019), and software based on general DNN models (Hu et al., 2019; Zhang et al., 2020a, b). Other researchers focus on testing deep learning libraries (Pham et al., 2019), assist the debugging process (Ma et al., 2018b), and detect adversarial examples online (Ma et al., 2019; Tao et al., 2018; Wang et al., 2019a; Xu et al., 2018). Unlike these papers that primarily focus on CV software, this paper focuses on NLP software.

2. Testing NLP Software

DNNs have boosted the performance of many NLP fields such as code analysis (Alon et al., 2019; Iyer et al., 2016; Pradel and Sen, 2018) and machine translation (Vaswani et al., 2017). In recent years, researchers have proposed a variety of metamorphic testing techniques for NLP software (Ribeiro et al., 2020; He et al., 2020; Gupta et al., 2020; He et al., 2021; Chen et al., 2021b, a; Sun et al., 2020). In addition to metamorphic testing techniques, another line of research for finding NLP software errors (Li et al., 2019; Jin et al., 2020; Zang et al., 2020; Jia et al., 2019; Li et al., 2021; Guo et al., 2021) is inspired by the adversarial attack concept in the CV field. Our work focuses on automatic quality evaluation of test cases generated by these testing techniques. Thus, we believe AEON complements with existing work.

3. Testing Criteria

Testing criterion, such as code coverage, has been widely utilized to measure how good a test suite is in traditional software (e.g., compilers). Inspired by code coverage in traditional software, DeepXplore (Pei et al., 2017) introduces the concept of neuron coverage for AI software: the percentage of neurons activated by the test cases. In recent years, researchers have proposed diverse variants of neuron coverage as testing criteria focusing on different activation magnitudes (Ma et al., 2018a; Xie et al., 2019). Researchers also develop neuron coverage specially designed for recurrent neuron networks (Du et al., 2019; Guo et al., 2019; Huang et al., 2019) to adopt the properties of sequence inputs. Different from neuron coverage metrics, which often act as test adequacy criteria of the test suite, our approach focuses on the quality evaluation of every test case. Thus, we think AEON can complement with existing testing criteria and contribute to research on NLP software testing.

Conclusion

This paper is the first to explore the quality of test cases generated by NLP software testing techniques. In an evaluation study, we surprisingly observe that 44% of the generated NLP test cases are of low quality, incurring Inconsistent or/and Unnatural issues. Thus, instead of improving NLP software, utilization of these test cases in model training could even degrade its accuracy and robustness. To this end, we introduce AEON, a novel, effective approach for automatic quality evaluation of NLP software testing cases. Given an original text and a generated test case, AEON returns two scores regarding similarity consistency and language naturalness. Our evaluation and user study show that AEON’s scores align well with the quality scores returned by humans. In particular, AEON achieves 69.4% and 68.5% AP on detecting Inconsistent and Unnatural test cases generated by four SOTA testing techniques on five widely-used datasets. We can conduct test selection or prioritization according to the scores returned by AEON. In our evaluation, models trained on the test cases selected by AEON consistently achieve better accuracy and robustness than models trained on randomly selected test cases. We believe that this work is the important first step toward systematic quality evaluation of NLP software test cases, which can further enhance the effectiveness of testing techniques for NLP software and complement existing testing criteria. We leave the work for automatically fixing Inconsistent and Unnatural test cases for future exploration.

References