On the Blind Spots of Model-Based Evaluation Metrics for Text Generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, Yulia Tsvetkov
Introduction
Automatic evaluation of machine-generated text (Celikyilmaz et al., 2020) has been a core research challenge in the field of natural language generation (NLG), as difficult as language generation itself. Encouraged by the phenomenal success of large-scale pretraining (Devlin et al., 2019), a recent series of work proposed to base evaluation metrics on pretrained language models (PLMs) (Zhang et al., 2020; Yuan et al., 2021; Pillutla et al., 2021). For example, BERTScore (Zhang et al., 2020) computes a similarity score between the contextualized embeddings of the hypothesis and the reference text. PLM-based metrics have been shown to have higher correlations with human annotations for various tasks (Yuan et al., 2021), and are becoming increasingly popular in practice.
However, PLMs have flaws. They could assign a high likelihood to degenerate, repetitive text (Holtzman et al., 2020) and could be insensitive to perturbations such as word order shuffling (Pham et al., 2021), negation (Ettinger, 2020), etc. These flaws, in combination with certain design choices, may lead to the metrics based on such PLMs being brittle and open to manipulation (Figure 1).
In this work, we develop a suite of stress tests with synthetic data for the robustness analysis of NLG metrics. In essence, we induce a variety of potential errors in clean text and examine the resulting drop in the metric scores. The tests are motivated by metric design choices, properties of PLMs, or general fluency/consistency errors. Our methodology facilitates full control over the synthesized error types, allowing us to test extreme or even adversarial scenarios that are not well covered in standard correlation-oriented evaluations.
Our tests are applied to a range of recently proposed and widely used PLM-based metrics for the tasks of open-ended generation, translation, and summarization. They reveal a number of glaring insensitivities, biases, and even loopholes in different metrics. Besides analyzing the reasons behind, we also provide practical suggestions and workarounds for a more reliable evaluation.
Methodology
We now discuss our methodology. For simplicity, in this section, let us assume a multi-reference translation dataset, where each sample has two reference translations produced by human translators, denoted by Ref-A and Ref-B. We will generalize our methodology to other tasks in §3.
We begin by computing a “base” metric score by considering Ref-A as hypotheses and Ref-B as references. Since Ref-A is produced by human translators, we assume that it is less likely to contain translation errors than machine-generated text, and it should be assigned a high score by the metric. Due to these two assumptions, and to disambiguate from the reference set (Ref-B), we term Ref-A as the gold hypothesis set.
For each test, we apply a synthesized error type (e.g., truncation) to the gold hypothesis set to construct a noised hypothesis set. We make sure that the amount or type of induced errors is sufficient to be distinctive from the original gold hypothesis (to be detailed in §5). The source texts and the references are left intact.
To determine whether a metric passes a test, a simple rank-based protocol is used: We claim that the metric fails the test for this dataset if the noised hypothesis set is not scored worse than the base score (from the gold set).As we will introduce in §4, all metrics except MAUVE are sample-level, and we compare the average score assigned to the gold/noised hypothesis set. This rank-based protocol can be easily extended to the comparison of different gradations of the same noise type (controlled by hyper-parameters). For example, a 20%-truncation is expected to rank lower than a 10%-truncation, as more information is lost.
Tasks and Datasets
Our tests cover three ubiquitous text generation tasks: open-ended generation, translation, and summarization. We now describe the dataset used for each task and the setting for gold hypotheses.
For open-ended generation, we use the WikiText-103 dataset (Merity et al., 2016). We randomly select 2000 paragraphs of length around 256 tokens from the dataset (preprocessing detailed in Appendix B.2). The samples typically contain seven or eight sentences. We divide them into two sets with 1000 samples each, and set one as the references and the other as the gold hypotheses. The reference set is only used for the MAUVE metric (more details given in Appendix A).
For summarization, we use the popular CNN-Dailymail (CNNDM) dataset (Hermann et al., 2015). Kryscinski et al. (2020) collected 10 additional human-annotated summaries (different from the original reference summary) for each of 100 samples in the test set. We set the CNNDM reference summaries to be the gold hypotheses, and use these 10 annotations as references. Correspondingly, the multi-reference version of metrics are used. The gold hypotheses typically contain three sentences.
For translation, we use the evaluation dataset from the WMT21 metrics shared task (Akhbardeh et al., 2021). We only use the source text and reference translations. We report results on the German-English (De-En) language pair, which contains 1000 translation pairs. There are two human-translated references (human-A and human-B) for each sample. We use human-A as the gold hypothesis and human-B as the reference. We also repeat key experiments on the Chinese-English (Zh-En) data and obtain very similar observations. Therefore, we omit the Zh-En results for brevity.
Most samples in WMT only contain one sentence, which makes some of our tests impossible (e.g., sentence switching). For this reason, we build a paragraph-level translation dataset based on the Zh-En part of the TED-Talks task (Duh, 2018). It contains 100 samples, where each sample has two human-translated references and on average contains 7 sentences. We name this dataset as TED-MT, and discuss how we build it in Appendix B.1.
Metrics
For open-ended text generation, we test MAUVE (Pillutla et al., 2021), GPT-PPL and MLM-PPL (Salazar et al., 2020). We report the negated GPT/MLM-PPL so that all metric scores are the higher the better.
MAUVE is a reference-based metric computed using contextualized embeddings from PLMs. We explore MAUVE with GPT2-large, RoBERTa-large, and ELECTRA-large (Clark et al., 2020) features. In Pillutla et al. (2021), the exploration is centered around the GPT-2 feature. However, in this work we find the choice of feature has a crucial impact on the metric’s robustness.
GPT-PPL denotes perplexity from the GPT2-large (Radford et al., 2019) model. MLM-PPL is the masked language model perplexity from a RoBERTa-large model (Liu et al., 2019). We use a definition similar to the formulation in Salazar et al. (2020) and provide details in Appendix A.
For translation and summarization, we test BERTScore (Zhang et al., 2020), MoverScore (Zhao et al., 2019), BARTScore (Yuan et al., 2021), UniEval (Zhong et al., 2022), COMET (Rei et al., 2020), PRISM (Thompson and Post, 2020), and BLEURT (Sellam et al., 2020). Among these metrics, PRISM and BLEURT are only applied for translation, and UniEval is only applied for summarization. While COMET was originally proposed for translation, Kasai et al. (2022b) showed it has superior human correlation for CNNDM. Therefore, we also include it for summarization. We also include the traditional metrics BLEU (for translation), and ROUGE-2/L (for summarization).
Both BERTScore and BARTScore have variants for precision (-p), recall (-r), and f-measure (-f). In addition, BARTScore has a faithfulness (-faithful) variant. We test two model options, namely BARTScore-cnn and BARTScore-para.For BARTS-cnn, the Bart model is finetuned on the CNNDM dataset (Hermann et al., 2015). For BARTS-para, it is further finetuned on the ParaBank2 dataset (Hu et al., 2019). UniEval reports scores on four aspects: coherence, consistency, fluency, and relevance, and the overall score is the average of the four.
By default, the metrics for translation and summarization are reference-based.There are two exceptions: The BARTScore-faithful and UniEval-relevance do not utilize reference. COMET and PRISM have a quality estimation (QE) variant (Specia et al., 2021), where users do not need to provide any reference.
In most cases, we directly use the released package or code for each metric and follow the recommended hyper-parameter or variant setting. We defer further implementation details and variant explanations to Appendix A.
Stress Tests and Results
We organize our findings into subsections each containing a set of tests with the corresponding motivation, description, results, and implications with practical workarounds. In general we perform each test for all metrics, and we primarily discuss metrics found to be problematic for brevity.
We group and order our tests by their motivations: The positioned-error (§5.1) and injection (§5.2) tests are mainly motivated by certain metric design choices; The freq-ngram (§5.3) and self-evaluation (§5.4) tests are motivated by certain PLM properties; Finally, the fluency/consistency (§5.5) tests mimic general errors that human or machine writers could make. See Table 1 for a catalogue along with the metrics affected.
For MAUVE, the features for reference/hypothesis texts are extracted using the PLM representation of the final token. Hence, it could be suboptimal if the PLM is biased to encode only the local context (Khandelwal et al., 2018; He et al., 2021).
To test for this bias, we create synthetic errors by replacing a span of 10 consecutive tokens in different positions of the gold hypothesis with (1) 10 random tokens from the vocabulary, or (2) randomly shuffled tokens of the original span. We experiment with three different error positions by replacing the tokens at the very start, the middle, and the very end of the gold hypotheses. A robust metric should give a significantly lower score to this clearly modified distribution of the hypotheses.
Shown in Table 2, MAUVE-GPT2 shows only a marginal drop (around 3%) for the random or shuffle errors in the start and middle positions. In comparison, MAUVE-RoBERTa penalizes errors in all positions severely, which aligns better with expectations. MAUVE-ELECTRA’s behavior is similar to the RoBERTa variant and is deferred to Appendix C.1.
We correlate this result with an attention pattern analysis. As shown in Figure 2, we observe that GPT2-large’s attention is concentrated on the diagonal of the plot, which indicates GPT-2 mostly attends to the near history. In contrast, RoBERTa-large attends heavily to specific (probably important) token positions regardless of the current token position. In summary, the attention patterns provide evidence that GPT-2 features encode less long-range context compared to RoBERTa.Besides this pattern, both GPT2-large and RoBERTa-large assign a large portion of attention to the very first token, which is also observed by Vig and Belinkov (2019). This pattern is typical across different data samples.
Currently, the default feature used by MAUVE is from GPT-2, which as we show, ignores errors at the start or the middle of the generations. Our analysis indicates that MLMs such as RoBERTa or ELECTRA could be a better choice. See §5.5 for results on MAUVE’s other blind spots.
2 The Injection Test
UniEval (Zhong et al., 2022) reframes NLG evaluation as a boolean question answering task. For example, a question such as “Is this a coherent summary? Summary: [HYPO] Document: ...” along with the hypothesis replacing the [HYPO] placeholder is inputted to a trained T5 model (Raffel et al., 2020), and the score is based on the output probability of answering “Yes”.
This test is inspired by a recent series of work teaching LMs to follow instructions (Wei et al., 2022; Mishra et al., 2022). We construct several valueless but misleading injection hypotheses, which attempt to “instruct” (via natural language) the underlying PLM to answer yes.To clarify, we do not modify the prompts in UniEval. The name “injection” is borrowed from the code injection hacking in software engineering. Results of two example injections are shown in Table 3.
We observe that UniEval is tricked to give a high score to the valueless injection hypotheses, and the more specific injection (Inj-1) gets a higher score. This is surprising because UniEval is trained with constructed positive/negative samples, and it is not trained to follow instructions. We surmise this result is more related to the PLM’s nature to make the output consistent with the context. More examples and discussion are given in Appendix F.
The injection test shows that the metric’s judgement can be misled by some valueless text span, which can be used for cheating. It can be detected by a low score from traditional metrics such as ROUGE (Table 3).
3 The Frequent n𝑛n-gram Test
Due to the statistical nature of LMs, they have been known to favor frequent -grams in the data. We now stress-test whether log-likelihood-based metrics would wrongly favor a random sequence of frequent -grams over the gold hypotheses.
For open-ended generation, we collect the top- most frequent -grams from the WikiText dataset. We then build synthetic hypotheses of length 256 by uniformly sampling -grams from this collection and concatenating them (see Table 12 in Appendix G for an example). To a human evaluator, these sequences are completely random and should get a lower score than the gold hypotheses.
Strikingly, as shown in Table 4 with 4-gram, we find that both GPT-PPL and MLM-PPL assign higher scores to the frequent -gram sequences than gold. This gap further increases when we concentrate on more frequent -grams. We present additional results with 3-gram in Appendix G.
To illustrate this issue, we plot step-wise next-token probability given by the underlying GPT2-large model. As shown in Figure 3, the probabilities exhibit a pattern that high-probability regions concentrate at the end of each 4-gram. We attribute this behavior to the LM’s utilization of local context (Khandelwal et al., 2018).
We conduct similar tests on translation or summarization but do not observe problematic behavior from the metrics. We surmise the reason could be due to the poor alignment between the random -gram sequence and the source/reference text.
This test shows that the affected metrics are biased towards frequent -gram rather than global coherence. This test strengthens the importance of diversity metrics such as rep-4gram.
4 The Self-Evaluation Bias
Log-probability-based metrics (e.g., GPT-PPL) are based on generative models such as GPT-2 (Radford et al., 2019) or BART (Lewis et al., 2019). At the same time, these PLMs are also used as base models for developing new NLG systems (Yang and Klein, 2021). Naturally, we wonder whether this could cause some level of bias in the evaluation. In the following tests, we demonstrate this bias for the case of GPT-PPL and BARTScore.
For GPT-PPL, we construct a setting that mimics how it is used in practice: For the generator, we finetune GPT-2 models of different sizes (small, medium, and large), and use the models to generate continuations of prompts from the WikiText dataset. The details of finetuning are available in Appendix H. We use top- sampling (Fan et al., 2018) with to decode. For evaluator, we use GPT-2 models off-the-shelf.
For different combinations of generator and evaluator, the results are shown in Table 5. Conventional wisdom in the community is that the larger GPT model should generate higher-quality text, which correlates with the scores from the OPT-2.7b (Zhang et al., 2022) model. However, perplexities from GPT2-small and -medium violate these expectations, ranking generations from their own base models higher than those of larger models. We term this as the self-evaluation bias.
BARTScore (Yuan et al., 2021) evaluates text generation quality as the log-probability of a seq2seq model. The default implementation relies on the finetuned BART-large model. Here, we test a hypothetical setting, where we base BARTScore on another popular PLM: T5 (Raffel et al., 2020). We use the BARTScore-cnn-faithful variant, and finetune all models on the CNNDM dataset (details in Appendix H). The results are shown in Table 6. For this experiment, we do not assume the supremacy of one model over the other, as that requires more rigorous human evaluation.
We observe an interesting but worrisome phenomenon: BART and T5 based evaluators strongly favor generators based on their own respective base models. This bias extends to different-sized variants of the base models as well. It is, however, less pronounced for the reference-based variant BARTScore-para.
Overall, these results show that the log-probability-based metrics could be unfairly biased towards their underlying PLMs. Basing the metric on different PLM could give inconsistent ranking for the same set of systems.
Hence, practitioners should avoid situations where the generation system and the metric are based on the exact same PLM, or where systems based on different types of PLMs are compared with a metric based on one of them. In such cases, the scores should be complemented with additional evaluations from reference-based metrics.While prior works follow this guideline by intuition (Liu et al., 2021), we show an explicit empirical analysis in support of this practice, which was previously lacking in the literature.
5 Fluency & Consistency Tests
The tests we discussed so far have been motivated by certain metric design choices or properties of the underlying PLMs. In this section, we move to more general tests, where we synthesize a range of perturbations that mimic human or machine errors.
Our tests cover two important aspects of natural language: fluency and consistency (some of our consistency tests are also related to coherence). Fluency tests focus on grammaticality, while consistency tests focus on temporal order, logic, or alignment with the source text.
Similar to previous sections, in each test we apply one type of noise to the gold hypothesis. The noise can be regarded as an exaggeration of the errors human or machine writers could make. In total, we design 10 fluency tests and 8 consistency tests. For brevity, we only discuss a subset of them in this section, which are listed in Table 7. The tests can generally be applied to all three tasks with a few exceptions (detailed in Appendix I).
Most tests involve a hyper-parameter influencing the amount of noise added. This enables us to test how the metric behaves as we induce different levels of noise. To quantify the noise level, we define noise-ratio, based on the Levenshtein distance:
where is the set of gold hypotheses, and is the noised hypothesis. We employ the noise-ratio as a crude proxy to quantify the amount of noise across different noise types.One shortcoming of the Levenshtein distance is that it does not allow the switching operation. Therefore, for switching-based noise types, we divide the noise-ratio by 2. For more details on the setup, please see Appendix I.
For each noise type, a robust metric should give monotonically decreasing scores with an increasing noise-ratio. We claim a metric fails the test if it deviates from this expectation.
5.2 Results
Results for a subset of metrics/tests are shown in Figure 4. Unsurprisingly, most tests are passed by the metrics. However, the truncation and sentence switching tests give striking results. We will focus on these two tests here, and defer more complete results and discussion to Appendix I.
A number of popular metrics fail the truncation test, including (some variants of) BARTScore, BERTScore, ROUGE, COMET, PRISM, UniEval, and MAUVE (Some figures are deferred to Appendix I), spanning across CNNDM, TED-MT, and WikiText datasets. This is undesirable because truncation not only makes the hypothesis disfluent but also causes a serious loss of information.
The analysis in Figure 5 offers an insight into the reason behind, where the values of three variants of BERTScore under the truncation test are plotted. We observe that precision increases with more truncation, canceling out the decrease in recall and leading to a non-decreasing f-measure. We conjecture that this happens due to the property of the dataset, where earlier parts of different summaries (of the same article) are more likely to overlap than the rear spans. In Figure 8 (Appendix I), we show a similar observation for BARTScore-para.
In comparison, all metrics pass the truncation test for WMT. We believe the reason is that in the WMT data, the gold hypothesis and the reference are highly similar (They mostly only differ by a few tokens). Therefore, it would be easier for the metrics to catch the loss of information.
Two metrics fail the sentence switching test: BARTScore-para-recall (Figure 12), and MAUVE-GPT2/RoBERTa (Figure 14). This result is more striking for MAUVE, as the hypotheses in WikiText typically contain a number of sentences, and the temporal or logical order is seriously disturbed by sentence switching (examples in Table 20, Appendix I). Note that considering the positioned error test of MAUVE, for the WikiText data, we intentionally do not switch the last sentence of the hypothesis paragraph.
Interestingly, MAUVE-ELECTRA passes sentence switching and other tests. We surmise this is due to the discriminative training of ELECTRA, making it sensitive to errors in the text. We also find that MAUVE-ELECTRA performs best in a human correlation evaluation (Appendix C.3). Therefore, within the scope of this work, ELECTRA is the best-performing feature for MAUVE. Appendix I contains more analysis on sentence switching.
However, also shown in Figure 4, MAUVE-ELECTRA penalizes some error types more drastically (e.g., article/preposition removal) compared to other metrics, which means it may benefit from some further calibration, and we leave it as future work.
Undesirable behaviors from the truncation test suggest that practitioners should either report all of the precision, recall, and f-measure for a complete picture or calibrate the f-measure to put more weight on recall than on precision.
The sentence switching test shows MAUVE-RoBERTa’s insensitivity to the temporal/logical disorder. We suggest use MAUVE-RoBERTa in combination with GPT-PPL.
Discussion
To save space, the copy-source test is deferred to Appendix D because its results are relatively unsurprising. We also defer the repetition test to Appendix E, as it is motivated by the well-known degeneration problem (Holtzman et al., 2020).
The tests we design rely on some level of understanding of the PLMs, or a detailed examination of the metric definitions. A natural next question is whether we can automate this process. As a case study, we focus on BERTScore and build a toy example, showing that one can design an adversarial attack algorithm (Cheng et al., 2018) to detect sample-level anomaly. We defer it to Appendix J.
We devote the rest of this section to prevent potential misunderstandings since this work contains negative results.
The results in this work should be regarded as complementary to the impressive human correlation results in the literature. For example, BLEU passes all our tests in translation, however, it is outperformed by PLM-based metrics in human correlation evaluations (Zhang et al., 2020). If a metric fails one of our tests, it only means the metric needs improvement on that particular aspect. Our main message is not to discourage the use of PLM-based metrics, nor to devalue existing work by metric developers or users. Instead, we suggest use the metrics with caution and with awareness of the blind spots.
While we have covered a large variety of stress tests in this work and we encourage future metric developers to use them for robustness analysis, the set is not exhaustive. Even if a metric passes all our tests, it does not guarantee that the metric is blind-spot-free. We also encourage developers to come up with novel tests targeting certain underlying property of their proposed metric (e.g., the positioned error test we design for MAUVE).
Related Work
In comparison to the vast literature on NLG metric development or benchmarking (Mathur et al., 2020; Celikyilmaz et al., 2020; Gehrmann et al., 2021; Kasai et al., 2022b; Hämäläinen and Alnajjar, 2021), the robustness analysis of PLM-based metrics is an under-explored area, where exisiting work focused on a relatively small subset of metrics or a limited definition of robustness. For example, Vu et al. (2022) explored BERTScore’s performance variation with changes in representation space and character perturbations. Kaster et al. (2021) propose a regression-based global explainability technique to disentangle metric scores along linguistic factors.
More related to our work, Hanna and Bojar (2021) conducted a fine-grained analysis of BERTScore on different error types. Caglayan et al. (2020) discussed some curious phenomena for a range of metrics. Chen et al. (2021) conducted diagnostic tests for factuality metrics with synthesized errors. Sun et al. (2022) found that some metrics are not robust to dialects. In comparison, this work is more comprehensive in that the design of our tests are inspired by a wider range of motivations, e.g., the properties of the underlying PLMs.
The use of synthetic data has been proven to be a powerful tool to analyze the capabilities of NLP models in tasks including natural language inference McCoy et al. (2019); Naik et al. (2018), question answering Ribeiro et al. (2019), reading comprehension Sugawara et al. (2020) and text classification Prabhakaran et al. (2019). Ribeiro et al. (2020) proposed a task-agnostic methodology, which synthesizes a large number of examinations for NLP models. Ruder et al. (2021) subsequently extended this methodology to a multilingual setting. Goel et al. (2021) built a more complete model evaluation system by integrating subpopulations, transformations, evaluation sets, and adversarial attacks. This work follows the same high-level spirit, while our focus is on NLG metrics.
This work takes inspiration from research analyzing the behavior of PLM’s representations (Belinkov and Glass, 2019). Masked LMs such as BERT have been shown to be insensitive to word order (Pham et al., 2021), negation (Ettinger, 2020), and named entities (Balasubramanian et al., 2020). GPT-like models were shown to prefer repetitive text (Holtzman et al., 2020). Staliūnaitė and Iacobacci (2020) studies what types of linguistic knowledge BERT acquires with a focus on compositional and lexical semantics. There are also important lines of work on layer representation probing (Belinkov, 2022), or attention analysis (Dong et al., 2021; Ji et al., 2022).
Conclusion
Using PLMs for NLG metrics is a double-edged sword. While the metrics benefit from the models’ powerful representations, their black-box nature may cause unexpected behavior. This work shows that stress tests, complementary to the standard human correlation tests, are powerful tools to cover corner cases, detect the metrics’ blind spots, and point out aspects where the metric could improve.
As a major implication for metric users, we suggest using combinations of metrics so that they can cover each other’s blind spots. While this has been an existing practice for a majority of work in the field, our results on the blind spots provide an explicit empirical argument for its importance. While we are still positive about the future of using PLM for NLG metrics, we call for more caution and awareness of potential blind spots from both metric users and developers. More generally speaking, a deeper understanding of the PLMs is in need.
Limitations
We have primarily focused our analysis on similarity or log-probability based metrics for NLG. There are other important and interesting metrics that future work could examine. For example, Deng et al. (2021) developed a family of interpretable metrics for various NLG tasks with the concept of information alignment. Xu et al. (2022) recently proposed a metric based on stratified error synthesis. In addition, there are several task-specific metrics for paraphrase generation (Shen et al., 2022), image captioning (Hessel et al., 2021; Kasai et al., 2022a), dialogue (Mehri and Eskenazi, 2020), controlled text generation (Ke et al., 2022), etc., which would be interesting to evaluate.
In §5.5, we design a number of fluency and consistency tests. It would be interesting to expand this set to be broader or more sophisticated (Ng et al., 2014). Also, there are other important aspects of text generation to consider, such as factuality (Wang et al., 2020; Pagnoni et al., 2021).
All of our diagnostic data are synthetically created. While it provides valuable insights on the metric’s behavior, it does not have a good coverage of errors in real-world settings. Expanding our analysis to real-world errors in a scalable way would be an important future direction.
Last but not least, we evaluate our proposed stress tests only on English texts. However, many language-specific properties can induce potential blind spots for metrics, especially for low-resource languages (Haddow et al., 2022) where PLMs may provide poor text representations. An important future direction is expanding the tests to multilingual settings (Thompson and Post, 2020; Pires et al., 2019).
Ethics Statement
Although the goal of our study is for more reliable evaluation, there is a risk of dual use of our tests: We investigate stress tests to identify blind spots in existing generation metrics, but a subset of the approaches (e.g., copy-source or injection) could be used for cheating in an evaluation. By an explicit discussion of how these blind spots can be utilized, we hope to increase awareness in the community of scenarios in which the metrics are not perfect and could be manipulated. Towards mitigating the risks, we have discussed countermeasures that can be adopted to cover or detect such blind spots.
Acknowledgements
We sincerely thank Jungo Kasai and Xiaochuang Han for useful discussions. This material is based upon work supported by the DARPA CMO under Contract No. HR001120C0124, by the National Science Foundation (NSF) under Grants No. IIS2203097, IIS2125201, IIS2040926, and NSF CAREER Grant No. IIS2142739. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the funding agencies.
References
Supplemental Materials
Appendix A Implementation Details of Metrics or Tests
The high-level motivation for MLM-PPL (Salazar et al., 2020) is using a bidirectional masked language model to compute a quantity similar to next-token perplexity in autoregressive models, by masking candidate tokens one by one and obtaining perplexity from masked token log probability. We follow a similar formulation of the “pseudo-perplexity” in Salazar et al. (2020). Given a sequence , we replace a token with the mask token [M], and predict it using all past and future tokens . Let denote the conditional log probability of predicting each token given its context. MLM-PPL is defined as below:
We use the default hyperparameter settings recommended in Pillutla et al. (2021). is set for the scaling constant. For the quantization algorithm, we use -means with 500 iterations and clusters, where is the number of generations.
We now explain why we set the reference set to be different from the gold set. According to the definition of MAUVE, if we set the gold and ref set to be exactly the same, then the score for the gold set will be 1.0 (full-score). In this setting, any stress test will be passed because the score of the perturbed set can only be lower. Since MAUVE is a distribution-based metric, in principle it is enough to ensure that the ref set is from the data distribution.
As suggested by Zhang et al. (2020), the f-measure variant of BERTScore is used for translation. However, the paper does not have recommendations for summarization. Therefore we test all three variants (precision, recall, f-measure).
As introduced in Yuan et al. (2021), BARTScore has four variants to tackle different scenarios, and each variant defines a pair of input-output for BART: precision (reference to hypothesis), recall (hypothesis to reference), f-measure, and faithfulness (source to hypothesis).
As suggested by the paper, for translation we use the f-measure. However, for summarization, the recommendations are a bit vague. In the main sections, we mainly report the faithfulness variant as it is used by the paper for the SummEval dataset (which is based on CNNDM). We also test the other three variants and defer their results to the appendix.
In addition to BARTScore-cnn and BARTScore-para, BARTScore also has a prompted modeling option which we currently do not have the capacity to test. We leave it as future work.
Following common practice, we use the f-measure of ROUGE-2 or ROUGE-L.
Our test code for translation or summarization is built upon the released code from BARTScore.https://github.com/neulab/BARTScore. We also benefit from the Hugging Face library (Wolf et al., 2020).https://github.com/huggingface/transformers. Some fluency and consistency tests are built using the spaCy library.https://github.com/explosion/spaCy. For the negation test, we utilize released code from the NLP CheckList (Ribeiro et al., 2020).https://github.com/marcotcr/checklist.
Appendix B More Information on Datasets
We find it hard to locate a public MT dataset satisfying: (1) Each sample has multiple references. (2) Each sample contains multiple sentences. Therefore, we decide to manually build one.
We build a paragraph-level translation dataset based on the Zh-En part of the Multitarget TED Talks Task (MTTT) (Duh, 2018). The original dataset contains consecutive sentences in a TED talk. We first manually form 100 coherent paragraphs by selecting spans of samples in the test and dev splits. Each paragraph contains at least 4 sentences and at most 10 sentences. Correspondingly, the English reference of the paragraph is the concatenation of the reference of each sentence.
One additional translation for each sample is needed. Two graduate students who are fluent in both English and Chinese help provide one additional translation for each paragraph. Each translator handles 50 samples. And then the translations are switched so that they can correct each other’s errors. An example is given in Table 19. In our experiments, the original reference is set to be the the gold hypothesis, and the added translation is used as reference for the metrics.
We will make this dataset available in the public version of this manuscript.
B.2 WikiText Preprocessing
For the gold/reference hypotheses of the WikiText-103 dataset, we sample paragraphs with more than 256 tokens and conduct preprocessing to clean up dataset artifacts and special symbols. First, we trim extra space around {’.’, ’,’, ’?’, ’!’, ’:’, ’;’, ’(’, ’)’, ”’s”, ’%’}. Next, we remove the special token ’@’ in the dot ’@.@’ and hyphen ’@-@’ tokens. We also remove extra space around quotation marks. Finally, the text is truncated to the last full sentence under a total length of 256, which is to ensure the gold hypotheses are of similar length.
Appendix C Details on the Positioned Error Test
The full set of results for the positioned error test is shown in Table 8. MAUVE-GPT2 is insensitive to errors at the start and middle positions. In contrast, both MAUVE-RoBERTa and MAUVE-ELECTRA give significantly lower scores for erroneous text compared to the gold hypothesis. We also observe MAUVE-ELECTRA is more sensitive compared to MAUVE-RoBERTa.
C.2 Attention Pattern Analysis
Here we provide details about the attention pattern analysis. We input two random samples (non-cherry-picked) from the WikiText dataset to GPT2-large and RoBERTa-large and visualize the attention distribution over the relative position in the text. The sample is truncated to length 200 for the convenience of this analysis.
As shown in Figure 11, we average the attention distribution over all transformer layers and attention heads and then group 20 x 20 (attention-from and attention-to) tokens into one attention block for ease of presentation. We also include a high-granularity version where we group 2 x 2 tokens into one attention block.
C.3 MAUVE Correlation with Human Judgment
We reproduce MAUVE’s correlation with human judgment in Pillutla et al. (2021) on the three MAUVE variants based on GPT2, RoBERTa, and ELECTRA, on the WebText dataset with the released code.https://github.com/krishnap25/mauve-experiments. Note that Pillutla et al. (2021) only considered MAUVE-GPT2, and the correlation scores for the RoBERTa/ELECTRA variants were not tested.
We follow their pairwise setup of evaluation: Each annotator receives the prompt and continuation from two different generation settings and selects the setting that is favored using a 5-point Likert scale. The annotators are asked about three aspects: whether the continuation is human-like, interesting, or sensible. There are 8 generation settings that consist of different (model, decoding) choices specified in Table 9 plus human written continuations. We use their provided human annotation directly. Also following Pillutla et al. (2021), we convert the pairwise preference scores into rankings by fitting a Bradley-Terry model (Marden, 1995), and compute the Spearman rank correlation between the MAUVE score and the fitted Bradley-Terry coefficients. We refer readers to Pillutla et al. (2021) for more details.
The results are shown in Table 10.Due to the stochastic nature of sampling, our reproduced generation is not guaranteed to be the exact replication of the ones used in Pillutla et al. (2021), which is currently not released. As a result, we observe slightly different correlation numbers for MAUVE-GPT2 compared to Pillutla et al. (2021). Compared to MAUVE-GPT2, although MAUVE-RoBERTa is slightly superior in the “interesting” aspect, it has a lower correlation on the human-like judgment. Nevertheless, MAUVE-ELECTRA shows a clearly superior correlation with human judgment on all three aspects compared to both the GPT-2 and RoBERTa variants. It also performs best in our stress tests.
Appendix D The Copy-Source Test
A number of metrics are based on the similarity between the hypothesis and the reference or source. Therefore, for tasks like summarization and translation, one could try to fool the metric by simply submitting a direct copy of the source text. We term it the copy-source test.
As reported in Table 11, for both translation and summarization datasets, we find that COMET-QE, BERTScore-r, several variants of BARTScore, and UniEval-overall not just fail to account for this simple trick but in fact obtain higher scores than gold hypotheses.
We attribute these behaviors to some of the metrics’ design choices. (1) COMET-QE relies on a cross-lingual RoBERTa encoder, but it does not check the language ID of the hypothesis. (2) BARTScore, computed as a length-averaged log-likelihood, fails to account for the length of the hypothesis, which in this case is the entire source article. While removing the average operation is a natural remedy and indeed leads to a lower score for the noised hypothesis (shown by BARTS-cnn-noavg in the table), it is not ideal as it would also favor overly short summaries. (3) BERTScore-r’s behavior on summarization, on the other hand, is not surprising since it is recall-oriented, and is alleviated by using the f-measure. (4) The take on UniEval is more nuanced. Strictly speaking, the copied source does not degrade the four aspects UniEval reports. However, they lead to a misleadingly high overall score.
The copy-source trick could be used to manipulate scores in a contest. Straightforward solutions can counter this trick. For example, contest organizers can implement checks for similarity between submitted hypotheses and the source text and reject the matches. For summarization, it would be useful to check whether the length of the hypothesis is within the expected range. For translation, a language ID check is helpful.
Appendix E The Repetition Test
It is well-known that GPT-like LMs suffer from a repetition problem—they tend to assign high likelihood to repetitive text (Holtzman et al., 2020).
For the repetition test, we append to each gold hypothesis copies of its last 4-gram to create a synthetic repetition problem (termed as Rep-), with an example available in Table 12. For this test, a robust metric should give a lower score for Rep- compared to gold, because synthetic repetition degrades quality.
The experimental results for the repetition test are shown in Table 13. The repetition problem plagues a wider range of models than expected. In addition to GPT-PPL, we find BARTScore, and MLM-PPL (based on RoBERTa) also prefer repetitive text.
As an illustrated example of the repetition test, Figure 6 shows the per-timestep next-token probability of a 4-gram repetitive text in the WikiText dataset, given by GPT-PPL. The first repetition of the 4-gram “hard to miss.” has a slightly higher probability compared to the original ending. As this 4-gram is repeated more times, the probability given by GPT-PPL becomes increasingly higher.
For metric users, it has been an established practice (especially for open-ended generation) to report diversity metrics like rep-4gram (Welleck et al., 2020) or -gram entropy (Zhang et al., 2018), as shown in Table 13. For metric developers, our results indicate that the degeneration issue can not be ignored even if the LM is not autoregressive.
Appendix F Auxiliary Results for the Injection Test
Table 14 contains auxiliary results of the injection test for UniEval on the summarization task. We note several additional interesting observations: (1) If we omit “And yes, it is relevant.”, the relevent score gets lower. (2) If we change the tone from positive to negative, the scores get lower. (3) Just repeating “Yes” is not effective.
In the lower part of the table, we also observe that the injection hypothesis can drastically increase the score of a random (irrelevant) reference summary.
Appendix G Auxiliary Results for the Frequent n𝑛n-gram Test
An example if the frequent -gram sequence is available in Table 12.
In Table 15, results of frequent 4-gram and 3-gram tests are shown. We observe that it is easier for the frequent 4-grams to confuse the log-probability-based metrics. Per-timestep next-token probability plots for examples of a 4-gram and a 3-gram test are shown in Figure 3 and Figure 7, respectively. In both cases, there are high probability regions concentrated at the end of each -gram. For example, “the” in the 3-gram “side of the” gets a higher probability than the first two tokens, and “of” in the 4-gram “in the middle of” gets a higher probability than the first three tokens.
Appendix H Details on the Finetuning (Self-Evaluation)
For GPT-PPL, we finetune the GPT-2 generators on the WikiText-103 training set for 2 epochs, with a learning rate of 1e-05 and a batch size of 16.
For BARTScore, we finetune the BART or T5 models on the CNNDM training set for 2 epochs, with a learning rate of 1e-05 and a batch size of 8. Beam search with beam size 5 is used for decoding.
Appendix I Auxiliary Description and Results of the Fluency and Consistency Tests
More details on the setup: Most noise types involve randomness. For each hyper-parameter, we report mean and standard-deviation over five runs with different random seeds. For each noise type and task, we set the hyper-parameters so that the gaps of noise-ratio between test points are close to or larger than 5%. The same set of random seeds and hyper-parameters are shared across all metrics.
The full set of tests is described by Table 16. For the detailed hyper-parameter setting, please refer to our to-be-released code.
In general, the tests can be applied to all three tasks. But there are exceptions due to the properties of the dataset: (1) We do not apply BERT-diverge to the WikiText data, as the task’s nature is open-ended. (2) We can not apply sentence switching to WMT as most samples only contain one sentence. (3) Due to similar reasons, we do not apply verb or named entity switching and sentence replacement to WMT. (4) Similarly, we do not apply named entity switching or generic named entity to TED-MT.
Compared to other tests, BERT-diverge is special in that its noise is generated automatically by an MLM, which is an interesting future direction for metric stress tests. One disadvantage of this approach is that we do not have a 100% guarantee that the perturbed hypothesis is indeed “diverged”. However, we do not observe empirical evidence of this weakness in the quantitative (Most metrics drop drastically with this noise) or qualitative examination.
The complete results for the fluency and consistency tests are shown in Figure 14 for open-ended generation, Figure 12 for summarization, and Figure 15/ Figure 16 for translation. For visibility, we plot fluency test and consistency tests separately for each metric. Failed tests are highlighted as bold lines.
We now discuss some interesting results which are not included in the main section.
For open-ended generation, both variants of MAUVE (-GPT2/-RoBERTa) fail the sentence switching test. Although MLM-PPL does not fail the test in terms of rank, the slope of the sentence switching curve is relatively much flatter than the other noise types, indicating an insensitivity.
Interestingly, while MAUVE-RoBERTa is robust to truncation, MAUVE-GPT2 only penalizes truncation in a binary manner. The score is much lower than gold for the first level of noise, but remains basically the same for other levels compared to the first level. This implies the GPT2 feature is not sensitive to the amount of information loss, which is problematic. From insights of the attention analysis (§5.1), we also attribute this to the locality of GPT2 embedding.
GPT-PPL and MLM-PPL are robust to truncation, but only penalize this error minimally as shown by the relatively flat slope of their truncation curves, which is not ideal.
For summarization, BARTScore-cnn/para-r fails a number of fluency tests involving stop-words, prepositions, etc. This suggests extra caution is needed when developing recall-orientated log-probability-based metrics.
ROUGE-2 and ROUGE-L fail the truncation and noised punctuation tests. ROUGE-2 also has a very marginal decrease in sentence switching, which is also undesirable.
Interestingly, BERT-diverge with COMET-QE is the only failure case for WMT (The same set of BERT-diverge noise is shared across metrics). A few examples are given in Table 17. We observe that the semantics of the hypotheses are clearly diverged, however, the scores from COMET-QE do not drop.
In addition, COMET-QE also fails article removal on summarization, while the reference-based COMET is more robust.
In Figure 8, we show how different variants of BARTScore-para behave under the truncation test. We also observe that the recall variant behaves well, while the precision and faithful variants are confused. But, BARTScore-para-recall fails the sentence switching test. Therefore, we recommend reporting the recall variant in combination with other variants.
In Figure 9, we test switching different units of the hypothesis. Interestingly, MAUVE-GPT2/RoBERTa drops drastically for all other types of units.We use {’,’,’.’,’?’,’!’} to deliminate sub-sentences.
Appendix J Can We Automate the Detection?
The tests we design rely on various intuitions including some level of understanding of the underlying PLM’s behavior, or a detailed examination of the metric definitions. A natural next question is whether we can automate this process. Ideally, we would like an algorithm to search for a noising transformation function of gold hypotheses that fools the targeted metric, while inducing perturbations visible to humans.
As a case study, we focus on BERTScore-f and build a toy example using a discrete-space adversarial attack algorithm (Cheng et al., 2018; Li et al., 2020; He and Glass, 2019) on WMT. Although it is only a preliminary attempt toward the ideal goal, the results show that it could be an interesting future direction.
On the high level, we design an enumeration-based algorithm that iteratively and greedily perturbs the hypothesis. Given a gold hypothesis and source text , the goal is to find a perturbed hypothesis that maximizes ,The notation means that is inputted as the hypothesis, and is inputted as the reference. subject to the noise-ratio being larger than a pre-specified value. i.e., the objective is to find a that BERTScore thinks is similar to and aligns with the source . The reference translations are not involved in this search.
In each perturbation step, we try two operations for each token in the current hypothesis: (1) Delete this token. (2) Replace this token with a token in a candidate set (detailed in Appendix J.1). Then, we select and apply the operation that maximizes . This iteration is repeated until the desired noise-ratio is reached. One disadvantage of this approach is that we do not have a 100% guarantee that the perturbed hypothesis is indeed “bad” (this problem is not crucial considering that we start from the gold hypothesis). However, we do not observe empirical evidence of this weakness in the quantitative or qualitative examination.
Figure 10 quantitatively demonstrates the effectiveness of the algorithm. Compared to BERTScore, the perturbations induce a large drop in a number of other metrics, implying that the perturbation is breaking the fluency/consistency of the gold hypotheses. In the meantime, the drop in BERTScore is marginal, which aligns with the objective.
We then inspect perturbed samples with high scores under BERTScore, with some examples shown in Table 18. The situation is especially common in articles (e.g., substitution of a and an), numbers (including the offset of date and time) and pronouns (e.g., substitution of he, she, it and they). While these substitutions are detrimental, they are not penalized by BERTScore. Incidentally, these patterns are not covered by our checks in Section 5.5, which demonstrates the value of this study.
Inspired by this, we attempt to design general noise transformation rules based on the observations (e.g., pronoun switching), and apply them to the dataset for BERTScore. However, we find that these patterns do not generalize to the whole WMT dataset. One key reason is that the transformation is only effective in confusing BERTScore for a subset of the hypotheses, which might not be surprising due to the nature of the adversarial attack. We conclude that more research is needed to make this framework practical and we leave it to future work.
We fix the targeted LM as RoBERTa since BERTScore is based on it.
In our iterative perturbation algorithm, for a hypothesis , we enumerate each token in it, and design the following perturbations: (1) Delete the token. The perturbed hypothesis becomes , (2) Substitute the token. We build the candidate token set in two ways: (a) Use [MASK] to replace , and employ the masked RoBERTa model to generate possible tokens with the highest scores (similar to BERT-diverge). (b) Utilize the word embedding in RoBERTa to find the possible tokens closest to . And (Some relatively meaningless substitutions, such as punctuation and uppercase/lowercase replacement will be filtered). In this way, we can get perturbed hypotheses . In our experiments, we set both and to eight.