InfoLM: A New Metric to Evaluate Summarization & Data2Text Generation
Pierre Colombo, Chloe Clavel, Pablo Piantanida
Introduction
A plethora of applications of natural language processing (NLP) performs text-to-text transformation (Mellish and Dale 1998; Belz and Reiter 2006; Specia, Scarton, and Paetzold 2018). Given an input, these systems are required to produce an output text that is coherent, readable and informative. Due to both high annotation costs and time, researchers tend to rely on automatic evaluation to compare the outputs of such systems. Reference-based automatic evaluation relies on comparing a candidate text produced by the NLG system and one or multiple reference texts (‘gold standard’) created by a human annotator. Generic automatic evaluation of NLG is a huge challenge as it requires building a metric that evaluates the similarity between a candidate and one or several gold-standard reference texts. However, the definition of success criteria is task-specific: as an example, evaluation of text summarization focuses on content, coherence, grammatically, conciseness, and readability (Mani 2001; Colombo et al. 2022), whereas machine translation focuses on fidelity, fluency and adequacy of the translation (Hovy 1999) and data2text generation (Gardent et al. 2017) consider criteria such as data coverage, correctness and text structure.
Automatic text evaluation metrics fall into two categories: metrics that are trained to maximise their correlations using human annotation (e.g., BLEND (Ma et al. 2017)) and untrained metrics (e.g., BLEU (Papineni et al. 2002), ROUGE (Lin 2004), BERTSCORE (Zhang et al. 2019), DepthScore (Staerman et al. 2021), BaryScore (Colombo et al. 2021c), MOVERSCORE (Zhao et al. 2019)). In this work, we focus on untrained metrics as trained metrics may not generalize well to new data (existing labelled corpora are of small size). Two categories of untrained metrics can be distinguished: word or character based-metrics that compute a score based on string representation and embedding-based metrics that rely on a continuous representation. String-based metrics (e.g., BLEU) often fail to robustly match paraphrases (Reiter and Belz 2009) as they mainly focus on the surface form as opposed to embedding-based metrics relying on continuous representations.
In this paper, we introduce InfoLM a family of new untrained metrics to evaluate text summarization and data2text generation. At the highest level InfoLM key components include: (1) a pre-trained masked language model (PMLM) that is used to compute two discrete probability distributions over the vocabulary. They represent the probability of observing each token of the vocabulary given the candidate and the reference sentence, respectively. (2) A contrast function that is used to measure the dissimilarity between aforementioned probability distributions. InfoLM differs from existing BERT-based metrics (e.g. BERTSCORE, MOVERSCORE) as it directly relies on the PMLM which outputs discrete probability distributions. Thus InfoLM does neither require to arbitrarily select one or several specific layers ( e.g. BERTSCORE relies on the 9th layer for bert-base-uncased), nor involves selecting arbitrary aggregations technics (e.g. see MOVERSCORE). As InfoLM relies on statistics on tokens it can also be seen as a string-based metric. However, it does not suffer from common pitfalls of string-based metrics (e.g. synonyms, need of an exact string-match) as the PMLM also allows ones to assign a high score to paraphrases and to capture distant dependencies. Contributions Our contributions are summarized below: (1) A set of novel metrics to automatically evaluate summarization and data2text generation. In this work, we introduce InfoLM which overcomes the common pitfall of string matching metrics and does not require to select a layer, nor to rely on a arbitrary aggregation function. InfoLM combines a pre-trained model and a contrast function denoted by between two discrete probability distributions. We explore the use of different choices of contrast functions such as -divergences, distances or Fisher-Rao distances. (2) Tasks. First, we demonstrate on both summarization and data2text that InfoLM is better suited than concurrent metrics. A comparison is conducted, using multiple correlation measures with human judgment both at the text and system level. Second, we dissect InfoLM to better understand the relative importance of each component (e.g. calibration, sensibility to the change of information measures).
Problem Statement and Related Work
In this section, we start by introducing notations and formulate the problem of both evaluating text generation and metrics. Then, we identify and present the most relevant related work and the existing approaches for the studied tasks.
Evaluating evaluation metrics. To assess the relevance of an evaluation metric , correlation with human judgment is considered to be one of the most important criteria (Koehn 2009; Chatzikoumi 2020). Debate on the relative merits of different correlations for the evaluation of automatic metrics is ongoing but classical correlation measures are Pearson (Leusch et al. 2003), Spearman (Melamed, Green, and Turian 2003) or Kendall (Kendall 1938) tests. Two meta-evaluation strategies are commonly used: (1) text-level correlation or (2) system-level correlation. Formally, the text-level correlation is computed as follows:
where the latter are the vectors composed of the averaged scores assigned by and the human, respectively. For the significance analysis, we follow Graham and Baldwin (2014) and use a William test to validate a significant improvement for dependent correlations (Steiger 1980).
2 Existing Metrics
Two types of string-based metrics exist: N-Grams matching and Edit distance-based metrics. N-Grams matching metrics count the number of N-grams in common between the candidate text and the reference text. The three most-used metrics are BLEU, ROUGE and METEOR (Banerjee and Lavie 2005). If no N-gram is in common between the input text candidate and the reference, these metrics fail to produce meaningful scores. The second category of metrics gathers edit distance-based metrics. They measure the number of basic operations such as edit/delete/insert to measure semantic equivalence. Variants include TER (Snover et al. 2006), CDER (Leusch, Ueffing, and Ney 2006), EED (Stanchev, Wang, and Ney 2019), CHARACTER (Wang et al. 2016). Edit distance-based metrics do not handle synonyms and focus on surface form. InfoLM can be seen as string-based but do not suffer from the aforementioned matching problem and can handle synonyms as it relies on a PMLM.
Embedding-based Metrics
Another class of metrics relies on word embeddings. These metrics either use static words embeddings such as word2vec (Mikolov et al. 2013) or contextualized embeddings such as ELMO (Peters et al. 2018), BERT (Devlin et al. 2018) and its variants (Sanh et al. 2019; Liu et al. 2019). Among the most popular metrics, we can mention MOVERSCORE, BERTSCORE, WMD (Kusner et al. 2015) , WMDO (Chow, Specia, and Madhyastha 2019). Different from these approaches InfoLM relies on a language model and work with discrete probability distributions instead of continuous representations.
Learning-based Metrics
Various trained metrics have been proposed such as BEER, BEND, RUSE, CIDER (Vedantam, Lawrence Zitnick, and Parikh 2015). These methods rely on train/dev/test sets composed of human evaluations. InfoLM does not require any training step and relies on a frozen PMLM.
PMLM as a Metric.
To the best of our knowledge, using a PMLM (i.e, without further training) as a reference-based automatic metric remains overlooked. The closest use we found was to rely on autoregressive models, such as GPT-2 (Radford et al. 2019), to compute the generated sentence perplexity and assess its fluency. Researchers mainly focused on the use of the learnt embedding of the PMLM. However, it remains an open question to find a reliable layer aggregation mechanism (e.g, BERTSCORE arbitrary selects a layer based on a chosen dataset, MOVERSCORE uses the 5 last layers). InfoLM addresses this aggregation issue by relying on the PMLM.
3 Masked Language Modeling
Language models based on masked language pre-training objectives Devlin et al. (2018); Liu et al. (2019) aim at reconstructing a corrupt version of an input text by minimizing a cross-entropy loss. This corrupted context corresponds to a ”local view” view of the sentence. To ensure fair comparison, we do not use existing alternatives (e.g. GPT-2 based models (Radford et al. 2019)) as concurrent works (Zhang et al. 2019; Zhao et al. 2019) rely on PMLM.
InfoLM
In this section, we first introduce a novel family of metrics called InfoLM and then detail the different components of these novel metrics. Notations We denote by , the parameter of the PMLM, its temperature and an information measure (see 3.3) which quantifies the similarity between two discrete distributions.
A PMLM has learnt the empirical distribution of a large text corpus. Given a text , the corrupted context with a mask at position j is denoted , the LM predicts a distribution over the vocabulary given the masked context. As an example, for a masked input sentence , a pretrained model could place high probabilities on tokens “food”, “meal” and low probability on ”the”. It is worth noting that represents the probability of observing each token of the vocabulary given the masked input .
Given , two masked contexts , from input texts , , with masks at positions j and k respectively, are equivalent (denoted ) if the two predicted discrete distributions given by the PMLM, namely and , are similar. Formally, if .
We have the intuition that two similar sentences will share several pairs of equivalent masked contexts. At this point, we make no claim on the relationship between equivalence and the masked context similarity.
In this work, we make the hypothesis that two similar sentences will share multiple equivalent masked contexts. However, pairwise comparisons of all the pairs of individual masked contexts are prohibitively expensive ( comparisons) when considering long texts. Motivated by efficiency, we instead propose to work with two ”global views” of the sentences that are well-formed probability distributions and are obtained through the aggregation of individual PMLM predictions. Aggregated probabilities for and are denoted and respectively.
Given , two texts are said to be similar (denoted ) if where and denotes the aggregated individual masked context predictions.
2 InfoLM
InfoLM uses the notion of similarity given in 3.2. Given a reference text together with a candidate text , InfoLM recursively masks each token position of both and to obtain individual masked contexts. By relying on a PMLM, InfoLM predicts one distribution for each individual masked contexts. The resulting distributions are then averaged (we refer to this operation ”bag of distributions”) to obtain and . The final step involves comparing two well formed discrete probability distributions and through a measure of information . InfoLM writes as:
It is worth to emphasize that and are two well formed discrete probability distributions. They represent the probability of observing each token of the vocabulary given the candidate and the reference sentence, respectively.
Aggregation Procedure
Rare tokens can be more indicative of text similarity than common tokens (Banerjee and Lavie 2005). Thus, for the aggregation of the individual masked contexts, we propose to compute a weighted ”bag of distributions” where the weights are normalized measures of the importance of each token. In practice, and write as:
LM Calibration.
Modern deep neural networks are overconfident (Guo et al. 2017). To re-calibrate language models several techniques have been proposed (e.g temperature scaling (Platt et al. 1999)). Here, we choose to study how calibration affects InfoLM by relying on temperatureWhen one token receive all the probability mass, and when , the probability becomes uniform. scaling motivated by simplicity and speed.
3 Information Measures
In this work, we focus on comparing a pair of discrete probability distributions through information measures (see Basseville (2013) for an exhaustive study). We rely on two types of information measures: divergences and distances. The divergence is a measure of dissimilarity that is always positive or equal to zero if (and only if) the two considered distributions are strictly identical. We call distance, a function that is symmetric, positive, respects the triangle inequality and is equal to zero if (and only if) the two considered distributions are strictly identical. We will use information measures that belong to either Csiszar -divergences or that are distances.
Distances
Connection with String Matching Metrics.
We adopt the following notations and . First, let us consider two texts such that . InfoLM with is closed to if . It means that all likely tokens (according to the PMLM) when considering are also likely when considering . For string matching metrics, it corresponds to a perfect match between and . Second, let us consider such that (dissimilar texts) and a measure of information that relies on product of and (e.g Fisher-Rao). In this case, thus all likely tokens when considering are unlikely when considering (the converse it true as well). For string matching metrics this corresponds to no match among the sub-strings of and .
Experimental Frameworks
In this section, we describe our experimental setting. We present the tasks and the baselines metrics use for each task.
Text summarization aims at compressing long texts into fluent, short sentences that preserve the salient information. Datasets. To compare the different metrics previous work (Bhandari et al. 2020) either relies on the TAC datasets (Dang and Owczarzak 2008; McNamee and Dang 2009) or on new summarization datasets extracted from CNN/DailyMail (Nallapati et al. 2016). As pointed out by Peyrard (2019); Bhandari et al. (2020), TAC datasets are old and contain flaws (e.g systems used to generate summaries were of poor quality), we choose to work with the newly assemble dataset from CNN/Daily News proposed in Bhandari et al. (2020). This dataset gathers 11,490 summaries and annotations are carried using the pyramid method (Nenkova and Passonneau 2004). Metrics. For text summarization, perhaps the most known metrics are ROUGE and its extensions (Ng and Abrecht 2015), or METEOR and its variants (Denkowski and Lavie 2014). Recently, a new set of metrics (e.g BERTSCORE, MOVERSCORE) have been applied to text summarization.
2 Data2Text Generation
Prior works mainly rely on two task-oriented dialogue datasets (i.e., BAGEL (Mairesse et al. 2010), SFHOTEL (Wen et al. 2015)). As sentence generated in these data-sets are unlikely to be representative of the progress of recent NLG systems we instead rely on a different dataset coming from the WebNLG2020 challenge (Gardent et al. 2017). Given the following example of triple: (John_Blaha birthDate 1942_08_26) (John_Blaha birthPlace San_Antonio) (John_E_Blaha job Pilot) the goal is to generate John Blaha, born in San Antonio on 1942-08-26, worked as a pilot. Annotations. The WebNLG task is evaluated by human annotators along four different axes: (1) Data Coverage: Are all the descriptions presented in the data included in the text? (2) Relevance: Does the text contains only predicates found in the data? (3) Correctness: Are predicates found in the data correctly mentioned and adequately introduced? (4) Text structure: Is the produced text well-structured, grammatically correct and written in acceptable English?, (5) Fluency: Does the text progress naturally? Is it easy to understand? Is it a coherent whole? Metrics. For this task, organisers rely on untrained metrics (e.g. BLEU, METEOR, TER, BERTSCORE) to compare the performance of the candidate systems. Thus, we will focus on system-level correlation.
Numerical Results
In this section, we study the performance of InfoLM on both text summarization and data2text generation.
General Analysis. 1 gathers the results of the correlation study between scores produced by different metrics and human judgement (i.e. pyramid score). We can reproduce results from Bhandari et al. (2020). We observe a different behavior depending on the type of systems to be evaluated (e.g., abstractive or extractive) and the chosen correlation coefficient. We observe that InfoLM with or with outperforms other BERT-based metrics such as MOVERSCORE or BERTSCORE (e.g., it is worth noting that both metrics perform poorly at the text or system-level when considering outputs from extractive systems). largely outperforms n-gram matching metrics (e.g., ROUGE metrics) on all datasets when measuring correlation with the Kendall and in almost all configurations (except when considering abstractive outputs at the system level) when using the Pearson . It is worth noting the overall good performance of the parameter-free Fisher-Rao distance. Choice of information geometric measure for InfoLM. In 1, we can observe two different types of groups depending on the global behaviour. First we notice that using , leads to poor performances in many configurations. Good performance of in some configurations is surprising as is extremely selective (i.e. computes ). As output produced by the PMLM is sparse, correspond to one likely word in one sentence and not likely at all in the other. The second group gathers , , , and and achieves the best performance overall. and achieve similar performance suggesting that the flexibility (e.g., robustness to outliers) introduced by the parameter in is not useful in our task. This observation is strengthened by the lower performance of . The difference of results between the two measures is due to the flexibility introduced by (i.e., controls the relative importance of the ration ) which can be interpreted in our case as the ability to control the importance attributed to less likely words (Jalalzai *). Takeaways. The best performing metric is obtained with . The Fisher-Rao distance, denoted by , achieves good performance in many scenarios and has the advantage to be parameter-free.
2 Results on Data2Text
Global Analysis: 2 gathers results of the correlation analysis of the metrics with human judgements following the five different axes. We observe that the five considered criteria of annotations are not independent: text structure and fluency achieve a strong correlation coefficient (). Additionally, all metrics achieve similar results when the correlation is computed on these two criteria. We observe that the best performing group of metric is based on InfoLM followed by metrics based on continuous representation from BERT (i.e., MOVERSCORE and BERTSCORE) followed by N-gram matching metrics. Regarding correctness, data coverage and relevance, we observe that both and achieve the best results on almost all correlation coefficients. On data coverage, InfoLM achieves improvement up to points in correlation compared to both BERT based or N-gram matching metrics. Regarding fluency and text structure, Fisher-Rao distance works better and slightly outperforms the second-best performing metric, namely BERTSCORE. Takeaways. Similar to summarisation, we observe very low correlation for , . We also observe that -divergences achieve lower results than both and divergences suggesting that, as noticed for summarisation, robustness to unlikely words (i.e., outliers) is less relevant for our task.
3 Further Analysis
In this experiment, we complete our global analysis by comparing the scores obtained by the different metrics with each other. We want to gain an understanding of how different our metric is from other metrics and how the choice of information geometric measures affects the predictions. 2 gathers the results of the experiment. We observe a high correlation () between , , and Note that these metrics consider the product of and .. Interestingly, we observe a lower correlation () with BERTSCORE and N-gram matching metrics, e.g., ROOGE) whereas BERTSCORE achieves a stronger correlation with ROUGE (). Takeaways. Through the correlation analysis in 2, we observe the impact of different geometry on InfoLM predictions. The correlation analysis shows that the prediction of InfoLM when using , , and are highly correlated and as illustrated by previous experience achieve high correlation scores which we believe validate our approach. It is worth noting that requires no tuning as it is parameter free.
Score Distributions
In 3, we study the text score distribution of different metrics on abstractive summary. The ideal metric would mimic the human score distribution (i.e. ) and be able to distinguish between good and bad quality summaries. The results show that ROUGE and BLEU struggle to distinguish between between good quality () low quality summaries () which has been reported in Peyrard (2019). We observe that , and metrics are able to make the distinction. Interestingly, as increases and the distances become more selective (i.e. focus on one word solely), the distances struggle to distinguish low from high scoring summaries. Takeaways. InfoLM when combined with , and is able to distinguish high-scoring from low scoring summaries.
Temperature Calibration
To study the impact of calibration, we choose to work on system-level correlation and report in 4 the achieved the correlation measured by the different coefficients. We limit our study to the Fisher-Rao distance as it is a parameter-free metric and is among the best-performing metrics of InfoLM. Due to space constraints, we report the result on extractive systems only. Takeaways. Fisher-Rao only considers product thus as increases and the predicted probability of the PMLM becomes more uniform more words are considered and the aggregated distributions become richer in terms of considered words. It is worth noting that when changing the temperature we observe a smooth change in correlation and observe an optimal temperature which is reached for . It suggests that InfoLM benefits from a PMLM that is not too selective (case ). For a specific application, the temperature of InfoLM can be tuned to improve correlation and InfoLM will likely benefit from well-calibrated PMLM.
Summary and Concluding Remarks
In this work, we presented InfoLM that does not require training and it is among the first metrics computing the similarity between two discrete probability distributions over the vocabulary (which is similar to string-based metrics) but also leverages the recent advance in language modeling thanks to a PMLM. Our experiments on both summarization and data2text generation demonstrate the validity of our approach. Among available contrast measures, the Fisher-Rao distance is parameter-free and thus, it is easy to use in practice while the -Divergence achieves better results but requires to select and . Future work includes extending our metrics to new tasks such as SLU (Chapuis et al. 2020, 2021; Dinkar et al. 2020; Colombo, Clavel, and Piantanida 2021), controlled sentence generation (Colombo et al. 2019, 2021b) and multi-modal learning (Colombo et al. 2021a; Garcia et al. 2019).
Acknowledgments
This work was also granted access to the HPC resources of IDRIS under the allocation 2021-AP010611665 as well as under the project 2021-101838 made by GENCI.