A Pseudo-Metric between Probability Distributions based on Depth-Trimmed Regions
Guillaume Staerman, Pavlo Mozharovskyi, Pierre Colombo, Stéphan Clémençon, Florence d'Alché-Buc
INTRODUCTION
Metrics or pseudo-metrics between probability distributions have attracted a long-standing interest in information theory (Kullback,, 1959; Rényi,, 1961; Csiszàr,, 1963; Stummer and Vajda,, 2012), probability theory and statistics (Billingsley,, 1999; Sriperumbudur et al.,, 2012; Panaretos and Zemel,, 2019; Rachev,, 1991). While they serve many purposes in machine learning (Cha and Srihari,, 2002; MacKay,, 2003), they are of crucial importance in automatic evaluation of natural language generation (see e.g. Kusner et al.,, 2015; Zhang et al.,, 2019), especially when leveraging deep contextualized embeddings such as the popular BERT (Devlin et al.,, 2018). Yet designing a measure to compare two probability distributions is a challenging research field. This is certainly due to the inherent difficulty in capturing in a single measure typical desired properties such as: (i) metric or pseudo metric properties, (ii) invariance under specific geometric transformations, (iii) efficient computation, and (iv) robustness to contamination.
One can find in the literature a vast collection of discrepancies between probability distributions that rely on different principles. The -divergences (Csiszàr,, 1963) are defined as the weighted average by a well-chosen function of the odds ratio between the two distributions. They are widely used in statistical inference but they are by design ill-defined when the supports of both distributions do not overlap, which appears to be a significant limitation in many applications. IPMs (Sriperumbudur et al.,, 2012) are based on a variational definition of the metric, i.e. the maximum difference in expectation for both distributions calculated over a class of measurable functions and give rise to various metrics (Maximum Mean Discrepancy (MMD), Dudley’s metric, -Wassertein Distance) depending on the choice of this class. However, except in the case of MMD, which appears to enjoy a closed-form solution, the variational definition raises issues in computation. From the side of Optimal transport (OT) (see Villani,, 2003; Peyré and Cuturi,, 2019), the -Wasserstein distance is based on a ground metric able to take into account the geometry of the space on which the distributions are defined. Its ability to handle non-overlapping support and appealing theoretical properties make OT a powerful tool, mainly when applied to generative models (Arjovsky et al.,, 2017) or automatic text evaluation (Zhao et al.,, 2019; Clark et al.,, 2019; Colombo et al., 2021a, ).
This paper presents a new discrepancy measure between probability distributions, well-defined for non-overlapping supports, that leverages the interesting features of data depths. This measure is studied through the lens of the previously stated properties, yielding the contributions listed below.
A new discrepancy measure between probability distributions involving the upper-level sets of data depth is introduced. We show that this measure defines a pseudo-metric in general. Its good behavior regarding major transformation groups, as well as its ability to factor out translations, are depicted. Its robustness is investigated through the concept of finite sample breakdown point.
An efficient approximation of the depth-trimmed regions-based pseudo-metric is proposed for convex depth functions such as halfspace, projection and integrated rank-weighted depths. This approximation relies on a nice feature of the Hausdorff distance when computed between convex bodies.
The behavior of this algorithm regarding its parameters is studied through numerical experiments, which also highlight the by-design robustness of the depth-trimmed regions based pseudo-metric. Applications to robust clustering of images and automatic evaluation of natural language generation (NLG) show the benefits of this approach when benchmarked with state-of-the-art probability metrics.
BACKGROUND ON DATA DEPTH
It follows that depth regions are nested, i.e. for any . These depth regions generalize the notion of quantiles to a multivariate distribution.
A depth function’s relevance to capturing information about a distribution relies on the statistical properties it satisfies. Such properties have been thoroughly investigated in Liu, (1990); Zuo and Serfling, (2000) and Dyckerhoff, (2004) with slightly different sets of axioms (or postulates) to be satisfied by a proper depth function. In this paper, we restrict to convex depth functions (Dyckerhoff,, 2004) mainly motivated by recent algorithmic developments including theoretical results (Nagy et al.,, 2020) as well as implementation guidelines (Dyckerhoff et al.,, 2021).
The general formulation (1) opens the door to various possible definitions. While these differ in theoretical and practically related properties such as robustness or computational complexity (see Mosler and Mozharovskyi,, 2021 for a detailed discussion), several postulates have been developed throughout the recent decades the “good” depth function should satisfy. Formally, a function is called a convex depth function if it satisfies the following postulates:
(Vanishing at infinity) .
Projection depth, being a monotone transform of the Stahel-Donoho outlyingness (Donoho and Gasko,, 1992; Stahel,, 1981), is defined as follows:
A PSEUDO-METRIC BASED ON DEPTH-TRIMMED REGIONS
In the remainder of this paper, when the quantity will be associated with depth regions of , the second argument of the function will be omitted, for notation simplicity. It is worth mentioning that for any , since is a monotone decreasing function. Thus, is the smallest depth region with probability larger than or equal to and can be defined in an identical way as:
where . The strict inequalities in (3) and in the definition of eliminate cases where the supremum does not exist. Indeed, when , the depth region is then an infinitesimal set with a probability higher than zero. To the best of our knowledge, the supremum exists (without necessarily being unique) in the case of the halfspace depth (Rousseeuw and Rutz,, 1999) and the projection depth (Zuo,, 2003) under mild assumptions. Still, no results have been derived for AI-IRW depth yet. The set where each region probability mass is equal to then defines quantile regions of .
Let and , for all pairs in , the depth-trimmed regions () discrepancy measure between and is defined as
Data depths provide robustness to (4) together with the -trimming. Indeed, data depths such as the three previously introduced in Section 2 exhibit attractive robustness properties. The asymptotic breakdown point of the halfspace and the integrated rank-weighted medians are higher than . In contrast, the projection median is known to have a breakdown point equal to (Donoho and Gasko,, 1992; Ramsay et al.,, 2019).
When , the -Wasserstein distance enjoys an explicit expression involving quantile and distribution functions. Let be two random variables where are univariate probability distributions. Denoting by the quantile function of , the -Wasserstein distance can be written as
Since data depth and its central regions are extensions of cdf and quantiles to dimension , is then a possible (center-outward) generalization of (5) to higher dimensions. When is associated with the halfspace depth, a simple calculus (see Lemma A.3 in the Appendix for mathematical details) leads to
Thus, in general where the equality holds for symmetric distributions.
We now investigate to which extent the proposed discrepancy measure satisfies the metric axioms. As a first go, we show that fulfills most conditions. However, it does not define distance in general.
For any convex data depth, is positive, symmetric and satisfies triangular inequality but the entailment does not hold in general.
Thus, defines a pseudo-metric rather than a distance. Based on distance, the proposed discrepancy measure preserves isometry invariance, as stated in the following proposition.
where is the push-forward of by . In particular, it ensures invariance of under translations and rotations.
Although formulas (4) and (5) are based on the same spirit, there are no apparent reasons why the proposed pseudo-metric should have the same behavior as the Wasserstein distance. It is the purpose of Proposition 3.6 to investigate the ability to factor out translations, for associated with the halfspace depth, giving a positive answer for the case of two Gaussian distributions with equal covariance matrices.
Consider two random variables following and with expectations and variance-covariance matrices respectively. Denoting by the centered versions of , it holds:
Now, let and . Then it holds:
Following Proposition 3.6: when , one has for any and providing a closed-form expression in the Gaussian case.
2 Robustness
In this part, we explore the robustness of the proposed distance, associated with the halfspace depth, given the finite sample breakdown point (BP; Donoho,, 1982; Donoho and Hubert,, 1983). This notion investigates the smallest contamination fraction under which the estimation breaks down in the worst case. Considering a sample composed of i.i.d. observations drawn from a distribution with empirical measure , the finite sample breakdown point of w.r.t. , denoted by is defined as
For the halfspace depth function, for any such that , it holds:
Thus, at least a proportion of outliers must be added to break down when considering larger regions, while central regions are robust independently of . For two datasets, breaks down if depth regions for at least one of the datasets do. The breakdown point is then the minimum between the breakdown points of each dataset. However, the breakdown point considers the worst case, i.e. the supremum over all possible contaminations, and is often pessimistic. Indeed the proposed pseudo-metric can handle more outliers in certain cases, as experimentally illustrated in Section 5.
EFFICIENT APPROXIMATE COMPUTATION
As we shall see in Section 5, mutual approximation of by points from the sample and of by taking maximum over a finite set of directions allows for stable estimation quality. Recently, motivated by their numerous applications, many algorithms have been developed for the (exact and approximate) computation of data depths; see, e.g., Section 5 of Mosler and Mozharovskyi, (2021) for a recent overview. Depths satisfying the projection property (which also include halfspace and projection depth, see Dyckerhoff, (2004)) can be approximated by taking minimum over univariate depths; see e.g. Rousseeuw and Struyf, (1998); Chen et al., (2013); Liu and Zuo, (2014), Nagy et al., (2020) for theoretical guarantees, and Dyckerhoff et al., (2021) for an experimental validation. The case of AI-IRW is easier since the expectation in Equation 2 can be approximated through Monte-Carlo approximation.
NUMERICAL EXPERIMENTS
In this section, we first measure the quality of the approximation introduced in Section 4 and explore its dependency on the number of projections. Further, we present two studies on robustness of the proposed pseudo-metric to outliers. On synthetic datasets, we investigate how behaves under the presence of outliers using two different settings. On a real image dataset extracted from Fashion-MNIST where images are seen as bags of pixels, we evaluate the robustness of spectral clustering based on . Finally, we analyze the relevance of using as an evaluation metric in natural language generation to compare the empirical distributions of words of a pair of texts. Where applicable, we include state-of-the-art methods for comparison. Due to space limitations, experiments on the influence of the parameters and , as well as on the statistical rates, are deferred to the Appendix section.
Approximation error in terms of the number of projections. Proposition 3.6 allows to derive a closed form expression for when are Gaussian distributions with the same variance-covariance matrix. In order to investigate the quality of the approximation on light-tailed and heavy-tailed distributions, we focus on computing (with , , and using the halfspace depth) for varying number of random projections between a sample of 1000 points stemming from for and two different samples. These two samples are constructed from 1000 observations stemming from Gaussian and symmetrical Cauchy distributions, both with a center equal to . Comparison with the approximation of max Sliced-Wasserstein (max-SW; see e.g. Kolouri et al.,, 2019), which shares the same closed-form as , is also provided. Denoting by max- the Monte-Carlo approximation of the max-SW, the relative approximation errors, i.e., and (max-, are computed investigating both the quality of the approximation and the robustness of these discrepancy measures. Results that report the averaged approximation error and the 25-75% empirical quantile intervals are depicted in Figure 1. They show that possesses the same behavior as max-SW when considering Gaussians while it behaves advantageously for Cauchy distribution. Computation times are depicted in Figure 2, highlighting a constant-multiple improvement compared to the max-SW, which is already computationally fast.
Robustness to outliers. We analyze the robustness of by measuring its ability to overcome outliers (its robustness regarding the influence of the parameter are given in the Section D.4 in the Appendix). In this benchmark, we naturally include existing robust extensions of the Wasserstein distance: Subspace Robust Wasserstein (SRW; Paty and Cuturi,, 2019) searching for a maximal distance on lower-dimensional subspaces, ROBOT (Mukherjee et al.,, 2020) and RUOT (Balaji et al.,, 2020) being robust modifications of the unbalanced optimal transport (Chizat et al.,, 2018). Medians-of-Means Wasserstein (MoMW; Staerman et al., 2021a, ) that replaces the empirical means in the Kantorivich duality formulae by the robust mean estimator MoM (see e.g. Lecué and Lerasle,, 2020; Laforgue et al.,, 2021), is not employed due to high computational burden. Further, for completeness, we add the standard Wasserstein distance (W) and its approximation, the Sliced-Wasserstein (Sliced-W; Rabin et al.,, 2012) distance, with the same number of projections () as . Since the scales of the compared methods differ, relative error is used as a performance metric, i.e., the ratio of the absolute difference of the computed distance with and without anomalies divided by the latter. Two settings for a pair of distributions are addressed: (a) Fragmented hypercube precedently studied in Paty and Cuturi, (2019), where the source distribution is uniform in the hypercube and the target distribution is transformed from the source via the map where is taken element-wisely. Outliers are drawn uniformly from . (b) Two multivariate standard Gaussian distributions, one shifted by , with outliers drawn uniformly from . Our analysis is conducted over 500 sampled points from the distributions described above.
To investigate the robustness of , we consider the following values of : computed with the projection depth. Thus, data depths are computed on source and target distributions such that , , of data with lower depth values w.r.t. each distribution are not used in computation of , respectively. Figure 3, which plots the relative error depending on the portion of outliers varying up to , illustrates advantageous behavior of (for ) for reasonable (starting with ) contamination. It also confirms the pessimism of the breakdown point provided in Proposition 3.7 since (represented by the blue curve) shows robustness to at least 20 % of outliers.
(Robust) Clustering on bags of pixels. We demonstrate the relevance of the proposed pseudo-metric through an application to (robust) clustering. To that end, we perform spectral clustering (Shi and Malik,, 2000) on two datasets derived from Fashion-MNIST (FM). Each grayscale image is seen as a bag of pixels (Jebara,, 2003), i.e. as an empirical probability distribution over a 3-dimensional space (the two first dimensions indicate the pixel position and the third one, its intensity). The first dataset (FM) is constructed by taking the 100 first images in each class of the Fashion-MNIST dataset. The second dataset (Cont. FM), considered contaminated, is designed by introducing white patches on the left corner of 50 images drawn uniformly in the first dataset, which yields 5% of contamination. We benchmark (using the projection depth) setting and with the Wasserstein (W), the Sliced-Wasserstein (Sliced-W) and the Maximum Mean Discrepancy (MMD; Gretton et al.,, 2007) distances. and the Sliced-Wasserstein are approximated by Monte-Carlo using 100 directions while the MMD distance is computed using a Gaussian kernel with a bandwidth equal to 1. As a baseline method, spectral clustering is also applied to images considered as vectors using Euclidean distance. Standard parameters of the scikit-learn spectral clustering implementation are employed with a number of clusters fixed to . Performances of the benchmarked metrics are assessed by measuring the normalized mutual information (NMI; Shannon,, 1948) and the adjusted rank index (ARI; Hubert and Arabie,, 1985), which are standard clustering evaluation measures when the ground truth class labels are available. Results presented in Table 1 show that for both cases, i.e. with or without contamination, spectral clustering based on outperforms spectral clustering based on the other metrics.
We follow previous BERT-based metrics and evaluate performances of (with , and using the AI-IRW depth) on two different NLG tasks namely: data2text generation (using the WebNLG 2020 dataset Ferreira et al.,, 2020) and summarization. For the sake of place, summarization results and additional experimental details are reported in Section E in the Appendix. For WebNLG, we follow standard methods to assess the performance of NLG metrics (see e.g. Zhao et al.,, 2019). We compute the correlation with the following annotation scores: correctness, data coverage, and relevance. We report in Table 2 correlation results on the WebNLG task using Pearson (), Spearman () and Kendall () correlation coefficients. When performing a fair comparison between metrics, i.e. when , W, Sliced-W, MMD are directly used on the output of BERT, we observe that achieves the best results on all configurations. It is worth noting that also compares favorably against existing state-of-the-art NLG methods in many different scenarios and shows promising results.
DISCUSSION
Leveraging the notion of statistical data depth function, a novel pseudo-metric between multivariate probability distributions—that meets the aforementioned requirements—was introduced. The developed framework exhibits inherent versatility due to numerous data depth variants. The linear approximation algorithm and the robustness property make a promising tool for a large spectrum of applications beyond clustering and NLG, e.g. in generative adversarial networks (GANs) or information retrieval. Moreover, recent works extending the notion of data depth to further types of data such as functional and time-series data (Nieto-Reyes and Battey,, 2016; Gijbels and Nagy,, 2017), directional (or spherical) data (Ley et al.,, 2014), random matrices (Paindaveine and Van Bever,, 2018), curves (or paths) data (Lafaye et al.,, 2020), and random sets (Cascos et al.,, 2021) shall allow for the use of the proposed pseudo-metric for a wide range of applications.
References
Appendix A PRELIMINARY RESULTS
First, we introduce additional notations and recall some lemmas, used in the subsequent proofs.
Furthermore, if and are convex bodies (i.e. non empty compact convex sets), the Hausdorff distance can be reformulated with support functions of :
where .
A.2 Quantile regions
and the upper quantile set of
A.3 Auxiliary results
We now recall useful results, so as to characterize the halfspace depth regions.
Let , for any , it holds: .
Let and be two random variables where are univariate probability distributions. Denoting by the quantile function of , then the depth-trimmed region based pseudo-metric (associated with the halfspace depth) is defined as
and for any , its upper-level sets to intervals
Now, the quantile function can be explicitly derived as function of :
Following the same reasoning, it holds . Further, by change of variables
Combining Equation 8 and the Hausdorff distance definition recalled in Section A.1 lead to the result.
Appendix B TECHNICAL PROOFS
We now prove the main results stated in the paper.
B.2 Proof of Proposition 3.5
where () holds because any data depth satisfies (D1) by definition. Furthermore,
where () holds by virtue of hypothesis . Replacing it in (B.2) yields the desired results.
B.3 Proof of Proposition 3.6
Combining (B.3) and (B.3) lead to the desired result.
Introducing and using triangular inequality, subadditivity of the supremum and linearity of the integral, we obtain:
B.4 Proof of Proposition 3.7
For to break down at , it needs to have at least one trimmed-region that breaks down. Then the breakdown point of is higher than the minimum of the breakdown point of each region. Indeed, we have
Now applying Lemma 3.1 in Donoho and Gasko, (1992) and Theorem 4 in Nagy and Dvořák, (2021), a lower bound of the breakdown point of each halfspace region, for every , is given by
Appendix C APPROXIMATION ALGORITHMS
In this part, we display the approximation algorithms of the halfspace depth (see Algorithm 2), the projection depth (see Algorithm 3) and the AI-IRW depth (see Algorithm 4, proposed in Staerman et al., 2021b, ) used in the first step of the Algorithm 1.
Appendix D ADDITIONAL EXPERIMENTS
Figure 4, which plots a family of AI-IRW (using MCD estimator) depth induced trimmed-contours for a dataset contaminated with outliers, illustrates its robustness.
D.2 Illustration of the depth trimmed-regions based pseudo-Metric
Figure 5, which plots a family of (approximated) AI-IRW depth induced trimmed-regions for two datasets contaminated with outliers, illustrates the key idea of our proposed pseudo-metric.
D.3 Empirical analysis of statistical rates
Deriving theoretical finite-sample analysis may appear to be challenging for the proposed pseudo-metric. Thus, we numerically investigate the statistical convergence speed of . To that end, we simulate two samples and from two standard Gaussian distributions in dimension two with varying sample sizes. We compute the between and with using the halfspace and the projection depths. Our proposed metric is computed with a high number of directions to isolate the statistical error. We report the estimation error (averaged over 10 runs, the true value of being equal to zero) in Figure 6. The experiment suggests that the statistical rates should be in .
D.4 The influence of the parameter ε𝜀\varepsilon
The parameter plays the role of the robust tuning parameter of . In this part, we complete our theoretical results provided in Section 3.2. We assess the robustness of our pseudo-metric making varying the parameter . Precisely, we simulate two normal samples and from two standard Gaussian distributions in dimension two with a sample size of . From that, we construct abnormal samples with a proportion of anomalies equal to . To that end, we choose a proportion of normal samples and replace their first (for ) and second (for ) coordinates as follows: and where follows a uniform distribution on $DR_{2,\varepsilon}\mathbf{X}\mathbf{Y}DR_{2,\varepsilon}\varepsilon\varepsilon1\%DR_{p,\varepsilon}\varepsilon$ provides robustness to our pseudo-metric when combined with a robust depth function.
Proposition 3.6 allows to derive a closed form expression for when are Gaussian distributions with the same variance-covariance matrix. In order to investigate the quality of the approximation on light-tailed and heavy-tailed distributions, we focus on computing (with ) for varying number of between a sample of 1000 points stemming from for , drawn from the Wishart distribution (with parameters ()) on the space of definite matrices and three different samples (which yields nine settings). These three samples are constructed from 1000 observations stemming from elliptically symmetric Cauchy, Student- and Gaussian distributions all centered at . Results that report the averaged approximation error and the 25-75% empirical quantile intervals are depicted in Figure 8. They show that converges slowly for Cauchy with growing , while it converges with small for Gaussian and Student- distributions.
D.6 Robustness to outliers
Datasets on which experiments regarding "Robustness to outliers" in Section 5 have been performed are displayed in Figure 9.
Appendix E APPLICATIONS TO NLP
In this section, we gather details on experimental settings and additional results on the automatic evaluation of natural language generation (NLG).
Many metrics have been recently introduced for the automatic evaluation of text generation. In this work, we rely on untrained metrics. These metrics can be grouped into two categories: string-based metrics that depend on the string representation of the input texts to compute the similarity score and embedding-based metrics that rely on a continuous representation of the texts.
String matching metrics can be divided into two categories: N-gram matching and edit distance-based metrics. Perhaps the most used N-gram matching metrics are BLEU, ROUGE and METEOR. Edit distance-based metrics (e.g. TER; Snover et al.,, 2006) measure the distance as the number of basic operations such as ‘edit’/‘delete’/‘insert’. Variants of TER include CHARACTERE (Wang et al.,, 2016), CDER (Leusch et al.,, 2006), EED (Stanchev et al.,, 2019). String-based metrics fail to produce meaningful scores in the case of paraphrases, especially if no common n-grams are found between the candidate and the reference text.
The second category of untrained metrics (namely embedding-based metrics) achieves state-of-the-art performance in many NLG evaluation tasks and has been introduced to address the issues mentioned above. Originally introduced for the widely used words embedding (Garcia et al.,, 2019; Colombo et al.,, 2019, 2020; Colombo et al., 2021b, ) such as Word2Vec (Mikolov et al.,, 2013) or Glove (Pennington et al.,, 2014), this class of metrics has leveraged recently introduced contextualized word representations (CWR). CWR such as BERT, ELMO (Peters et al.,, 2018), HILAMOD (Chapuis et al.,, 2020, 2021) or ROBERTA (Liu et al., 2019b, ) are popular in NLP (Witon et al.,, 2018) as they achieve SOTA performance on many tasks. The two most popular metrics are MoverScore and BertScore.
E.2 Evaluation
For the task of evaluation of text generation, we assume that we have access to a dataset where represents the -th generated text by the -th natural generation system, and represents score assigned by the human annotatorIn practice an averaged score is considered as each sentence is annotated by 3 different annotators. The considered datasets directly provide the aggregated score. to , and is the reference text. is the number of available texts, and is the number of different systems.
To assess the relevance of an evaluation metric , the correlation with the human judgment is considered one of the most important criteria (Banerjee and Lavie,, 2005; Koehn,, 2009; Chatzikoumi,, 2020). To measure this correlation, two evaluation strategies are commonly adopted and built on top of a classical correlation measure, denoted , e.g. Kendall (; Kendall,, 1938), Pearson (; Leusch et al.,, 2003) or Spearman (; Melamed et al.,, 2003).
The text level correlation () measures the ability of the metric to distinguish between badly and well generated text. Formally, is defined as follows:
The system level correlation () assesses the ability of a metric to distinguish between good and bad systems. Formally, is defined as follows:
We refer the reader to Bhandari et al., (2020) for further details on the evaluation of text generation.
E.3 Results on Data2text
In this section, we gather further details and results on data2text automatic evaluation.
In WebNLG 2020, the goal is to create new efficient generation algorithms that can verbalise knowledge-based fragments. These algorithms are called Knowledge Base Verbalizers (Gardent et al.,, 2017) and are used during the micro-planning phase of NLG systems (Ferreira et al.,, 2018). WebNLG has been gathered to be more representative of the progress of recent NLG systems than previously existing task-oriented dialogue datasets (see e.g. SFHOTEL (Wen et al.,, 2015) and BAGEL (Mairesse et al.,, 2010)). As previously mentioned for the data2text task we work on the WebNLG2020 challenge (Gardent et al.,, 2017; Perez-Beltrachini et al.,, 2016). Data and system performances can be found in https://webnlg-challenge.loria.fr/. The task consists in mapping RDF triples to natural language (RDF format is used for many application including FOAF (see http://www.foaf-project.org/). For WebNLG 2020, the triplets are extracted from DBpedia (Auer et al.,, 2007). Data have been made freely available from the authors at https://gitlab.com/shimorina/webnlg-dataset/-/tree/master/release_v3.0. To compose this dataset, 15 systems (both symbolic and neural-based) have been used. The final dataset is composed of over 3k samples of human annotations (see https://webnlg-challenge.loria.fr/files/WebNLG-2020-Presentation.pdf for more details).
Example: Given the following triplet (John_Blaha birthDate 1942_08_26) (John_Blaha birthPlace San_Antonio) (John_Blaha job Pilot) the ground-truth reference is John Blaha, born in San Antonio on 1942-08-26, worked as a pilot.
E.3.2 Results
We gather in Table 3 complete results on the WebNLG tasks including results on ROUGE-2. To compare (with , , ) with the different metrics (i.e. Wasserstein, Sliced-Wasserstein, MMD), we work on Roberta-based model from the HuggingFace hub (Wolf et al.,, 2019) and extract representation from the 11th layer. From Table 3, we observe a similar behavior from BertScore and MoverScore. This similarity has also been reported in a different setting in the previous work of Zhao et al., (2019). Overall, we observe that is always among its group’s top-scoring metrics and achieves the best overall results on several configurations. It is worth noticing that only relies on information available in the candidate and the reference text. In contrast, BertScore and MoverScore use IDF information computed on every dataset.
E.4 Results on summarization
In this section, we gather experimental details and results on the automatic evaluation of the text summarization task.
Text summarization has attracted much attention in recent years (Zhang et al.,, 2020). Two types of models exist: extractive and abstractive. In extractive summarization, the system copies chunks of informative fragments from the input texts, whereas, in abstractive summarization, the system generates novel words. In this section, we describe our experimental setting. We present the tasks and the baseline metrics used for the automatic evaluation of summarization. We work with the dataset from Bhandari et al., (2020) for this task. This dataset has been introduced to solve several flaws (Rankel et al.,, 2013) present in existing summarization datasets such as TAC (Dang and Owczarzak,, 2008; McNamee and Dang,, 2009). The dataset has been annotated using the pyramid score (Nenkova et al.,, 2007; Nenkova and Passonneau,, 2004) and automatically built from the CNN/Daily News (Bhandari et al.,, 2020). It gathers 11 490 summaries coming from 11 extractive systems (See et al.,, 2017; Chen and Bansal,, 2018; Raffel et al.,, 2019; Gehrmann et al.,, 2018; Dong et al.,, 2019; Liu and Lapata,, 2019; Lewis et al.,, 2019; Yoon et al.,, 2020) and 14 abstractive systems (Zhou et al.,, 2018; Narayan et al.,, 2018; Kedzie et al.,, 2018; Zhong et al.,, 2019; Liu and Lapata,, 2019; Dong et al.,, 2019; Wang et al.,, 2020; Zhong et al.,, 2020).
Example: The goal is to assign a similarity score between a reference text: “Manchester United take on Manchester City on Sunday. Match will begin at 4 pm local time at United’s Old Trafford home. Police have no objections to kick-off being so late in the afternoon. Last late afternoon weekend kick-off in the Manchester derby saw 34 fans arrested at Wembley in 2011 fa cup semi-final” and the text generated by a NLG system: “Manchester Derby takes place at Old Trafford on Sunday afternoon police have no objections to the late afternoon kick-off both sides are challenging for a top-four spot in the Premier League the man in charge of patrolling the sell-out clash has no such fears”.
E.4.2 Results
We gather in Table 4, the results on the summarization task. We use a bert-based uncased model and rely on the representations extracted from the 9th layer (similarly to BertScore). For this experiment the following parameters are used: , , . For this task, we can reproduce results from Bhandari et al., (2020) where the different behavior regarding the extractive and the abstractive systems is also observed. In this experiment, we observe that can achieve stronger results than other metrics based on Wasserstein, Sliced-Wasserstein and MMD. We also observe that outperforms MoverScore and BertScore on extractive systems (on and ). We believe these results support our approach.