Evaluation of Text Generation: A Survey
Asli Celikyilmaz, Elizabeth Clark, Jianfeng Gao
Introduction
Natural language generation (NLG), a sub-field of natural language processing (NLP), deals with building software systems that can produce coherent and readable text (Reiter & Dale, 2000a) NLG is commonly considered a general term which encompasses a wide range of tasks that take a form of input (e.g., a structured input like a dataset or a table, a natural language prompt or even an image) and output a sequence of text that is coherent and understandable by humans. Hence, the field of NLG can be applied to a broad range of NLP tasks, such as generating responses to user questions in a chatbot, translating a sentence or a document from one language into another, offering suggestions to help write a story, or generating summaries of time-intensive data analysis.
The evaluation of NLG model output is challenging mainly because many NLG tasks are open-ended. For example, a dialog system can generate multiple plausible responses for the same user input. A document can be summarized in different ways. Therefore, human evaluation remains the gold standard for almost all NLG tasks. However, human evaluation is expensive, and researchers often resort to automatic metrics for quantifying day-to-day progress and for performing automatic system optimization. Recent advancements in deep learning have yielded tremendous improvements in many NLP tasks. This, in turn, presents a need for evaluating these deep neural network (DNN) models for NLG.
In this paper we provide a comprehensive survey of NLG evaluation methods with a focus on evaluating neural NLG systems. We group evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics that require no training, and (3) machine-learned metrics. For each category, we discuss the progress that has been made, the challenges still being faced, and proposals for new directions in NLG evaluation.
NLG is defined as the task of building software systems that can write (e.g., producing explanations, summaries, narratives, etc.) in English and other human languagesFrom Ehud Reiter’s Blog (Reiter, 2019).. Just as people communicate ideas through writing or speech, NLG systems are designed to produce natural language text or speech that conveys ideas to its readers in a clear and useful way. NLG systems have been used to generate text for many real-world applications, such as generating weather forecasts, carrying interactive conversations with humans in spoken dialog systems (chatbots), captioning images or visual scenes, translating text from one language to another, and generating stories and news articles.
NLG techniques range from simple template-based systems that generate natural language text using rules to machine-learned systems that have a complex understanding of human grammarFor an extensive survey on the evolution of NLG techniques, please refer to Gatt & Krahmer (2018).. The first generation of automatic NLG systems uses rule-based or data-driven pipeline methods. In their book, Reiter & Dale (2000b) presented a classical three-stage NLG architecture. The first stage is document planning, which determines the content and its order and generates a text plan outlining the structure of messages. The second is the micro-planning stage, when referring expressions that identify objects like entities or places are generated, along with the choice of words to be used and how they are aggregated. Collating similar sentences to improve readability with a natural flow also occurs in this stage. The last stage is realization, in which the actual text is generated, using linguistic knowledge about morphology, syntax, semantics, etc. Earlier work has focused on modeling discourse structures and learning representations of relations between text units for text generation (McKeown, 1985; Marcu, 1997; Ono et al., 1994; Stede & Umbach, 1998), for example using Rhetorical Structure Theory (Mann & Thompson, 1987) or Segmented Discourse Representation Theory (Asher & Lascarides, 2003). There is a large body of work that is based on template-based models and has used statistical methods to improve generation by introducing new methods such as sentence compression, reordering, lexical paraphrasing, and syntactic transformation, to name a few (Sporleder, 2005; Steinberger, 2006; Knight, 2000; Clarke & Lapata, 2008; Quirk et al., 2004).
These earlier text generation approaches and their extensions play an important role in the evolution of NLG research. Following this earlier work, an important direction that several NLG researchers have focused on is data-driven representation learning, which has gained attention with the availability of more data sources. Availability of large datasets, treebanks, corpora of referring expressions, as well as shared tasks have been beneficial in the progress of several NLG tasks today (Gkatzia et al., 2015; Gatt et al., 2007; Mairesse et al., 2010; Konstas & Lapata, 2013; Konstas & Lapara, 2012).
The last decade has witnessed a paradigm shift towards learning representations from large textual corpora in an unsupervised manner using deep neural network (DNN) models. Recent NLG models are built by training DNN models, typically on very large corpora of human-written texts. The paradigm shift starts with the use of recurrent neural networks (Graves, 2013) (e.g., long short-term memory networks (LSTM) (Hochreiter & Schmidhuber, 1997), gated recurrent units (GRUs) (Cho et al., 2014), etc.) for learning language representations, and later sequence-to-sequence learning (Sutskever et al., 2014), which opens up a new chapter characterised by the wide application of the encoder-decoder architecture. Although sequence-to-sequence models were originally developed for machine translation, they were soon shown to improve performance across many NLG tasks. These models’ weakness of capturing long-span dependencies in long word sequences motivated the development of attention networks (Bahdanau et al., 2015) and pointer networks (Vinyals et al., 2015). The Transformer architecture (Vaswani et al., 2017), which incorporates an encoder and a decoder, both implemented using the self-attention mechanism, is being adopted by new state-of-the-art NLG systems. There has been a large body of research in recent years that focuses on improving the performance of NLG using large-scale pre-trained language models for contextual word embeddings (Peters et al., 2018; Devlin et al., 2018; Sun et al., 2019; Dong et al., 2019), using better sampling methods to reduce degeneration in decoding (Zellers et al., 2019; Holtzman et al., 2020), and learning to generate tex t with better discourse structures and narrative flow (Yao et al., 2018; Fan et al., 2019b; Dathathri et al., 2020; Rashkin et al., 2020).
Neural models have been applied to many NLG tasks, which we will discuss in this paper, including:
summarization: common tasks include single or multi-document tasks, query-focused or generic summarization, and summarization of news, meetings, screen-plays, social blogs, etc.
machine translation: sentence- or document-level.
dialog response generation: goal-oriented or chit-chat dialog.
long text generation: most common tasks are story, news, or poem generation.
data-to-text generation: e.g., table summarization.
caption generation from non-textual input: input can be tables, images, or sequences of video frames (e.g., in visual storytelling), to name a few.
2 Why a Survey on Evaluation of Natural Language Generation
Text generation is a key component of language translation, chatbots, question answering, summarization, and several other applications that people interact with everyday. Building language models using traditional approaches is a complicated task that needs to take into account multiple aspects of language, including linguistic structure, grammar, word usage, and reasoning, and thus requires non-trivial data labeling efforts. Recently, Transformer-based neural language models have been shown to be very effective in leveraging large amounts of raw text corpora from online sources (such as Wikipedia, search results, blogs, Reddit posts, etc.). For example, one of most advanced neural language models, GPT-2/3 (Radford et al., 2019; Brown et al., 2020), can generate long texts that are almost indistinguishable from human-generated texts (Zellers et al., 2019; Brown et al., 2020). Empathetic social chatbots, such as XiaoIce (Zhou et al., 2020), seem to understand human dialog well and can generate interpersonal responses to establish long-term emotional connections with users.
Many NLG surveys have been published in the last few years (Gatt & Krahmer, 2018; Zhu et al., 2018; Zhang et al., 2019a). Others survey specific NLG tasks or NLG models, such as image captioning (Bernardi et al., 2016; Kilickaya et al., 2017; Hossain et al., 2018; Li et al., 2019; Bai & An, 2018), machine translation (Dabre et al., 2020; Han & Wong, 2016; Wong & Kit, 2019), summarization (Deriu et al., 2009; Shi et al., 2018), question generation (Pan et al., 2019), extractive key-phrase generation (Cano & Bojar, 2019), deep generative models (Pelsmaeker & Aziz, 2019; Kim et al., 2018), text-to-image synthesis (Agnese et al., 2020), and dialog response generation (Liu et al., 2016; Novikova et al., 2017; Deriu et al., 2019; Dusek et al., 2019; Gao et al., 2019), to name a few.
There are only a few published papers that review evaluation methods for specific NLG tasks, such as image captioning (Kilickaya et al., 2017), machine translation (Goutte, 2006), online review generation (Garbacea et al., 2019), interactive systems (Hastie & Belz, 2014a), and conversational dialog systems (Deriu et al., 2019), and for human-centric evaluations (Lee et al., 2019; Amidei et al., 2019b). The closest to our paper is the NLG survey paper of Gkatzia & Mahamood (2015), which includes a section on NLG evaluation metrics.
Different from this work, our survey is dedicated to NLG evaluation, with a focus on the evaluation metrics developed recently for neural text generation systems, and provides an in-depth analysis of existing metrics to-date. To the best of our knowledge, our paper is the most extensive and up-to-date survey on NLG evaluation.
3 Outline of The Survey
We review NLG evaluation methods in three categories in Sections 2-4:
Human-Centric Evaluation. The most natural way to evaluate the quality of a text generator is to involve humans as judges. Naive or expert subjects are asked to rate or compare texts generated by different NLG systems or to perform a Turing test (Turing, 1950) to distinguish machine-generated texts from human-generated texts.
Some human evaluations may require the judging of task-specific criteria (e.g., evaluating that certain entity names appear correctly in the text, such as in health report summarization), while other human evaluation criteria can be generalized for most text generation tasks (e.g., evaluating the fluency or grammar of the generated text).
Untrained Automatic Metrics. This category, also known as automatic metrics, is the most commonly used in the research community. These evaluation methods compare machine-generated texts to human-generated texts (reference texts) based on the same input data and use metrics that do not require machine learning but are simply based on string overlap, content overlap, string distance, or lexical diversity, such as -gram match and distributional similarity. For most NLG tasks, it is critical to select the right automatic metric that measures the aspects of the generated text that are consistent with the original design goals of the NLG system.
Machine-Learned Metrics. These metrics are often based on machine-learned models, which are used to measure the similarity between two machine-generated texts or between machine-generated and human-generated texts. These models can be viewed as digital judges that simulate human judges. We investigate the differences among these evaluations and shed light on the potential factors that contribute to these differences.
To see how these evaluation methods are applied in practice, we look at the role NLG shared tasks have played in NLG model evaluation (Section 5) and at how evaluation metrics are applied in two NLG subfields (Section 6): automatic document summarization and long-text generation. Lastly, we conclude the paper with future research directions for NLG evaluation (Section 7).
Human-Centric Evaluation Methods
Whether a system is generating an answer to a user’s query, a justification for a classification model’s decision, or a short story, the ultimate goal in NLG is to generate text that is valuable to people. For this reason, human evaluations are typically viewed as the most important form of evaluation for NLG systems and are held as the gold standard when developing new automatic metrics. Since automatic metrics still fall short of replicating human decisions (Reiter & Belz, 2009; Krahmer & Theune, 2010; Reiter, 2018), many NLG papers include some form of human evaluation. For example, Hashimoto et al. (2019) report that 20 out of 26 generation papers published at ACL2018 presented human evaluation results.
While human evaluations give the best insight into how well a model performs in a task, it is worth noting that human evaluations also pose several challenges. First, human evaluations can be expensive and time-consuming to run, especially for the tasks that require extensive domain expertise. While online crowd-sourcing platforms such as Amazon Mechanical Turk have enabled researchers to run experiments on a larger scale at a lower cost, they come with their own problems, such as maintaining quality control (Ipeirotis et al., 2010; Mitra et al., 2015). Second, even with a large group of annotators, there are some dimensions of generated text quality that are not well-suited to human evaluations, such as diversity (Hashimoto et al., 2019). Thirdly, there is a lack of consistency in how human evaluations are run, which prevents researchers from reproducing experiments and comparing results across systems. This inconsistency in evaluation methods is made worse by inconsistent reporting on methods; details on how the human evaluations were run are often incomplete or vague. For example, Lee et al. (2021) find that in a sample of NLG papers from ACL and INLG, only 57% of papers report the number of participants in their human evaluations.
In this section, we describe common approaches researchers take when evaluating generated text using only human judgments, grouped into intrinsic (§2.1) and extrinsic (§2.2) evaluations (Belz & Reiter, 2006). However, there are other ways to incorporate human subjects into the evaluation process, such as training models on human judgments, which will be discussed in Section 4.
An intrinsic evaluation asks people to evaluate the quality of generated text, either overall or along some specific dimension (e.g., fluency, coherence, correctness, etc.). This is typically done by generating several samples of text from a model and asking human evaluators to score their quality.
The simplest way to get this type of evaluation is to show the evaluators the generated texts one at a time and have them judge their quality individually. They are asked to vote whether the text is good or bad, or to make more fine-grained decisions by marking the quality along a Likert or sliding scale (see Figure 1(a)). However, judgments in this format can be inconsistent and comparing these results is not straightforward; Amidei et al. (2019b) find that analysis on NLG evaluations in this format is often done incorrectly or with little justification for the chosen methods.
To more directly compare a model’s output against baselines, model variants, or human-generated text, intrinsic evaluations can also be performed by having people choose which of two generated texts they prefer, or more generally, rank a set of generated texts. This comparative approach has been found to produce higher inter-annotator agreement (Callison-Burch et al., 2007) in some cases. However, while it captures models’ relative quality, it does not give a sense of the absolute quality of the generated text. One way to address this is to use a method like RankME (Novikova et al., 2018), which adds magnitude estimation (Bard et al., 1996) to the ranking task, asking evaluators to indicate how much better their chosen text is over the alternative(s) (see Figure 1(b)). Comparison-based approaches can become prohibitively costly (by requiring lots of head-to-head comparisons) or complex (by requiring participants to rank long lists of output) when there are many models to compare, though there are methods to help in these cases. For example, best-worst scaling (Louviere et al., 2015) has been used in NLG tasks (Kiritchenko & Mohammad, 2016; Koncel-Kedziorski et al., 2019) to simplify comparative evaluations; best-worst scaling asks participants to choose the best and worst elements from a set of candidates, a simpler task than fully ranking the set that still provides reliable results.
Almost all text generation tasks today are evaluated with intrinsic human evaluations. Machine translation is one of the text generation tasks in which intrinsic human evaluations have made a huge impact on the development of more reliable and accurate translation systems, as automatic metrics are validated through correlation with human judgments. One metric that is most commonly used to judge translated output by humans is measuring its adequacy, which is defined by the Linguistic Data Consortium as “how much of the meaning expressed in the gold-standard translation or source is also expressed in the target translation”https://catalog.ldc.upenn.edu/docs/LDC2003T17/TransAssess02.pdf. The annotators must be bilingual in both the source and target languages in order to judge whether the information is preserved across translation. Another dimension of text quality commonly considered in machine translation is fluency, which measures the quality of the generated text only (e.g., the target translated sentence), without taking the source into account. It accounts for criteria such as grammar, spelling, choice of words, and style. A typical scale used to measure fluency is based on the question “Is the language in the output fluent?”. Fluency is also adopted in several text generation tasks including document summarization (Celikyilmaz et al., 2018; Narayan et al., 2018), recipe generation (Bosselut et al., 2018), image captioning (Lan et al., 2017), video description generation (Park et al., 2018), and question generation (Du et al., 2017), to name a few.
While fluency and adequacy have become standard dimensions of human evaluation for machine translation, not all text generation tasks have an established set of dimensions that researchers use. Nevertheless, there are several dimensions that are common in human evaluations for generated text. As with adequacy, many of these dimensions focus on the content of the generated text. Factuality is important in tasks that require the generated text to accurately reflect facts described in the context. For example, in tasks like data-to-text generation or summarization, the information in the output should not contradict the information in the input data table or news article. This is a challenge to many neural NLG models, which are known to “hallucinate” information (Holtzman et al., 2020; Welleck et al., 2019); Maynez et al. (2020) find that over 70% of generated single-sentence summaries contained hallucinations, a finding that held across several different modeling approaches. Even if there is no explicit set of facts to adhere to, researchers may want to know how well the generated text follows rules of commonsense or how logical it is. For generation tasks that involve extending a text, researchers may ask evaluators to gauge the coherence or consistency of a text—how well it fits the provided context. For example, in story generation, do the same characters appear throughout the generated text, and do the sequence of actions make sense given the plot so far?
Other dimensions focus not on what the generated text is saying, but how it is being said. As with fluency, these dimensions can often be evaluated without showing evaluators any context. This can be something as basic as checking for simple language errors by asking evaluators to rate how grammatical the generated text is. It can also involve asking about the overall style, formality, or tone of the generated text, which is particularly important in style-transfer tasks or in multi-task settings. Hashimoto et al. (2019) ask evaluators about the typicality of generated text; in other words, how often do you expect to see text that looks like this? These dimensions may also focus on how efficiently the generated text communicates its point by asking evaluators how repetitive or redundant it is.
Note that while these dimensions are common, they may be referred to by other names, explained to evaluators in different terms, or measured in different ways (Lee et al., 2021). Howcroft et al. (2020) found that 2̃5% of generation papers published in the last twenty years failed to mention what the evaluation dimensions were, and less than half included definitions of these dimensions. More consistency in how user evaluations are run, especially for well-defined generation tasks, would be useful for producing comparable results and for focused efforts for improving performance in a given generation task. One way to enforce this consistency is by handing over the task of human evaluation from the individual researchers to an evaluation platform, usually run by people hosting a shared task or leaderboard. In this setting, researchers submit their models or model outputs to the evaluation platform, which organizes and runs all the human evaluations. For example, GENIE (Khashabi et al., 2021) and GEM (Gehrmann et al., 2021) both include standardized human evaluations for understanding models’ progress across several generation tasks. ChatEval is an evaluation platform for open-domain chatbots based on both human and automatic metrics (Sedoc et al., 2019), and TuringAdvice (Zellers et al., 2020) tests models’ language understanding capabilities by having people read and rate the models’ ability to generate advice.
Of course, as with all leaderboards and evaluation platforms, with uniformity and consistency come rigidity and the possibility of overfitting to the wrong objectives. Discussions of how to standardize human evaluations should take this into account. A person’s goal when producing text can be nuanced and diverse, and the ways of evaluating text should reflect that.
2 Extrinsic Evaluation
An extrinsic evaluation measures how successful the system is in a downstream task. Extrinsic evaluations are the most meaningful evaluation as they show how a system actually performs in a downstream setting, but they can also be expensive and difficult to run (Reiter & Belz, 2009). For this reason, intrinsic evaluations are more common than extrinsic evaluations (Gkatzia & Mahamood, 2015; van der Lee et al., 2019) and have become increasingly so, which van der Lee et al. (2019) attribute to a recent shift in focus on NLG subtasks rather than full systems.
An NLG system’s success can be measured from two different perspectives: a user’s success in a task and the system’s success in fulfilling its purpose (Hastie & Belz, 2014b). Extrinsic methods that measure a user’s success at a task look at what the user is able to take away from the system, e.g., improved decision making or higher comprehension accuracy (Gkatzia & Mahamood, 2015). For example, Young (1999), which Reiter & Belz (2009) point to as one of the first examples of extrinsic evaluation of generated text, evaluate automatically generated instructions by the number of mistakes subjects made when they followed them. System success-based extrinsic evaluations, on the other hand, measure an NLG system’s ability to complete the task for which it has been designed. For example, Reiter et al. (2003) generate personalized smoking cessation letters and report how many recipients actually gave up smoking. Post-editing, most often seen in machine translation (Aziz et al., 2012; Denkowski et al., 2014), can also be used to measure a system’s success by measuring how many changes a person makes to a machine-generated text.
Extrinsic human evaluations are commonly used in evaluating the performance of dialog systems (Deriu et al., 2019) and have made an impact on the development of the dialog modeling systems. Various approaches have been used to measure the system’s performance when talking to people, such as measuring the conversation length or asking people to rate the system. The feedback is collected by real users of the dialog system (Black et al., 2011; Lamel et al., 2000; Zhou et al., 2020) at the end of the conversation. The Alexa Prizehttps://developer.amazon.com/alexaprize follows a similar strategy by letting real users interact with operational systems and gathering the user feedback over a span of several months. However, the most commonly used human evaluations of dialog systems is still via crowdsourcing platforms such as Amazon Mechanical Turk (AMT) (Serban et al., 2016a; Peng et al., 2020; Li et al., 2020; Zhou et al., 2020). Jurcicek et al. (2011) suggest that using enough crowdsourced users can yield a good quality metric, which is also comparable to the human evaluations in which subjects interact with the system and evaluate afterwards.
3 The Evaluators
For many NLG evaluation tasks, no specific expertise is required of the evaluators other than a proficiency in the language of the generated text. This is especially true when fluency-related aspects of the generated text are the focus of the evaluation. Often, the target audience of an NLG system is broad, e.g., a summarization system may want to generate text for anyone who is interested in reading news articles or a chatbot needs to carry out a conversation with anyone who could access it. In these cases, human evaluations benefit from being performed on as wide a population as possible.
Evaluations can be performed either in-person or online. An in-person evaluation could simply be performed by the authors or a group of evaluators recruited by the researchers to come to the lab and participate in the study. The benefits of in-person evaluation are that it is easier to train and interact with participants, and that it is easier to get detailed feedback about the study and adapt it as needed. Researchers also have more certainty and control over who is participating in their study, which is especially important when trying to work with a more targeted set of evaluators. However, in-person studies can also be expensive and time-consuming to run. For these reasons, in-person evaluations tend to include fewer participants, and the set of people in proximity to the research group may not accurately reflect the full set of potential users of the system. In-person evaluations may also be more susceptible to response biases, adjusting their decisions to match what they believe to be the researchers’ preferences or expectations (Nichols & Maner, 2008; Orne, 1962).
To mitigate some of the drawbacks of in-person studies, online evaluations of generated texts have become increasingly popular. While researchers could independently recruit participants online to work on their tasks, it is common to use crowdsourcing platforms that have their own users whom researchers can recruit to participate in their task, either by paying them a fee (e.g., Amazon Mechanical Turkhttps://www.mturk.com/) or rewarding them by some other means (e.g., LabintheWildhttp://www.labinthewild.org/, which provides participants with personalized feedback or information based on their task results). These platforms allow researchers to perform large-scale evaluations in a time-efficient manner, and they are usually less expensive (or even free) to run. They also allow researchers to reach a wider range of evaluators than they would be able to recruit in-person (e.g., more geographical diversity). However, maintaining quality control online can be an issue (Ipeirotis et al., 2010; Oppenheimer et al., 2009), and the demographics of the evaluators may be heavily skewed depending on the user base of the platform (Difallah et al., 2018; Reinecke & Gajos, 2015). Furthermore, there may be a disconnect between what evaluators online being paid to complete a task would want out of an NLG system and what the people who would be using the end product would want.
Not all NLG evaluation tasks can be performed by any subset of speakers of a given language. Some tasks may not transfer well to platforms like Amazon Mechanical Turk where workers are more accustomed to dealing with large batches of microtasks. Specialized groups of evaluators can be useful when testing a system for a particular set of users, as in extrinsic evaluation settings. Researchers can recruit people who would be potential users of the system, e.g., students for educational tools or doctors for bioNLP systems. Other cases that may require more specialized human evaluation are projects where evaluator expertise is important for the task or when the source texts or the generated texts consist of long documents or a collection of documents. Consider the task of citation generation (Luu et al., 2020): given two scientific documents A and B, the task is to generate a sentence in document A that appropriately cites document B. To rate the generated citations, the evaluator must be able to read and understand two different scientific documents and have general expert knowledge about the style and conventions of academic writing. For these reasons, Luu et al. (2020) choose to run human evaluations with expert annotators (in this case, NLP researchers) rather than crowdworkers.
4 Inter-Evaluator Agreement
While evaluatorsIn text generation, ‘judges’ are also commonly used. often undergo training to standardize their evaluations, evaluating generated natural language will always include some degree of subjectivity. Evaluators may disagree in their ratings, and the level of disagreement can be a useful measure to researchers. High levels of inter-evaluator agreement generally mean that the task is well-defined and the differences in the generated text are consistently noticeable to evaluators, while low agreement can indicate a poorly defined task or that there are not reliable differences in the generated text.
Nevertheless, measures of inter-evaluator agreement are not frequently included in NLG papers. Only 18% of the 135 generation papers reviewed in Amidei et al. (2019a) include agreement analysis (though on a positive note, it was more common in the most recent papers they studied). When agreement measures are included, agreement is usually low in generated text evaluation tasks, lower than what is typically considered “acceptable” on most agreement scales (Amidei et al., 2018, 2019a). However, as Amidei et al. (2018) point out, given the richness and variety of natural language, pushing for the highest possible inter-annotator agreement may not be the right choice when it comes to NLG evaluation.
While there are many ways to capture the agreement between annotators (Banerjee et al., 1999), we highlight the most common approaches used in NLG evaluation. For an in-depth look at annotator agreement measures in natural language processing, refer to Artstein & Poesio (2008).
A simple way to measure agreement is to report the percent of cases in which the evaluators agree with each other. If you are evaluating a set of generated texts by having people assign a score to each text , then let be the agreement in the scores for (where if the evaluators agree and if they don’t). Then the percent agreement for the task is:
So means the evaluators did not agree on their scores for any generated text, while means they agreed on all of them.
However, while this is a common way people evaluate agreement in NLG evaluations (Amidei et al., 2019a), it does not take into account the fact that the evaluators may agree purely by chance, particularly in cases where the number of scoring categories are low or some scoring categories are much more likely than others (Artstein & Poesio, 2008). We need a more complex agreement measure to capture this.
4.2 Cohen’s κ𝜅\kappa
Cohen’s (Cohen, 1960) is an agreement measure that can capture evaluator agreements that may happen by chance. In addition to , we now consider , the probability that the evaluators agree by chance. So, for example, if two evaluators ( and ) are scoring texts with a score from the set , then would be the odds of them both scoring a text the same:
For Cohen’s , is estimated using the frequency with which the evaluator assigned each of the scores across the task.There are other related agreement measures, e.g., Scott’s (Scott, 1955), that only differ from Cohen’s in how to estimate . These are well described in Artstein & Poesio (2008), but we do not discuss these here as they are not commonly used for NLG evaluations (Amidei et al., 2019a). Once we have both and , Cohen’s can then be calculated as:
4.3 Fleiss’ κ𝜅\kappa
As seen in Equation 2, Cohen’s measures the agreement between two annotators, but often many evaluators have scored the generated texts, particularly in tasks that are run on crowdsourcing platforms. Fleiss’ (Fleiss, 1971) can measure agreement between multiple evaluators. This is done by still looking at how often pairs of evaluators agree, but now considering all possible pairs of evaluators. So now , which we defined earlier to be the agreement in the scores for a generated text , is calculated across all evaluator pairs:
Then we can once again define , the overall agreement probability, as it is defined in Equation 1—the average agreement across all the texts.
To calculate , we estimate the probability of a judgment by the frequency of the score across all annotators. So if is the proportion of judgments that assigned a score , then the likelihood of two annotators assigning score by chance is . Then our overall probability of chance agreement is:
With these values for and , we can use Equation 3 to calculate Fleiss’ .
4.4 Krippendorff’s α𝛼\alpha
Each of the above measures treats all evaluator disagreements as equally bad, but in some cases, we may wish to penalize some disagreements more harshly than others. Krippendorff’s (Krippendorff, 1970), which is technically a measure of evaluator disagreement rather than agreement, allows different levels of disagreement to be taken into account.Note that there are other measures that permit evaluator disagreements to be weighted differently. For example, weighted (Cohen, 1968) extends Cohen’s by adding weights to each possible pair of score assignments. In NLG evaluation, though, Krippendorff’s is the most common of these weighted measures; in the set of NLG papers surveyed in Amidei et al. (2019a), only one used weighted .
Like the measures above, we again use the frequency of evaluator agreements and the odds of them agreeing by chance. However, we will now state everything in terms of disagreement. First, we find the probability of disagreement across all the different possible score pairs , which are weighted by whatever value we assign the pair. So:
(Note that when , i.e., the pair of annotators agree, should be 0.)
Next, to calculate the expected disagreement, we make a similar assumption as in Fleiss’ : the random likelihood of an evaluator assigning a score can be estimated from the overall frequency of . If is the proportion of all evaluation pairs that assign scores and , then we can treat it as the probability of two evaluators assigning scores and to a generated text at random. So is now:
Finally, we can calculate Krippendorff’s as:
Untrained Automatic Evaluation Metrics
With the increase of the numbers of NLG applications and their benchmark datasets, the evaluation of NLG systems has become increasingly important. Arguably, humans can evaluate most of the generated text with little effortFor some domains that require domain knowledge (e.g., factual correctness of a scientific article) or background knowledge (e.g., knowledge of a movie or recipe) might be necessary to evaluate the generated recipe obeys the actual instructions if the instructions are not provided. However, human evaluation is costly and time-consuming to design and run, and more importantly, the results are not always repeatable (Belz & Reiter, 2006). Thus, automatic evaluation metrics are employed as an alternative in both developing new models and comparing them against state-of-the-art approaches. In this survey, we group automatic metrics into two categories: untrained automatic metrics that do not require training (this section), and machine-learned evaluation metrics that are based on machine-learned models (Section 4).
Untrained automatic metrics for NLG evaluation are used to measure the effectiveness of the models that generate text, such as in machine translation, image captioning, or question generation. These metrics compute a score that indicates the similarity (or dissimilarity) between an automatically generated text and human-written reference (gold standard) text. Untrained automatic evaluation metrics are fast and efficient and are widely used to quantify day-to-day progress of model development, e.g., comparing models trained with different hyperparameters. In this section we review untrained automatic metrics used in different NLG applications and briefly discuss the advantages and drawbacks of commonly used metrics. We group the untrained automatic evaluation methods, as in Table 1, into five categories:
grammatical feature based metricsNo specific metric is defined, mostly syntactic parsing methods are used as metric. See section 3.5 for more details.
We cluster some of these metrics in terms of different efficiency criteria (where applicable) in Table 2Recent work reports that with limited human studies most untrained automatic metrics have weaker correlation with human judgments and the correlation strengths would depend on the specific human’s evaluation criteria (Shimorina, 2021). The information in Table 2 relating to correlation with human judgments is obtained from the published work which we discuss in this section. We suggest the reader refer to the model-based evaluation metrics in the next section, in which we survey evaluation models that have reported tighter correlation with the human judgments on some of the evaluation criteria.. Throughout this section, we will provide a brief description of the selected untrained metrics as depicted in Table 1, discuss about how they are used in evaluating different text generation tasks and provide references for others for further read. We will highlight some of their strengths and weaknesses in bolded sentences.
n-gram overlap metrics are commonly used for evaluating NLG systems and measure the degree of “matching” between machine-generated and human-authored (ground-truth) texts. In this section we present several n-gram match features and the NLG tasks they are used to evaluate.
f-score. Also called F-measure, the f-score is a measure of accuracy. It balances the generated text’s precision and recall by the harmonic mean of the two measures. The most common instantiation of the f-score is the f1-score (). In text generation tasks such as machine translation or summarization, f-score gives an indication as to the quality of the generated sequence that a model will produce (Melamed et al., 2003; Aliguliyev, 2008). Specifically for machine translation, f-score-based metrics have been shown to be effective in evaluating translation quality.
bleu. The Bilingual Evaluation Understudy (bleu) is one of the first metrics used to measure the similarity between two sentences (Papineni et al., 2002). Originally proposed for machine translation, it compares a candidate translation of text to one or more reference translations. bleu is a weighted geometric mean of n-gram precision scores.
It has been argued that although bleu has significant advantages, it may not be the ultimate measure for improved machine translation quality (Callison-Burch & Osborne, 2006). While earlier work has reported that bleu correlates well with human judgments (Lee & Przybocki, 2005; Denoual & Lepage, 2005), more recent work argues that although it can be a good metric for the machine translation task (Zhang et al., 2004) for which it is designed, it doesn’t correlate well with human judgments for other generation tasks (such as image captioning or dialog response generation). Reiter (2018) reports that there’s not enough evidence to support that bleu is the best metric for evaluating NLG systems other than machine translation. Caccia et al. (2018) found that generated text with perfect bleu scores was often grammatically correct but lacked semantic or global coherence, concluding that the generated text has poor information content.
Outside of machine translation, bleu has been used for other text generation tasks, such as document summarization (Graham, 2015), image captioning (Vinyals et al., 2014), human-machine conversation (Gao et al., 2019), and language generation (Semeniuta et al., 2019). In Graham (2015), it was concluded that bleu achieves strongest correlation with human assessment, but does not significantly outperform the best-performing rouge variant. A more recent study has demonstrated that n-gram matching scores such as bleu can be an insufficient and potentially less accurate metric for unsupervised language generation (Semeniuta et al., 2019).
Text generation research, especially when focused on short text generation like sentence-based machine translation or question generation, has successfully used bleu for benchmark analysis with models since it is fast, easy to calculate, and enables a comparison with other models on the same task. However, bleu has some drawbacks for NLG tasks where contextual understanding and reasoning is the key (e.g., story generation (Fan et al., 2018; Martin et al., 2017) or long-form question answering (Fan et al., 2019a)). It considers neither semantic meaning nor sentence structure. It does not handle morphologically rich languages well, nor does it map well to human judgments (Tatman, 2019). Recent work by Mathur et al. (2020) investigated how sensitive the machine translation evaluation metrics are to outliers. They found that when there are outliers in tasks like machine translation, metrics like bleu lead to high correlations yielding false conclusions about reliability of these metrics. They report that when the outliers are removed, these metrics do not correlate as well as before, which adds evidence to the unreliablity of bleu.
We will present other metrics that address some of these shortcomings throughout this paper.
rouge. Recall-Oriented Understudy for Gisting Evaluation (rouge) (Lin, 2004) is a set of metrics for evaluating automatic summarization of long texts consisting of multiple sentences or paragraphs. Although mainly designed for evaluating single- or multi-document summarization, it has also been used for evaluating short text generation, such as machine translation (Lin & Och, 2004), image captioning (Cui et al., 2018), and question generation (Nema & Khapra, 2018; Dong et al., 2019). rouge includes a large number of distinct variants, including eight different n-gram counting methods to measure n-gram overlap between the generated and the ground-truth (human-written) text: rouge-{1/2/3/4} measures the overlap of unigrams/bigrams/trigrams/four-grams (single tokens) between the reference and hypothesis text (e.g., summaries); rouge-l measures the longest matching sequence of words using longest common sub-sequence (LCS); rouge-s (less commonly used) measures skip-bigramA skip-gram (Huang et al., 1992) is a type of n-gram in which tokens (e.g., words) don’t need to be consecutive but in order in the sentence, where there can be gaps between the tokens that are skipped over. In NLP research, they are used to overcome data sparsity issues.-based co-occurrence statistics; rouge-su (less commonly used) measures skip-bigram and unigram-based co-occurrence statistics.
Compared to bleu, rouge focuses on recall rather than precision and is more interpretable than bleu (Callison-Burch & Osborne, 2006). Additionally, rouge includes the mean or median score from individual output text, which allows for a significance test of differences in system-level rouge scores, while this is restricted in bleu (Graham & Baldwin, 2014; Graham, 2015). However, rouge’s reliance on n-gram matching can be an issue, especially for long-text generation tasks (Kilickaya et al., 2017), because it doesn’t provide information about the narrative flow, grammar, or topical flow of the generated text, nor does it evaluate the factual correctness of the text compared to the corpus it is generated from.
meteor. The Metric for Evaluation of Translation with Explicit ORdering (meteor) (Lavie et al., 2004; Banerjee & Lavie, 2005) is a metric designed to address some of the issues found in bleu and has been widely used for evaluating machine translation models and other models, such as image captioning (Kilickaya et al., 2017), question generation (Nema & Khapra, 2018; Du et al., 2017), and summarization (See et al., 2017; Chen & Bansal, 2018; Yan et al., 2020). Compared to bleu, which only measures precision, meteor is based on the harmonic mean of the unigram precision and recall, in which recall is weighted higher than precision. Several metrics support this property since it yields high correlation with human judgments (Lavie & Agarwal, 2007).
meteor has several variants that extend exact word matching that most of the metrics in this category do not include, such as stemming and synonym matching. These variants address the problem of reference translation variability, allowing for morphological variants and synonyms to be recognized as valid translations. The metric has been found to produce good correlation with human judgments at the sentence or segment level (Agarwal & Lavie, 2008). This differs from bleu in that meteor is explicitly designed to compare at the sentence level rather than the corpus level.
cider. Consensus-based Image Description Evaluation (cider) is an automatic metric for measuring the similarity of a generated sentence against a set of human-written sentences using a consensus-based protocol. Originally proposed for image captioning (Vedantam et al., 2014), cider shows high agreement with consensus as assessed by humans. It enables a comparison of text generation models based on their “human-likeness,” without having to create arbitrary calls on weighing content, grammar, saliency, etc. with respect to each other.
The cider metric presents three explanations about what a hypothesis sentence should contain: (1) n-grams in the hypothesis sentence should also occur in the reference sentences, (2) If an n-gram does not occur in a reference sentence, it shouldn’t be in the hypothesis sentence, (3) n-grams that commonly occur across all image-caption pairs in the dataset should be assigned lower weights, since they are potentially less informative. While cider has been adopted as an evaluation metric for image captioning and has been shown to correlate well with human judgments on some datasets (PASCAL-50S and ABSTRACT-50S datasets) (Vedantam et al., 2014), recent studies have shown that metrics that include semantic content matching such as spice can correlate better with human judgments (Anderson et al., 2016; Liu et al., 2017).
nist. Proposed by the US National Institute of Standards and Technology, nist (Martin & Przybocki, 2000) is a method similar to bleu for evaluating the quality of text. Unlike bleu, which treats each n-gram equally, nist heavily weights n-grams that occur less frequently, as co-occurrences of these n-grams are more informative than common n-grams (Doddington, 2002).
gtm. The gtm metric (Turian & Melamed, 2003) measures n-gram similarity between the model-generated hypothesis translation and the reference sentence by using precision, recall and F-score measures.
hlepor. Harmonic mean of enhanced Length Penalty, Precision, n-gram Position difference Penalty, and Recall (hlepor), initially proposed for machine translation, is a metric designed for morphologically complex languages like Turkish or Czech (Han et al., 2013a). Among other factors, hlepor uses part-of-speech tags’ similarity to capture syntactic information.
ribes. Rank-based Intuitive Bilingual Evaluation Score (ribes) (Isozaki et al., 2010) is another untrained automatic evaluation metric for machine translation. It was developed by NTT Communication Science Labs and designed to be more informative for Asian languages–―like Japanese and Chinese—since it doesn’t rely on word boundaries. Specifically, ribes is based on how the words in generated text are ordered. It uses the rank correlation coefficients measured based on the word order from the hypothesis (model-generated) translation and the reference translation.
dice and masi. Used mainly for referring expression generation evaluation, dice (Gatt et al., 2008) measures the overlap of a set of attributes between the human-provided referring expression and the model-generated expressions. The expressions are based on an input image (e.g., the large chair), and the attributes are extracted from the expressions, such as the type or color (Chen & van Deemter, 2020). The masi metric (Gatt et al., 2008), on the other hand, adapts the Jaccard coefficient, which biases it in favour of similarity when a set of attributes is a subset of the other attribute set.
2 Distance-Based Evaluation Metrics for Content Selection
A distance-based metric in NLG applications uses a distance function to measure the similarity between two text units (e.g., words, sentences). First, we represent two text units using vectors. Then, we compute the distance between the vectors. The smaller the distance, the more similar the two text units are. This section reviews distance-based similarity measures where text vectors can be constructed using discrete tokens, such as bag of words (§3.2.1) or embedding vectors (§3.2.2). We note that even though the embeddings that are used by these metrics to represent the text vectors are pre-trained, these metrics are not trained to mimic the human judgments, as in the machine-learned metrics that we summarize in Section 4.
Edit distance, one of the most commonly used evaluation metrics in natural language processing, measures how dissimilar two text units are based on the minimum number of operations required to transform one text into the other. We summarize some of the well-known edit distance measures below.
Word error rate (wer) has been commonly used for measuring the performance of speech recognition systems, as well as to evaluate the quality of machine translation systems (Tomás et al., 2003). Specifically, wer is the percentage of words that need to be inserted, deleted, or replaced in the translated sentence to obtain the reference sentence, i.e., the edit distance between the reference and hypothesis sentences.
wer has some limitations. For instance, while its value is lower-bounded by zero, which indicates a perfect match between the hypothesis and reference text, its value is not upper-bounded, making it hard to evaluate in an absolute manner (Mccowan et al., 2004). It is also reported to suffer from weak correlation with human evaluation. For example, in the task of spoken document retrieval, the wer of an automatic speech recognition system is reported to poorly correlate with the retrieval system performance (Kafle & Huenerfauth, 2017).
Translation edit rate (ter) (Snover et al., 2006) is defined as the minimum number of edits needed to change a generated text so that it exactly matches one of the references, normalized by the average length of the references. While ter has been shown to correlate well with human judgments in evaluating machine translation quality, it suffers from some limitations. For example, it can only capture similarity in a narrow sense, as it only uses a single reference translation and considers only exact word matches between the hypothesis and the reference. This issue can be partly addressed by constructing a lattice of reference translations, a technique that has been used to combine the output of multiple translation systems (Rosti et al., 2007).
2.2 Vector Similarity-Based Evaluation Metrics
In NLP, embedding-based similarity measures are commonly used in addition to n-gram-based similarity metrics. Embeddings are real-valued vector representations of character or lexical units, such as word-tokens or n-grams, that allow tokens with similar meanings to have similar representations. Even though the embedding vectors are learned using supervised or unsupervised neural network models, the vector-similarity metrics we summarize below assume the embeddings are pre-trained and simply used as input to calculate the metric.
The vector-based similarity measure meant uses word embeddings and shallow semantic parses to compute lexical and structural similarity (Lo, 2017). It evaluates translation adequacy by measuring the similarity of the semantic frames and their role fillers between the human references and the machine translations.
Inspired by the meant score, yisiYiSi, is the romanization of the Cantonese word 意思 , which translates as ‘meaning’ in English. (Lo, 2019) is proposed to evaluate the accuracy of machine translation model outputs. It is based on the weighted distributional lexical semantic similarity, as well as shallow semantic structures. Specifically, it extracts the longest common character sub-string from the hypothesis and reference translations to measure the lexical similarity.
Earth mover’s distance (emd), also known as the Wasserstein metric (Rubner et al., 1998), is a measure of the distance between two probability distributions. Word mover’s distance (wmd; Kusner et al., 2015) is a discrete version of emd that calculates the distance between two sequences (e.g., sentences, paragraphs, etc.), each represented with relative word frequencies. It combines item similarityThe similarity can be defined as cosine, Jaccard, Euclidean, etc. on bag-of-word (BOW) histogram representations of text (Goldberg et al., 2018) with word embedding similarity. In short, wmd has several intriguing properties:
It is hyperparameter-free and easy to use.
It is highly interpretable as the distance between two documents can be broken down and explained as the sparse distances between few individual words.
It uses the knowledge encoded within the word embedding space, which leads to high retrieval accuracy.
Empirically, wmd has been instrumental to the improvement of many NLG tasks, specifically sentence-level tasks, such as image caption generation (Kilickaya et al., 2017) and natural language inference (Sulea, 2017). However, while wmd works well for short texts, its cost grows prohibitively as the length of the documents increases, and the BOW approach can be problematic when documents become large as the relation between sentences is lost. By only measuring word distances, the metric cannot capture information conveyed in the group of words, for which we need higher-level document representations (Dai et al., 2015).
Sentence Mover’s Distance (smd) is an automatic metric based on wmd to evaluate text in a continuous space using sentence embeddings (Clark et al., 2019; Zhao et al., 2019). smd represents each document as a collection of sentences or of both words and sentences (as seen in Figure 2), where each sentence embedding is weighted according to its length. smd measures the cumulative distance of moving the sentences embeddings in one document to match another document’s sentence embeddings. On a summarization task, smd correlated better with human judgments than rouge (Clark et al., 2019).
Zhao et al. (2019) proposed a new version of smd that attains higher correlation with human judgments. Similar to smd, they used word and sentence embeddings by taking the average of the token-based embeddings before the mover’s distance is calculated. They also investigated different contextual embeddings models including ELMO and BERT by taking the power mean (which is an embedding aggregation method) of their embeddings at each layer of the encoding model.
3 n-gram-Based Diversity Metrics
The lexical diversity score measures the breadth and variety of the word usage in writing (Inspector, 2013). Lexical diversity is desirable in many NLG tasks, such as conversational bots (Li et al., 2018), story generation (Rashkin et al., 2020), question generation (Du et al., 2017; Pan et al., 2019), and abstractive question answering (Fan et al., 2019). Nevertheless, diversity-based metrics are rarely used on their own, as text diversity can come at the cost of text quality (Montahaei et al., 2019a; Hashimoto et al., 2019; Zhang et al., 2021), and some NLG tasks do not require highly diverse generations. For example, Reiter et al. (2005) reported that a weather forecast system was preferred over human meteorologists as the system produced report has a more consistent use of certain classes of expressions relating to reporting weather forecast.
In this section we review some of the metrics designed to measure the quality of the generated text in terms of lexical diversity.
is a measure of lexical diversity (Richards, 1987), mostly used in linguistics to determine the richness of a writer’s or speaker’s vocabulary. It is computed as the number of unique words (types) divided by the total number of words (tokens) in a given segment of language.
Although intuitive and easy to use, ttr is sensitive to text length because the longer the document, the lower the likelihood that a token will be a new type. This causes the ttr to drop as more words are added. To remedy this, a diversity metric, hd-d (hyper-geometric distribution function), was proposed to compare texts of different lengths (McCarthy & Jarvis, 2010).
Measuring diversity using n-gram repetitions is a more generalized version of ttr, which has been use for text generation evaluation. Li et al. (2016) has shown that modeling mutual information between source and targets significantly decreases the chance of generating bland responses and improves the diversity of responses. They use bleu and distinct word unigram and bigram counts to evaluate the proposed diversity-promoting objective function for dialog response generation.
Zhu et al. (2018) proposed self-bleu as a diversity evaluation metric, which calculates a bleu score for every generated sentence, treating the other generated sentences as references. The average of these bleu scores is the self-bleu score of the text, where a lower self-bleu score implies higher diversity. Several NLG papers have reported that self-bleu achieves good generation diversity (Zhu et al., 2018; Chen et al., 2018a; Lu et al., 2018). However, others have reported some weakness of the metric in generating diverse output (Caccia et al., 2018) or detecting mode collapse (Semeniuta et al., 2019) in text generation with GAN (Goodfellow et al., 2014) models.
4 Explicit Semantic Content Match Metrics
Semantic content matching metrics define the similarity between human-written and model-generated text by extracting explicit semantic information units from text beyond n-grams. These metrics operate on semantic and conceptual levels and are shown to correlate well with human judgments. We summarize some of them below.
The pyramid method is a semi-automatic evaluation method (Nenkova & Passonneau, 2004) for evaluating the performance of document summarization models. Like other untrained automatic metrics that require references, this untrained metric also requires human annotations. It identifies summarization content units (SCUs) to compare information in a human-generated reference summary to the model-generated summary. To create a pyramid, annotators select sets of text spans that express the same meaning across summaries, and each SCU is a weighted according to the number of summaries that express the SCU’s meaning.
The pyramid metric relies on manual human labeling effort, which makes it difficult to automate. peak: Pyramid Evaluation via Automated Knowledge Extraction (Yang et al., 2016) was presented as a fully automated variant of pyramid model, which can automatically assign the pyramid weights and was shown to correlate well with human judgments.
Semantic propositional image caption evaluation (spice) (Anderson et al., 2016) is an image captioning metric that measures the similarity between a list of reference human written captions of an image and a hypothesis caption generated by a model. Instead of directly comparing the captions’ text, spice parses each caption to derive an abstract scene graph representation, encoding the objects, attributes, and relationships detected in image captions (Schuster et al., 2015), as shown in Figure 3. spice then computes the f-score using the hypothesis and reference scene graphs over the conjunction of logical tuples representing semantic propositions in the scene graph to measure their similarity. spice has been shown to have a strong correlation with human ratings.
Compared to n-gram matching methods, spice can capture a broader sense of semantic similarity between a hypothesis and a reference text by using scene graphs. However, even though spice correlates well with human evaluations, a major drawback is that it ignores the fluency of the generated captions (Sharif et al., 2018).
Liu et al. (2017) proposed spider, which is a linear combination of spice and cider. They show that optimizing spice alone often results in captions that are wordy and repetitive because while scene graph similarity is good at measuring the semantic similarity between captions, it does not take into account the syntactical aspects of texts. Thus, a combination of semantic graph similarity (like spice) and n-gram similarity measure (like cider) yields a more complete quality evaluation metric. However, the correlation of spider and human evaluation is not reported.
4.1 Semantic Similarity Models used as Evaluation Metrics
Other text generation work has used the confidence scores obtained from semantic similarity methods as an evaluation metric. Such models can evaluate a reference and a hypothesis text based on their task-level semantics. The most commonly used methods based on the sentence-level similarity are as follows:
Semantic Textual Similarity (STS) is concerned with the degree of equivalence in the underlying semantics of paired text (Agirre et al., 2016). STS is used as an evaluation metric in text generation tasks such as machine translation, summarization, and dialogue response generation in conversational systems. The official score is based on weighted Pearson correlation between predicted similarity and human-annotated similarity. The higher the score, the better the the similarity prediction result from the algorithm (Maharjan et al., 2017; Cer et al., 2017b).
Paraphrase identification (PI) considers if two sentences express the same meaning (Dolan & Brockett, 2005; Barzilay & Lee, 2003). PI is used as a text generation evaluation score based on the textual similarity (Kauchak & Barzilay, 2006) of a reference and hypothesis by finding a paraphrase of the reference sentence that is closer in wording to the hypothesis output. For instance, given the pair of sentences:
reference: “However, Israel’s reply failed to completely clear the U.S. suspicions.” hypothesis: “However, Israeli answer unable to fully remove the doubts.”
PI is concerned with learning to transform the reference sentence into:
paraphrase: “However, Israel’s answer failed to completely remove the U.S. suspicions.”
which is closer in wording to the hypothesis. In Jiang et al. (2019), a new paraphrasing evaluation metric, tiger, is used for image caption generation evaluation. Similarly, Liu et al. (2019a) introduce different strategies to select useful visual paraphrase pairs for training by designing a variety of scoring functions.
Textual entailment (TE) is concerned with whether a hypothesis can be inferred from a premise, requiring understanding of the semantic similarity between the hypothesis and the premise (Dagan et al., 2006; Bowman et al., 2015). It has been used to evaluate several text generation tasks, including machine translation (Padó et al., 2009), document summarization (Long et al., 2018), language modeling (Liu et al., 2019b), and video captioning (Pasunuru & Bansal, 2017).
Machine Comprehension (MC) is concerned with the sentence matching between a passage and a question, pointing out the text region that contains the answer (Rajpurkar et al., 2016). MC has been used for tasks like improving question generation (Yuan et al., 2017; Du et al., 2017) and document summarization (Hermann et al., 2015).
5 Syntactic Similarity-Based Metrics
A syntactic similarity metric captures the similarity between a reference and a hypothesis text at a structural level to capture the overall grammatical or sentence structure similarity.
In corpus linguistics, part of speech (POS) tagging is the process of assigning a part-of-speech tag (e.g., verb, noun, adjective, adverb, and preposition, etc.) to each word in a sentence, based on its context, morphological behaviour, and syntax. POS tags have been commonly used in machine translation evaluation to evaluate the quality of the generated translations. tesla (Dahlmeier et al., 2011) was introduced as an evaluation metric to combine the synonyms of bilingual phrase tables and POS tags, while others use POS n-grams together with a combination of morphemes and lexicon probabilities to compare the target and source translations (Popovic et al., 2011; Han et al., 2013b). POS tag information has been used for other text generation tasks such as story generation (Agirrezabal et al., 2013), summarization (Suneetha & Fatima, 2011), and question generation (Zerr, 2014).
Syntactic analysis studies the arrangement of words and phrases in well-formed sentences. For example, a dependency parser extracts a dependency tree of a sentence to represent its grammatical structure. Several text generation tasks have enriched their evaluation criteria by leveraging syntactic analysis. In machine translation, Liu & Gildea (2005) used constituent labels and head-modifier dependencies to extract structural information from sentences for evaluation, while others use shallow parsers (Lo et al., 2012) or dependency parsers (Yu et al., 2014, 2015). Yoshida et al. (2014) combined a sequential decoder with a tree-based decoder in a neural architecture for abstractive text summarization.
Machine-Learned Evaluation Metrics
Many of the untrained evaluation metrics described in Section 3 assume that the generated text has significant word (or n-gram) overlap with the ground-truth text. However, this assumption does not hold for NLG tasks that permit significant diversity and allow multiple plausible outputs for a given input (e.g., a social chatbot). Table 3 shows two examples from the dialog response generation and image captioning tasks, respectively. In both tasks, the model-generated outputs are plausible given the input, but they do not share any words with the ground-truth output.
One solution to this problem is to use embedding-based metrics, which measure semantic similarity rather than word overlap, as in Section 3.2.2. But embedding-based methods cannot help in situations when the generated output is semantically different from the reference, as in the dialog example. In these cases, we can build machine-learned models (trained on human judgment data) to mimic human judges to measure many quality metrics of output, such as factual correctness, naturalness, fluency, coherence, etc. In this section we survey the NLG evaluation metrics that are computed using machine-learned models, with a focus on recent neural models.
Neural approaches to sentence representation learning seek to capture semantic meaning and syntactic structure of sentences from different perspectives and topics and to map a sentence onto an embedding vector using neural network models. As with word embeddings, NLG models can be evaluated by embedding each sentence in the generated and reference texts.
Extending word2vec (Mikolov et al., 2013) to produce word or phrase embeddings, one of the earliest sentence embeddings models, Deep Semantic Similarity Model (dssm) (Huang et al., 2013) introduced a series of latent semantic models with a deep structure that projects two or more text streams (such as a query and multiple documents) into a common low-dimensional space where the relevance of one text towards the other text can be computed via vector distance. The skip-thought vectors model (Kiros et al., 2015) exploits the encoder-decoder architecture to predict context sentences in an unsupervised manner (see Figure 4). Skip-thought vectors allow us to encode rich contextual information by taking into account the surrounding context, but are slow to train. fastsent (Hill et al., 2016) makes training efficient by representing a sentence as the sum of its word embeddings, but also dropping any knowledge of word order. A simpler weighted sum of word vectors (Arora et al., 2019) weighs each word vector by a factor similar to the tf-idf score, where more frequent terms are weighted less. Similar to fastsent, it ignores word order and surrounding sentences. Extending dssm models, infersent (Conneau et al., 2017) is an effective model, which uses lstm-based Siamese networks, with two additional advantages over the fastsent. It encodes word order and is trained on a high-quality sentence inference dataset. On the other hand, quick-thought (Logeswaran & Lee, 2018) is based on an unsupervised model of universal sentence embeddings trained on consecutive sentences. Given an input sentence and its context, a classifier is trained to distinguish a context sentence from other contrastive sentences based on their embeddings.
The recent large-scale pre-trained language models (PLMs) such as elmo and bert use contextualized word embeddings to represent sentences. Even though these PLMs outperform the earlier models such as dssms, they are more computationally expensive to use for evaluating NLG systems. For example, the sentence similarity metrics that use Transformer-based encoders, such as bert model (Devlin et al., 2018) and its extension roberta (Liu et al., 2019c), to obtain sentence representations are designed to learn textual similarities in sentence-pairs using distance-based similarity measures at the top layer as learning signal, such as cosine similarity similar to dssm. But both are much more computationally expensive than dssm due to the fact that they use a much deeper NN architecture, and need to be fine-tuned for different tasks. To remedy this, Reimers & Gurevych (2019) proposed sentbert, a fine-tuned bert on a “general” task to optimize the BERT parameters, so that a cosine similarity between two generated sentence embeddings is strongly related to the semantic similarity of the two sentences. Then the fine-tuned model can be used to evaluate various NLG tasks. Focusing on machine translation task, esim also computes sentence representations from bert embeddings (with no fine-tuning), and later computes the similarity between the translated text and its reference using metrics such as the average recall of its reference (Chen et al., 2017; Mathur et al., 2019).
2 Regression-Based Evaluation
Shimanaka et al. (2018) proposed a segment-level machine translation evaluation metric named ruse. They treat the evaluation task as a regression problem to predict a scalar value to indicate the quality of translating a machine-translated hypothesis to a reference translation . They first do a forward pass on the GRU (gated-recurrent unit) based on an encoder to generate and represent as a -dimensional vector. Then, they apply different matching methods to extract relations between and by (1) concatenating ; (2) getting the element-wise product (); (3) computing the absolute element-wise distance (see Figure 5). ruse is demonstrated to be an efficient metric in machine translation shared tasks in both segment-level (how well the metric correlates with human judgments of segment quality) and system-level (how well a given metric correlates with the machine translation workshop official manual ranking) metrics.
3 Evaluation Models with Human Judgments
For more creative and open-ended text generation tasks, such as chit-chat dialog, story generation, or online review generation, current evaluation methods are only useful to some degree. As we mentioned in the beginning of this section, word-overlap metrics are ineffective as there are often many plausible references in these scenarios and collecting them all is impossible. Even though human evaluation methods are useful in these scenarios for evaluating aspects like coherency, naturalness, or fluency, aspects like diversity or creativity may be difficult for human judges to assess as they have no knowledge about the dataset that the model is trained on (Hashimoto et al., 2019). Language models can learn to copy from the training dataset and generate samples that a human judge will rate as high in quality, but may fail in generating diverse samples (i.e., samples that are very different from training samples), as has been observed in social chatbots (Li et al., 2016; Zhou et al., 2020). A language model optimized only for perplexity may generate coherent but bland responses. Such behaviours are observed when generic pre-trained language models are used for downstream tasks ‘as-is’ without fine-tuning on in-domain datasets of related downstream tasks. A commonly overlooked issue is that conducting human evaluation for every new generation task can be expensive and not easily generalizable.
To calibrate human judgments and automatic evaluation metrics, model-based approaches that use human judgments as attributes or labels have been proposed. Lowe et al. (2017) introduced a model-based evaluation metric, adem, which is learned from human judgments for dialog system evaluation, specifically response generation in a chatbot setting. Using Twitter data (each tweet response is a reference, and its previous dialog turns are its context), they have different models (such as RNNs, retrieval-based methods, or other human responses) generate responses and ask humans to judge the appropriateness of the generated response given the context. For evaluation they use a higher quality labeled Twitter dataset (Ritter et al., 2011), which contains dialogs on a variety of topics.
Using this score-labeled dataset, the adem evaluation model is trained as follows: First, a latent variational recurrent encoder-decoder model (vhred) (Serban et al., 2016b) is pre-trained on a dialog dataset to learn to represent the context of a dialog. vhred encodes the dialog context into a vector representation, from which the model generates samples of initial vectors to condition the decoder model to generate the next response. Using the pre-trained vhred model as the encoder, they train adem as follows (see Figure 6). First, the dialog context, , the model generated response , and the reference response are fed to vhred to get their embedding vectors, , and . Then, each embedding is linearly projected so that the model response can be mapped onto the spaces of the dialog context and the reference response to calculate a similarity score. The similarity score measures how close the model responses are to the context and the reference response after the projection, as follows:
adem is optimized for squared error loss between the predicted score and the human judgment score with L-2 regularization in an end-to-end fashion. The trained evaluation model is shown to correlate well with human judgments. adem is also found to be conservative and give lower scores to plausible responses.
With the motivation that a good evaluation metric should capture both the quality and the diversity of the generated text, Hashimoto et al. (2019) proposed a new evaluation metric named Human Unified with Statistical Evaluation (huse), which focuses on more creative and open-ended text generation tasks, such as dialog and story generation. Unlike the adem metric, which relies on human judgments for training the model, huse combines statistical evaluation and human evaluation metrics in one model, as shown in Figure 7.
huse considers the conditional generation task that, given a context sampled from a prior distribution , outputs a distribution over possible sentences . The evaluation metric is designed to determine the similarity of the output distribution and a human generation reference distribution . This similarity is scored using an optimal discriminator that determines whether a sample comes from the reference or hypothesis (model) distribution (Figure 7). For instance, a low-quality text is likely to be sampled from the model distribution. The discriminator is implemented approximately using two probability measures: (i) the probability of a sentence under the model, which can be estimated using the text generation model, and (ii) the probability under the reference distribution, which can be estimated based on human judgment scores. On summarization and chitchat dialog tasks, huse has been shown to be effective to detect low-diverse generations that humans fail to detect.
4 BERT-Based Evaluation
Given the strong performance of bert (Devlin et al., 2018) across many tasks, there has been work that uses bert or similar pre-trained language models for evaluating NLG tasks, such as summarization and dialog response generation. Here, we summarize some of the recent work that fine-tunes bert to use as evaluation metrics for downstream text generation tasks.
One of the bert-based models for semantic evaluation is bertscore (Zhang et al., 2020a). As illustrated in Figure 8, it leverages the pre-trained contextual embeddings from bert and matches words in candidate and reference sentences by cosine similarity. bertscore has been shown to correlate well with human judgments on sentence-level and system-level evaluations. Moreover, bertscore computes precision, recall, and F1 measures, which are useful for evaluating a range of NLG tasks.
Kané et al. (2019) presented a bert-based evaluation method called roberta-sts to detect sentences that are logically contradictory or unrelated, regardless whether they are grammatically plausible. Using roberta (Liu et al., 2019c) as a pre-trained language model, roberta-sts is fine-tuned on the STS-B dataset (Cer et al., 2017a) to learn the similarity of sentence pairs on a Likert scale. Another evaluation model is fine-tuned on the Multi-Genre Natural Language Inference Corpus (Williams et al., 2018) in a similar way to learn to predict logical inference of one sentence given the other. Both model-based evaluators, roberta-sts and its extension, have been shown to be more robust and correlate better with human evaluation than automatic evaluation metrics such as bleu and rouge.
Another recent bert-based machine-learned evaluation metric is bleurt (Sellam et al., 2020), which was proposed to evaluate various NLG systems. The evaluation model is trained as follows: A checkpoint from bert is taken and fine-tuned on synthetically generated sentence pairs using automatic evaluation scores such as bleu or rouge, and then further fine-tuned on system-generated outputs and human-written references using human ratings and automatic metrics as labels. The fine-tuning of bleurt on synthetic pairs is an important step because it improves the robustness to quality drifts of generation systems. As shown in the plots in Figure 9, as the NLG task gets more difficult, the ratings get closer as it is easier to discriminate between “good” and “bad” systems than to rank “good” systems. To ensure the robustness of their metric, they investigate with training datasets with different characteristics, such as when the training data is highly skewed or out-of-domain. They report that the training skew has a disastrous effect on bleurt without pre-training; this pre-training makes bleurt significantly more robust to quality drifts.
As discussed in Section 2, humans can efficiently evaluate the performance of two models side-by-side, and most embedding-based similarity metrics reviewed in the previous sections are based on this idea. Inspired by this, the comparator evaluator (Zhou & Xu, 2020) was proposed to evaluate NLG models by learning to compare a pair of generated sentences by fine-tuning bert. A text pair relation classifier is trained to compare the task-specific quality of a sample hypothesis and reference based on the win/loss rate. Using the trained model, a skill rating system is built. This system is similar to the player-vs-player games in which the players are evaluated by observing a record of wins and losses of multiple players. Then, for each player, the system infers the value of a latent, unobserved skill variable that indicates the records of wins and losses. On story generation and open domain dialogue response generation tasks, the comparator evaluator metric demonstrates high correlation with human evaluation.
5 Evaluating Factual Correctness
An important issue in text generation systems is that the model’s generation could be factually inconsistent, caused by distorted or fabricated facts about the source text. Especially in document summarization tasks, the models that abstract away salient aspects, have been shown to generate text with up to 30% factual inconsistencies (Kryscinski et al., 2019b; Falke et al., 2019; Zhu et al., 2020). There has been a lot of recent work that focuses on building models to verify the factual correctness of the generated text, focusing on semantically constrained tasks such as document summarization or image captioning, some of which we summarize here.
Some recent evaluation metrics have addressed factual correctness via entailment-based models (Falke et al., 2019; Maynez et al., 2020; Dušek & Kasner, 2020). However, these sentence-level, entailment-based approaches do not capture which part of the generated text is non-factual. Goyal & Durrett presented a new localized entailment-based approach using dependency trees to reformulate the entailment problem at the dependency arc level. Specifically, they align the the semantic relations yielded by the dependency arcs (see Figure 11) in the generated output summary to the input sentences. Their dependency arc entailment model improves factual consistency and shows stronger correlations with human judgments in generation tasks such as summarization and paraphrasing.
Models adhering to the facts in the source have started to gain more attention in “conditional” or “grounded” text generation tasks, such as document summarization (Kryscinski et al., 2019b) and data-to-text generation (Reiter, 2007; Lebret et al., 2016a; Sha et al., 2017; Puduppully et al., 2018; Wang, 2019; Nan et al., 2021). In one of the earlier works on structured data-to-text generation, Wiseman et al. (2017) dealt with the coherent generation of multi-sentence summaries of tables or database records. In this work, they first trained an auxiliary model as relation extraction classifier (entity-mention pairs) based on information extraction to evaluate how well the text generated by the model can capture the information in a discrete set of records. Then the factual evaluation is based on the alignment between the entity-mention predictions of this classifier against the source database records. Their work was limited to a single domain (basketball game tables and summaries) and assumed that the tables has similar attributes, which can be limiting for open-domain data-to-text generation systems.
Dhingra et al. (2019) extended this approach and introduced the parent measure. Their evaluation approach first aligns the entities in the table and the reference and generated text with a neural attention-based model and later measures similarities on word overlap, entailment and other metrics over the alignment. They conduct a large scale human evaluation study which yielded that parent correlates with human judgments better than several n-gram match and information extraction based metrics they used for evaluation. Parikh et al. proposed a new controllable text generation task, totto, which generates a sentence to describe a highlighted cell in a given table and extended the parent to adapt to their tasks so the metric takes into account the highlighted cell in the table.
Factual consistency evaluations have also appeared in multi-modal generation tasks, such as image captioning. In one such work (Chen et al., 2018b), a new style-focused factual rnn-type decoder is constructed to allow the model to preserve factual information in longer sequences without requiring additional labels. In this model, they query a reference model to adaptively learn to add factual information into the model.
Zhang et al. (2019b) proposed a way to tackle the problem of factual correctness in summarization models. Focusing on summarizing radiology reports, they extend pointer networks for abstractive summarization by introducing a reward-based optimization that trains the generators to obtain more rewards when they generate summaries that are factually aligned with the original document. Specifically, they design a fact extractor module so that the factual accuracy of a generated summary can be measured and directly optimized as a reward using policy gradient, as shown in Figure 12. This fact extractor is based on an information extraction module and extracts and represents the facts from generated and reference summaries in a structured format. The summarization model is updated via reinforcement learning using a combination of the NLL (negative log likelihood) loss, a rouge-based loss, and a factual correctness-based loss (Loss=++). Their work suggests that for domains in which generating factually correct text is crucial, a carefully implemented information extraction system can be used to improve the factual correctness of neural summarization models via reinforcement learning.
To evaluate the factual consistency of the text generation models, Eyal et al. (2019b) presented a question-answering-based parametric evaluation model named Answering Performance for Evaluation of Summaries (apes) (see Figure 13). Their evaluation model is designed to evaluate document summarization and is based on the hypothesis that the quality of a generated summary is associated with the number of questions (from a set of relevant ones) that can be answered by reading the summary.
To build such an evaluator to assess the quality of generated summaries, they introduce two components: (a) a set of relevant questions for each source document and (b) a question-answering system. They first generate questions from each reference summary by masking each of the named entities present in the reference based on the method described in Hermann et al. (2015). For each reference summary, this results in several triplets in the form (generated summary, question, answer), where question refers to the sentence containing the masked entity, answer refers to the masked entity, and the generated summary is generated by their summarization model. Thus, for each generated summary, metrics can be derived based on the accuracy of the question answering system in retrieving the correct answers from each of the associated triplets. This metric is useful for summarizing documents for domains that contain lots of named entities, such as biomedical or news article summarization.
6 Composite Metric Scores
The quality of many NLG models like machine translation and image captioning can be evaluated for multiple aspects, such as adequacy, fluency, and diversity. Many composite metrics have been proposed to capture a multi-dimensional sense of quality. Sharif et al. (2018) presented a machine-learned composite metric for evaluating image captions. The metric incorporates a set of existing metrics such as meteor, wmd, and spice to measure both adequacy and fluency. They evaluate various combinations of the metrics they chose to compose and and show that their composite metrics correlate well with human judgments.
Li & Chen (2020) propose a composite reward function to evaluate the performance of image captions. The approach is based on refined Adversarial Inverse Reinforcement Learning (rAIRL), which eases the reward ambiguity (common in reward-based generation models) by decoupling the reward for each word in a sentence. The proposed composite reward is shown on MS COCO data to achieve state-of-the-art performance on image captioning. Some examples generated from this model that uses the composite reward function are shown in Figure 14. They have shown that their metric not only generates grammatical captions but also correlates well with human judgments.
Shared Tasks for NLG Evaluation
Shared tasks in NLG are designed to boost the development of sub-fields and continuously encourage researchers to improve upon the state-of-the-art. With shared tasks, the same data and evaluation metrics are used to efficiently benchmark models. NLG shared tasks are common not only because language generation is a growing research field with numerous unsolved research challenges, but also because many NLG generation tasks do not have an established evaluation pipeline. NLG researchers are constantly proposing new shared tasks as new datasets and tasks are introduced to support efficient evaluation of novel approaches in language generation. Even though shared tasks are important for NLG research and evaluation, there are potential issues that originate from the large variability and a lack of standardisation in the organisation of shared tasks, not just for language generation but for language processing in general. In (Parra Escartín et al., 2017), some of these ethical concerns are discussed.
In this section we survey some of the shared tasks that focus on the evaluation of text generation systems that are aimed at comparing and validating different evaluation measures.
The Attribute Selection for Generating Referring Expressions (GRE) (asgre) Challenge (Gatt & Belz, 2008) was one of the first shared-task evaluation challenges in NLG. It was designed for the content determination of the GRE task, selecting the properties to describe an intended referent. The goal of this shared task was to evaluate the submitted systems on minimality (the proportion of descriptions in the system-generated output that are maximally brief compared to the original definition),uniqueness and humanlikeness.
2 Embedded Text Generation
To spur research towards human-machine communication in situated settings, Generating Instructions in Virtual Environments (GIVE) has been introduced as a challenge and an evaluation testbed for NLG (Koller et al., 2009). In this challenge a human player is given a task to solve in a simulated 3D space. A generation module’s task is to guide the human player, using natural language instructions. Only the human user can effect any changes in the world, by moving around, manipulating objects, etc. This challenge evaluates NLG models on referring expression generation, aggregation, grounding, realization, and user modeling. This challenge has been organized in four consecutive years (Striegnitz et al., 2011).
3 Regular Expression Generation (REG) in Context
The goal in this task is to map a representation of an intended referent in a given textual context to a full surface form. The representation of the intended referring expression maybe one from possible list of referring expressions for that referent and/or a set of semantic and syntactic properties. This challenge has been organized under different sub-challenges: GREC-Full has focused on improving the referential clarity and fluency of the text in which systems were expected to replace regular expressions and where necessary to produce as clear, fluent and coherent a text as possible (Belz & Kow, 2010). The GREC-NEG Task at Generation Challenges 2009 (Belz et al., 2009) evaluated models in select correct coreference chains for all people entities mentioned in short encyclopaedic texts about people collected from Wikipedia.
4 Regular Expression Generation from Attribute Sets
This task tries to answer the following question: Given a symbol corresponding to an intended referent, how do we work out the semantic content of a referring expression that uniquely identifies the entity in question? (Bohnet & Dale, 2005). The input to these models consists of sets of attributes (e.g., {type=lamp, colour=blue, size=small}), where at least one attribute set is labelled the intended referent, and the remainder are the distractors. Then the task is to build a model that can output a set of attributes for the intended referent that uniquely distinguishes it from the distractors. Gatt et al. (2008) have introduced the tune Corpus and the tuna Challenge based on this corpus that covered a variety of tasks, including attribute selection for referring expressions, realization and end-to-end referring expression generation.
5 Deep Meaning Representation to Text (SemEval)
SemEval is a series of NLP workshops organized around the goal of advancing the current state of the art in semantic analysis and to help create high-quality annotated datasets to approach challenging problems in natural language semantics. Each year a different shared task is introduced for the teams to evaluate and benchmark models. For instance, Task 9 of the SemEval 2017 challenge was (sem, 2017) on text generation from AMR (Abstract Meaning Representation), which has focused on generating valid English sentences given AMR (Banarescu et al., 2013) annotation structure.
6 WebNLG
The WebNLG challenge introduced a text generation task from RDF triples to natural language text, providing a corpus and common benchmark for comparing the microplanning capacity of the generation systems that deal with resolving and using referring expressions, aggregations, lexicalizations, surface realizations and sentence segmentations (Gardent et al., 2017). A second challenge has taken place in 2020 (Zhou & Lampouras, 2020), three years after the first one, in which the dataset size increased (as did the coverage of the verbalisers) and more categories and an additional language were included to promote the development of knowledge extraction tools, with a task that mirrors the verbalisation task.
7 E2E NLG Challenge
Introduced in 2018, E2E NLG Challenge (Dušek et al., 2018) provided a high quality and large quantity training dataset for evaluating response generation models in spoken dialog systems. It introduced new challenges such that models should jointly learn sentence planning and surface realisation, while not requiring costly alignment between meaning representations and corresponding natural language reference texts.
8 Data-to-Text Generation Challenge
Most existing work in data-to-text (or table-to-text) generation focused on introducing datasets and benchmarks rather than organizing challanges. Some of these earlier works include: eathergov (Liang et al., 2009), robocup (Chen & Mooney, 2008), rotowire (Wiseman et al., 2017), e2e (Novikova et al., 2016), wikibio (Lebret et al., 2016b) and recently totto (Parikh et al., 2020). Banik et al. introduced a text generation from knowledge basehttp://www.kbgen.org challenge in 2013 to benchmark various systems on the content realization stage of generation. Given a set of relations which form a coherent unit, the task is to generate complex sentences that are grammatical and fluent in English.
9 GEM Benchmark
Introduced in ACL 2021, the gem benchmarkhttps://gem-benchmark.com (Gehrmann et al., 2021) aims to measure the progress in NLG, while continuously adding new datasets, evaluation metrics and human evaluation standards. gem provides an environment by providing easy testing of different NLG tasks and evaluation strategies.
Examples of Task-Specific NLG Evaluation
In the previous sections, we reviewed a wide range of NLG evaluation metrics individually. However, these metrics are constantly evolving due to rapid progress in more efficient, reliable, scalable and sustainable neural network architectures for training neural text generation models, as well as ever growing compute resources. Nevertheless, it is not easy to define what really is an “accurate,” “trustworthy” or even “efficient” metric for evaluating an NLG model or task. Thus, in this section we present how these metrics can be jointly used in research projects to more effectively evaluate NLG systems for real-world applications. We discuss two NLG tasks, automatic document summarization and long-text generation, that are sophisticated enough that multiple metrics are required to gauge different aspects of the generated text’s quality.
A text summarization system aims to extract useful content from a reference document and generate a short summary that is coherent, fluent, readable, concise, and consistent with the reference document. There are different types of summarization approaches, which can be grouped by their tasks into (i) generic text summarization for broad topics; (ii) topic-focused summarization, e.g., a scientific article, conversation, or meeting summarization; and (iii) query-focused summarization, such that the summary answers a posed query. These approaches can also be grouped by their method: (i) extractive, where a summary is composed of a subset of sentences or words in the input document; and (ii) abstractive, where a summary is generated on-the-fly and often contains text units that do not occur in the input document. Depending on the number of documents to be summarized, these approaches can also be grouped into single-document or multi-document summarization.
Evaluation of text summarization, regardless of its type, measures the system’s ability to generate a summary based on: (i) a set of criteria that are not related to references (Dusek et al., 2017), (ii) a set of criteria that measure its closeness to the reference document, or (iii) a set of criteria that measure its closeness to the reference summary. Figure 15 shows the taxonomy of evaluation metrics (Steinberger & Jezek, 2009) in two categories: intrinsic and extrinsic, which will be explained below.
Intrinsic evaluation of generated summaries can focus on the generated text’s content, text quality, and factual consistency, each discussed below.
Content evaluation compares a generated summary to a reference summary using automatic metrics. The most widely used metric for summarization is rouge, though other metrics, such as bleu and f-score, are also used. Although rouge has been shown to correlate well with human judgments for generic text summarization, the correlation is lower for topic-focused summarization like extractive meeting summarization (Liu & Liu, 2008). Meetings are transcripts of spontaneous speech, and thus usually contain disfluencies, such as pauses (e.g., ‘um,’ ‘uh,’ etc.), discourse markers (e.g., ‘you know,’ ‘i mean,’ etc.), repetitions, etc. Liu & Liu (2008) find that after such disfluencies are cleaned, the rouge score is improved. They even observed fair amounts of improvement in the correlation between the rouge score and human judgments when they include the speaker information of the extracted sentences from the source meeting to form the summary.
Evaluating generated summaries based on quality has been one of the challenging tasks for summarization researchers. As basic as it sounds, since the definition of a “good quality summary” has not been established and finding the most suitable metrics to evaluate quality remains an open research area. Below are some criteria of text, which are used in recent papers as human evaluation metrics to evaluate the quality of generated text in comparison to the reference text.
Coherence and Cohesion measure how clearly the ideas are expressed in the summary (Lapata & Barzilay, 2005). In particular, the idea that, in conjunction with cohesion, which is to hold the context as a whole, coherence should measure how well the text is organised and “hangs together.” Consider the examples in Table 5, from the scientific article abstract generation task. The models must include factual information, but it must also be presented in the right order to be coherent.
Readability and Fluency, associated with non-redundancy, are linguistic quality metrics used to measure how repetitive the generated summary is and how many spelling and grammar errors there are in the generated summary (Lapata, 2003).
Focus indicates how many of the main ideas of the document are captured, while avoiding superfluous details.
Informativeness, which is mostly used to evaluate question-focused summarization, measures how well the summary answers a question. Auto-regressive generation models trained to generate a short summary text given a longer document(s) may yield shorter summaries due to reasons relating to bias in the training data or type of the decoding method (e.g., beam search can yield more coherent text compared to top-k decoding but can yield shorter text if a large beam size is used) (Huang et al., 2017). Thus, in comparing different model generations, the summary text length has also been used as an informativeness measure since a shorter text typically preserves less information (Singh & Jin, 2016).
These quality criterion are widely used as evaluation metrics for human evaluation in document summarization. They can be used to compare a system-generated summary to a source text, a human-generated summary, or to another system-generated summary.
One thing that is usually overlooked in document summarization tasks is evaluating the generated summaries’ factual correctness. It has been shown in many recent work on summarization that models frequently generate factually incorrect text. This is partially because the models are not trained to be factually consistent and can generate about anything related to the prompt, Table 6 shows a sample summarization model output, in which the claims made are not consistent with the source document (Kryscinski et al., 2019b). Zhang et al.
It is imperative that the summarization models are factually consistent and that any conflicts between a source document and its generated summary (commonly referred to as faithfulness (Durmus et al., 2020; Wang et al., 2020b)) can be easily measured, especially for domain-specific summarization tasks like patient-doctor conversation summarization or business meeting summarization. As a result, factual-consistency-aware and faithful text generation research has drawn a lot of attention in the community in recent years (Kryscinski et al., 2019a, b; Zhang et al., 2019b; Wang et al., 2020a; Durmus et al., 2020; Wang et al., 2020b). A common approach is to use a model-based approach, in which a separate component is built on top of a summarization engine that can evaluate the generated summary based on factual consistency, as discussed in Section 4.5.
1.2 Extrinsic Summarization Evaluation Methods
Extrinsic evaluation metrics test the generated summary text by how it impacts the performance of downstream tasks, such as relevance assessment, reading comprehension, and question answering. Cohan & Goharian (2016) propose a new metric, sera (Summarization Evaluation by Relevance Analysis), for summarization evaluation based on the content relevance of the generated summary and the human-written summary. They find that this metric yields higher correlation with human judgments compared to rouge, especially on the task of scientific article summarization. Eyal et al. (2019a) and Wang et al. (2020a) measure the performance of a summary by using it to answer a set of questions regarding the salient entities in the source document.
2 Long Text Generation Evaluation
A long text generation system aims to generate multi-sentence text, such as a single paragraph or a multi-paragraph document. Common applications of long-form text generation are document-level machine translation, story generation, news article generation, poem generation, summarization, and image description generation, to name a few. This research area presents a particular challenge to state-of-the-art approaches that are based on statistical neural models, which are proven to be insufficient to generate coherent long text. As an example, in Figure 16 and 17 we show two generated text from two long-text generation models, Grover (Zellers et al., 2019) and PlotMachines (Rashkin et al., 2020). Both of these controlled text models are designed to generate a multi-paragraph story given a list of attributes (in these examples a list of outline points are provided), and the models should generate a coherent long story related to the outline points. These examples demonstrate some of the cohesion issues with these statistical models. For instance, in the Grover output, the model often finishes the story and then starts a new story partway through the document. In contrast, PlotMachines adheres more to a beginning-middle-ending structure. For example, GPT-2 (Radford et al., 2018) can generate remarkably fluent sentences, and even paragraphs, for a given topic or a prompt. However, as more sentences are generated and the text gets longer, it starts to wander, switching to unrelated topics and becoming incoherent (Rashkin et al., 2020).
Evaluating long-text generation is a challenging task. New criteria need to be implemented to measure the quality of long generated text, such as inter-sentence or inter-paragraph coherence in language style and semantics. Although human evaluation methods are commonly used, we focus our discussion on automatic evaluation methods in this section.
Text with longer context (e.g., documents, longer conversations, debates, movies scripts, etc.) usually consist of sections (e.g., paragraphs, sets, topics, etc.) that constitute some structure, and in natural language generation such structures are referred to as discourse (Jurafsky & Martin, 2009). Considering the discourse structure of the generated text is crucial in evaluating the system. Especially in open-ended text generation, the model needs to determine the topical flow, structure of entities and events, and their relations in a narrative flow that is coherent and fluent. One of the major tasks in which discourse plays an important role is document-level machine translation (Gong et al., 2015). Hajlaoui & Popescu-Belis (2013) present a new metric called Accuracy of Connective Translation (ACT) (Meyer et al., 2012) that uses a combination of rules and automatic metrics to compare the discourse connection between the source and target documents. Joty et al. (2017), on the other hand, compare the source and target documents based on the similarity of their discourse trees.
2.2 Evaluation via Lexical Cohesion
Lexical cohesion is a surface property of text and refers to the way textual units are linked together grammatically or lexically. Lexical similarity (Lapata & Barzilay, 2005) is one of the most commonly used metrics in story generation. Roemmele et al. (2017) filter the -grams based on lexical semantics and only use adjectives, adverbs, interjections, nouns, pronouns, proper nouns, and verbs for lexical similarity measure. Other commonly used metrics compare reference and source text on word- (Mikolov et al., 2013) or sentence-level (Kiros et al., 2015) embedding similarity averaged over the entire document. Entity co-reference is another metric that has been used to measure coherence (Elsner & Charniak, 2008). An entity should be referred to properly in the text and should not be used before introduced. Roemmele et al. (2017) capture the proportion of the entities in the generated sentence that are co-referred to an entity in the corresponding context as a metric of entity co-reference, in which a higher co-reference score indicates higher coherence.
In machine translation, Wong & Kit (2019) introduce a feature that can identify lexical cohesion at the sentence level via word-level clustering using WordNet (Miller, 1995) and stemming to obtain a score for each word token, which is averaged over the sentence. They find that this new score improves correlation of bleu and ter with human judgments. Other work, such as Gong et al. (2015), uses topic modeling together with automatic metrics like bleu and meteor to evaluate lexical cohesion in machine translation of long text. Chow et al. (2019) investigate the position of the word tokens in evaluating the fluency of the generated text. They modify wmd by adding a fragmentation penalty to measure the fluency of a translation for evaluating machine translation systems.
2.3 Evaluation via Writing Style
Gamon (2004) show that an author’s writing is consistent in style across a particular work. Based on this finding, Roemmele et al. (2017) propose to measure the quality of generated text based on whether it presents a consistent writing style. They capture the category distribution of individual words between the story context and the generated following sentence using their part-of-speech tags of words (e.g., adverbs, adjectives, conjunctions, determiners, nouns, etc.).
Text style transfer reflects the creativity of the generation model in generating new content. Style transfer can help rewrite a text in a different style, which is useful in creative writing such as poetry generation (Ghazvininejad et al., 2016). One metric that is commonly used in style transfer is the classification score obtained from a pre-trained style transfer model (Fu et al., 2018). This metric measures whether a generated sentence has the same style as its context.
2.4 Evaluation with Multiple References
One issue of evaluating text generation systems is the diversity of generation, especially when the text to evaluate is long. The generated text can be fluent, valid given the input, and informative for the user, but it still may not have lexical overlap with the reference text or the prompt that was used to constrain the generation. This issue has been investigated extensively (Li et al., 2016; Montahaei et al., 2019b; Holtzman et al., 2020; Welleck et al., 2019; Gao et al., 2019). Using multiple references that cover as many plausible outputs as possible is an effective solution to improving the correlation of automatic evaluation metrics (such as adequacy and fluency) with human judgments, as demonstrated in machine translation (Han, 2018; Läubli et al., 2020) and other NLG tasks.
Conclusions and Future Directions
Text generation is central to many NLP tasks, including machine translation, dialog response generation, document summarization, etc. With the recent advances in neural language models, the research community has made significant progress in developing new NLG models and systems for challenging tasks like multi-paragraph document generation or visual story generation. With every new system or model comes a new challenge of evaluation. This paper surveys the NLG evaluation methods in three categories:
Human-Centric Evaluation. Human evaluation is the most important for developing NLG systems and is considered the gold standard when developing automatic metrics. But it is expensive to execute, and the evaluation results are difficult to reproduce.
Untrained Automatic Metrics. Untrained automatic evaluation metrics are widely used to monitor the progress of system development.
A good automatic metric needs to correlate well with human judgments. For many NLG tasks, it is desirable to use multiple metrics to gauge different aspects of the system’s quality.
Machine-Learned Evaluation Metrics. In the cases where the reference outputs are not complete, we can train an evaluation model to mimic human judges. However, as pointed out in Gao et al. (2019), any machine-learned metrics might lead to potential problems such as overfitting and ‘gaming of the metric.’
We conclude this paper by summarizing some of the challenges of evaluating NLG systems:
Detecting machine-generated text and fake news. As language models get stronger by learning from increasingly larger corpora of human-written text, they can generate text that is not easily distinguishable from human-authored text. Due to this, new systems and evaluation methods have been developed to detect if a piece of text is machine- or human-generated. A recent study (Schuster et al., 2019) reports the results of a fact verification system to identify inherent bias in training datasets that cause fact-checking issues. In an attempt to combat fake news, Vo & Lee (2019) present an extensive analysis of tweets and a new tweet generation method to identify fact-checking tweets (among many tweets), which were originally produced to persuade posters to stop tweeting fake news. Gehrmann et al. introduced GLTR, which is a tool that helps humans to detect if a text is written by a human or generated by a model. Other research focuses on factually correct text generation, with a goal of providing users with accurate information. Massarelli et al. (2019) introduce a new approach for generating text that is factually consistent with the knowledge source. Kryscinski et al. (2019b) investigate methods of checking the consistency of a generated summary against the document from which the summary is generated. Zellers et al. (2019) present a new controllable language model that can generate an article with respect to a given headline, yielding more trustworthy text than human-written text of fake information. Nevertheless, large-scale language models (even controllable ones), have a tendency to hallucinate and generate nonfactual information, which the model designers should measure and prevent. Future work should focus on the analysis of the text generated from large-scale language models, emphasize careful examination of such models in terms of how they learn and reproduce potential biases that in the training data (Sheng et al., 2020; Bender et al., 2021).
Making evaluation explainable. Explainable AI refers to AI and machine learning methods that can provide human-understandable justifications for their behaviour (Ehsan et al., 2019). Evaluation systems that can provide reasons for their decisions are beneficial in many ways. For instance, the explanation could help system developers to identify the root causes of the system’s quality problems such as unintentional bias, repetition, or factual inconsistency. The field of explainable AI is growing, particularly in generating explanations of classifier predictions in NLP tasks (Ribeiro et al., 2016, 2018; Thorne et al., 2019). Text generation systems that use evaluation methods that can provide justification or explanation for their decisions will be more trusted by their users. Future NLG evaluation research should focus on developing easy-to-use, robust, and explainable evaluation tools.
Improving corpus quality. Creating high-quality datasets with multiple reference texts is essential for not only improving the reliability of evaluation but also for allowing the development of new automatic metrics that correlate well with human judgments (Belz & Reiter, 2006). Among many important critical aspects of building corpora for natural language generation tasks, the accuracy, timeliness, completeness, cleanness and unbiasedness of the data plays a very important role. The collected corpus (whether created manually or automatically through retrieval or generation) must be accurate so the generation models can serve for the downstream tasks more efficiently. The corpora used for language generation tasks should be relevant to the corresponding tasks so intended performance can be achieved. Missing information, information biased toward certain groups, ethnicities, religions, etc. could prevent the models from gathering accurate insights and could damage the efficiency of the task performance (Eckart et al., 2012; McGuffie & Newhouse, 2020; Barbaresi, 2015; Bender et al., 2021; Gehrmann et al., 2021).
Standardizing evaluation methods. Most untrained automatic evaluation metrics are standardized using open source platforms like Natural Language Toolkit (NLTK)nltk.org or spaCyspacy.io. Such platforms can significantly simplify the process of benchmarking different models. However, there are still many NLG tasks that use task-specific evaluation metrics, such as metrics to evaluate the contextual quality or informativeness of generated text. There are also no standard criteria for human evaluation methods for different NLG tasks.
It is important for the research community to collaborate more closely to standardize the evaluation metrics for NLG tasks that are pursued by many research teams. One effective way to achieve this is to organize challenges or shared tasks, such as the Evaluating Natural Language Generation Challengehttps://framalistes.org/sympa/info/eval.gen.chal and the Shared Task on NLG Evaluationhttps://github.com/evanmiltenburg/Shared-task-on-NLG-Evaluation.
Developing effective human evaluations. For most NLG tasks, there is little consensus on how human evaluations should be conducted. Furthermore, papers often leave out important details on how the human evaluations were run, such as who the evaluators are and how many people evaluated the text (van der Lee et al., 2019). Clear reporting of human evaluations is very important, especially for replicability purposes.
We encourage NLG researchers to design their human evaluations carefully, paying attention to best practices described in NLG and crowdsourcing research, and to include the details of the studies and data collected from human evaluations, where possible, in their papers. This will allow new research to be consistent with previous work and enable more direct comparisons between NLG results. Human evaluation-based shared tasks and evaluation platforms can also provide evaluation consistency and help researchers directly compare how people perceive and interact with different NLG systems.
Evaluating ethical issues. There is still a lack of systematic methods for evaluating how effectively an NLG system can avoid generating improper or offensive language. The problem is particularly challenging when the NLG system is based on neural language models whose output is not always predictable. As a result, many social chatbots, such as XiaoIce (Zhou et al., 2020), resort to hand-crafted policies and editorial responses to make the system’s behavior predictable. However, as pointed out by Zhou et al. (2020), even a completely deterministic function can lead to unpredictable behavior. For example, a simple answer “Yes” could be perceived as offensive in a given context. For these reasons and others, NLG evaluations should also consider the ethical implications of their potential responses and applications. We should also note that the landscape and focus of ethics in AI in general is constantly changing due to new advances in neural text generation, and as such, continuing development of ethical evaluations of the machine-generated content is crucial for new advances in the field.
We encourage researchers working in NLG and NLG evaluation to focus on these challenges moving forward, as they will help sustain and broaden the progress we have seen in NLG so far.