Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation
Yixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev
Introduction
Human evaluation plays an essential role in both assessing the rapid development of summarization systems in recent years Lewis et al. (2020a); Zhang et al. (2020a); Brown et al. (2020); Sanh et al. (2022); He et al. (2022) and in assessing the ability of automatic metrics to evaluate such systems as a proxy for manual evaluation Bhandari et al. (2020); Fabbri et al. (2022a); Gao and Wan (2022). However, while human evaluation is regarded as the gold standard for evaluating both summarization systems and automatic metrics, as suggested by Clark et al. (2021) an evaluation study does not become “gold” automatically without proper practices. For example, achieving a high inter-annotator agreement among annotators can be difficult Goyal et al. (2022), and there can be a near-zero correlation between the annotations of crowd-workers and expert annotators Fabbri et al. (2022a). Also, a human evaluation study without a large enough sample size can fail to find statistically significant results due to insufficient statistical power Card et al. (2020).
Therefore, we believe it is important to ensure that human evaluation can indeed serve as a solid foundation for evaluating summarization systems and automatic metrics. For this, we propose using a robust human evaluation protocol for evaluating the salience of summaries that is more objective by dissecting the summaries into fine-grained content units and defining the annotation task based on those units. Specifically, we introduce the Atomic Content Unit (ACU) protocol for summary salience evaluation (§3), which is modified from the Pyramid Nenkova and Passonneau (2004) and LitePyramid Shapira et al. (2019) protocols. We demonstrate that with the ACU protocol, a high inter-annotator agreement can be established among crowd-workers, which leads to more stable system evaluation results and better reproducibility.
We then collect, through both in-house annotation and crowdsourcing, RoSE, a large human evaluation benchmark of human-annotated summaries with the ACU evaluation protocol on recent state-of-the-art summarization systems, which yields higher statistical power (§4). To support evaluation across datasets and domains, our benchmark consists of test sets over three summarization datasets, CNN/DailyMail (CNNDM) Nallapati et al. (2016), XSum Narayan et al. (2018), and SamSum Gliwa et al. (2019), and annotations on the validation set of CNNDM to facilitate automatic metric training. To gain further insights into the characteristics of different evaluation protocols, we conduct human evaluation with three other protocols (§5). Specifically, we analyze protocol differences in the context of both fine-tuned models and large language models (LLMs) in a zero-shot setting such as GPT-3 Brown et al. (2020). We find that different protocols can lead to drastically different results, which can be affected by annotators’ prior preferences, highlighting the importance of aligning the protocol with the summary quality intended to be evaluated. We note that our benchmark enables a more trustworthy evaluation of automatic metrics (§6), as shown by statistical characteristics such as tighter confidence intervals and more statistically significant comparisons (§6.2). Our evaluation includes recent methods based on LLMs Fu et al. (2023); Liu et al. (2023), and we found that they cannot outperform traditional metrics despite their successes on related benchmarks such as SummEval Fabbri et al. (2022a).
We summarize our key findings in Tab. 1. Our contributions are the following: (1) We propose the ACU protocol for high-agreement human evaluation of summary salience. (2) We curate the RoSE benchmark, consisting of 22000 summary-level annotations and requiring over 150 hours of in-house annotation, across three summarization datasets, which can lay a solid foundation for training and evaluating automatic metrics.We release our benchmark and evaluation scripts at https://github.com/Yale-LILY/ROSE. (3) We compare four human evaluation protocols for summarization and show how they can lead to drastically different model preferences. (4) We evaluate automatic metrics across different human evaluation protocols and call for human evaluation to be conducted with a clear evaluation target aligned with the evaluated systems or metrics, such that task-specific qualities can be evaluated without the impact of general, input-agnostic preferences of annotators. We note that the implications of our findings can become even more critical with the progress of LLMs trained with human preference feedback (Ouyang et al., 2022) and call for a more rigorous human evaluation of LLM performance.
Related Work
Human Evaluation Benchmarks Human annotations are essential to the analysis of summarization research progress. Thus, recent efforts have focused on aggregating model outputs and annotating them according to specific quality dimensions Huang et al. (2020); Bhandari et al. (2020); Stiennon et al. (2020); Zhang and Bansal (2021); Fabbri et al. (2022a); Gao and Wan (2022). The most relevant work to ours is Bhandari et al. (2020), which annotates summaries according to semantic content units, motivated by the Pyramid Nenkova and Passonneau (2004) and LitePyramid Shapira et al. (2019) protocols. However, this benchmark only covers a single dataset (CNNDM) without a focus on similarly-performing state-of-the-art systems, which may skew metric analysis Tang et al. (2022a) and not fully reflect realistic scenarios Deutsch et al. (2022). In contrast, our benchmark consists only of outputs from recently-introduced models over three datasets.
Summarization Meta-Evaluation With a human evaluation dataset, there exist many directions of meta-evaluation, or re-evaluation of the current state of evaluation, such as metric performance analyses, understanding model strengths, and human evaluation protocol comparisons.
Within metric meta-analysis, several studies have focused on the analysis of ROUGE Lin (2004b), and its variations Rankel et al. (2013); Graham (2015), across domains such as news Lin (2004a), meeting summarization (Liu and Liu, 2008), and scientific articles (Cohan and Goharian, 2016). Other studies analyze a broader set of metrics Peyrard (2019); Bhandari et al. (2020); Deutsch and Roth (2020); Fabbri et al. (2022a); Gabriel et al. (2021); Kasai et al. (2022b), including those specific to factual consistency evaluation Kryscinski et al. (2020); Durmus et al. (2020); Wang et al. (2020); Maynez et al. (2020); Laban et al. (20d); Fabbri et al. (2022b); Honovich et al. (2022); Tam et al. (2022).
Regarding re-evaluating model performance, a recent line of work has focused on evaluating zero-shot large language models Goyal et al. (2022); Liang et al. (2022); Tam et al. (2022), noting their high performance compared to smaller models.
As for the further understanding of human evaluation, prior work has compared approaches to human evaluation Hardy et al. (2019), studied annotation protocols for quality dimensions such as linguistic quality Steen and Markert (2021) and factual consistency Tang et al. (2022b), and noted the effects of human annotation inconsistencies on system rankings Owczarzak et al. (2012). The unreliability and cost of human evaluation in certain settings have been emphasized Chaganty et al. (2018); Clark et al. (2021), with some work noting that thousands of costly data points may need to be collected in order to draw statistically significant conclusions Wei and Jia (2021). Our meta-analysis focuses on this latter aspect, and we further analyze potential confounding factors in evaluation such as length and protocol design, with respect to both small and large zero-shot language models.
Atomic Content Units for Summarization Evaluation
We now describe our Atomic Content Unit (ACU) annotation protocol for reference-based summary salience evaluation, including the procedure of writing ACUs based on reference summaries and matching the written ACUs with system outputs.
In this work, we focus on a specific summarization meta-evaluation study on summary salience. Salience is a desired summary quality that requires the summary to include all and only important information of the input article. The human evaluation of summary salience can be conducted in either reference-free or reference-based manners. The former asks the annotators to assess the summary directly based on the input article Fabbri et al. (2022a), while the latter requires the annotators to assess the information overlap between the system output and reference summary Bhandari et al. (2020), under the assumption that the reference summary is the gold standard of summary salience.We note salience can be an inherently subjective quality, and the reference summary of common datasets may not always be the actual “gold standard”, discussed more in §7. Given that reference-based protocols are more constrained, we focus on reference-based evaluation for our human judgment dataset collection, and we conduct a comparison of protocols in §5.
2 ACU Annotation Protocol
Inspired by the Pyramid Nenkova and Passonneau (2004) and LitePyramid Shapira et al. (2019) protocols and subsequent annotation collection efforts Bhandari et al. (2020); Zhang and Bansal (2021), the ACU protocol is designed to reduce the subjectivity of reference-based human evaluation by simplifying the basic annotation unit – the annotators only need to decide on the presence of a single fact, extracted from one text sequence, in another text sequence, to which a binary label can be assigned with more objectivity. Specifically, the evaluation process is decomposed into two steps: (1) ACU Writing – extracting facts from one text sequence, and (2) ACU Matching – checking for the presence of the extracted facts in another sequence. We formulate the ACU protocol as a recall-based protocol, such that the first step only needs to be performed once for the reference summary, allowing for reproducibility and reuse of these units when performing matching on new system outputs.
ACU Writing While the LitePyramid approach defines its basic content unit as a sentence containing a brief fact, we follow Bhandari et al. (2020) to emphasize a shorter, more fine-grained information unit. Specifically, we define the ACU protocol with the concept of atomic facts – elementary information units in the reference summaries, which no longer need to be further split for the purpose of reducing ambiguity in human evaluation.We note that it can be impossible to provide a practical definition of atomic facts. Instead, we use it as a general concept for fine-grained information units. Then, ACUs are constructed based on one atomic fact and other minimal, necessary information.
Fig. 1 shows an example of the written ACUs. To ensure annotation quality, we (the authors) write all the ACUs used in this work. We define guidelines to standardize the annotation process; for each summary sentence the annotator creates an ACU constituting the main information from the subject of the main clause (e.g., root), followed by additional ACUs for other facts while including the minimal necessary information from the root. We provide rules for dealing with quotations, extraneous adjectives, noisy summaries, and additional cases. We note that there can still be inherent subjectivity in the written ACUs among different annotators even with the provided guidelines. However, such subjectivity should be unbiased in summary comparison because all the candidate summaries are evaluated by the same set of written ACUs.
ACU Matching Given ACUs written for a set of reference summaries, our protocol evaluates summarization system performance by checking the presence of the ACUs in the system-generated summaries as illustrated in Fig. 1. For this step, we recruit annotators on Amazon Mechanical Turkhttps://www.mturk.com/ (MTurk). The annotators must pass a qualification test, and additional requirements are specified in Appendix A. Besides displaying the ACUs and the system outputs, we also provide the reference summaries to be used as context for the ACUs.
Scoring Summaries with ACU ACU matching annotations can be aggregated into summary scores. We first define an un-normalized ACU score of a candidate summary given a set of ACUs as:
3 ACU Annotation Collection
We collect ACU annotations on three summarization datasets: CNNDM Nallapati et al. (2016), XSum Narayan et al. (2018), and SamSum Gliwa et al. (2019). To reflect the latest progress in text summarization, we collect and annotate the generated summaries of pre-trained summarization systems proposed in recent years.We release all of the system outputs with a unified, cased, untokenized format to facilitate future research. Detailed information about the summarization systems we used can be found in Appendix A.2.
Table 2 shows the statistics of the collected annotations. The annotations are collected from the test set of the above datasets, and additionally from the validation set of CNNDM to facilitate the training of automatic evaluation metrics. In total, we collect around 21.8k ACU-level annotations and around 22k summary-level annotations, aggregated over around 50k individual summary-level judgments.
To calculate inter-annotator agreement, we use Krippendorff’s alpha (Krippendorff, 2011). The aggregated summary-level agreement score of ACU matching is 0.7571, and the ACU-level agreement score is 0.7528. These agreement scores are higher than prior collections, such as RealSumm Bhandari et al. (2020) and SummEval Fabbri et al. (2022a), which have an average agreement score of crowd-workers 0.66 and 0.49, respectively.
RoSE Benchmark Analysis
We first analyze the robustness of our collected annotations and a case study on the system outputs.
We analyze the statistical power of our collected human annotations to study whether it can yield stable and trustworthy results Card et al. (2020). Statistical power is the probability that the null hypothesis of a statistical significance test is rejected given there is a real effect. For example, for a human evaluation study that compares the performance of two genuinely different systems, a statistical power of 0.80 means there is an 80% chance that a significant difference will be observed. Further details can be found in Appendix B.1.
We conduct the power analysis for pair-wise system comparisons with ACU scores (Eq. 1) focusing on two factors, the number of test examples and the observed system difference. Specifically, we run the power analysis with varying sample sizes, and group the system pairs into buckets according to their performance difference, as determined by ROUGE1 recall scores (Fig.2).We note that these scores are proxies of the true system differences, and the power analysis is based on the assumption that the systems have significantly different performance. We observe the following: (1) A high statistical powerAn experiment is usually considered sufficiently powered if the statistical power is over 0.80. is difficult to reach when the system performance is similar. Notably, while the sample size of the human evaluation performed in recent work is typically around 50-100,We provide a brief survey of the practices of human evaluation in recent text summarization research in Appendix F. such sample size can only reach a power of 0.80 when the ROUGE1 recall score difference is above 5. (2) Increasing the sample size can effectively raise the statistical power. For example, when the system performance difference is within the range of 1-2 points, the power of a 500-sample set is around 0.50 while a 100-sample set only has a power of around 0.20. The results of power analysis on three datasets with both ROUGE and ACU score differences are provided in Appendix B.2 with the same patterns, which indicates that our dataset can provide more stable summarization system evaluation thanks to its higher statistical power.
2 Summarization System Analysis
As a case study, in Tab. 3 we analyze the summary characteristics of the recent summarization systems we collected on the CNNDM test set. XSum and SamSum results are shown in Appendix A.3. Apart from the ACU scores, we note that the average summary length of different systems can greatly vary, and such differences are not always captured by the widely-used ROUGE F1. For example, the length of GSum Dou et al. (2021) is around 40% longer than GLOBAL Ma et al. (2021) while they have very similar ROUGE1 F1 scores. Besides, we note all systems in Tab. 3 have longer summaries than the reference summaries, whose average length is only 54.93. This can be a potential risk to users who may prefer shorter, more concise summaries. Meanwhile, the systems that generate longer summaries may be favored by users who prefer more informative summaries. Therefore, we join the previous work Sun et al. (2019); Song et al. (2021); Gehrmann et al. (2022); Goyal et al. (2022) in advocating treating summary lengths as a separate aspect of summary quality in evaluation, as in earlier work in summarization research.For example, the DUC evaluation campaigns set a pre-specified maximum summary length, or summary budget.
Evaluating Annotation Protocols
Apart from ACU annotations, we collect human annotations with three different protocols to better understand their characteristics. Specifically, two reference-free protocols are investigated: Prior protocol evaluates the annotators’ preferences of summaries without the input document, while Ref-free protocol evaluates if summaries cover the salient information of the input document. We also consider one reference-based protocol, Ref-based, which evaluates the content similarity between the generated and reference summaries. Appendix D.1 provides detailed instructions for each protocol.
We collected three annotations per summary on a 100-example subset of the above CNNDM test set using the same pool of workers from our ACU qualification. Except for ACU, all of the summaries from different systems are evaluated within a single task with a score from 1 (worst) to 5 (best), similar to the EASL protocol (Sakaguchi and Van Durme, 2018). We collect (1) annotations of the 12 above systems, with an inter-annotator agreement (Krippendorff’s alpha) of 0.3455, 0.2201, 0.2741 on Prior, Ref-free, Ref-based protocols respectively; (2) annotations for summaries from GPT-3 Brown et al. (2020),We use the “text-davinci-002” version of GPT-3. T0 Sanh et al. (2022), BRIO, and BART to better understand annotation protocols with respect to recently introduced large language models applied to zero-shot summarization.
2 Results Analysis
We investigate both the summary-level and system-level correlations of evaluation results of different protocols to study their inherent similarity. Details of correlation calculation are in Appendix C.
Results on Fine-tuned Models We show the system-level protocol correlation when evaluating the fine-tuned models in Tab. 4, and the summary-level correlation can be found in Appendix D.2. We use the normalized ACU score (Eq. 2) because the other evaluation protocols are supposed to resemble an F1 score, while the ACU score is by definition recall-based. We have the following observations:
(1) The Ref-free protocol has a strong correlation with the Prior protocol, suggesting that the latter may have a large impact on the annotator’s document-based judgments.
(2) Both the Prior and Ref-free protocols have a strong correlation with summary length, showing that annotators may favor longer summaries.
(3) The Ref-free protocol and the Ref-based protocol have a negative correlation while ideally they are supposed to measure similar quality aspects.
We perform power analysis on the results following the procedure in §4.1 and found that ACU protocol can yield higher statistical power than the Ref-based protocol, suggesting that the ACU protocol leads to more robust evaluation results. We also found that the reference-free Prior and Ref-free protocols have higher power than the reference-based protocols. However, we note that they are not directly comparable because they have different underlying evaluation targets, as shown by the near-zero correlation between them. Further details are provided in Appendix D.2.
Results on Large Language Models The results are shown in Tab. 5. Apart from the system outputs, we also annotate reference summaries for reference-free protocols. We found that under the Ref-free protocol, GPT-3 receives the highest score while the reference summary is the least favorite one, similar to the findings of recent work Goyal et al. (2022); Liang et al. (2022). However, we found the same pattern with the Prior protocol, showing that the annotators have a prior preference for GPT-3. We provide an example in Appendix D.2 comparing GPT-3 and BRIO summaries under different protocols. Given the strong correlation between the Prior and Ref-free protocols, we note that there is a risk that the annotators’ decisions are affected by their prior preferences that are not genuinely related to the task requirement. As a further investigation, we conduct an annotator-based case study including 4 annotators who annotated around 20 examples in this task, in which we compare two summary-level correlations (Eq. 3) given a specific annotator: (1) the correlation between their own Ref-free protocol scores and Prior scores; (2) the correlation between their Ref-free scores and the Ref-free scores averaged over the other annotations on each example. We found that the average value of the former is 0.404 while the latter is only 0.188, suggesting that the annotators’ own Prior score is a better prediction of their Ref-free score than the Ref-free score of other annotators.
Evaluating Automatic Metrics
We analyze several representative automatic metrics, with additional results in Appendix E on 50 automatic metric variants. We focus the metric evaluation on ACU annotations because of two insights from §5: (1) Reference-based metrics should be evaluated with reference-based human evaluation. (2) ACU protocol provides higher statistical power than the summary-level Ref-based protocol.
We use the correlations between automatic metric scores and ACU annotation scores of system outputs to analyze and compare automatic metric performance. The following metrics are evaluated:
(1) lexical overlap based metrics, ROUGE (Lin, 2004b), METEOR (Lavie and Agarwal, 2007), CHRF (Popović, 2015); (2) pre-trained language model based metrics, BERTScore (Zhang et al., 2020c), BARTScore (Yuan et al., 2021); (3) question-answering based metrics, SummaQA (Scialom et al., 2019), QAEval (Deutsch et al., 2021a); (4) Lite3Pyramid (Zhang and Bansal, 2021), which automates the LitePyramid evaluation process; (5) evaluation methods based on large language models, GPTScore Fu et al. (2023) and G-Eval Liu et al. (2023), with two variants that are based on GPT-3.5OpenAI’s gpt-3.5-turbo-0301: https://platform.openai.com/docs/models/gpt-3-5. (G-Eval-3.5) and GPT-4OpenAI’s gpt-4-0314: https://platform.openai.com/docs/models/gpt-4. OpenAI (2023) (G-Eval-4) respectively. We note that for LLM-based evaluation we require the metric to calculate the recall score. For G-Eval-3.5 we report two variants that are based on greedy decoding (G-Eval-3.5) and sampling (G-Eval-3.5-S) respectively, Details of the LLM-based evaluation are in Appendix E.2.
Tab. 6 shows the results, with additional results of more metrics in Appendix E.3. We note:
(1) Several automatic metrics from the different families of methods (e.g., ROUGE, BARTScore) are all able to achieve a relatively high correlation with the ACU scores, especially at the system level.
(2) Metric performance varies across different datasets. In particular, metrics tend to have stronger correlations on the SamSum dataset and weaker correlations on the XSum dataset. We hypothesize that one reason is that the reference summaries of the XSum dataset contain more complex structures.
(3) Despite their successes Fu et al. (2023); Liu et al. (2023) in other human evaluation benchmarks such as SummEval, LLM-based automatic evaluation cannot outperform traditional methods such as ROUGE on RoSE. Moreover, their low summary-level correlation with ACU scores suggests that their predicted scores may not be well-calibrated.
Following Deutsch et al. (2022), we further investigate metric performance when evaluating system pairs with varying performance differences. Specifically, we group the system pairs based on the difference of their ACU scores into different buckets and calculate the modified Kendall’s correlation (Deutsch et al., 2022) on each bucket. The system pairs in each bucket are provided in Appendix E.4. Tab. 7 shows that the automatic metrics generally perform worse when they are used to evaluate similar-performing systems.
2 Analysis of Metric Evaluation
We analyze the metric evaluation with respect to the statistical characteristics and the impact of different human evaluation protocols on metric evaluation.
Confidence Interval We select several representative automatic metrics and calculate the confidence intervals of their system-level correlations with the ACU scores using bootstrapping. Similar to Deutsch et al. (2021b), we find that the confidence intervals are large. However, we found that having a larger sample size can effectively reduce the confidence interval, which further shows the importance of increasing the statistical power of the human evaluation dataset as discussed in §4.1. We provide further details in Appendix E.5.
Power Analysis of Metric Comparison We conduct a power analysis of pair-wise metric comparison with around 200 pairs, which corresponds to the chance of a statistical significance result being found. More details can be found in Appendix E.6. The results are in Fig.3, showing similar patterns as in the power analysis of summarization system comparison (§4.1):
(1) Significant results are difficult to find when the metric performance is similar;
(2) Increasing the sample size can effectively increase the chance of finding significant results.
Correlations under Different Human Evaluation Protocols We analyze the metric correlations under different human evaluation protocols (§5). The results are shown in Tab. 8, with more results in Appendix E.7. We note: (1) Metric performance differs greatly under different protocols, likely because the protocols can have weak correlations with each other (§5.2). (2) The reference-based automatic metrics generally perform better under reference-based evaluation protocols, but can have negative correlations with reference-free protocols.
Conclusion and Implications
We introduce RoSE, a benchmark whose underlying protocol and scale allow for more robust summarization evaluation across three datasets. With our benchmark, we re-evaluate the current state of human evaluation and its implications for both summarization system and automatic metric development, and we suggest the following:
(1) Alignment in metric evaluation. To evaluate automatic metrics, it is important to use an appropriate human evaluation protocol that captures the intended quality dimension to be measured. For example, reference-based automatic metrics should be evaluated by reference-based human evaluation, which disentangles metric performance from the impact of reference summaries.
(2) Alignment in system evaluation. We advocate for targeted evaluation, which clearly defines the intended evaluation quality. Specifically, text summarization, as a conditional generation task, should be defined by both the source and target texts along with pre-specified, desired characteristics. Clearly specifying characteristics to be measured can lead to more reliable and objective evaluation results. This will be even more important for LLMs pre-trained with human preference feedback for disentangling annotators’ prior preferences for LLMs with the task-specific summary quality.
(3) Alignment between NLP datasets and tasks. Human judgments for summary quality can be diverse and affected by various factors such as summary lengths, and reference summaries are not always favored. Therefore, existing summarization datasets (e.g. CNNDM) should only be used for the appropriate tasks. For example, they can be used to define a summarization task with specific requirements (e.g. maximum summary lengths), and be important for studying reference-based metrics.
Limitations
Biases may be present in the data annotator as well as in the data the models were pretrained on. Furthermore, we only include English-language data in our benchmark and analysis. Recent work has noted that language models may be susceptible to learning such data biases Lucy and Bamman (2021), thus we request that the users be aware of potential issues in downstream use cases.
As described in Appendix D.1, we take measures to ensure a high quality benchmark. There will inevitably be noise in the dataset collection process, either in the ACU writing or matching step, and high agreement of annotations does not necessarily coincide with correctness. However, we believe that the steps taken to spot check ACU writing and filter workers for ACU matching allow us to curate a high-quality benchmark. Furthermore, we encourage the community to analyze and improve RoSE in the spirit of evolving, living benchmarks Gehrmann et al. (2021).
For reference-based evaluation, questions about reference quality arise naturally. We also note that the original Pyramid protocol was designed for multi-reference evaluation and weighting of semantic content units, while we do not weight ACUs during aggregation. As discussed above, we argue that our benchmark and analysis are still valuable given the purpose of studying conditional generation and evaluating automatic metrics for semantic overlap in targeted evaluation. We view the collection of high-quality reference summaries as a valuable, orthogonal direction to this work, and we plan to explore ACU weighting in future work.
Acknowledgements
We thank the anonymous reviewers for their constructive comments. We are grateful to Arman Cohan for insightful discussions and suggestions, Daniel Deutsch for the initial discussions, Richard Yuanzhe Pang for sharing system outputs, and Philippe Laban for valuable comments.
References
Appendix A Benchmark Data Collection
We discuss the detailed settings of ACU collection in §3.2. To ensure the consistency of written ACUs among different annotators, we require each annotator to be familiar with the annotation protocol and proofread each other’s annotations to resolve any differences in initial annotations. After establishing a consistent understanding of the task, we have each reference summary annotated by one annotator. We note that there are multiple valid ways of writing the same atomic fact. In preliminary protocol analysis, we had multiple annotators write ACUs for the same reference summaries and did not find large differences in downstream inter-annotator agreement for ACU matching. The average time to write ACUs of one summary ranges from 2 to 5 minutes, and the overall annotation time for ACU writing is around 150 hours.
We use the following qualifications, in addition to a qualification test, to recruit MTurk workers with good track records: HIT approval rate greater than or equal to 98%, number of HITs approved greater than or equal to 10000, and located in either the United Kingdom or the United States. Workers were compensated between 0.55 per summary-level ACU HITs, with HITs bucketed according to the number of ACUs to be matched. For protocol comparison HITs, workers were compensated between 3. All HITs were carefully calibrated to equal a $12/hour pay rate.
The datasets we used for the collection are CNNDM, XSum and SamSum. The data release licenses are the Apache License for CNNDM and XSum, and CC BY-NC-ND 4.0 for SamSum. Our collected benchmark will be released under the 3-Clause BSD license.
A.2 Summarization Models
We list the summarization models for ACUs annotations on CNNDM, XSum, and SamSum in §3.3.
BART (Lewis et al., 2020b) introduce a denoising autoencoder for pretraining sequence to sequence tasks which is applicable to both natural language understanding and generation tasks.
Pegasus (Zhang et al., 2020b) introduce a model pretrained with a novel objective function designed for summarization by which important sentences are removed from an input document and then generated from the remaining sentences.
MatchSum Zhong et al. (2020) propose a summary-level extractive system using semantic match between the extracted summary and the source document.
CTRLSum He et al. (2020) introduce a method for controllable summarization based on keyword or descriptive prompt control tokens.
CLIFF Cao and Wang (2021) propose to use contrastive learning to improve factual consistency. We use the CLIFF output that uses an underlying BART model.
GOLD Pang and He (2021) frames text generation as an offline reinforcement learning problem, using importance weighting and assigning weights to examples that receive a higher probability from the generation model.
GSum Dou et al. (2021) is a framework for incorporating forms of summarization guidance.
SimCLS Liu and Liu (2021) is a two-stage summarization model where candidates from BART are reranking by a RoBERTa Liu et al. (2019) scoring model trained using contrastive learning.
FROST Narayan et al. (20d) propose to do content planning in both pretraining and finetuning summarization models with plans in the form of entity chains.
GLOBAL Ma et al. (2021) propose a variation of beam search that takes into account the global attention distribution.
BRIO Liu et al. (2022) proposes to train a summarization model both as a token-level generator and an evaluator of sequence candidates through contrastive reranking.
BRIO-Ext Liu et al. (2022) uses BRIO’s reranker on candidate extractive summaries from MatchSum.
The following models were included in protocol-comparison annotations.
T0 Sanh et al. (2022) introduces a prompt-based model that is fine-tuned on multiple tasks, including summarization.
GPT-3 Brown et al. (2020) is the davinci-002 model trained on human demonstrations and model outputs highly rated by humans.https://beta.openai.com/docs/model-index-for-researchers
XSum Systems
For XSum we reuse several of the above models with their XSum-trained checkpoints as well as several variations from the above paper due to the scarcity of widely-available, easily-reproducible XSum outputs.
CLIFF-Pegasus Cao and Wang (2021) is the CLIFF algorithm applied with Pegasus as the underlying model.
BRIO-ranking Liu et al. (2022) is the paper’s reranking model.
SamSum Systems
We use system outputs from Gao and Wan (2022).
UniLM (Dong et al., 2019) is a model pretrained on unidirection, bidirection, and sequence-to-sequence language modeling tasks.
Ctrl-DiaSumm (Liu and Chen, 2021) propose controlled generation using named entity plans.
PLM-BART (Feng et al., 2021) use DialogGPT Zhang et al. (2020d) to annotate input dialogues before finetuning.
CODS (Wu et al., 2021) propose a two-stage generation model that first generates a sketch that is then used as a signal to the second-stage summarizer.
MV-BART (Chen and Yang, 2020) propose a multi-view encoder and a decoder that attends to these conversation views.
S-BART (Chen and Yang, 2020) encodes utterances as well as action and discourse graphs and introduces a decoder that attends to these different levels of granularity.
A.3 ACU Scores of Summarization Models
We report the ACU scores of the summarization systems we annotated on the XSum and SamSum datasets (§3) in Tab. 9 and Tab. 10, respectively. The results on CNNDM can be found in Tab. 3. For the normalized ACU score (Eq. 2), we set the normalization strength to 2, 5, 0.5, on CNNDM, XSum, SamSum, respectively, by a grid search for de-correlating the summary length and the normalized score at the summary level.
Appendix B Power Analysis
We describe the algorithm for the power analysis in §4.1 in Alg.1. While prior work Card et al. (2020); Wei and Jia (2021) uses parametric methods to estimate statistical power, we conduct the power analysis with the bootstrapping test Tibshirani and Efron (1993) as recent work Deutsch et al. (2021b) has shown that the assumptions of the parametric methods do not always hold for human evaluation of text summarization. The process involves (1) iteratively sampling a set of examples with a certain sample size from an existing dataset, (2) running the significance test on the sampled set, and (3) estimating the power by averaging across the trials.
The essence of the test is to have a series of simulated datasets sampled from the existing dataset and run the significance test on the sampled sets. Here the existing dataset consists of human-annotated scores of system outputs. We use paired bootstrapping for the significance test. The power analysis is conducted over all the system pairs.
B.2 Powers of ACU Annotations
Fig.4, Fig.5, and Fig.6 show the power analysis results on CNNDM, XSum and SamSum respectively in §4.1, where the system pairs are grouped by their performance difference in either ACU or ROUGE1 recall scores. Similar to our findings on CNNDM in §4.1, we observe that increasing the sample size can effectively raise the statistical power.
Appendix C Calculating Correlations
We use correlations to analyze the inherent similarity between different human evaluation protocols, and the performance of automatic metrics, which is evaluated based on the correlations between the metric-calculated summary scores and the human-annotated summary scores. Specifically, given system outputs on each of the data samples and two different evaluation methods (e.g., human evaluation and an automatic metric) resulting in two -row, -column score matrices and , the summary-level correlation is an average of sample-wise correlations:
where , are the evaluation results on the -th data sample and is a function calculating a correlation coefficient (e.g., the Pearson correlation coefficient). In contrast, the system-level correlation is calculated on the aggregated system scores:
where and contain entries which are the system scores from the two evaluation methods averaged across data samples, e.g., .
Appendix D Protocol Comparison
The 100 examples chosen for annotation in §5 are a subset of the CNNDM ACU test set, and as here we aim to analyze trends among protocols as opposed to observing statistically significant differences among systems, we believe 100 examples suffice for this collection.
We summarize and compare different protocols in Tab. 11. We provide the following instructions to annotators for non-ACU annotations. We will release the full interface and instructions.
Prior: We ask the annotator to imagine each of the candidate summaries to be evaluated as a summary of a longer news article and answer the following question: how good do you think this summary is?
Ref-free: The rating measures how well the summary captures the key points of the news article. Consider whether all and only the important aspects are contained in the summary.
Ref-based: The rating measures how similar two summaries are. The similarity depends on if the summaries contain similar information, not if they use the same words.
D.2 Results Analysis
We present the result analysis of §5.2 here.
We show the summary-level Pearson’s Correlation Coefficients among different protocols in Tab. 12.
Power Analysis
The power analysis results on the Prior, Ref-free, Ref-based, and ACU protocols are shown in Fig. 7.
Case Study
We show a case study in Tab. 13 comparing the summaries generated by BRIO and GPT-3. GPT-3 scores higher on Prior and Ref-free (3.33/3.33 for BRIO and 3.66/4.00 for GPT-3). However, the BRIO summary scores 0.77 on un-normalized ACU annotations while GPT-3 scores 0.33. Also, Ref-based annotations favor BRIO over GPT-3 (3.66 vs. 3.33).
Appendix E Metric Analysis
We provide additional metric details as well as results for other metrics in §6. Note that for ROUGE, we use the Python implementation. https://pypi.org/project/ROUGE-score/
BLEU (Papineni et al., 2002) is a corpus-level precision-focused metric that calculates n-gram overlap and includes a brevity penalty.
CIDEr (Vedantam et al., 2015) computes {1-4}-gram co-occurrences, down-weighting common n-grams and calculating cosine similarity between the n-grams of the candidate and reference texts.
Statistics (Grusky et al., 2018) reports summary statistics such as the length, novel and repeated n-grams in the summary, the compression ratio between the summary and article, and measures of the level of extraction. Coverage is the percentage of words that are part of an extractive fragment and density is the average length of the extractive fragment each summary word belongs to.
MoverScore (Zhao et al., 2019) measures semantic distance with Word Mover’s Distance Kusner et al. (2015) on pooled BERT n-gram embeddings.
SUPERT (Gao et al., 2020) measures the semantic similarity of summaries with pseudo-reference summaries created by extracting salient sentences from the source documents.
BLANC (Vasilyev et al., 2020) measures the performance gains of a pre-trained language model on language understanding tasks on the input document when given access to a document summary.
QAEval (Deutsch et al., 2021a) reports both an F1 and exact match (em) score. We do not report the learned answer overlap metric.
SummaQA (Scialom et al., 2019) reports an F1 score and model confidence. We plan to report QuestEval Scialom et al. (2021) in a future version.
Lite3Pyramid includes four variations of the metric depending on the entailment model (two vs three-class entailment model) and how the output is used (as a probability vs a 0/1 label).
CTC (Deng et al., 2021) proposes metrics for Compression, transduction, and creation tasks as variations of textual alignment. Relevance is scored as the average bi-directional alignment between generated and reference summaries.
SimCSE (Gao et al., 2021) apply contrastive learning to learn improved sentence representations, which can then be used to compare generated and reference summary similarity.
UniEval (Zhong et al., 2022) frames text evaluation as the answer to yes or no questions, in our case whether the summary is relevant or not, and constructs pseudo-data to fine-tune language models for this setting.
E.2 Metrics based on Large Language Models
In §6.1 we evaluate two different LLM-based automatic evaluation methods.
GPTScore (Fu et al., 2023) formulates the text evaluation as the text-filling task and takes the token probability predicted by the LLMs as the quality score. We use the following prompt for calculating the recall score of the system outputs:
Answer the question based on the following reference summary and candidate summary.
Question: Can all of the information in the reference summary be found in the candidate summary? (a). Yes. (b). No.
The LLM-predicted probability of the last token, “Yes”, is used as the recall score. We use the OpenAI’s text-davinci-003 as the LLM.
G-Eval Liu et al. (2023) introduces a similar task as GPTscore, but has the LLM to predict a numerical score directly instead of using the LLM-predicted probability. We use the following prompt for the task:
You will receive a reference summary and a candidate summary. Your task is to compare these two summaries and assess the extent to which the candidate summary covers the information presented in the reference summary.
Please indicate your agreement with the following statement: “All of the information in the reference summary can be found in the candidate summary.”
Use the following 5-point scale when determining your response:
We note that we set the sampling temperature to 0 to ensure more deterministic behavior for G-Eval-3.5 and G-Eval-4. We also experiment with a sampling strategy with GPT-3.5 (G-Eval-3.5-S), where we sample 5 outputs with a temperate 1 and take the average score as the final prediction.
E.3 Metric Correlation with ACU Scores
We collect in total 50 different automatic metrics (including different variations of the same metric), and evaluate their performance using our collected ACU benchmark on CNNDM, XSum and SamSum datasets with three different correlation coefficients (§6.1). Tab. 15 reports the system-level correlation with the un-normalized ACU score (Eq. 1). Tab. 16 reports the summary-level correlation with the un-normalized ACU score (Eq. 1). Tab. 17 reports the system-level correlation with the normalized ACU score (Eq. 2). Tab. 18 reports the summary-level correlation with the normalized ACU score (Eq. 2).
E.4 System Pairs for Fine-grained Metric Evaluation
For metric elevation in §6.1, we provide the system pairs in the six different buckets grouped by their performance differences below.
Bucket 1: CLIFF V.S. FROST, CTRLSUM V.S. GSUM, BART V.S. CLIFF, GOLD V.S. FROST, BART V.S. FROST, CLIFF V.S. GOLD, BRIO V.S. GSUM, GOLD V.S. PEGASUS, BRIO V.S. CTRLSUM, BART V.S. GOLD, BRIO-EXT V.S. MATCHSUM.
Bucket 2: FROST V.S. PEGASUS, CLIFF V.S. PEGASUS, PEGASUS V.S. GLOB, BRIO-EXT V.S. SIMCLS, BART V.S. PEGASUS, BRIO V.S. MATCHSUM, BART V.S. SIMCLS, GOLD V.S. GLOB, CLIFF V.S. SIMCLS, MATCHSUM V.S. GSUM, SIMCLS V.S. FROST.
Bucket 3: MATCHSUM V.S. SIMCLS, FROST V.S. GLOB, MATCHSUM V.S. CTRLSUM, CLIFF V.S. GLOB, BRIO V.S. BRIO-EXT, GOLD V.S. SIMCLS, BART V.S. GLOB, BRIO-EXT V.S. GSUM, BRIO-EXT V.S. CTRLSUM, BART V.S. BRIO-EXT, SIMCLS V.S. PEGASUS.
Bucket 4: CLIFF V.S. BRIO-EXT, BRIO-EXT V.S. FROST, BRIO V.S. SIMCLS, GOLD V.S. BRIO-EXT, BART V.S. MATCHSUM, CLIFF V.S. MATCHSUM, SIMCLS V.S. GSUM, MATCHSUM V.S. FROST, SIMCLS V.S. GLOB, SIMCLS V.S. CTRLSUM, BRIO-EXT V.S. PEGASUS.
Bucket 5: GOLD V.S. MATCHSUM, MATCHSUM V.S. PEGASUS, BART V.S. BRIO, BRIO-EXT V.S. GLOB, BRIO V.S. CLIFF, BRIO V.S. FROST, BART V.S. GSUM, BART V.S. CTRLSUM, BRIO V.S. GOLD, CLIFF V.S. GSUM, FROST V.S. GSUM.
Bucket 6: CLIFF V.S. CTRLSUM, MATCHSUM V.S. GLOB, CTRLSUM V.S. FROST, GOLD V.S. GSUM, BRIO V.S. PEGASUS, GOLD V.S. CTRLSUM, PEGASUS V.S. GSUM, CTRLSUM V.S. PEGASUS, BRIO V.S. GLOB, GLOB V.S. GSUM, CTRLSUM V.S. GLOB.
E.5 Confidence Interval
We select several automatic metrics and calculate the confidence intervals of their system-level correlations with the ACU scores (§6.2). The results are in Fig. 8. Similar to Deutsch et al. (2021b), we found that the confidence intervals are large. However, having a larger sample size can effectively reduce the confidence interval. Specifically, we use re-sampling to generate a series of synthetic sample sets with several different sizes and calculate the confidence interval by averaging over the sampled sets with the same size. As shown in Fig. 9, larger sample sizes lead to more stable results.
E.6 Power Analysis of Metric Comparison
We use Alg.1 to conduct a power analysis of metric comparison based on their Kendall’s correlations with ACU scores (§6.2). We choose 20 metrics for comparison, resulting in 190 metric pairs in total, which are (1) BARTScore-r-parabank, (2) BERTScore-r-deberta, (3) BERTScore-r-roberta, (4) BLANC, (5) CHRF, (6) CTC, (7) Meteor, (8) Lite2Pyramid-p2c, (9) QAEval-em, (10) QAEval-f1, (11) ROUGE1, (12) ROUGE1r, (13) ROUGE2, (14) ROUGE2r, (15) ROUGEL, (16) ROUGELr, (17) SimCSE, (18) SummaQA, (19) SummaQA-prob, (20) SUPERT. We note that we use the permutation test instead of the paired bootstrapping test to calculate the statistical significance for metric comparison, since Deutsch et al. (2021b) found that the permutation test works better for detecting significant results in metric comparison.
E.7 Metric Correlation with Different Human Evaluation Protocols
We present the correlations between automatic metrics and different human evaluation protocols in Tab. 19 as discussed in §6.2.
Appendix F Human Evaluation Practices in Recent Text Summarization Research
We provide a brief survey for the human evaluation practices of 55 selected papers on text summarization published at NAACLhttps://aclanthology.org/events/naacl-2022/, ACLhttps://aclanthology.org/events/acl-2022/, and EMNLPhttps://preview.aclanthology.org/emnlp-22-ingestion/volumes/2022.emnlp-main/ from 2022. We follow the design of a similar study in Gehrmann et al. (2022) as described below. The results are shown in Tab. 14.
Performed Human Evaluation: Report “yes”, if a human evaluation of any kind is done. We report that 71% of analyzed papers did human evaluation.
Significance Test: Report “yes”, if a significance test is done on the human annotation results. Of the 39 papers that conducted human evaluation, a total of 27 papers reported the result of a significance test (68%), which is much higher compared to the 25% reported in the previous survey Gehrmann et al. (2022).
Power Analysis: Report “yes”, if a power analysis of any kind is mentioned. Of the 39 papers that conducted human evaluation, none of the papers did power analysis, the same as the result provided in the previous survey of Gehrmann et al. (2022). Inter-annotator Agreement: Report “yes”, if any kind of agreement test is conducted to evaluate the quality of human annotation themselves. Overall, we report a total of 12 papers (28%) that did agreement tests and documented specific agreement values. 9 out of the 12 papers recorded the specific agreement test, with Krippendorff’s alpha as the most commonly used measurement.
Participants (crowd-worker, expert, etc.) Report “yes”, if at least the number of human evaluators, document sample size, annotators per document, or their demographics is mentioned. We show the sample size of the conducted human evaluation study in Fig. 10, and note that around 93% of them are less or equal to 200.
Released Human Evaluation Data: Report “yes”, if the authors release the human evaluation data.