News Summarization and Evaluation in the Era of GPT-3

Tanya Goyal, Junyi Jessy Li, Greg Durrett

Introduction

Fine-tuning pre-trained models on domain-specific datasets has been the leading paradigm in text summarization research in recent years Lewis et al. (2020); Zhang et al. (2020); Raffel et al. (2020). These models generate high-quality summaries on standard benchmarks, but still require sizeable training datasets to adapt to new settings, e.g., summarizing data from a new source domain or producing a summary in a different style. The success of prompting large language models (GPT-3 Brown et al. (2020), T0 Sanh et al. (2022), PaLM Chowdhery et al. (2022), etc.) provides an alternative approach, namely learning from natural language task instructions and/or a few demonstrative examples in the context without updating model parameters. While recent work Zhao et al. (2021); Min et al. (2022); Ye and Durrett (2022) has evaluated this paradigm across a number of tasks, it has only been studied for text summarization with unreliable automatic metrics He et al. (2022b); Chowdhery et al. (2022); Ouyang et al. (2022) or in non-standard settings Saunders et al. (2022).

In this paper, we conduct the first systematic study of the impact of prompt-based models on the text summarization research space, using an Instruct-tuned 175B GPT-3 model (text-davinci-002) Brown et al. (2020); Ouyang et al. (2022) as a case study. Figure 1 shows that GPT-3 summaries are extremely high-quality and adaptable to different summarization settings. Starting from these observations, we aim to answer three main questions. First, how do prompt-based GPT-3 summaries compare to those obtained from state-of-the-art fine-tuned summarization models Zhang et al. (2020); Liu et al. (2022)? We compare these approaches using A/B testing on a new corpus of recent news articles, and find that our study participants overwhelmingly prefer GPT-3 summaries across two different “styles” with different prompts (three-sentence and single-sentence). Moreover, these summaries do not suffer from limitations due to low-quality training data that plague fine-tuned generic summarization models Maynez et al. (2020); Goyal et al. (2022).

Second, are existing automatic metrics well-suited to evaluating prompt-based summaries? Recent work has shown that classic reference-based such as ROUGE Lin (2004) and BERTScore Zhang* et al. (2020) are unreliable when small improvements are reported Peyrard (2019); Fabbri et al. (2021); however large differences, on the order of say 55 Rouge points or greater, are considered to be correlated with human preferences Bhandari et al. (2020); Deutsch et al. (2022). However, we find that the same is no longer true when evaluating GPT-3 summaries. These summaries score much lower on automatic metrics (77 ROUGE-L points on average) than all prior state-of-the-art models while comfortably outperforming them on human evaluation. Furthermore, we show that recent reference-free metrics, e.g. QA-based metrics Fabbri et al. (2022); Durmus et al. (2020) and trained factuality models Kryscinski et al. (2020); Goyal and Durrett (2020), similarly fail to adapt to this shift from the fine-tuned to prompting, and need to be re-visited.

Finally, how can prompting be used beyond generic summarization? We focus on keyword-based and aspect-based summarization. For keyword-based summarization, we find that GPT-3 consistently generates more coherent and keyword-relevant summaries compared to current fine-tuned alternatives: crowd annotators prefer GPT-3 summaries over a baseline model He et al. (2022a) 70% of the time. We observe mixed results for the aspect-based setting, where GPT-3 summaries show frequent failure cases with simple prompts.

Taken together, this evidence suggests that GPT-3 represents a fundamental paradigm shift in summarization, changing what data we need (or don’t need) and what approaches we can now explore. Evaluating these systems will require a new framework distinct from the automatic metrics that have dominated the last decade of summarization research.

Models and Setup

Recent zero- and few-shot prompting based models Brown et al. (2020); Sanh et al. (2022), have shown impressive generalization capabilities on unseen tasks specified using prompts alone and without performing any gradient updates Mishra et al. (2022). In this work, we want to compare their text summarization performance against the current state-of-the-art models.

Figure 2 shows the broad categories of all available summarization approaches, including current SOTA models and prompting-based models. The former set consists of fine-tuned language models, trained on a large number of article-summary pairs (e.g. BART Lewis et al. (2020), PEGASUS Zhang et al. (2020), BRIO Liu et al. (2022)) to obtain dataset-specific systems. This category also includes models aimed at tasks beyond generic summarization, such as keyword- or query-based summarization, that still rely on standard datasets for training He et al. (2022a).

On the other extreme are zero- or few-shot models, (e.g. GPT3 Brown et al. (2020), PaLM Chowdhery et al. (2022)), that are not explicitly trained for any particular task, as discussed above. Recent work Ouyang et al. (2022); Wei et al. (2022); Sanh et al. (2022) has improved on these models by introducing instruction-tuned models. Here, pre-trained language models are fine-tuned on multiple tasks (which may include summarization) using instruction templates in order to align their training with inference time usage.

In this work, we compare the summarization performance of three models that are representative of this space of options:

OpenAI’s text-davinci-002, a GPT-3 model Brown et al. (2020) from the Instruct series Ouyang et al. (2022). While we do not know the exact training details for this release of the model, the previous in the series (text-davinci-001) was fine-tuned on a combination of prompts submitted to their API and labeler prompts spanning multiple tasks. These tasks include summarization but not (to our knowledge) standard summarization datasets like CNN/DM Hermann et al. (2015); Nallapati et al. (2016) or XSum Narayan et al. (2018).

We choose the text-davinci-002 version for our experiments in order to benchmark the best available prompt-based model.We did not observe obvious quality differences in generated summaries between text-davinci-001 and text-davinci-002. Examples are included in Appendix C. We refer to this approach as GPT3-D2.

BRIO Liu et al. (2022), a fine-tuned summarization model that reports state-of-the art results on both CNN/DM and XSum. We will use versions of this model fine-tuned on each of these two datasets.

T0 Sanh et al. (2022), a prompt-based model fine-tuned on multiple tasks including standard summarization datasets. This provides a useful point of comparison between task-specific fine-tuned (BRIO) and bigger instruction-tuned models (GPT3-D2).

2 Using GPT3-D2 for summarization

Fine-tuned models largely follow the “style” of reference summaries in their training data, and hence, generated summaries show large variance between datasets (see Table 1 for basic summary statistics of standard summarization datasets). To ensure fair comparison between these and GPT3-D2, we adapt the latter’s prompt to align with dataset-specific styles.

Specifically, we follow prior work Sanh et al. (2022) and use sentence-count length prompts to adapt to each dataset. Although these datasets also differ along other attributes, e.g. CNN/DM is lead-biased whereas XSum requires drawing inferences from a whole article, we do not attempt to control any other attributed of the summary. Figure 3 shows an example of different length GPT3-D2 summaries for the same news article, using the following prompt format:

Summarize the above article in N sentences.

We found that GPT3-D2 summaries faithfully follow the given length constraint in 98% of the test instances used in our human study data in Section 3.

Given this setup, we first compare the summary quality of the three summarization models through a human annotation study (Section 3). Then, we evaluate the current suite of summarization metrics for prompt-based summarization (Section 4). Finally, in Section 5, we briefly discuss GPT3-D2 performance on summarization tasks beyond generic summarization and new challenges.

Human evaluation of GPT3-D2 summaries

Generated summaries of fine-tuned models Lewis et al. (2020); Zhang et al. (2020); Liu et al. (2022) emulate gold-standard summaries in their training datasets. In contrast, prompt-based GPT3-D2 models generate summaries based on how the given task description surfaces behavior learned during pre-training or instruction-tuning. In this section, we ask: how do these paradigms compare? Does learning from gold summaries lead to a better summarization model? To answer this, we conduct a human study to compare outputs of our 3 representative models and collect human preferences of quality.

We choose two standard fine-tuning datasets whose summaries differ along multiple dimensions such as length and abstractiveness:

CNN/DM Hermann et al. (2015); Nallapati et al. (2016) contains reference summaries that are approximately 3-4 sentences long. Summaries in this dataset are highly extractive and lead-biased.

XSum Narayan et al. (2018) contains 1 sentence summaries of BBC news articles. In this dataset, references summaries, and consequently generated summaries from fine-tuned models are highly abstractive.

Datasets for evaluation

Because GPT3-D2’s pre-training and instruction-tuning datasets are unknown, it may have been trained on existing articles and summaries in the test splits of these standard benchmarks. We therefore run our human study on 100 recent articles from CNNAlthough the BRIO’s CNN/DM model also includes DailyMail data in its training, we do not use this news source in our study as it is now widely considered to be unreliable. E.g. according to Media Bias / Fact Check site, DM’s factual reporting is rated ‘low’ https://mediabiasfactcheck.com/daily-mail/. and BBC, collected between March 1, 2022 and June 31, 2022. We call these CNN-2022 and BBC-2022 respectively.

Model details

We use the publicly released BRIO-XSum and BRIO-CNN/DM models to generate summaries.Models at: https://github.com/yixinL7/BRIO For T0, we use a prompt we selected from its prompt repository for CNN/DM and XSum datasets.Repository with T0 prompts: https://github.com/bigscience-workshop/promptsource Finally, to generate GPT3-D2 summaries, we set N=3N=3 for CNN and N=1N=1 for BBC in our standard sentence-count prompt template from Section 2.

For a maximally fair comparison in this “realistic” setting, we take some additional steps to improve the output of BRIO-XSum. In order to automate dataset creation, XSum removes the first sentence from news articles to use as the gold summary for training, then treats the rest of the sentences as the article to summarize. This setup differs from the real world usage of summarization systems where the complete article is summarized. Due to this mismatch, BRIO-XSum often generates very low quality outputs, e.g. All images: Strule Shared Education Campus in Figure 4, for around 30% of the articles. We manually identify these examples and first attempt to fix them by selecting a summary without such obvious failures from further down the beam (we use beam size =10=10). However, if we cannot find a “better” summary, we remove the first sentence of the article and re-sample a new summary to align with its noisy training. This latter strategy often results in factually incorrect summary generations, as is well documented in prior research Maynez et al. (2020); Goyal and Durrett (2021).

Design of the human study

We design an A/B test to collect preference annotations. For each given article, annotators are shown summaries from all three summarization systems (BRIO, T0 and GPT3-D2). They are then asked to select their most and least preferred summary or summaries. In addition to these multiple choice questions, we also ask for a free-text justification of both choices.

We make two design decisions for our human study: first, we do not provide annotators with specific definitions of summary quality to avoid introducing our own biases. It is also quite challenging to produce a unified definition of quality for the very different “styles” of summaries evaluated in this study. Instead, we ask them to rely on their own preferences based on summaries they would like to see if they were browsing the web, which we believe to be a representative scenario for non-expert consumers of news summaries. Detailed task instructions are included in Appendix F.

Second, we allow multiple selections for both the best and worst summary questions to cater to scenarios in which different summarization systems output similar quality summaries without meaningful differences.

We hire crowd annotators through Prolific. For both CNN and BBC, we recruit 60 unique participants to annotate the 100 summaries in each dataset. Each annotator was asked to annotate 5 articles and each article was annotated by 3 annotators. Additionally, we use the Prolific’s demographic filters to restrict participation to USA (or UK) residents for CNN (or BBC). We anticipate that residents from these respective countries are better positioned to understand country-specific news events and evaluate their summaries. Participants were paid approximately $11/hr for their work.

2 Results

Figure 4 shows examples of generated summaries from all three summarization systems for both CNN and BBC articles. For CNN, we observe that fine-tuned BRIO summaries tend to be highly extractive and generally include a high number of named entities (dates, percentages, names), reflecting the data it was trained on. In contrast, GPT3-D2 summaries are more abstractive and less specific, but provide a more exhaustive overview of the article content. Table 2 provides quantitative evidence of this; we use percentage of novel n-grams to measure abstractiveness, and number of named entities per 100 words to measure specificity.

For BBC, we observe inverse trends where BRIO and T0 are more abstractive compared to GPT3-D2. Again, this can be attributed to the XSum training data used to train both these prior models. For GPT3-D2 summaries, on the other hand, the level of abstractiveness does not differ between datasets. Finally, Table 2 shows that GPT3-D2 summaries tend to have longer sentences, and therefore similar number of summary sentences often results in a longer summary for both datasets. We study the effect of this length difference on human preference judgments in Appendix B.

Which systems do humans prefer?

Results of our human study are summarized in Table 3. We report the percentage of times a particular system is the most/least preferred model according to majority vote combining all three annotator’s choices.As we allow multiple system selections, note that more that one system could be the majority. However, this is rare after majority vote: only 2% of the articles in CNN and 7% in BBC have multiple best summaries. Across both datasets and styles, we observe a clear preference for GPT3-D2 summaries compared to the other two models. In fact, in both scenarios, the GPT3-D2 outperforms the next best model by at least 20 percentage points. This improvement is statistically significant according to a paired bootstrap test (CNN p−p-value =2×10−3=2\times 10^{-3}, BBC p−p-value =6×10−4=6\times 10^{-4}).

Note that the next best model differs between the two datasets. For BBC, annotators prefer T0 summaries over BRIO. Annotator rationales often mentioned misleading or incorrect information as the primarily reason for selecting BRIO as the worst summary, confirming the issues that have been observed with XSum-trained models Maynez et al. (2020); Pagnoni et al. (2021); Goyal and Durrett (2021). Although T0 also includes XSum training data, we hypothesize that its multi-task framework helps offset the noisy signal from XSum.

In contrast, annotators rate T0 as the worst summarization system for CNN. The most common rationales for these were shorter length and inclusion of irrelevant details, e.g. long quotes, while missing key points. Some annotators also commented that these T0 summaries were less coherent compared to the other models. Interestingly, we did not observe similar complaints for the single-sentence T0 summaries for BBC.

Do annotators agree with each other?

To study this, we plot the distribution of annotator votes for each summarization system and dataset in Figure 5. Additionally, we report the inter-annotator agreement, measured using Krippendorff’s alpha with MASI distance Passonneau (2006), to account for multiple selections of best or worst summary allowed in our study design.

The vote distribution shows that although more annotators prefer GPT3-D2 summaries, this choice is only unanimous, i.e. supported by all three annotators, for less that 30% of the annotated articles. Conversely, although BRIO (or T0) summaries are less preferred than GPT3-D2 for the CNN (or BBC) dataset on aggregate, they were voted as the best summary by at least one annotator for more than 60% of the articles. This demonstrate two things: first, when comparing summaries from two strong models, the choice is inherently ambiguous (similar observations in Clark et al. (2021)). Second, these results and the diversity in the written rationales, show that there does not exist a universal definition of a “good” summary and that different summary properties appeal to different annotators. Regardless, the aggregate preference for GPT3-D2 is high enough across the board to give us confidence in its strength.

How do these results impact the field?

Progress in text summarization research in the last five years has been enabled by the construction of large-scale text summarization datasets that involved scraping news articles and pairing them with any available summary-like data Hermann et al. (2015); Narayan et al. (2018); Grusky et al. (2018). The CNN/DM dataset considers bullet points accompanying news articles as its summary. These “gold” standard summaries provided useful training signal to train impressive supervised models Lewis et al. (2020); Zhang et al. (2020); Liu et al. (2022) and hence, their quality or alignment with human preferences was largely ignored.

We found that, despite its popularity, XSum is largely unsuitable for fine-tuning models like BRIO for realistic summarization settings. Even though a CNN/DM-trained BRIO model performed better, the results of our human study question the continued utility of hill-climbing on this dataset, as it seems users may simply prefer a different style of summary altogether. In fact, this preference for GPT3-D2 is much larger than incremental improvements reported in other human evaluation settings, e.g. improvements on XSum on the GENIE leaderboard Khashabi et al. (2022). Furthermore, as we we will see in Section 5, the greater flexibility of GPT3-D2 compared to these systems makes it more suitable for news summarization tasks beyond generic summarization.

If a system designer collects a large-scale dataset of high-quality summaries that they wish to emulate, we believe a fine-tuned system may outperform GPT3-D2. However, better-trained models on datasets collected via “incidental” supervision are less likely to help.

Can current automatic metrics evaluate GPT3-D2 summaries?

Automatic metrics proposed for summarization evaluation can be broadly divided into two categories: (1) reference-based, that compare generated summaries against available gold summaries, and (2) reference-free that only rely on the input document. Here, we compare their performance at evaluating prompt-based GPT3-D2 summaries.

We evaluate automatic metrics using summaries from 4 different summarization datasets, listed in Table 1. For each dataset, we construct our evaluation sets by randomly sampling 500This size is chosen to give sufficient statistical power Card et al. (2020) while keeping costs for GPT3-D2 evaluation low to enable others to compare on this subset. We outline costs in Appendix D. articles from the standard test split.Note that these standard datasets were released before 2020. Therefore, it is possible that some article-summary pairs in our test set overlap with GPT3-D2’s training data. However, we do not observe a qualitative difference in GPT3-D2’s performance on these older articles. We compare the same 3 summarization systems from Section 3 in our analysis. Additionally, we also report results using the fine-tuned PEGASUS model Zhang et al. (2020), as BRIO fine-tuned models are not available for all datasets.

We publicly release this corpus of summarization outputs to standardize the test sets and support future research into GPT3-D2 based summarization. Link: https://tagoyal.github.io/zeroshot-news-annotations.html.

1 Reference-based metrics

Here, we study if the gold summaries of the standard datasets are useful for evaluation, especially when evaluating prompt-based summaries that are not trained to emulate the gold. We benchmark the performance of 3 different summarization metrics: (1) overlap-based metrics, specifically ROUGE Lin (2004) METEOR Banerjee and Lavie (2005) and BLEU Papineni et al. (2002). (2) similarity-based metrics, that compute similarity between embeddings representations of generated and reference summaries. Specifically, we report BERTScore Zhang* et al. (2020) and MoverScore Zhao et al. (2019). (3) a QA-based metric, specifically QAEval Deutsch et al. (2021). Although most QA-metrics are reference-free (discussed in Section 4.2), QAEval uses the reference summaries to indicate saliency. We report both exact match (EM) and F1 components of QAEval.

Table 4 outlines the results. It shows that BRIO and PEGASUS models, fine-tuned to emulate the reference summaries, outperform GPT3-D2 summaries according to all reference-based automatic metrics. The difference in their assigned scores is very high, e.g. >7 ROUGE-L points between GPT3-D2 and BRIO. For comparison, these reported scores for GPT3-D2 are even lower than the trivial Lead-3 baseline reported in prior work Fabbri et al. (2021); Grusky et al. (2018). This clearly demonstrates that current automatic reference-based metrics cannot be used to reliably measure summary quality under the prompting paradigm.

Amongst prompting-based models, we observe that T0 summaries report better metric scores than GPT3-D2 for all datasets except Newsroom. Interestingly, out of the four datasets evaluated here, Newsroom is the only one not used to train the T0 model. This further shows that access to dataset-specific reference summaries during training improves performance according to these metrics, rendering them unsuitable for evaluating prompt-based models.

2 Reference-free metrics

Table 5 outlines the scores for each summarization system according to the above reference-free metrics. Ideally, we want the relative rankings of different systems according to these metrics to correspond to human preferences, i.e. GPT3-D2 > BRIO > T0 for CNN/DMAlthough the human study in Section 3 is only run on CNN articles, the underlying fine-tuned model is same for both CNN and DM. Therefore, it we can reasonably expect it to display similar quality differences with respect to GPT3-D2. and GPT3-D2 > T0 > BRIO for XSum.Note that while annotators were not explicitly asked to rate factuality, we instructed them to carefully check factuality and appropriately downvote non-factual summaries.

Overall, we observe that none of the reference-free metrics we evaluate follow these trends for both CNN/DM and XSum datasets. In particular, we observe that GPT3-D2 summaries report low factuality scores (except XSum) even though we rarely found any factual errors in our qualitative analysis of its generated summaries.

Interestingly, we noticed a roughly inverse relation to abstractiveness; summarization systems that generated more abstractive summaries (see Table 2) were generally scored lower by all automatic reference-based metrics. For instance, GPT3-D2 is scored lower than BRIO by both quality metrics for all datasets except XSum; the latter is the only dataset for which GPT3-D2 summaries are less abstractive. Such shortcomings of reference-free evaluation metrics due to spurious correlations have also been studied in prior work Durmus et al. (2022). These issues become more exaggerated when the summarization systems being compared exhibit very different properties.

Discussion

On the surface, the failure of reference-free metrics at evaluating GPT3-D2 summaries is more surprising that reference-based metrics as the later explicitly compares generated summaries with references that GPT3-D2 is not trained to imitate. Therefore, GPT3-D2 understandably scores lower than fine-tuned systems.

However, we note two different issues with reference-free metrics: (1) Some of these, e.g. FactCC and DAE, use reference summaries as positive examples to train the metric. Therefore, although “reference-free” at test time, they are still trained to reward the summary properties seen in the standard summarization benchmarks. (2) Even completely reference-free metrics, e.g. QuestEval and QAFactEval, have only been evaluated on reference-based benchmarks and fine-tuned models. Therefore, the choice of different components, such as question answering or question generation models to use, etc. has been dictated by the error space of prior fine-tuned models Tang et al. (2023). These decisions also now need to be re-visited to incorporate GPT3-D2 evaluation; we leave this for future work.

Beyond Generic Summarization

Previously, we observed that GPT3-D2 models faithfully follow simple “style” instructions in the given prompts. This provides a promising direction to tackle other use cases in news summarization beyond the generic summarization task from Section 3.

Different users can have very different information needs from the same article, all of which cannot be satisfied with a single generic summary. Prior work has introduced several task formulations to address this gap, including keyword-focused He et al. (2022a), query-focused Baumel et al. (2014); He et al. (2022a), or aspect-focused summarization Krishna and Srinivasan (2018); Ahuja et al. (2022), amongst others. Here, we evaluate GPT3-D2 performance at two of these use cases.

In keyword-based summarization, the output summaries must succinctly summarize the input document focusing on a given keyword; these generally correspond to specific entities or events directly mentioned in the document. In contrast, the control units in aspect-based summarization are high-level topics that can be common across multiple similar types of documents. For e.g., for the input article in Figure 1, Donald Trump or Russian interference in 2016 elections are keyword controls whereas charges against the defendants is a higher-level aspect that can serve as the query for any news article discussing a lawsuit or investigation.

We use the recently proposed CTRLSum He et al. (2022a), a fine-tuned BART model, as our baseline. It can be flexibly adapted for both keyword- and aspect-based settings by including a prompt as additional input to the encoder. We use the prompt template recommended in the original paper.Trained model publicly released at: https://github.com/salesforce/ctrl-sum.

Control Units

For the keyword-focused setting, we use named entities extracted from the input article as the control units. For aspect-focused summarization, we directly use the aspects introduced in the guided summarization task from TAC 2011.https://tac.nist.gov/2011/Summarization/Guided-Summ.2011.guidelines.html It defined 5 broad categories of newswire articles, such as accidents and natural disasters, investigations and trial, etc., and multiple aspects for each category. For example, the “investigations and trials” category includes aspects such as “who is the defendant or under trial?”, “who is investigating, prosecuting, judging?”, and so on.

Qualitative Analysis

Figure 6 shows examples of keyword- and aspect-focused summaries using GPT3-D2 and the baseline CTRLSum model. The keywords or aspects are highlighted in bold within the GPT3-D2 prompt displayed on the left.

In this example, representative of average GPT3-D2 quality, the keyword-focused GPT3-D2 summary first gives a brief overview of the article setting before providing keyword-relevant information. In contrast, the CTRLSum summary exhibits poor discourse structure and reads like a list of facts stapled together.

The figure also shows aspect-focused summaries for two aspects associated with the “investigations and trial” category most appropriate for the chosen article. We see mixed results here for GPT3-D2; it generates a factually incorrect summary for the first aspect, listing multiple people from the input article as defendants instead of only “Donald Trump”. For the second aspect, it correctly maps the high-level concept “defendant” to “Donald Trump” in the input article and generates the correct answer to the input query: “The defendant’s reaction to charges in the above article is denial of charges”.

On the other hand, CTRLSum fails to generate aspect-focused summaries for both cases. We believe that it struggles to align high-level concepts and explicit entities in the article due to a lack of such aspect-specific examples in its training data. Instead, it generates summaries focusing on lexically similar words, i.e. “defenders” for both cases.

Based off of GPT3-D2’s promising keyword-focused summarization capabilities observed above, we next conduct a human study to systematically compare it against the CTRLSum baseline. We leave further explorations of aspect-based summarization to future work, given the mixed to poor results for both models at this task.

2 Human Study: Keyword-focused summarization

Similar to Section 3, we design an A/B test to compare the two models. We use the same set of 100 CNNWe run this study using only CNN articles as the baseline CTRLSum model is trained on CNN/DM. articles as Section 3. We randomly extract 2 distinct named entities from each article. In the study interface, the annotator is shown the article-keyword pair and GPT3-D2 and CTRLSum summaries corresponding to it. They are asked to select the summary that best summarizes the input article while focusing on the given keyword. Exact task instructions are included in Appendix F.

Again, we run this study using the Prolific platform. We recruit 60 participants to annotate the 100 articles; each article is annotated by 3 annotators which includes annotations for 2 separate keywords. Each annotator evaluates 5 articles.

Results

Figure 7 shows the distribution of annotator votes between the GPT3-D2 and CTRLSum models. Annotators show a clear preference for GPT3-D2. In fact, for nearly 70% of all article-keyword pairs, GPT3-D2 is preferred over CTRLSum by a majority of the annotators. The main rationales given for this choice were better contextualization of keyword-related information and better coherence in GPT3-D2 summaries.

Impact

These results show that prompting GPT-3 models present a promising alternative to fine-tuned models for such specialized summarization tasks that can be easily described using textual prompts. One of the major drawbacks of fine-tuned models is that they are constrained by what data is available and how it can be transformed to create new task-specific training data. CTRLSum relied on the SQuAD question answering dataset Rajpurkar et al. (2016) because the required “queries” or “questions” were unavailable at scale for summaries in standard summarization datasets. In contrast, prompt-based models are not constrained by the availability of task-specific data and can flexibly adapt to new tasks. Future research should focus on further exploring these capabilities and possible improvements on currently “unsolved” tasks such as aspect-based or plan-based summarization.

Discussion and Related Work

In recent years, research in text summarization Rush et al. (2015); Nallapati et al. (2016); See et al. (2017); Lewis et al. (2020); Zhang et al. (2020); Liu et al. (2022) has typically relied on comparisons with gold test sets for evaluation, possibly augmented with reference-free metrics for dimensions like factuality. This paper shows that all these metrics are completely ineffective at evaluating GPT-3 summaries. Although issues with these metrics, particularly low correlation with human judgments, have also been studied earlier Fabbri et al. (2021); Deutsch and Roth (2021), they are considered reliable when comparing systems in different score ranges Peyrard (2019); Deutsch et al. (2022). However, GPT-3 challenges these established practices and evaluation protocols, and poses an urgent need for better evaluation.

This brings us to manual evaluation, generally considered to be the gold standard for generation evaluation. The majority of summarization research now reports results from a human study in addition to automatic metrics, but there is a general lack of consensus on what dimensions to evaluate, task design, and other factors Hardy et al. (2019). This presents difficulties in conducting reliable and reproducible comparisons between systems Karpinska et al. (2021), another factor contributing to the popularity of automatic metrics. Although recent efforts like GENIE Khashabi et al. (2022) have taken steps to standardize manual evaluation protocols across systems, its annotation is not universally affordable and the quality is not strictly monitored. We hope that future work addresses these challenges and democratizes human evaluations.

The ultimate test of summarization systems is with actual users using the systems in practice. Jones (2007) discusses the need to align task formulations with actual applications scenarios (“purpose factors”). However, the research in text summarization until now has been constrained to certain problems or domains by the heavy dependence on large-scale training data: for example, producing a bullet-point summary of a news article has emerged as standard due to availability of data from CNN, not because it is shown to be the best way to present information.

Now, the success of prompt-based models can allow realistic use-cases to drive research in a more top-down way. We already show that GPT3-D2 improves upon prior keyword-focused summarization systems that were trained on artificially adapted training data. In future research, we are interested in tackling other real world use cases, such as update summarization and plan- or aspect-based summarization. Additionally, adapting GPT3-D2 to documents longer than the allowed context, or structured inputs such as tables, presents research challenges beyond the current capabilities of GPT-3 and would be interesting to study.We very briefly discuss long document summarization with GPT-3 in Appendix E.

Conclusion

In this work, we performed the first systematic study comparing prompt-based GPT-3 and fine-tuned models at the news summarization task. We analyzed the impact of prompting on the summarization field, including training paradigms and evaluation practices. Finally, to support further research in this direction, we release a large corpus of generated summaries for multiple prompt-based and fine-tuned models, as well as human preference judgments comparing these systems.

Limitations

In the text generation evaluation literature, there does not exist a standardized task design for comparing different system generations. In our work, we chose a human evaluation workflow that directly asks annotators to compare systems, while other prior work has opted for Likert-scale judgments and/or evaluation along multiple quality dimensions Gehrmann et al. (2022). The latter strategy of evaluating different dimensions could surface more insights into which “style” properties of GPT-3 summaries provide them an edge over fine-tuned models; however, such analysis is outside the scope of this paper. Our experiments comparing overall quality reveal that current summarization datasets are not well-aligned with user preferences. We leave more fine-grained analysis into these preference judgments for future work.

The experiments in this paper are run on English-language news summarization datasets as these serve as common benchmarks in the summarization literature. However, user rankings of system outputs might be different when evaluating other domains, e.g., summaries of scientific text. While we believe that automatic metrics would fail to evaluate GPT-3 summaries on these domains also (generated summaries would still look different from the reference summaries), users may prefer models that are specifically fine-tuned on domain-specific data for niche domains.

Finally, we do not know exact datasets or tasks used to train GPT3-D2. It is possible that its RLHF training Ouyang et al. (2022) included summarization examples, and therefore, preference judgments from human annotators for its different outputs. However, our arguments in this paper do not rely on the specifics of the GPT3-D2 system, merely that such a system exists. If anything, the existence of potentially better data underscores that further work should collect new data for summarization model tuning, and our claims about metrics still hold regardless of the details of how the GPT3-D2 summaries were produced.

References

Appendix A Implementation Details

To generate GPT3-D2 summaries for all experiments in this paper, we use the standard prompt format outlined in Section 2. We set N=3N=3 for CNN and DailyMail, N=2N=2 for Newsroom, and N=1N=1 for XSum/BBC. For the latter, the prompt is slightly modified to “Summarize the above article briefly in 1 sentence.”

For T0, we use the following prompts: a) CNN/DM: “Summarize the article below in 3 to 4 sentences?”, b) Newsroom: “Summarize the article below in 2 to 3 sentences?”, and c) XSum/BBC: “Summarize the article below in 1 sentence?”

Factuality Metrics

In Section 4.2, we evaluated several recently proposed factuality metrics. We note that multiple versions have been released for some of these models in recent years. Here, we specify the versions used in our experiments to ensure reproducibility of results:

QuestEval: We use version 0.2.4 of the questeval python package and report numbers using the precision-only setting.

DAE: We use the updated version of the DAE model trained for document-level factuality. Latest code and model released at https://github.com/tagoyal/factuality-datasets.

SummaC: We use the SummaC-Conv model (model_name = ‘vitc’) and sentence-level granularity in our experiments.

Keyword-based data

For our keyword-based human study, we extracted two named entities per article, as discussed in Section 5. In practice, we constrained the first keyword to be lead-biased, i.e. it was extracted from the first three sentences of the article, and the second keyword was extracted from the remaining article. As CNN-based summarization models are generally lead-biased, this allowed us to benchmark models under both settings.

Appendix B Are annotator judgments of quality correlated with length?

In Section 3, results of the human study showed that annotators provide shorter length as one of the main reasons for selecting T0 summaries as the worst for the CNN dataset. Here, we investigate if the choice between GPT3-D2 and BRIO is similarly influenced by their length differences; GPT3-D2 summaries are on average 9 words longer.

To study this, we plot the difference in summary length against the difference in annotator score (measured as the no. of votes for a summarization system) between the best summarization system (GPT3-D2) and the next best system (BRIO for CNN and T0 for BBC). The resulting plot is shown in Figure 8. In general we observe low correlation between these; Pearson’s ρ\rho is 0.17 for CNN and .02 for the BBC dataset. These correlation values cannot solely explain the large differences in annotator judgments reported in the human study results of Section 3; additional quality factors must have influenced this choice. Anecdotally, we observe that the GPT summaries are slightly less information dense; our impression is that these contain a similar level of information content, but are easier to read and understand despite being a bit more verbose.

Appendix C Qualitative differences between GPT-3 versions

Figure 9 shows examples comparing summaries from text-davinci-001 (GPT3-D1) to those from GPT3-D2. For BBC-style single sentence summaries, we observed that the two models generated very similar summaries with high content and lexical overlap. More variance is observed for CNN-style summaries. In our anecdotal assessment, GPT3-D1 generated more detailed summaries while those from GPT3-D2 are less information dense.

Appendix D Human study and API costs

At the time of running our experiments, GPT-3 API’s text-davinci-002 version was priced at $0.06 per 1K tokens. New pricing information is available at: https://openai.com/api/pricing/.

In our experiments, we generated around 2600 GPT3-D2 summaries across all experiments in Section 3 (human study), Section 4 (evaluation of metrics) and Section 5 (keyword-based human study). We spent a total of approximately $150 on API requests.

For the human study, we paid participants 4pertask(eachtaskinvolvedannotationfor5articles).Onaverage,thistranslatedto4 per task (each task involved annotation for 5 articles). On average, this translated to11/hr of work. The combined cost for the generic summarization (Section 3) and the keyword-based summarization (Section 5) studies was $1020, including platform costs and bonus payments.

Appendix E Long document summarization using GPT3-D2

Summarization of long documents has attracted significant interest in recent years Cohan et al. (2018); Kryscinski et al. (2021). Here, we study how naive prompting of GPT-3 performs at long-document summarization.

First, we extract text from a long input article from the CNN website.Article link: https://www.cnn.com/2021/09/07/opinions/covid-19-good-and-bad-news-ranney/index.html Next, we follow the commonly used segment-then-summarize procedure from prior work Zhao et al. (2020); Zhang et al. (2022). We divide the input article into 3 disjoint segments, summarize each segment separately and concatenate these outputs to form the final summary.

Figure 10 shows the prompt used and the generated summaries for each segment. While individual segment summaries are high quality, we can see that the concatenated summary is not coherent and includes repeated “introductory” sentences outlining similar content. Related to this, it also does not cover all important aspects of the input article as a majority of its ‘length budget’ is spent on a high-level overview. We also observed that the generated summaries for long documents often focus on less unimportant parts of the document, e.g. “…everyone should take the precaution of … opening windows to let the fresh air in” in the illustrated example. This is, in part, due to the segmentation of the input article: GPT3-D2 still exhibits some lead bias and treats the beginning of the input segment as more salient. Therefore, the exact segmentation of the article also dictates the quality of the final summary, and cannot be readily fixed by altering the prompt.

These observations show that while GPT3-D2 produces superior segment-level summaries, it is more difficult to adapt it to “non-natural” text inputs without fine-tuning. Therefore, techniques that have shown promising results for fine-tuned models, e.g. segment-then-summarize or extract-then-abstract Zhang et al. (2021) approaches, are not as effective when directly applied with prompting-based models.

Appendix F Task Instructions

Task instructions provided to crowd annotators for the generic summarization task setting are shown in Figure 14 and those for the keyword-based setting are shown in Figure 15.

Appendix G Examples of generated summaries

We show examples of generated summaries for articles for generic summarization for CNN-2022 and BBC-2022 in Figures 11 and 12. It includes summaries from the 3 different summarization models evaluated in the human study in Section 3.

Examples of keyword-focused summaries are shown in Figure 13 for CNN. It includes summaries generated by GPT3-D2 and CTRLSum models.