Human-like Summarization Evaluation with ChatGPT
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, Xiaojun Wan
Introduction
Text summarization is a task that involves generating a condensed version of one or multiple documents. Thanks to the advancements in deep learning-based techniques, automatic summarization has made significant strides. Specifically, the emergence of large language models such as InstructGPT has resulted in comparable performance to reference summaries written by humans, even in zero-shot settings (Zhang et al., 2023).
Evaluating text summarization, like other text generation tasks, is a challenging problem. While human evaluation is considered the gold standard, it is expensive and time-consuming. As a result, automatic evaluation metrics play a crucial role. ROUGE (Lin, 2004) and its variants, which are based on reference summaries and n-gram matching, are widely accepted and used in various types of summarization. However, surface-level word matching cannot accurately reflect the quality of the summary. Additionally, it is challenging to evaluate the factual accuracy of the summary without utilizing the source document. Recently, evaluation metrics based on pre-trained models such as BERTScore (Zhang et al., 2020) and BARTScore (Yuan et al., 2021) have achieved better correlation with human judgments. Factuality evaluation methods based on entailment classification, such as FactCC (Kryscinski et al., 2020), and question answering, such as FEQA (Durmus et al., 2020), have also been used to evaluate the factual consistency of summaries. Despite the existence of advanced automatic evaluation metrics, their performance, usability, and interpretability are still far from satisfactory.
Large language models (LLMs) offer completely different possibilities for the automatic evaluation of summarization. GPT-3 (Brown et al., 2020) has the ability of in-context learning, and instruction tuning allows LLMs to align with human evaluation (Ouyang et al., 2022). These two abilities make it possible for LLMs to mimic the behavior of human evaluators, who generally evaluate summaries by understanding examples and instructions. We refer to this automatic evaluation method that views large models as human evaluators as human-like automatic evaluation. The most prominent feature of this evaluation method is its flexibility, which unifies all types of automatic evaluation in form and can simulate many of the practices of human evaluators. Unlike previous automatic evaluation metrics that give one or more numerical values as evaluation results, the evaluation results of this human-like automatic evaluation are fully reflected in the generated responses, which may include scoring, comparison, labels, and explanations.
We conducted an evaluation of the evaluation ability of ChatGPT, a recently popular LLM, using four commonly used human evaluation methods for summarization. The methods include Likert scale scoring, pairwise comparison, Pyramid (Nenkova and Passonneau, 2004), and binary factuality evaluation. Our findings indicate that ChatGPT is capable of completing annotations relatively smoothly using these methods. In addition, our results demonstrate that ChatGPT outperforms commonly used automatic evaluation metrics on some datasets. Furthermore, we analyzed the impact of different prompts, compared the performance of ChatGPT with human evaluation, and examined the quality of the generated explanations and invalid responses.
Preliminary
We select several evaluation metrics that are commonly used in summarization:
ROUGE (Lin, 2004), which is the dominant automatic evaluation metric in summarization, is widely used by researchers. The most commonly used ROUGE measures are ROUGE-1, ROUGE-2, and ROUGE-L, which evaluate the similarity between two texts based on the overlap of unigrams, bigrams, and the longest common sequence.
BERTScore (Zhang et al., 2020) assesses the similarity between two texts at the token level by measuring the soft overlap using contextual embeddings from BERT. Similarly, MoverScore (Zhao et al., 2019) uses n-gram embeddings that are pooled from BERT to compute the semantic distance between two texts at the n-gram level.
BARTScore (Yuan et al., 2021) https://github.com/neulab/BARTScore, also for ROUGE, BERTScore, and MoverScore. views evaluation as a natural language generation task and considers that when the quality of the generated text is higher, BART is more likely to generate it from the source text or the reference, or to generate the reference from it. BARTScore can be flexibly applied to evaluate text from various perspectives.
FactCC https://github.com/salesforce/factCC and DAE https://github.com/tagoyal/factuality-datasets are two factuality metrics based on classification. When evaluating a summary, we use NLTK version 3.7, https://www.nltk.org/ to split it into individual sentences and classify each one as factually correct or not. The factual score of the summary is then calculated as the ratio of sentences that are factually correct.
2 Human Evaluation Methods
There are several commonly used methods for human evaluation, including the Likert scale scoring and pairwise comparison for general text generation, as well as Pyramid and binary factuality evaluation specifically designed for summarization. After introducing each method, we will list the datasets we used that were annotated in this way.
Likert scale scoring is the most common method for human evaluation. Specifically, given a source document and a generated summary, annotators rate the summary on several dimensions. Typically, this is an absolute evaluation, meaning each summary is evaluated individually without explicit comparison to other summaries. Dimensions usually include factual consistency, informativeness, fluency, etc. The rating scale is usually 1 (worst) to 5 (best). We used SummEval (Fabbri et al., 2021) and Newsroom datasets (Grusky et al., 2018).
Pairwise comparison is a relative human evaluation method. Given a source document and two generated summaries, annotators choose the one that is of higher quality. This method is used in reinforcement learning based human feedback for summarization. We used the TLDR dataset (Stiennon et al., 2022).
Pyramid (Nenkova and Passonneau, 2004) is a human evaluation method designed for summarization that is based on reference summaries. Prior to human annotation, several semantic content units (SCUs) are extracted from the reference summary. For each SCU, annotators judge whether it presents in the generated summary. For single-document summarization, the final score of the summary is the proportion of SCUs it contains. We used the REALSumm dataset (Bhandari et al., 2020).
Binary factuality evaluation is a method for evaluating the factual correctness of summaries. Given a source document and a sentence in the generated summary, annotators judge whether the sentence is faithful to the source document. We used the QAGS dataset (Wang et al., 2020).
Experiments
We used the ChatGPT API (gpt-3.5-turbo-0301) provided by OpenAI for our experiments. To reduce randomness, we set temperature to 0. In addition, we set max_tokens to 256. We kept the default values for other parameters.
2 Prompt Design
When designing prompts, we made it as identical as possible to the original instructions of human evaluations.
Figure 1 shows the template for Likert scale scoring. ChatGPT is asked to rate four dimensions at a time. For SummEval, the four dimensions are relevance, faithfulness The original term used in SummEval was ”consistency”. Since we did not add definitions of the dimensions in the prompt, we used ”faithfulness”, which is more representative of its actual meaning, fluency, and coherence. For Newsroom, the four dimensions are relevance, informativeness, fluency, and coherence. Figure 2 shows the template for pairwise comparison.
Figure 3 shows the template for Pyramid. The number of SCUs depends on the content of the reference summary, up to 16.
Figure 4 shows the template for binary factuality evaluation. The sentences are from the generated summaries.
3 Post-processing of Results
The vast majority of ChatGPT responses contained annotation results, which can be extracted by some simple rules. For invalid responses, we considered them as failing to complete the tagging successfully and marked them as NAN (not a number).
4 Evaluation
For Likert scale scoring, we computed sample-level, system-level, and dataset-level correlation with human judgments. For the other human evaluation methods, we calculated the accuracy of the responses generated by ChatGPT using human annotation as the answer.
5 Results
Tables 1 and 2 show that ChatGPT has a good ability to evaluate summaries with Likert scale scoring. On SummEval, it performs substantially better than the existing evaluation metrics. On Newsroom, it performs second only to BARTScore_s_h and BARTScore_cnn_s_h.
Tables 3, 4 and 5 illustrate that ChatGPT can also perform relatively smoothly on pairwise comparisons, Pyramid, and binary factuality evaluation. Nevertheless, from the current experimental results, ChatGPT has not yet shown a very large advantage except on QAGS_XSUM.
Analysis and Discussion
We tried several different prompts on SummEval. As shown in Figure 5, more detailed step instructions and dimension definitions are added. These instructions and definitions are from the original human evaluation. In addition, we consider setting the system prompt as "You are a human annotator that rates the quality of summaries." when using ChatGPT API.
Table 6 shows that changing prompts result in a significant change in the performance of the human-like automatic evaluation using ChatGPT, especially in terms of system-level correlations. From the current results, these changes do not make it to achieve higher correlations with human judgments, except for a modest improvement in a few cases by adding dimension definitions alone.
2 Comparison with human evaluation
In terms of accuracy, there is still an overall gap between the current automatic human-like evaluations using ChatGPT compared to human experts. Table 6 illustrates that in most cases, the correlation between scores given by a human expert and the average of scores given by human experts is substantially better than ChatGPT at all levels. However, the correlation between ChatGPT and human evaluations (0.889) is already higher than that of a particular human expert (0.843) in terms of system-level correlation of fluency.
For variance and reproducibility, automatic human-like evaluations using ChatGPT are more controllable. It is easy to know from Table 6 that the scores of the same samples will not be identical between different human annotators. Belz et al. (2021) pointed out that reproducing the manual evaluation was difficult. In contrast, we can make the human-like manual evaluation based on ChatGPT reproducible by setting randomness parameters (e.g., temperature) at decoding time.
In terms of cost, it is cheaper to perform the human-like automatic evaluation. Taking SummEval as an example, in our experiments, the assessment of one summary consumed about 1000 tokens, and it took about USD https://openai.com/pricing to finish the evaluation on the whole dataset. Assuming that a single annotator spends 5 hours annotating the whole dataset. It costs USD. It is estimated that the cost of human evaluation is about 10 to 20 times higher than human-like automatic evaluation using ChatGPT.
3 The quality of generated explanations
We sampled and examined the responses generated by ChatGPT on SummEval, and found the following characteristics of the explanations given by ChatGPT:
ChatGPT sometimes provides scores or labels followed by an explanation, even if it is not explicitly asked to provide the explanation in the prompt. Of course, it is possible to add a request such as "You do not need to explain." to the prompt so that it does not generate an explanation, but the impact of this on the evaluation scores is unknown.
The explanations generated by ChatGPT are generally self-consistent but not necessarily correct. The generated explanations generally coincide with its scoring. For example, Table 7 shows that ChatGPT and ChatGPT+def both scored low for the faithfulness of the summary, and they both pointed out factual errors in the summary. However, the correctness of these explanations still needs further testing.
The combination of ChatGPT’s explanations and scoring can better confirm whether it understands the requirements of the evaluation, for example, the dimension definitions. Without providing dimension definitions (see Figure 5), ChatGPT’s understanding of fluency and coherence converged. After examining multiple samples we found that its explanations of the scoring of these two dimensions are placed together and the dataset-level correlation between the scoring of these two dimensions is 0.960. ChatGPT is better able to distinguish between these two dimensions when dimension definitions are provided. Its explanations of the scoring of the two dimensions are separated and the dataset-level correlation between the two dimensions drops to 0.843.
4 Invalid responses
ChatGPT sometimes generates invalid responses, but this fraction is only about 1% at most (see Table 9). As shown in Table 8, common invalid responses were refusing to evaluate, not evaluating as required, writing a new summary, and continuing to write the summary. The reason why invalid responses are generated needs to be further explored.
Related Work
There are some concurrent studies using LLMs for human-like NLG evaluation. According to Kocmi and Federmann (2023), LLMs are currently the most advanced evaluators of translation quality. Wang et al. (2023) tested ChatGPT’s ability to be an evaluator on three NLG meta-evaluation datasets. Ji et al. (2023) explored the effectiveness of ChatGPT in ranking model-generated content. Luo et al. (2023) investigated ChatGPT’s ability to evaluate factual consistency in summarization.. Liu et al. (2023) utilized ChatGPT and GPT-4 to assess the quality of NLG outputs with chain-of-thoughts.
Conclusion
From the above experiments using ChatGPT for human-like summarization evaluation, the key findings are as follows:
ChatGPT has the ability to perform summarization evaluation using various human evaluation methods. In some instances, it attains a higher correlation with human judgments than existing evaluation metrics.
The performance of ChatGPT on summarization evaluation is highly dependent on prompt design.
Human-like evaluation with ChatGPT is more cost-effective and reproducible than human evaluation.
The explanation generated by ChatGPT is consistent with its scoring.