Exploring the Limits of ChatGPT for Query or Aspect-based Text Summarization

Xianjun Yang, Yan Li, Xinlu Zhang, Haifeng Chen, Wei Cheng

Introduction

Text summarization has long been a pivotal challenge in the field of Natural Language Processing (NLP). The main objective of this task is to succinctly condense a lengthy document into a shorter version while ensuring that the most crucial information is preserved. With the recent rise of advanced language models like ChatGPT, there has been a heightened interest in leveraging these models for text summarization tasks. However, it is noteworthy that the majority of existing research studies Goyal et al. (2022); Zhang et al. (2023) have primarily concentrated on generating a general summary for news-related content.

Aspect- or query-based summarization represents a more diverse and nuanced form of text summarization that has garnered significant attention within the NLP community. Unlike generic summarization, these tasks involve generating summaries that are customized to particular aspects or queries, rather than a single condensed version of the entire document. Consequently, this approach demands a deeper level of comprehension of the document, with respect to the specific interests and needs of the users.

In this paper, we present a comprehensive evaluation of ChatGPT’s performance on four distinct aspect-based and query-based text summarization tasks. Our experimental analysis indicates that ChatGPT’s summarization capabilities are on par with traditional fine-tuning methods, based on Rouge scores. The outcomes of this study offer valuable perspectives on the potential of ChatGPT for text summarization tasks, and emphasize the necessity for innovative approaches in this field. The achievement of ChatGPT in text summarization tasks holds promising implications for the development of practical and effective summarization systems.

Recently, a study by Goyal et al. (2022) demonstrated that, while GPT-3 generated summarizations achieved lower Rouge scores compared to traditional fine-tuning methods, human annotators favored the text generated by GPT-3. In addition, a thorough analysis of large language models for news summarization by Zhang et al. (2023) concluded that the summarizations produced by these models were already comparable to those generated by humans, which they attributed to instruction tuning. Notably, Bang et al. (2023) conducted a comprehensive investigation of ChatGPT’s multi-task, multilingual, and multimodal evaluation, including text summarization as a case study, and arrived at similar conclusions. However, the evaluation datasets used in the News domain for general summarization were not explicitly designed for text summarization and focused solely on general summarization. As such, we aim to explore how ChatGPT performs in the diverse summarization of lengthy articles across multiple domains, using high-quality data.

This work makes several significant contributions, including:

Being the first systematic attempt to extend the usage of LLMs beyond generic summarization and examining the performance of ChatGPT in aspect or query-based summarization.

Demonstrating that ChatGPT-generated diverse specific summaries are highly comparable to traditional fine-tuning methods in terms of Rouge scores.

Conducting an in-depth analysis of the LLM-generated summaries and identifying several potential future research directions that could leverage the strengths of LLMs.

Together, these contributions provide novel insights into the capabilities of ChatGPT for diverse text summarization tasks and underscore the potential of LLMs as a powerful tool for NLP research.

Related Work

Aspect- and query-based summarization are two critical forms of text summarization that differ significantly from general summaries in that they are not input-agnostic. These tasks aim to generate summaries that are tailored to specific aspects or queries for various types of content, such as news articles Kulkarni et al. (2020), meetings Zhong et al. (2021), stories Wang et al. (2022), and Wikipedia articles Yang et al. (2022). This approach contrasts with previous methods such as CNN/DM Hermann et al. (2015) or XSUM Narayan et al. (2018), which focus on developing a single generic summary of the entire document. By leveraging aspect- or query-based summarization, it is possible to create more targeted and personalized summaries that cater to the specific interests and needs of different users.

Aspect-based and query-based summarization can be accomplished using a variety of methods, including end-to-end and extract-then-summarize approaches. End-to-end summarization directly produces the summaries without manipulating the original inputs. In contrast, the extract-then-summarize approach involves identifying and extracting the most important sentences or phrases from the original document to form a shorter document, which is then summarized to fit the input limit of language models, such as BART, which has a token limit of 1024 Lewis et al. (2020). Additionally, abstractive summarization methods aim to generate new sentences or phrases that summarize the original document’s content, rather than simply extracting and rephrasing existing text. There is no one-size-fits-all approach to aspect- and query-based summarization, and the choice of method depends on factors such as the size and complexity of the input, the desired length and level of detail of the summary, and the target audience’s needs and preferences.

2 Large Language Models

In recent years, large language models such as GPT-3 Brown et al. (2020) and ChatGPT have garnered substantial interest in the field of natural language processing. These models are trained on vast quantities of text data and have achieved remarkable performance on a range of NLP tasks, including text classification, question answering, and machine translation.

Several studies have investigated using large language models for text summarization tasks. For instance, Goyal et al. Goyal et al. (2022) observed that while GPT-3-generated summaries obtained slightly lower Rouge scores than traditional fine-tuning methods, human evaluators preferred the former. Similarly, Zhang et al. Zhang et al. (2023) reported that LMM-generated summaries were considered as good as human-written summaries in the News domain. Besides, Qin et al. (2023) also competitively examine the performance of ChatGPT and GPT-3.5 for various tasks, including dialogue summarization dataset SAMSUm Gliwa et al. (2019)

As recent studies have highlighted the potential of large language models for text summarization, it is essential to further investigate their performance on diverse summarization tasks in various domains. Our work aims to contribute to this ongoing research by evaluating the capabilities of ChatGPT on aspect-based and query-based summarization tasks and providing insights into its strengths and limitations.

Task Formulation

Aspect- and query-based summarization are essential tasks for text summarization because they are considered more challenging and valuable for real-world production. These tasks aim to generate a summary tailored to specific aspects or queries, rather than generating a single generic summary of the entire document.

In this study, we evaluated the performance of ChatGPT on a series of aspect- and query-based text summarization benchmarks. The main steps of our experimental methodology are as follows:

Data collection: We selected publicly available datasets as listed in Tabel 1, ensuring that they are consistent with previous finetuning methods.

Model evaluation: We conducted an evaluation of ChatGPT’s performance on question and answer pairs, utilizing Rouge scores as our evaluation metric. Due to the lack of an API provided by ChatGPT for processing large amounts of input data, we manually evaluated 100 examples selected at random from each test set on the ChatGPT platform. Previous research efforts Goyal et al. (2022); Zhang et al. (2023) have also been limited in their testing of GPT-3 on a small number of instances.

Here we list prompts used in our experiments for generated summaries.

SQuALITY The prompt is Q: Query. Answer the question in around 200 words. Article: story. for a specific question, while Q: Query. Answer the question in around 450 words. Article: story. for a general question. In the second case, if the generated summaries are much shorter than 450 words, we will specify Your response is too short. Please answer it in around 450 words. to get a second-round conversation.

QMSum The prompt is Q: Query. Article: meeting or Q: Query. Article: golden meeting, where meeting is the initial meeting, while golden meeting is the provided golden spans of sentences in the original long meeting.

CovidET The prompt is Q: Summarize this article with respect to Aspect within one short sentence. Article0. A: Answer0. Q: Summarize this article with respect to Aspect within one short sentence. Article. A: , where Article0 and Answer0 are randomly picked from the training instances to serve as the in-context one-shot example. When within one short sentence is omitted, we observed the summaries are much longer and even on par with the input article. We observe a significantly lower performance for zero-shot from preliminary experiments, likely due to the concise input and output. Thus we adopt 1-shot experiments for CovidET.

NEWTS The prompt is Article. Summarize this article with respect to Aspect: , where the Aspect is some continuous words serving as certain topics.

Unless expressly stated, we did not conduct any further conversation with ChatGPT to correct the answer. We test zero-shot performance for all datasets except for CovidET.

Experiments and Analysis

We use the ChatGPT https://chat.openai.com/chat platform for conducting our experiments between February 10 to Feb. 15, 2022. To eliminate the effects of historical chats, we clear each conversation after generating each summary. For all datasets, we use their originally released corpus as testing examples.

2 Analysis

The overall results are shown in Table 2. As we can see, ChatGPT achieves comparable performance with traditional finetuning methods in all datasets. Surprisingly, when provided with golden annotation of meeting spans in QMSum, ChatGPT even outperforms finetuning in terms of Rouge-1 and Rouge-2, though Rouge-L lags. The worst performance of ChatGPT is observed in CovidET, where the inputs are usually around 128 words, and the summary is almost always comprised of only one sentence with around 20 words. We attribute this low performance to the untypical length in CovidET, compared with most summarization datasets where the inputs and outputs are always much longer. Regarding the news domain, ChatGPT outperforms finetuning in terms of all Rouge scores. This finding is consistent with the previous conclusion that Instruct-GPT could achieve near SOTA performance for a general summary in the news domain. We suspect this is due to the large availability of news corpus for pre-training.

In the context of meeting dialogue summarization in QMSum, we examine two scenarios where the input length exceeds the maximum token limit of ChatGPT. In the first scenario, we extract and summarize the input by splitting it into two parts, asking ChatGPT to extract salient information about the question, and then combining the extracted parts and performing a second round of summarization. We use finetuning as the comparing baseline for the extraction and summarization. The results demonstrate that ChatGPT outperforms finetuning in terms of Rouge-2 and exhibits comparable performance on Rouge-1, but performs significantly worse on Rouge-L. However, when given golden spans, ChatGPT performs slightly better on Rouge-1 and Rouge-2 in the zero-shot setting than fine-tuning on the golden inputs. In both cases, we observe a significant gap in Rouge-L, likely caused by the data feature of oral dialogues. Rouge-L stands for Longest Common Subsequence (LCS), and the finetuning could bias toward the close-to-oral dialogues summaries, but ChatGPT tends to make more formal summaries. As a result, lower Rouge-L does not necessarily imply worse performance, and we intend to evaluate it using human evaluations in the future.

Lastly, since the inputs in the English story summaries dataset SQuaLITY are usually longer than 3000 words, we directly truncate them to fit into ChatGPT, following finetuning baseline that also truncates them. Notice that for some queries that can not be answered by the truncated paragraphs, ChatGPT will directly return the message that it could not be answered. We abandon such instances since the summaries are meaningless. We also use two prompts for a general summary of the story’s plot or some specific questions, observing that they correspond to very different summaries as detailed in 3.1. From the results, we again see similar performance on all evaluations of Rouge scores, with ChatGPT lagging by only 1 point.

Following Grusky et al. (2018), we use the Coverage, Density, and Compression to measure to what extent the summary is derivative of a text, how well the word sequence of a summary can be described as a series of extractions, and word ratio between the article and summary. We also use unique n-grams(n=1,2,3,41,2,3,4) to denote how many unique words are presented in the summaries. The results are calculated for golden references and ChatGPT-generated summaries in Table 3. As we can see, ChatGPT-generated text consistently achieves a lower compression ratio, indicating that it prefers generating more extended summaries. While for coverage and density, there is no apparent difference for all scenarios. For articles with long inputs like QMSum and SQuaLITY, the unique n-grams(n=1,21,2) are usually higher for ChatGPT while lower for unique n-grams(n=3,43,4), suggesting that ChatGPT-summaries are more abstractive in terms of short words. While for the remaining datasets, ChatGPT almost always generates less fraction of unique n-grams.

3 Insights

In the above, we see the exceptional summarization ability of ChatGPT across Reddit posts, news, dialogue, and meeting domains toward various aspects and queries. From some case studies as shown in Table 4 and 5 in the Appendix, we can tell the ChatGPT-generated summaries are surprisingly good and even better than the given references. We leave the complete human evaluation for future work. Considering the zero-shot performance of ChatGPT does not involve any additional labeling efforts for training data, and ChatGPT could even achieve better results given more appropriate prompts or multiple conversations for self-correction, we believe it is time for rethinking future directions for various text summarization tasks. Theoretically speaking, our experiments merely establish the lower threshold of ChatGPT’s capabilities for aspect or query-based summarization. We are of the conviction that in the near future (possibly within a few months), ChatGPT could conceivably exceed the performance achieved through fine-tuning, owing to the utilization of superior prompts, the incorporation of multiple conversations involving self-correction, and the self-enhancement of ChatGPT itself.

Conclusion

In this paper, we evaluated the performance of ChatGPT on aspect- and query-based text summarization tasks across diverse domains. These results demonstrate the super ability of ChatGPT for various controllable text summarization tasks. However, the Rouge score might not be a good indicator for evaluating the performance of ChatGPT in text summarization tasks. We will conduct human evaluations of the generated text shortly to provide a more comprehensive assessment of ChatGPT’s performance. In conclusion, our findings suggest that ChatGPT holds promise as a powerful tool for text summarization and lays the insights for future research in this area.

In the era of ChatGPT, we conclude some future directions that are worthy of investigating: 1. Retrieval module: In view of the fact that the training of large language models such as ChatGPT is constantly challenged by constraints on input length, the solution lies in the adoption of a lighter model such as LED Beltagy et al. (2020), which is adept at swiftly retrieving significant sentences from lengthy inputs. By integrating LED, ChatGPT can effectively tackle the processing of lengthy documents. 2. GPT-generated text detection: Although ChatGPT-generated summaries are already very fluent and consistent, they might still include nonfactual or biased summaries. Thus it is important to develop tools for detecting such ChatGPT-generated summaries before they are widely deployed for real applications. 3. Better prompts: Given that our preliminary experiments have not thoroughly explored the possibility space of prompts, and we have yet to examine multiple conversations to refine our summaries, we are of the opinion that improving summary quality through enhanced prompting would be a topic of independent interest.

Limitations

In this study, we evaluated the performance of ChatGPT on aspect- and query-based text summarization tasks. However, our experiments were limited by the maximum input sequence length of ChatGPT, which is currently set to around 5000 tokens. This limitation could impact the generalizability of our results to other text summarization tasks and datasets, as the length of documents can vary widely in real-world applications.

It is important to note that the results of this study should not be used to make decisions that could negatively impact individuals or groups. Further research is needed to thoroughly assess the ethical implications of using language models such as ChatGPT for text summarization tasks, particularly regarding fairness, bias, and factuality.

References

Appendix

Appendix A Generated examples

Here we show some ChatGPT-generated summaries in Table 4 and 5, together with their golden references.