Abstractive Text Summarization by Incorporating Reader Comments
Shen Gao, Xiuying Chen, Piji Li, Zhaochun Ren, Lidong Bing, Dongyan Zhao, Rui Yan
Introduction
Abstractive summarization can be regarded as a sequence mapping task that the source text is mapped to the target summary, and has drawn much attention since the deep neural networks are widely applied in natural language processing field. Recently, sequence-to-sequence (seq2seq) framework (?) has been proved effective for the task of abstractive summarization (?; ?) and other text generation tasks (?; ?). In this paper, we use “aspect” to denote the topic described in a specific paragraph or a sentence of a news document, and use “main aspect” to denote the central topic which the author tends to convey to the readers. Although a document may describe an event in many different aspects, the summary of this document should always focus on the main aspect. As shown in Table 1, the good summary describes the main aspect and the bad summary describes another trivial aspect that is not the main point of the document. To focus on the main aspect, some summarization methods (?; ?; ?) first select several sentences about the main aspect and then generate the summary. However, it is very challenging to discover which is the main aspect of the news document.
Nowadays, a great number of news comments are generated by readers to express their opinions about the event. Some comments may mention the main aspect of the document for several times. Take the case in Table 1 as an example, the focused aspect of the reader is “investment of Toyota” which is also the main aspect of this document. To be specific, we define “reader focused aspect” to denote the focused aspect by a reader through the comments. Intuitively, these reader comments may help the summary generator capture the main aspect of document, thereby improving the quality of the generated summary. Therefore, in this paper, we investigate a new problem setting of the task of abstractive text summarization. We name such paradigm of extension as reader-aware abstractive text summarization.
The effect of comments or social contexts in document summarization have been explored by several previous works (?; ?; ?; ?). Unlike these approaches that directly extract sentences from the original document (?; ?; ?), we aim to generate a natural-sounding summary from scratch instead of extracting words from the document.
Generally, existing text summarization approaches confront two challenges when addressing reader-aware summarization task. The first challenge is that reader comments are very noisy and informative. Not all the information provided by the comments is useful when modeling the reader focused aspects. Therefore, it is crucial to make the model own the ability of capturing main aspect and filtering noisy information when incorporating reader comments. The second challenge is how to generate summaries by jointly modeling the main aspect of document and the reader focused aspect revealed by comments. Meanwhile, the model should not be sensitive to the diverse unimportant aspects introduced by some reader comments. Thus, simply absorbing all the reader aspect information to directly guide the model to generate summary is not feasible, as it will make the generator lose the ability of modeling the main aspect.
In this paper, we propose a summarization framework named reader-aware summary generator (RASG) that incorporates reader comments to improve the summarization performance. Specifically, a seq2seq architecture with attention mechanism is employed as the basic summary generator. We first calculate alignment between the reader comments words and document words, and this alignment information is regarded as reader attention representing the “reader focused aspect”. Then, we treat the decoder attention weights as the focused aspect of the generated summary, a.k.a., “decoder focused aspect”. After each decoding step, a supervisor is designed to measure the distance between the reader focused aspect and the decoder focused aspect. Given this distance, a goal tracker provides the goal to the decoder to induce it to reduce this distance. The training of our framework RASG is conducted in an adversarial way. To evaluate the performance of our model, we collect a large amount of document-summary pairs associated with several reader comments from social media website. Extensive experiments conducted on this dataset show that RASG significantly outperforms the state-of-the-art baselines in terms of ROUGE metrics and human evaluations.
To sum up, our contributions can be summarized as follows:
We propose a reader-aware abstractive text summarization task. To solve this task, we propose an end-to-end learning framework to conduct the reader attention modeling and reader-aware summary generation.
We design a supervisor as well as a goal tracker to guide the generator to focus on the main aspect of the document.
To reduce the noisy information introduced by the reader comments, we propose a denoising module to identify which comments are helpful for summary generation automatically.
We release a large scale abstractive text summarization dataset associated with reader comments. Experimental results on this dataset demonstrate the effectiveness of our proposed framework.
Related Work
Text summarization can be classified into extractive and abstractive methods. Extractive methods (?; ?) read the article and get the representations of the sentences and article to select sentences. However, summaries generated by extractive methods always suffer from redundancy problem. Recently, with the emergence of neural network models for text generation, a vast majority of the literature on summarization is dedicated to abstractive summarization (?; ?; ?). On the text summarization benchmark dataset CNN/DailyMail, the state-of-the-art abstractive methods outperform the best extractive method in terms of ROUGE score. Most methods for abstractive text summarization are based on the sequence-to-sequence model (?), which encodes the source texts into the semantic representation with an encoder, and generates the summaries from the representation with a decoder. To tackle the out-of-vocabulary problem, some researchers employ the copy mechanism to copy some words from the input document to summary (?; ?). To capture the main aspect of document, Chen et al. (?) propose to select salient sentences and then rewrite these sentences to a concise summary. This approach achieves the state-of-the-art of text summarization on CNN/DailyMail benchmark dataset. Unlike document summarization that needs to encode a long text, social media summarization usually reads short and noisy text and has become a popular task these days. After Hu et al. (?) propose a short text summarization dataset on social media and many researchers follow this task. Lin et al. (?) propose a seq2seq based model which uses an CNN to refine the representation of source context. Wang et al. (?) use convolutional seq2seq model to summarize text and use the policy gradient algorithm to directly optimize the ROUGE score. However, these summarization models do not utilize the reader’s comments in generating summaries.
To consider the reader’s comments into text summarization, the reader-aware summarization is proposed and it mainly takes the form of extractive approaches. Graph-based method has been used for comment oriented summarization task such as (?; ?), where they identify three relations (topic, quotation, and mention) by which comments can be linked to one another. Recently, Nguyen et al. (?) publish a small extractive sentence-comment dataset which can not be used to train neural models due to its small size. Li et al. (?) propose an unsupervised compressive multi-document summarization model using sparse coding method. Following previous work, there are some models (?; ?) using variational auto-encoder to model the latent semantic of original article and reader comments. Different from our abstractive summarization task, these related works are all based on extractive or compressive approaches.
Problem Formulation
Before presenting our approach for the reader-aware summarization, we first introduce our notations and key concepts.
To begin with, for a document , we assume there is a comment set where is the -th comment, denotes the -th word in document , and denotes the -th word in -th comment sentence . Given the document , the summary generator reads the comments , then generates a summary . Finally, we use the difference between generated summary and ground truth summary as the training signal to optimize the model parameters.
The Proposed RASG Model
In this section, we propose our reader-aware summary generator, abbreviated as RASG. The overview of RASG is shown in Figure 1 which can be split into four main parts:
Summary generator is a seq2seq based architecture with attention and copy mechanisms.
Reader attention module learns a semantic alignment between each word in document and comments, thus captures the reader focused aspect.
Supervisor measures the semantic gap between decoder focused aspect and reader focused aspect. There is also a discriminator which uses convolutional neural network to extract features and then distinguishes how similar is decoder focused aspect to reader focused aspect.
Goal tracker utilizes the semantic gap learned by supervisor and the features extracted learned by the discriminator to set a goal, which is further utilized as a more specific guidance for summary generator to produce better summary.
Summary generator
At the beginning, we use an embedding matrix to map one-hot representation of each word in the document and comments to a high-dimensional vector space. We denote as the embedding representation of word . From these embedding representations, we employ a bi-directional recurrent neural network (Bi-RNN) to model the temporal interactions between words:
where denotes the hidden state of -th step in Bi-RNN for document . We denote the final hidden state of as the vector representation of the document . Following (?; ?), we choose the long short-term memory (LSTM) as the Bi-RNN cell.
Then we apply a linear transform layer on the input document vector representation and use the output of this layer as the initial state of decoder LSTM, shown in Equation 2. In order to reduce the burden of compressing document information into initial state , we use the attention mechanism (?) to summarize the input document into context vector dynamically and we will show the detail of these in the following sections. We then concatenate the context vector with the embedding of previous step output and feed this into decoder LSTM, shown in Equation 3. We use the notion as the concatenation of two vectors.
Finally, an output projection layer is applied to get the final generating distribution over vocabulary, as shown in Equation 7. We concatenate goal vector , gap content , and the output of decoder LSTM as the input of the output projection layer. The goal vector represents the goal of current generation step, the gap content denotes the semantic gap between generated summary and reader focused document and we will show the details of these variables in the following sections.
In order to handle the out-of-vocabulary (OOV) problem, we equip the pointer network (?; ?; ?) with our decoder, which makes our decoder capable to copy words from the source text. The design of the pointer network is the same as the model used in (?), thus we omit this procedure in our paper due to the limited space. We use the negative log-likelihood as the loss function:
Denoising module
Due to the fact that reader comments are a kind of informal text, they may consist of many noisy information, and not all the comments are helpful for generating better summaries. Consequently, we employ a denoising module to distinguish which comments are helpful. First, we employ an to model the comment word embeddings:
where denotes the hidden state of -th word in -th comment . Next, we use average-pooling operation over these hidden states to produce a vector representation of -th comment, shown in Equation 10. Finally, we apply a linear transform with sigmoid function to predict whether the comment is useful, and the sigmoid output also can be seen as a salience score of -th comment given the document representation .
To train the denoising module, we use the cross entropy loss to supervise this procedure.
where is the ground truth salience score of comments. denotes the -th comment is helpful for generating summary and vice-versa.
Reader attention modeling
To model the reader focused aspect, we first calculate the word alignment of reader comments towards the document. We use the embeddings of words in document and comments to calculate the semantic alignment score. Precisely, is the alignment socre between the -th document word and the -th word in the -th comment , as shown in Equation 13:
In Equation 14, we use a max-operation over the alignment to signify whether the -th word of document is focused by the -th comment. We regard the alignment score as the reader attention weight for the -th reader comment to the -th document word.
In order to reduce the interference caused by the noisy comments, we employ the comment salience score obtained from the denoising module to weighted combine the -th reader attention , as shown in Equation 15. It means that noisy comments will contribute less in the procedure of reader attention modeling.
Supervisor
where represents the focused aspect by the latest decoding steps, a.k.a., decoder focused aspect.
Next, we use the reader attention to weighted sum the document hidden states :
where represents the reader focused aspect.
For encouraging the decoder focused aspect become similar to the reader focused aspect, we employ an CNN based discriminator to signify the difference between the decoder focused aspect and the reader focused aspect . Then we can use this difference to guide the decoder focus on the reader focused aspect. Typically, the discriminator is a binary classifier which can be decomposed into a convolutional feature extractor shown in Equation 20 and a sigmoid classification layer shown in Equation 21 and 22.
where denotes the convolutional operation, trainable parameter denotes the convolutional kernel, and and are both the classification probabilities.
Note that a token generated at time will influence not only the gradient received at that time but also the gradient at subsequent time steps. Intuitively, the decoding attention of latter decoding step is more similar to the attention of final summary than the earlier steps. Thus we propose to define the cumulative loss with a discount factor as the loss functions. Note that the training objective for discriminator can be interpreted as maximizing the log-likelihood for classification, whether the input in Equation 20 comes from reader focused aspect or from decoder focused aspect.
where denotes the semantic of unfocused document aspects by summary generator, a.k.a., gap content. To encourage the summary generator focus on the unfocused document aspects, we feed the gap content to the generator, as shown in Equation 7.
Goal tracker
Since the discriminator only provides a scalar guiding signal at each decoding step, it becomes relatively less informative when the sentence length goes larger. Inspired by LeakGAN (?), the proposed RASG framework allows discriminator to provide additional information, denoted as goal vector . In view of there is certain relationship between the goal of current decoding step and previous steps, we need to model the temporal interactions between the goal of each step. More specifically, we introduce a goal tracker module, an LSTM that takes the extracted feature vector and gap content as its input at each step and outputs a goal vector :
In order to achieve higher consistency of reader focused aspect, we feed the goal vector into the generator to guide the generation of the next word, as shown in Equation 7.
Model training
As our model is trained in an adversarial manner, we re-split the parameters in our model into two parts: (1) generation module including the parameters of summary generator, reader attention module and goal tracker; (2) discriminator module including the parameters of CNN classifier. As for training generation module, we sum up the loss function of denoising module , cross entropy between ground truth and the result of discriminator , as shown in Equation 27. We use the to optimize the parameters of generation module.
Next, we train the discriminator module to maximize the probability of assigning the correct label to both generated aspect and reader focused aspect . More specifically, we optimize the parameters of discriminator module according to the loss function calculated in Equation 23.
Experimental Setup
We list four research questions that guide the experiments: RQ1: Does RASG outperform other baselines? RQ2: What is the effect of each module in RASG? RQ3: Does RASG capture useful information from noisy comments? RQ4: Can goal tracker give a helpful guidance to decoder?
Dataset
We collect the document-summary-comments pair data from Weibo which is the largest social network website in China, and users can read a document and post a comment about the document on this website. Each sample of data contains a document, a summary and several reader comments. Most comments are about the readers’ opinion of their focused aspect in the document. In order to train the denoising module, we should give a ground truth label for -th comment. When there is at least one common word in summary and comment, we regard such comment is helpful for generating summary. Accordingly, we give the to -th comment when it contains at least one common word and give when it does not. In total, our training dataset contains 863826 training samples. The average length of document is 67.08 words, average length of comment is 16.61 words and average length of summary is 16.56 words. The average comments number of a document is 9.11.
Evaluation metrics
For evaluation metrics, we adopt ROUGE score (?) which is widely applied for summarization evaluation (?; ?). The ROUGE metrics compare generated summary with the reference summary by computing overlapping lexical units, including ROUGE-1 (unigram), ROUGE-2 (bi-gram) and ROUGE-L (longest common subsequence).
Comparison methods
In order to prove the effectiveness of each module in RASG, we conduct some ablation models introduced in Table 2.
To evaluate the performance of our proposed dataset and model, we compare it with the following baselines:
(1) S2S: Sequence-to-sequence framework (?) has been proposed for language generation task. (2) S2SR: We simply add the reader attention on attention distribution in each decoding step. (3) CGU: Lin et al. (?) propose to use the convolutional gated unit to refine the source representation, which achieves the state-of-the-art performance on social media text summarization dataset. (4) LEAD1: LEAD1 is a commonly used baseline (?; ?), which selects the first sentence of document as the summary. (5) TextRank: Mihalcea et al. (?) propose to build a graph, then add each sentence as a vertex and use link to represent semantic similarity. Sentences are sorted based on final scores and a greedy algorithm is employed to select summary sentences.
Implementation details
We implement our experiments in TensorFlow (?) on an NVIDIA P40 GPU. The word embedding dimension is set to 256 and the number of hidden units is 512. We set the in the Equation 17 and in Equation 23 and 24. We use Adagrad optimizer (?) as our optimizing algorithm. We employ beam search with beam size 5 to generate more fluency summary sentence.
Experimental Results
For research question RQ1, we examine the performance of our model in terms of ROUGE. Table 3 lists performances of all comparisons in terms of ROUGE score. We see that RASG achieves a 11.0%, 9.1% and 6.6% increment over the state-of-the-art method CGU in terms of ROUGE-1, ROUGE-2 and ROUGE-L respectively. It is worth noticing that the baseline model S2SR achieves better performance than S2S which demonstrates the effectiveness of incorporating reader focused aspect in summary generation. However when compared with RASG, S2SR achieves lower performance in terms of all ROUGE score. Thus, simply adding the reader focused aspect into generation procedure is not a good reader-aware summarization method.
Ablation study
Next, we turn to research question RQ2. We conduct ablation tests on the usage of denoising module, supervisor as well as the goal tracker and the ROUGE score result is shown in Table 4. The discriminator provides the scalar training signal for generator training and the feature vector for goal tracker. Consequently, there is an increment of 17.51% from RASG w/o GTD to RASG w/o GT in terms of ROUGE-L, which demonstrates the effectiveness of discriminator. As for the effectiveness of goal tracker, compared with RASG and RASG w/o GT, RASG w/o GTD offers a decrease of 45.23% and 17.88% in terms of ROUGE-1, respectively. This demonstrates that the goal tracker with the feature from discriminator plays an important role in producing better summary. However, using the goal tracker without the feature extracted by the discriminator does not help improve the performance of the summary generator, shown by the performance of RASG w/o GTD. Finally, RASG w/o DM offers a decrease of 10.22% compared with RASG in terms of ROUGE-L, which demonstrates the effectiveness of denoising module.
Denoising ability
Next, we turn to research question RQ3. Due to the fact that the denoising module is learned in a supervised way, there is a ground truth label associated with each comment. Thus when the predict salience score we classify it as a helpful comment and vice-versa. As the denoising module can be regarded as a binary classifier to classify each comment to or , we calculate the classification recall score of comments to measure the performance of this module. The recall curve is shown in Figure 2. As the training progresses, the recall score is on a steady upward curve which proves the improved performance of denoising module. To conclude, the denoising module can give a meaningful salience score for the subsequent process.
Analysis of goal tracker
Human evaluation
We ask three highly-educated Ph.D. students to rate 100 generated summaries of different models according to consistency and fluency. These annotators are all native speakers. The rating score ranges from 1 to 3 and 3 is the best. We take the average score of all summaries as the final score of each model, as shown in Table 5. It can be seen that RASG outperforms other baseline models in both sentence fluency and consistency by a large margin. We calculate the kappa statistics in terms of fluency and consistency, and the score is 0.33 and 0.29 respectively. To prove the significance of the above results, we also do the paired student t-test between our model and CGU model (row with shaded background), the p-value are 0.0017 and 0.0012 for fluency and consistency respectively.
Case analysis
Figure 3 shows a document and its corresponding summaries generated by different methods. We can observe that S2S does generate fluent summary. However, the generated aspect is contradictory to the focused aspect of reader or ground truth summary. Meanwhile, RASG overcomes this shortcoming by using goal vector and gap content given by goal tracker and supervisor at training stage, and produces the summary that is not only fluent but also consistent with main aspect of document.
Conclusion
In this paper, we propose a new framework named reader-aware summary generator (RASG) which aims to generate summaries for document from social media incorporating the reader comments. In order to capture the reader focused aspect, we design a reader attention component with a denoising module to capture the alignment between comments and document. We employ a supervisor to measure the semantic gap between generated summary and reader focused aspect. A goal tracker uses the information of semantic gap and the feature extracted by the discriminator to produce a goal vector to guide the summary generator. In our experiments, we have demonstrated the effectiveness of RASG and have found significant improvements over state-of-the-art baselines in terms of ROUGE and human evaluations. Moreover, we have verified the effectiveness of each module in RASG for improving the summarization performance.
Acknowledgments
We would like to thank the anonymous reviewers for their constructive comments. We would also like to thank Zhujun Zhang, Sicong Jiang for their helps on this project. This work was supported by the National Key Research and Development Program of China (No. 2017YFC0804001), the National Science Foundation of China (NSFC No. 61876196, No. 61672058), Alibaba Innovative Research (AIR) Fund. Rui Yan was sponsored by CCF-Tencent Open Research Fund and Microsoft Research Asia (MSRA) Collaborative Research Program.