A Unified Model for Extractive and Abstractive Summarization using Inconsistency Loss

Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, Min Sun

Introduction

Text summarization is the task of automatically condensing a piece of text to a shorter version while maintaining the important points. The ability to condense text information can aid many applications such as creating news digests, presenting search results, and generating reports. There are mainly two types of approaches: extractive and abstractive. Extractive approaches assemble summaries directly from the source text typically selecting one whole sentence at a time. In contrast, abstractive approaches can generate novel words and phrases not copied from the source text. Hence, abstractive summaries can be more coherent and concise than extractive summaries.

Extractive approaches are typically simpler. They output the probability of each sentence to be selected into the summary. Many earlier works on summarization Cheng and Lapata (2016); Nallapati et al. (2016a, 2017); Narayan et al. (2017); Yasunaga et al. (2017) focus on extractive summarization. Among them, Nallapati et al. (2017) have achieved high ROUGE scores. On the other hand, abstractive approaches Nallapati et al. (2016b); See et al. (2017); Paulus et al. (2017); Fan et al. (2017); Liu et al. (2017) typically involve sophisticated mechanism in order to paraphrase, generate unseen words in the source text, or even incorporate external knowledge. Neural networks Nallapati et al. (2017); See et al. (2017) based on the attentional encoder-decoder model Bahdanau et al. (2014) were able to generate abstractive summaries with high ROUGE scores but suffer from inaccurately reproducing factual details and an inability to deal with out-of-vocabulary (OOV) words. Recently, See et al. (2017) propose a pointer-generator model which has the abilities to copy words from source text as well as generate unseen words. Despite recent progress in abstractive summarization, extractive approaches Nallapati et al. (2017); Yasunaga et al. (2017) and lead-3 baseline (i.e., selecting the first 3 sentences) still achieve strong performance in ROUGE scores.

We propose to explicitly take advantage of the strength of state-of-the-art extractive and abstractive summarization and introduced the following unified model. Firstly, we treat the probability output of each sentence from the extractive model Nallapati et al. (2017) as sentence-level attention. Then, we modulate the word-level dynamic attention from the abstractive model See et al. (2017) with sentence-level attention such that words in less attended sentences are less likely to be generated. In this way, extractive summarization mostly benefits abstractive summarization by mitigating spurious word-level attention. Secondly, we introduce a novel inconsistency loss function to encourage the consistency between two levels of attentions. The loss function can be computed without additional human annotation and has shown to ensure our unified model to be mutually beneficial to both extractive and abstractive summarization. On CNN/Daily Mail dataset, our unified model achieves state-of-the-art ROUGE scores and outperforms a strong extractive baseline (i.e., lead-3). Finally, to ensure the quality of our unified model, we conduct a solid human evaluation and confirm that our method significantly outperforms recent state-of-the-art methods in informativity and readability.

To summarize, our contributions are twofold:

We propose a unified model combining sentence-level and word-level attentions to take advantage of both extractive and abstractive summarization approaches.

We propose a novel inconsistency loss function to ensure our unified model to be mutually beneficial to both extractive and abstractive summarization. The unified model with inconsistency loss achieves the best ROUGE scores on CNN/Daily Mail dataset and outperforms recent state-of-the-art methods in informativity and readability on human evaluation.

Related Work

Text summarization has been widely studied in recent years. We first introduce the related works of neural-network-based extractive and abstractive summarization. Finally, we introduce a few related works with hierarchical attention mechanism.

Extractive summarization. Kågebäck et al. (2014) and Yin and Pei (2015) use neural networks to map sentences into vectors and select sentences based on those vectors. Cheng and Lapata (2016), Nallapati et al. (2016a) and Nallapati et al. (2017) use recurrent neural networks to read the article and get the representations of the sentences and article to select sentences. Narayan et al. (2017) utilize side information (i.e., image captions and titles) to help the sentence classifier choose sentences. Yasunaga et al. (2017) combine recurrent neural networks with graph convolutional networks to compute the salience (or importance) of each sentence. While some extractive summarization methods obtain high ROUGE scores, they all suffer from low readability.

Abstractive summarization. Rush et al. (2015) first bring up the abstractive summarization task and use attention-based encoder to read the input text and generate the summary. Based on them, Miao and Blunsom (2016) use a variational auto-encoder and Nallapati et al. (2016b) use a more powerful sequence-to-sequence model. Besides, Nallapati et al. (2016b) create a new article-level summarization dataset called CNN/Daily Mail by adapting DeepMind question-answering dataset Hermann et al. (2015). Ranzato et al. (2015) change the traditional training method to directly optimize evaluation metrics (e.g., BLEU and ROUGE). Gu et al. (2016), See et al. (2017) and Paulus et al. (2017) combine pointer networks Vinyals et al. (2015) into their models to deal with out-of-vocabulary (OOV) words. Chen et al. (2016) and See et al. (2017) restrain their models from attending to the same word to decrease repeated phrases in the generated summary. Paulus et al. (2017) use policy gradient on summarization and state out the fact that high ROUGE scores might still lead to low human evaluation scores. Fan et al. (2017) apply convolutional sequence-to-sequence model and design several new tasks for summarization. Liu et al. (2017) achieve high readability score on human evaluation using generative adversarial networks.

Hierarchical attention. Attention mechanism was first proposed by Bahdanau et al. (2014). Yang et al. (2016) proposed a hierarchical attention mechanism for document classification. We adopt the method of combining sentence-level and word-level attention in Nallapati et al. (2016b). However, their sentence attention is dynamic, which means it will be different for each generated word. Whereas our sentence attention is fixed for all generated words. Inspired by the high performance of extractive summarization, we propose to use fixed sentence attention.

Our model combines state-of-the-art extractive model Nallapati et al. (2017) and abstractive model See et al. (2017) by combining sentence-level attention from the former and word-level attention from the latter. Furthermore, we design an inconsistency loss to enhance the cooperation between the extractive and abstractive models.

Our Unified Model

We propose a unified model to combine the strength of both state-of-the-art extractor Nallapati et al. (2017) and abstracter See et al. (2017). Before going into details of our model, we first define the tasks of the extractor and abstracter.

Problem definition. The input of both extractor and abstracter is a sequence of words w=[w1,w2,...,wm,...]\mathbf{w}=\left[w_{1},w_{2},...,w_{m},...\right], where mm is the word index. The sequence of words also forms a sequence of sentences s=[s1,s2,...,sn,...]\mathbf{s}=\left[s_{1},s_{2},...,s_{n},...\right], where nn is the sentence index. The mthm^{th} word is mapped into the n(m)thn(m)^{th} sentence, where n(⋅)n(\cdot) is the mapping function. The output of the extractor is the sentence-level attention β=[β1,β2,...,βn,...]\bm{\beta}=\left[\beta_{1},\beta_{2},...,\beta_{n},...\right], where βn\beta_{n} is the probability of the nthn^{th} sentence been extracted into the summary. On the other hand, our attention-based abstractor computes word-level attention αt=[α1t,α2t,...,αmt,...]\bm{\alpha}^{t}=\left[\alpha_{1}^{t},\alpha_{2}^{t},...,\alpha_{m}^{t},...\right] dynamically while generating the ttht^{th} word in the summary. The output of the abstracter is the summary text y=[y1,y2,...,yt,...]\textbf{y}=\left[y^{1},y^{2},...,y^{t},...\right], where yty^{t} is ttht^{th} word in the summary.

In the following, we introduce the mechanism to combine sentence-level and word-level attentions in Sec. 3.1. Next, we define the novel inconsistency loss that ensures extractor and abstracter to be mutually beneficial in Sec. 3.2. We also give the details of our extractor in Sec. 3.3 and our abstracter in Sec. 3.4. Finally, our training procedure is described in Sec. 3.5.

Pieces of evidence (e.g., Vaswani et al. (2017)) show that attention mechanism is very important for NLP tasks. Hence, we propose to explicitly combine the sentence-level βn\beta_{n} and word-level αmt\alpha^{t}_{m} attentions by simple scalar multiplication and renormalization. The updated word attention α^mt\hat{\alpha}^{t}_{m} is

The multiplication ensures that only when both word-level αmt\alpha^{t}_{m} and sentence-level βn\beta_{n} attentions are high, the updated word attention α^mt\hat{\alpha}^{t}_{m} can be high. Since the sentence-level attention βn\beta_{n} from the extractor already achieves high ROUGE scores, βn\beta_{n} intuitively modulates the word-level attention αmt\alpha^{t}_{m} to mitigate spurious word-level attention such that words in less attended sentences are less likely to be generated (see Fig. 2). As highlighted in Sec. 3.4, the word-level attention α^mt\hat{\alpha}^{t}_{m} significantly affects the decoding process of the abstracter. Hence, an updated word-level attention is our key to improve abstractive summarization.

2 Inconsistency Loss

Instead of only leveraging the complementary nature between sentence-level and word-level attentions, we would like to encourage these two-levels of attentions to be mostly consistent to each other during training as an intrinsic learning target for free (i.e., without additional human annotation). Explicitly, we would like the sentence-level attention to be high when the word-level attention is high. Hence, we design the following inconsistency loss,

where K\mathcal{K} is the set of top K attended words and TT is the number of words in the summary. This implicitly encourages the distribution of the word-level attentions to be sharp and sentence-level attention to be high. To avoid the degenerated solution for the distribution of word attention to be one-hot and sentence attention to be high, we include the original loss functions for training the extractor ( LextL_{ext} in Sec. 3.3) and abstracter (LabsL_{abs} and LcovL_{cov} in Sec. 3.4). Note that Eq. 1 is the only part that the extractor is interacting with the abstracter. Our proposed inconsistency loss facilitates our end-to-end trained unified model to be mutually beneficial to both the extractor and abstracter.

3 Extractor

Our extractor is inspired by Nallapati et al. (2017). The main difference is that our extractor does not need to obtain the final summary. It mainly needs to obtain a short list of important sentences with a high recall to further facilitate the abstractor. We first introduce the network architecture and the loss function. Finally, we define our ground truth important sentences to encourage high recall.

Architecture. The model consists of a hierarchical bidirectional GRU which extracts sentence representations and a classification layer for predicting the sentence-level attention βn\beta_{n} for each sentence (see Fig. 3).

Extractor loss. The following sigmoid cross entropy loss is used,

where gn∈{0,1}g_{n}\in\{0,1\} is the ground-truth label for the nthn^{th} sentence and NN is the number of sentences. When gn=1g_{n}=1, it indicates that the nthn^{th} sentence should be attended to facilitate abstractive summarization.

Ground-truth label. The goal of our extractor is to extract sentences with high informativity, which means the extracted sentences should contain information that is needed to generate an abstractive summary as much as possible. To obtain the ground-truth labels g={gn}n\mathbf{g}=\{g_{n}\}_{n}, first, we measure the informativity of each sentence sns_{n} in the article by computing the ROUGE-L recall score Lin (2004) between the sentence sns_{n} and the reference abstractive summary y^={y^t}t\mathbf{\hat{y}}=\{\hat{y}^{t}\}_{t}. Second, we sort the sentences by their informativity and select the sentence in the order of high to low informativity. We add one sentence at a time if the new sentence can increase the informativity of all the selected sentences. Finally, we obtain the ground-truth labels g\mathbf{g} and train our extractor by minimizing Eq. 3. Note that our method is different from Nallapati et al. (2017) who aim to extract a final summary for an article so they use ROUGE F-1 score to select ground-truth sentences; while we focus on high informativity, hence, we use ROUGE recall score to obtain as much information as possible with respect to the reference summary y^\hat{\mathbf{y}}.

4 Abstracter

The second part of our model is an abstracter that reads the article; then, generate a summary word-by-word. We use the pointer-generator network proposed by See et al. (2017) and combine it with the extractor by combining sentence-level and word-level attentions (Sec. 3.1).

Pointer-generator network. The pointer-generator network See et al. (2017) is a specially designed sequence-to-sequence attentional model that can generate the summary by copying words in the article or generating words from a fixed vocabulary at the same time. The model contains a bidirectional LSTM which serves as an encoder to encode the input words w\mathbf{w} and a unidirectional LSTM which serves as a decoder to generate the summary y\mathbf{y}. For details of the network architecture, please refer to See et al. (2017). In the following, we describe how the updated word attention α^t\bm{\hat{\alpha}}^{t} affects the decoding process.

Notations. We first define some notations. hmeh^{e}_{m} is the encoder hidden state for the mthm^{th} word. htdh^{d}_{t} is the decoder hidden state in step tt. h∗(α^t)=∑mMα^mt×hmeh^{*}(\bm{\hat{\alpha}}^{t})=\sum_{m}^{M}\hat{\alpha}_{m}^{t}\times h^{e}_{m} is the context vector which is a function of the updated word attention α^t\bm{\hat{\alpha}}^{t}. Pvocab(h∗(α^t))\mathbf{P}^{vocab}(h^{*}(\bm{\hat{\alpha}}^{t})) is the probability distribution over the fixed vocabulary before applying the copying mechanism.

where W1W_{1}, W2W_{2}, b1b_{1} and b2b_{2} are learnable parameters. Pvocab={Pwvocab}w\mathbf{P}^{vocab}=\{P^{vocab}_{w}\}_{w} where Pwvocab(h∗(α^t))P^{vocab}_{w}(h^{*}(\bm{\hat{\alpha}}^{t})) is the probability of word ww being decoded. pgen(h∗(α^t))∈p^{gen}(h^{*}(\bm{\hat{\alpha}}^{t}))\in is the generating probability (see Eq.8 in See et al. (2017)) and 1−pgen(h∗(α^t))1-p^{gen}(h^{*}(\bm{\hat{\alpha}}^{t})) is the copying probability.

Final word distribution. Pwfinal(α^t)P_{w}^{final}(\bm{\hat{\alpha}}^{t}) is the final probability of word ww being decoded (i.e., yt=wy^{t}=w). It is related to the updated word attention α^t\bm{\hat{\alpha}}^{t} as follows (see Fig. 4),

Note that Pfinal={Pwfinal}w\mathbf{P}^{final}=\{P_{w}^{final}\}_{w} is the probability distribution over the fixed vocabulary and out-of-vocabulary (OOV) words. Hence, OOV words can be decoded. Most importantly, it is clear from Eq. 3.4 that Pwfinal(α^t)P_{w}^{final}(\bm{\hat{\alpha}}^{t}) is a function of the updated word attention α^t\bm{\hat{\alpha}}^{t}. Finally, we train the abstracter to minimize the negative log-likelihood:

where y^t\hat{y}^{t} is the ttht^{th} token in the reference abstractive summary.

Coverage mechanism. We also apply coverage mechanism See et al. (2017) to prevent the abstracter from repeatedly attending to the same place. In each decoder step tt, we calculate the coverage vector ct=∑t′=1t−1α^t′\mathbf{c}^{t}=\sum_{t^{\prime}=1}^{t-1}\bm{\hat{\alpha}}^{t^{\prime}} which indicates so far how much attention has been paid to every input word. The coverage vector ct\mathbf{c}^{t} will be used to calculate word attention α^t\bm{\hat{\alpha}}^{t} (see Eq.11 in See et al. (2017)). Moreover, coverage loss LcovL_{cov} is calculated to directly penalize the repetition in updated word attention α^t\bm{\hat{\alpha}}^{t}:

The objective function for training the abstracter with coverage mechanism is the weighted sum of negative log-likelihood and coverage loss.

5 Training Procedure

We first pre-train the extractor by minimizing LextL_{ext} in Eq. 3 and the abstracter by minimizing LabsL_{abs} and LcovL_{cov} in Eq. 6 and Eq. 7, respectively. When pre-training, the abstracter takes ground-truth extracted sentences (i.e., sentences with gn=1g_{n}=1) as input. To combine the extractor and abstracter, we proposed two training settings : (1) two-stages training and (2) end-to-end training.

End-to-end training. For end-to-end training, the sentence-level attention β\bm{\beta} is soft attention and will be combined with the word-level attention αt\bm{\alpha}^{t} as described in Sec. 3.1. We end-to-end train the extractor and abstracter by minimizing four loss functions: LextL_{ext}, LabsL_{abs}, LcovL_{cov}, as well as LincL_{inc} in Eq. 2. The final loss is as below:

where λ1\lambda_{1}, λ2\lambda_{2}, λ3\lambda_{3}, λ4\lambda_{4} are hyper-parameters. In our experiment, we give LextL_{ext} a bigger weight (e.g., λ1=5\lambda_{1}=5) when end-to-end training with LincL_{inc} since we found that LincL_{inc} is relatively large such that the extractor tends to ignore LextL_{ext}.

Experiments

We introduce the dataset and implementation details of our method evaluated in our experiments.

We evaluate our models on the CNN/Daily Mail dataset Hermann et al. (2015); Nallapati et al. (2016b); See et al. (2017) which contains news stories in CNN and Daily Mail websites. Each article in this dataset is paired with one human-written multi-sentence summary. This dataset has two versions: anonymized and non-anonymized. The former contains the news stories with all the named entities replaced by special tokens (e.g., @entity2); while the latter contains the raw text of each news story. We follow See et al. (2017) and obtain the non-anonymized version of this dataset which has 287,113 training pairs, 13,368 validation pairs and 11,490 test pairs.

2 Implementation Details

We train our extractor and abstracter with 128-dimension word embeddings and set the vocabulary size to 50k for both source and target text. We follow Nallapati et al. (2017) and See et al. (2017) and set the hidden dimension to 200 and 256 for the extractor and abstracter, respectively. We use Adagrad optimizer Duchi et al. (2011) and apply early stopping based on the validation set. In the testing phase, we limit the length of the summary to 120.

Pre-training. We use learning rate 0.15 when pre-training the extractor and abstracter. For the extractor, we limit both the maximum number of sentences per article and the maximum number of tokens per sentence to 50 and train the model for 27k iterations with the batch size of 64. For the abstracter, it takes ground-truth extracted sentences (i.e., sentences with gn=1g_{n}=1) as input. We limit the length of the source text to 400 and the length of the summary to 100 and use the batch size of 16. We train the abstracter without coverage mechanism for 88k iterations and continue training for 1k iterations with coverage mechanism (Labs:Lcov=1:1L_{abs}:L_{cov}=1:1).

Two-stages training. The abstracter takes extracted sentences with βn>\beta_{n}> 0.5, where β\bm{\beta} is obtained from the pre-trained extractor, as input during two-stages training. We finetune the abstracter for 10k iterations.

End-to-end training. During end-to-end training, we will minimize four loss functions (Eq. 8) with λ1=5\lambda_{1}=5 and λ2=λ3=λ4=1\lambda_{2}=\lambda_{3}=\lambda_{4}=1. We set K to 3 for computing LincL_{inc}. Due to the limitation of the memory, we reduce the batch size to 8 and thus use a smaller learning rate 0.01 for stability. The abstracter here reads the whole article. Hence, we increase the maximum length of source text to 600. We end-to-end train the model for 50k iterations.

Results

Our unified model not only generates an abstractive summary but also extracts the important sentences in an article. Our goal is that both of the two types of outputs can help people to read and understand an article faster. Hence, in this section, we evaluate the results of our extractor in Sec. 5.1 and unified model in Sec. 5.2. Furthermore, in Sec. 5.3, we perform human evaluation and show that our model can provide a better abstractive summary than other baselines.

To evaluate whether our extractor obtains enough information for the abstracter, we use full-length ROUGE recall scoresAll our ROUGE scores are reported by the official ROUGE script. We use the pyrouge package. https://pypi.org/project/pyrouge/0.1.3/ between the extracted sentences and reference abstractive summary. High ROUGE recall scores can be obtained if the extracted sentences include more words or sequences overlapping with the reference abstractive summary. For each article, we select sentences with the sentence probabilities β\beta greater than 0.50.5. We show the results of the ground-truth sentence labels (Sec. 3.3) and our models on the test set of the CNN/Daily Mail dataset in Table 1. Note that the ground-truth extracted sentences can’t get ROUGE recall scores of 100 because reference summary is abstractive and may contain some words and sequences that are not in the article. Our extractor performs the best when end-to-end trained with inconsistency loss.

2 Results of Abstractive Summarization

We use full-length ROUGE-1, ROUGE-2 and ROUGE-L F-1 scores to evaluate the generated summaries. We compare our models (two-stage and end-to-end) with state-of-the-art abstractive summarization models Nallapati et al. (2016b); Paulus et al. (2017); See et al. (2017); Liu et al. (2017) and a strong lead-3 baseline which directly uses the first three article sentences as the summary. Due to the writing style of news articles, the most important information is often written at the beginning of an article which makes lead-3 a strong baseline. The results of ROUGE F-1 scores are shown in Table 2. We prove that with help of the extractor, our unified model can outperform pointer-generator (the third row in Table 2) even with two-stages training (the fifth row in Table 2). After end-to-end training without inconsistency loss, our method already achieves better ROUGE scores by cooperating with each other. Moreover, our model end-to-end trained with inconsistency loss achieves state-of-the-art ROUGE scores and exceeds lead-3 baseline.

where TT is the length of the summary. The average inconsistency rates on test set are shown in Table 4. Our inconsistency loss significantly decrease RincR_{inc} from about 20%20\% to 4%4\%. An example of inconsistency improvement is shown in Fig. 5.

3 Human Evaluation

We perform human evaluation on Amazon Mechanical Turk (MTurk)https://www.mturk.com/ to evaluate the informativity, conciseness and readability of the summaries. We compare our best model (end2end with inconsistency loss) with pointer-generator See et al. (2017), generative adversarial network Liu et al. (2017) and deep reinforcement model Paulus et al. (2017). For these three models, we use the test set outputs provided by the authorshttps://github.com/abisee/pointer-generator and https://likicode.com for the first two. For DeepRL, we asked through email.. We randomly pick 100 examples in the test set. All generated summaries are re-capitalized and de-tokenized. Since Paulus et al. (2017) trained their model on anonymized data, we also recover the anonymized entities and numbers of their outputs.

We show the article and 6 summaries (reference summary, 4 generated summaries and a random summary) to each human evaluator. The random summary is a reference summary randomly picked from other articles and is used as a trap. We show the instructions of three different aspects as: (1) Informativity: how well does the summary capture the important parts of the article? (2) Conciseness: is the summary clear enough to explain everything without being redundant? (3) Readability: how well-written (fluent and grammatical) the summary is? The user interface of our human evaluation is shown in the supplementary material.

We ask the human evaluator to evaluate each summary by scoring the three aspects with 1 to 5 score (higher the better). We reject all the evaluations that score the informativity of the random summary as 3, 4 and 5. By using this trap mechanism, we can ensure a much better quality of our human evaluation. For each example, we first ask 5 human evaluators to evaluate. However, for those articles that are too long, which are always skipped by the evaluators, it is hard to collect 5 reliable evaluations. Hence, we collect at least 3 evaluations for every example. For each summary, we average the scores over different human evaluators.

The results are shown in Table 3. The reference summaries get the best score on conciseness since the recent abstractive models tend to copy sentences from the input articles. However, our model learns well to select important information and form complete sentences so we even get slightly better scores on informativity and readability than the reference summaries. We show a typical example of our model comparing with other state-of-the-art methods in Fig. 6. More examples (5 using CNN/Daily Mail news articles and 3 using non-news articles as inputs) are provided in the supplementary material.

Conclusion

We propose a unified model combining the strength of extractive and abstractive summarization. Most importantly, a novel inconsistency loss function is introduced to penalize the inconsistency between two levels of attentions. The inconsistency loss enables extractive and abstractive summarization to be mutually beneficial. By end-to-end training of our model, we achieve the best ROUGE-recall and ROUGE while being the most informative and readable summarization on the CNN/Daily Mail dataset in a solid human evaluation.

Acknowledgments

We thank the support from Cheetah Mobile, National Taiwan University, and MOST 107-2634-F-007-007, 106-3114-E-007-004, 107-2633-E-002-001. We thank Yun-Zhu Song for assistance with useful survey and experiment on the task of abstractive summarization.

References