Deep Recurrent Generative Decoder for Abstractive Text Summarization
Piji Li, Wai Lam, Lidong Bing, Zihao Wang
Introduction
Automatic summarization is the process of automatically generating a summary that retains the most important content of the original text document Edmundson (1969); Luhn (1958); Nenkova and McKeown (2012). Different from the common extraction-based and compression-based methods, abstraction-based methods aim at constructing new sentences as summaries, thus they require a deeper understanding of the text and the capability of generating new sentences, which provide an obvious advantage in improving the focus of a summary, reducing the redundancy, and keeping a good compression rate Bing et al. (2015); Rush et al. (2015); Nallapati et al. (2016).
Some previous research works show that human-written summaries are more abstractive Jing and McKeown (2000). Moreover, our investigation reveals that people may naturally follow some inherent structures when they write the abstractive summaries. To illustrate this observation, we show some examples in Figure 1, which are some top story summaries or headlines from the channel “Technology” of CNN. After analyzing the summaries carefully, we can find some common structures from them, such as “What”, “What-Happened” , “Who Action What”, etc. For example, the summary “Apple sues Qualcomm for nearly 20 million] for misleading drivers”, and “[Bipartisan bill] aims to [reform] [H-1B visa system]” also follow the structure of “Who Action What”. The summary “The emergence of the ‘cyber cold war”’ matches with the structure of “What”, and the summary “St. Louis’ public library computers hacked” follows the structure of “What-Happened”.
Intuitively, if we can incorporate the latent structure information of summaries into the abstractive summarization model, it will improve the quality of the generated summaries. However, very few existing works specifically consider the latent structure information of summaries in their summarization models. Although a very popular neural network based sequence-to-sequence (seq2seq) framework has been proposed to tackle the abstractive summarization problem Lopyrev (2015); Rush et al. (2015); Nallapati et al. (2016), the calculation of the internal decoding states is entirely deterministic. The deterministic transformations in these discriminative models lead to limitations on the representation ability of the latent structure information. Miao and Blunsom (2016) extended the seq2seq framework and proposed a generative model to capture the latent summary information, but they did not consider the recurrent dependencies in their generative model leading to limited representation ability.
To tackle the above mentioned problems, we design a new framework based on sequence-to-sequence oriented encoder-decoder model equipped with a latent structure modeling component. We employ Variational Auto-Encoders (VAEs) Kingma and Welling (2013); Rezende et al. (2014) as the base model for our generative framework which can handle the inference problem associated with complex generative modeling. However, the standard framework of VAEs is not designed for sequence modeling related tasks. Inspired by Chung et al. (2015), we add historical dependencies on the latent variables of VAEs and propose a deep recurrent generative decoder (DRGD) for latent structure modeling. Then the standard discriminative deterministic decoder and the recurrent generative decoder are integrated into a unified decoding framework. The target summaries will be decoded based on both the discriminative deterministic variables and the generative latent structural information. All the neural parameters are learned by back-propagation in an end-to-end training paradigm.
The main contributions of our framework are summarized as follows: (1) We propose a sequence-to-sequence oriented encoder-decoder model equipped with a deep recurrent generative decoder (DRGD) to model and learn the latent structure information implied in the target summaries of the training data. Neural variational inference is employed to address the intractable posterior inference for the recurrent latent variables. (2) Both the generative latent structural information and the discriminative deterministic variables are jointly considered in the generation process of the abstractive summaries. (3) Experimental results on some benchmark datasets in different languages show that our framework achieves better performance than the state-of-the-art models.
Related Works
Automatic summarization is the process of automatically generating a summary that retains the most important content of the original text document Nenkova and McKeown (2012). Traditionally, the summarization methods can be classified into three categories: extraction-based methods Erkan and Radev (2004); Goldstein et al. (2000); Wan et al. (2007); Min et al. (2012); Nallapati et al. (2017); Cheng and Lapata (2016); Cao et al. (2016); Song et al. (2017), compression-based methods Li et al. (2013); Wang et al. (2013); Li et al. (2015, 2017), and abstraction-based methods. In fact, previous investigations show that human-written summaries are more abstractive Barzilay and McKeown (2005); Bing et al. (2015). Abstraction-based approaches can generate new sentences based on the facts from different source sentences. Barzilay and McKeown (2005) employed sentence fusion to generate a new sentence. Bing et al. (2015) proposed a more fine-grained fusion framework, where new sentences are generated by selecting and merging salient phrases. These methods can be regarded as a kind of indirect abstractive summarization, and complicated constraints are used to guarantee the linguistic quality.
Recently, some researchers employ neural network based framework to tackle the abstractive summarization problem. Rush et al. (2015) proposed a neural network based model with local attention modeling, which is trained on the Gigaword corpus, but combined with an additional log-linear extractive summarization model with handcrafted features. Gu et al. (2016) integrated a copying mechanism into a seq2seq framework to improve the quality of the generated summaries. Chen et al. (2016) proposed a new attention mechanism that not only considers the important source segments, but also distracts them in the decoding step in order to better grasp the overall meaning of input documents. Nallapati et al. (2016) utilized a trick to control the vocabulary size to improve the training efficiency. The calculations in these methods are all deterministic and the representation ability is limited. Miao and Blunsom (2016) extended the seq2seq framework and proposed a generative model to capture the latent summary information, but they do not consider the recurrent dependencies in their generative model leading to limited representation ability.
Some research works employ topic models to capture the latent information from source documents or sentences. Wang et al. (2009) proposed a new Bayesian sentence-based topic model by making use of both the term-document and term-sentence associations to improve the performance of sentence selection. Celikyilmaz and Hakkani-Tur (2010) estimated scores for sentences based on their latent characteristics using a hierarchical topic model, and trained a regression model to extract sentences. However, they only use the latent topic information to conduct the sentence salience estimation for extractive summarization. In contrast, our purpose is to model and learn the latent structure information from the target summaries and use it to enhance the performance of abstractive summarization.
Framework Description
As shown in Figure 2, the basic framework of our approach is a neural network based encoder-decoder framework for sequence-to-sequence learning. The input is a variable-length sequence representing the source text. The word embedding is initialized randomly and learned during the optimization process. The output is also a sequence , which represents the generated abstractive summaries. Gated Recurrent Unit (GRU) Cho et al. (2014) is employed as the basic sequence modeling component for the encoder and the decoder. For latent structure modeling, we add historical dependencies on the latent variables of Variational Auto-Encoders (VAEs) and propose a deep recurrent generative decoder (DRGD) to distill the complex latent structures implied in the target summaries of the training data. Finally, the abstractive summaries will be decoded out based on both the discriminative deterministic variables and the generative latent structural information .
2 Recurrent Generative Decoder
where is a recurrent neural network such as vanilla RNN, Long Short-Term Memory (LSTM) Hochreiter and Schmidhuber (1997), and Gated Recurrent Unit (GRU) Cho et al. (2014). No matter which one we use for , the common transformation operation is as follows:
For the inference stage, the variational-encoder can map the observed variable and the previous latent structure information to the posterior probability distribution of the latent structure variable . It is obvious that this is a recurrent inference process in which contains the historical dynamic latent structure information. Compared with the variational inference process of the typical VAEs model, the recurrent framework can extract more complex and effective latent structure features implied in the sequence data.
For the generation process, based on the latent structure variable , the target word at the time step is drawn from a conditional probability distribution . The target is to maximize the probability of each generated summary based on the generation process according to:
For the purpose of solving the intractable integral of the marginal likelihood as shown in Equation 3, a recognition model is introduced as an approximation to the intractable true posterior . The recognition model parameters and the generative model parameters can be learned jointly. The aim is to reduce the Kulllback-Leibler divergence (KL) between and :
Let represent the last two terms from the right part of Equation 4:
Since the first KL-divergence term of Equation 4 is non-negative, we have meaning that is a lower bound (the objective to be maximized) on the marginal likelihood. In order to differentiate and optimize the lower bound , following the core idea of VAEs, we use a neural network framework for the probabilistic encoder for better approximation.
3 Abstractive Summary Generation
We also design a neural network based framework to conduct the variational inference and generation for the recurrent generative decoder component similar to some design in previous works Kingma and Welling (2013); Rezende et al. (2014); Gregor et al. (2015). The encoder component and the decoder component are integrated into a unified abstractive summarization framework. Considering that GRU has comparable performance but with less parameters and more efficient computation, we employ GRU as the basic recurrent model which updates the variables according to the following operations:
where is the reset gate, is the update gate. denotes the element-wise multiplication. is the hyperbolic tangent activation function.
As shown in the left block of Figure 2, the encoder is designed based on bidirectional recurrent neural networks. Let be the word embedding vector of the -th word in the source sequence. GRU maps and the previous hidden state to the current hidden state in feed-forward direction and back-forward direction respectively:
The discriminative deterministic decoding is an improved attention modeling based recurrent sequence decoder. The first hidden state is initialized using the average of all the source input states: , where is the source input hidden state. is the input sequence length. The deterministic decoder hidden state is calculated using two layers of GRUs. On the first layer, the hidden state is calculated only using the current input word embedding and the previous hidden state :
The final deterministic hidden state is the output of the second decoder GRU layer, jointly considering the word , the previous hidden state , and the attention context :
For the component of recurrent generative model, inspired by some ideas in previous works Kingma and Welling (2013); Rezende et al. (2014); Gregor et al. (2015), we assume that both the prior and posterior of the latent variables are Gaussian, i.e., and , where and denote the variational mean and standard deviation respectively, which can be calculated via a multilayer perceptron. Precisely, given the word embedding , the previous latent structure variable , and the previous deterministic hidden state , we first project it to a new hidden space:
To generate summaries precisely, we first integrate the recurrent generative decoding component with the discriminative deterministic decoding component, and map the latent structure variable and the deterministic decoding hidden state to a new hidden variable:
Given the combined decoding state at the time , the probability of generating any target word is given as follows:
4 Learning
Although the proposed model contains a recurrent generative decoder, the whole framework is fully differentiable. As shown in Section 3.3, both the recurrent deterministic decoder and the recurrent generative decoder are designed based on neural networks. Therefore, all the parameters in our model can be optimized in an end-to-end paradigm using back-propagation. We use and to denote the training source and target sequence. Generally, the objective of our framework consists of two terms. One term is the negative log-likelihood of the generated summaries, and the other one is the variational lower bound mentioned in Equation 5. Since the variational lower bound also contains a likelihood term, we can merge it with the likelihood term of summaries. The final objective function, which needs to be minimized, is formulated as follows:
Experimental Setup
We train and evaluate our framework on three popular datasets. Gigawords is an English sentence summarization dataset prepared based on Annotated Gigawordshttps://catalog.ldc.upenn.edu/ldc2012t21 by extracting the first sentence from articles with the headline to form a source-summary pair. We directly download the prepared dataset used in Rush et al. (2015). It roughly contains 3.8M training pairs, 190K validation pairs, and 2,000 test pairs. DUC-2004http://duc.nist.gov/duc2004 is another English dataset only used for testing in our experiments. It contains 500 documents. Each document contains 4 model summaries written by experts. The length of the summary is limited to 75 bytes. LCSTS is a large-scale Chinese short text summarization dataset, consisting of pairs of (short text, summary) collected from Sina Weibohttp://www.weibo.com Hu et al. (2015). We take Part-I as the training set, Part-II as the development set, and Part-III as the test set. There is a score in range labeled by human to indicate how relevant an article and its summary is. We only reserve those pairs with scores no less than 3. The size of the three sets are 2.4M, 8.7k, and 725 respectively. In our experiments, we only take Chinese character sequence as input, without performing word segmentation.
2 Evaluation Metrics
We use ROUGE score Lin (2004) as our evaluation metric with standard options. The basic idea of ROUGE is to count the number of overlapping units between generated summaries and the reference summaries, such as overlapped n-grams, word sequences, and word pairs. F-measures of ROUGE-1 (R-1), ROUGE-2 (R-2), ROUGE-L (R-L) and ROUGE-SU4 (R-SU4) are reported.
3 Comparative Methods
We compare our model with some baselines and state-of-the-art methods. Because the datasets are quite standard, so we just extract the results from their papers. Therefore the baseline methods on different datasets may be slightly different.
TOPIARY Zajic et al. (2004) is the best on DUC2004 Task-1 for compressive text summarization. It combines a system using linguistic based transformations and an unsupervised topic detection algorithm for compressive text summarization.
MOSES+ Rush et al. (2015) uses a phrase-based statistical machine translation system trained on Gigaword to produce summaries. It also augments the phrase table with “deletion” rulesto improve the baseline performance, and MERT is also used to improve the quality of generated summaries.
ABS and ABS+ Rush et al. (2015) are both the neural network based models with local attention modeling for abstractive sentence summarization. ABS+ is trained on the Gigaword corpus, but combined with an additional log-linear extractive summarization model with handcrafted features.
RNN and RNN-context Hu et al. (2015) are two seq2seq architectures. RNN-context integrates attention mechanism to model the context.
CopyNet Gu et al. (2016) integrates a copying mechanism into the sequence-to-sequence framework.
RNN-distract Chen et al. (2016) uses a new attention mechanism by distracting the historical attention in the decoding steps.
RAS-LSTM and RAS-Elman Chopra et al. (2016) both consider words and word positions as input and use convolutional encoders to handle the source information. For the attention based sequence decoding process, RAS-Elman selects Elman RNN Elman (1990) as decoder, and RAS-LSTM selects Long Short-Term Memory architecture Hochreiter and Schmidhuber (1997).
LenEmb Kikuchi et al. (2016) uses a mechanism to control the summary length by considering the length embedding vector as the input.
ASC+FSC1 Miao and Blunsom (2016) uses a generative model with attention mechanism to conduct the sentence compression problem. The model first draws a latent summary sentence from a background language model, and then subsequently draws the observed sentence conditioned on this latent summary.
lvt2k-1sent and lvt5k-1sent Nallapati et al. (2016) utilize a trick to control the vocabulary size to improve the training efficiency.
4 Experimental Settings
For the experiments on the English dataset Gigawords, we set the dimension of word embeddings to 300, and the dimension of hidden states and latent variables to 500. The maximum length of documents and summaries is 100 and 50 respectively. The batch size of mini-batch training is 256. For DUC-2004, the maximum length of summaries is 75 bytes. For the dataset of LCSTS, the dimension of word embeddings is 350. We also set the dimension of hidden states and latent variables to 500. The maximum length of documents and summaries is 120 and 25 respectively, and the batch size is also 256. The beam size of the decoder was set to be 10. Adadelta Schmidhuber (2015) with hyperparameter and is used for gradient based optimization. Our neural network based framework is implemented using Theano Theano Development Team (2016).
Results and Discussions
We first depict the performance of our model DRGD by comparing to the standard decoders (StanD) of our own implementation. The comparison results on the validation datasets of Gigawords and LCSTS are shown in Table 1. From the results we can see that our proposed generative decoders DRGD can obtain obvious improvements on abstractive summarization than the standard decoders. Actually, the performance of the standard decoders is similar with those mentioned popular baseline methods.
The results on the English datasets of Gigawords and DUC-2004 are shown in Table 2 and Table 3 respectively. Our model DRGD achieves the best summarization performance on all the ROUGE metrics. Although ASC+FSC1 also uses a generative method to model the latent summary variables, the representation ability is limited and it cannot bring in noticeable improvements. It is worth noting that the methods lvt2k-1sent and lvt5k-1sent Nallapati et al. (2016) utilize linguistic features such as parts-of-speech tags, named-entity tags, and TF and IDF statistics of the words as part of the document representation. In fact, extracting all such features is a time consuming work, especially on large-scale datasets such as Gigawords. lvt2k and lvt5k are not end-to-end style models and are more complicated than our model in practical applications.
The results on the Chinese dataset LCSTS are shown in Table 4. Our model DRGD also achieves the best performance. Although CopyNet employs a copying mechanism to improve the summary quality and RNN-distract considers attention information diversity in their decoders, our model is still better than those two methods demonstrating that the latent structure information learned from target summaries indeed plays a role in abstractive summarization. We also believe that integrating the copying mechanism and coverage diversity in our framework will further improve the summarization performance.
2 Summary Case Analysis
In order to analyze the reasons of improving the performance, we compare the generated summaries by DRGD and the standard decoders StanD used in some other works such as Chopra et al. (2016). The source texts, golden summaries, and the generated summaries are shown in Table 5. From the cases we can observe that DRGD can indeed capture some latent structures which are consistent with the golden summaries. For example, our result for S(1) “Wuhan wins men’s soccer title at Chinese city games” matches the “Who Action What” structure. However, the standard decoder StanD ignores the latent structures and generates some loose sentences, such as the results for S(1) “Results of men’s volleyball at Chinese city games” does not catch the main points. The reason is that the recurrent variational auto-encoders used in our framework have better representation ability and can capture more effective and complicated latent structures from the sequence data. Therefore, the summaries generated by DRGD have consistent latent structures with the ground truth, leading to a better ROUGE evaluation.
Conclusions
We propose a deep recurrent generative decoder (DRGD) to improve the abstractive summarization performance. The model is a sequence-to-sequence oriented encoder-decoder framework equipped with a latent structure modeling component. Abstractive summaries are generated based on both the latent variables and the deterministic states. Extensive experiments on benchmark datasets show that DRGD achieves improvements over the state-of-the-art methods.