A Semantic Relevance Based Neural Network for Text Summarization and Text Simplification

Shuming Ma, Xu Sun

Introduction

Text summarization and text simplification is to make the text easier to read and understand, especially for poor readers, including children, non-native speakers, and the functionally illiterate. Text summarization is to simplify the texts at the document level. The source texts are often consist of many sentences or paragraphs, and the simplified texts are some brief sentences of the main ideas of the source texts. Text simplification is to simplify the texts at the sentence level. It aims to simplify the sentences to reduce the lexical and structure complexity. Unlike text summarization, it does not require the simplified sentences are shorter, but requires the words are simple to understand.

In some previous work, extractive summarization achieves satisfying performance by selecting a few sentences from source texts Radev et al. (2004); Cheng and Lapata (2016); Cao et al. (2015). By extracting the sentences, the generated texts are grammatical, and retain the same meaning with the source texts. However, it does not simplify the texts but only shorten the texts. Some previous related work regards text simplification as a combination of three operations: splitting, deletion and paraphrasing, which requires some rule-based models or heavy syntactic features Zhu, Bernhard, and Gurevych (2010); Woodsend and Lapata (2011); Filippova et al. (2015).

Most recent approaches use sequence-to-sequence model for text summarization Rush, Chopra, and Weston (2015); Hu, Chen, and Zhu (2015) and text simplification Nisioi et al. (2017); Cao et al. (2017); Zhang and Lapata (2017). Sequence-to-sequence model is a widely used end-to-end framework for text generation, such as machine translation. It compresses the source text information into dense vectors with the neural encoder, and the neural decoder generates the target text using the compressed vectors.

For both text summarization and text simplification, the simplified texts must have high semantic relevance to the source texts. However, current sequence-to-sequence models tend to produce grammatical and coherent simplified texts regardless of the semantic relevance to source texts. Table 1 shows that the summary generated by a LSTM sequence-to-sequence model (Seq2seq) is similar to the source text literally, but it has low semantic relevance.

In this work, our goal is to improve the semantic relevance between source texts and generated simplified texts for text summarization and text simplification. To achieve this goal, we propose a Semantic Relevance Based neural network model (SRB). In our model, we compress the source texts into dense vectors with the encoder, and decode the dense vector into simplified texts with the decoder. The encoder produces the representation of source texts, and the decoder produces the representation of the generated texts. A similarity evaluation component is introduced to measure the relevance of source texts and generated texts. During training, it maximizes the similarity score to encourage high semantic relevance between source texts and simplified texts. In order to better represent a long source text, we introduce a self-gated attention encoder to memory the input text. We conduct the experiments on three corpus, namely LCSTS, PWKP, and EW-SEW. Experiments show that our proposed model has better performance than the state-of-the-art systems on two benchmark corpus.

The contributions of this work are as follow:

We propose a Semantic Relevance Based neural network model (SRB) to improve the semantic relevance between source texts and generated simplified texts for text summarization and text simplification. A similarity evaluation component is introduced to measure the relevance of source texts and generated texts, and the similarity score is maximized to encourage high semantic relevance between source texts and simplified texts.

We introduce a self-gated encoder to better represent a long redundant text. We perform the experiments on three corpus, namely LCSTS, PWKP, and EW-SEW. Experiments show that our proposed model outperforms the state-of-the-art systems on two benchmark corpus.

Background: Sequence-to-sequence Model

Most recent models for text summarization and text simplification are based on the sequence-to-sequence model. The sequence-to-sequence model is able to compress source texts x={x1,x2,...,xN}x=\{x_{1},x_{2},...,x_{N}\} into a continuous vector representation with an encoder, and then generates the simplified text y={y1,y2,...,yM}y=\{y_{1},y_{2},...,y_{M}\} with a decoder. In the previous work Nisioi et al. (2017); Hu, Chen, and Zhu (2015), the encoder is a two layer Long Short-term Memory Network (LSTM) Hochreiter and Schmidhuber (1997), which maps source texts into the hidden vector {h1,h2,...,hN}\{h_{1},h_{2},...,h_{N}\}. The decoder is a uni-directional LSTM, producing the hidden output sts_{t}, which is the dense representation of the words at the ttht^{th} time step. Finally, the word generator computes the distribution of output words yty_{t} with the hidden state sts_{t} and the parameter matrix WW:

Attention mechanism is introduced to better capture context information of source texts Bahdanau, Cho, and Bengio (2014). Attention vector ctc_{t} is calculated by the weighted sum of encoder hidden states:

where g(st,hi)g(s_{t},h_{i}) is an attentive score between the decoder hidden state sts_{t} and the encoder hidden state hih_{i}. When predicting an output word, the decoder takes account of the attention vector, which contains the alignment information between source texts and simplified texts. With the attention mechanism, the word generator computes the distribution of output words yty_{t}:

Proposed Model

Our goal is to improve the semantic relevance between source texts and simplified texts, so our proposed model encourages high similarity between their representations. Figure 1 shows our proposed model. The model consists of three components: encoder, decoder and a similarity function. The encoder compresses source texts into semantic vectors, and the decoder generates summaries and produces semantic vectors of the generated summaries. Finally, the similarity function evaluates the relevance between the sematic vectors of source texts and generated summaries. Our training objective is to maximize the similarity score so that the generated summaries have high semantic relevance to source texts.

The goal of the complex text encoder is to provide a series of dense representation of source texts for the decoder and the semantic relevance component. In the previous work Nisioi et al. (2017), the complex text encoder is a two-layer uni-directional Long Short-term Memory Network (LSTM), which produces the dense representation {h1,h2,...,hN}\{h_{1},h_{2},...,h_{N}\} from the source text {x1,x2,...,xN}\{x_{1},x_{2},...,x_{N}\}.

However, in text summarization and text simplification, source texts are usually very long and noisy. Therefore, some encoding information in the beginning of the texts will vanish until the end of the texts, which leads to bad representations of the texts. Bi-directional LSTM is an alternative to deal with the problem, but it needs double time to encoder the source texts, and it does not represents the middle of the texts well when the texts are too long. To solve the problem, we propose a self-gated encoder to better represent a long text.

In text summarization and text simplification, some words or information in the source texts are unimportant, so they need to be simplified or discarded. Therefore, we introduce a self-gated encoder, which can reduce the unnecessary information and enhance the important information to represent a long text.

Self-gated encoder try to measure the importance of each word, and decide how much information is reserved as the representation of the texts. At each time step, every upcoming word xtx_{t} is fed into the LSTM cell, which outputs the dense vector hth_{t}:

where ff is the LSTM function, and hth_{t} is the output vector of the LSTM cell. A feed-forward neural network is used to measure the importance and decide how much information is reversed:

where gg is the feed-forward neural network function, and βt\beta_{t} measures the proportion of the reserved information. Finally, the reversed information is computed by multiplying βt\beta_{t}:

where h^t\hat{h}_{t} is the representation at the ttht_{th} time step, and e^t+1\hat{e}_{t+1} is the input embedding of xt+1x_{t+1} at the t+1tht+1_{th} time step.

2 Simplified Text Decoder

The goal of the simplified text decoder is to generate a series of simplified words from the dense representation of source texts. In our model, the dense representations of the source texts are fed into an attention layer to generate the context vector ctc_{t}:

where sts_{t} is the dense representation of generated simplified computed by a two-layer LSTM.

In this way, ctc_{t} and sts_{t} respectively represent the context information of source texts and the target texts at the ttht^{th} time step. To predict the ttht^{th} word, the decoder uses ctc_{t} and sts_{t} to generate the probability distribution of the candidate words:

where WW and WcW_{c} are the parameter matrix of the output layer. Finally, the word with the highest probability is predicted:

3 Semantic Relevance

Our goal is to compute the semantic relevance of source texts and generated texts given the source semantic vector VtV_{t} and the generated sementic vector VsV_{s}. Here, we use cosine similarity to measure the semantic relevance, which is represented with a dot product and magnitude:

Source texts and generated texts share the same language, so it is reasonable to assume that their semantic vectors are distributed in the same space. Cosine similarity is a good way to measure the distance between two vectors in the same space.

With the semantic relevance metric, the problem is how to get the semantic vector VsV_{s} and VtV_{t}. There are several methods to represent a text or a sentence, such as mean pooling of LSTM output or reserving the last state of LSTM. In our model, we select the last state of the encoder as the representation of source texts:

A natural idea to get the semantic vector of a summary is to feed it into the encoder as well. However, this method wastes much time because we encode the same sentence twice. Actually, the last output of the decoder s^M\hat{s}_{M} contains information of both source text and generated summaries. We simply compute the semantic vector of the summary by subtracting h^N\hat{h}_{N} from s^M\hat{s}_{M}:

Previous work has proved that it is effective to represent a span of words without encoding them once more Wang and Chang (2016).

4 Training

Given the model parameter θ\theta and input text xx, the model produces corresponding summary yy and semantic vector VsV_{s} and VtV_{t}. The objective is to minimize the loss function:

where p(y∣x;θ)p(y|x;\theta) is the conditional probability of summaries given source texts, and is computed by the encoder-decoder model. cos(Vs,Vt)cos(V_{s},V_{t}) is cosine similarity of semantic vectors VsV_{s} and VtV_{t}. This term tries to maximize the semantic relevance between source input and target output.

We use Adam optimization method to train the model, with the default hyper-parameters: the learning rate α=0.001\alpha=0.001, and β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, ϵ=1e−8\epsilon=1e-8.

Experiments

In this section, we present the evaluation of our model and show its performance on three popular corpus. Besides, we perform a case study to explain the semantic relevance between generated summary and source text.

We introduce a Chinese text summarization dataset and two popular text simplification datasets. The simplification datasets are both from the alignments between English Wikipedia websitehttp://en.wikipedia.org and Simple English Wikipedia websitehttp://simple.wikipedia.org. The Simple English Wikipedia is built for “the children and adults who are learning the English language”, and the articles are composed with “easy words and short sentences”. Therefore, Simple English Wikipedia is a natural public simplified text corpus. Most of the text simplification benchmark datasets are constructed from Simple English Wikipedia.

Large Scale Chinese Short Text Summarization Dataset (LCSTS). LCSTS is constructed by \namecitelcsts. The dataset consists of more than 2.4 million text-summary pairs, constructed from a famous Chinese social media website called Sina Weiboweibo.sina.com. It is split into three parts, with 2,400,591 pairs in PART I, 10,666 pairs in PART II and 1,106 pairs in PART III. All the text-summary pairs in PART II and PART III are manually annotated with relevant scores ranged from 1 to 5, and we only reserve pairs with scores no less than 3. Following the previous work, we use PART I as training set, PART II as development set, and PART III as test set.

Parallel Wikipedia Simplification Corpus (PWKP). PWKP Zhu, Bernhard, and Gurevych (2010) is a widely used benchmark for evaluating text simplification systems. It consists of aligned complex text from English WikiPedia (as of Aug. 22nd, 2009) and simple text from Simple Wikipedia (as of Aug. 17th, 2009). The dataset contains 108,016 sentence pairs, with 25.01 words on average per complex sentence and 20.87 words per simple sentence. Following the previous work Zhang and Lapata (2017), we remove the duplicate sentence pairs, and split the corpus with 89,042 pairs for training, 205 pairs for development and 100 pairs for test.

English Wikipedia and Simple English Wikipedia (EW-SEW). EW-SEW is a publicly available dataset provided by \nameciteHwangEA2015. To build the corpus, they first align the complex-simple sentence pairs, score the semantic similarity between the complex sentence and the simple sentence, and classify each sentence pair as a good, good partial, partial, or bad match. Following the previous work Nisioi et al. (2017), we discard the unclassified matches, and use the good matches and partial matches with a scaled threshold greater than 0.45. The corpus contains about 150K good matches and 130K good partial matches. We use this corpus as the training set, and the dataset provided by Xu et al. Xu et al. (2016) as the development set and the test set. The development set consists of 2,000 sentence pairs, and the test set contains 359 sentence pairs. Besides, each complex sentence is paired with 8 reference simplified sentences provided by Amazon Mechanical Turk workers.

2 Settings

We describe the experimental details of text summarization and text simplification respectively.

Text Summarization. To alleviate the risk of word segmentation mistakes Xu and Sun (2016); Sun, Wang, and Li (2012), we use Chinese character sequences as both source inputs and target outputs. We limit the model vocabulary size to 4000, which covers most of the common characters. Each character is represented by a random initialized word embedding. We tune our parameter on the development set. In our model, the embedding size is 400, the hidden state size of encoder-decoder is 500, and the size of gated attention network is 1000. We use Adam optimizer to learn the model parameters, and the batch size is set as 32. The parameter λ\lambda is 0.0001. Both the encoder and decoder are based on LSTM unit. Following the previous work Hu, Chen, and Zhu (2015), our evaluation metric is F-score of ROUGE: ROUGE-1, ROUGE-2 and ROUGE-L Lin and Hovy (2003).

Text Simplification. The text simplification datasets contain a lot of named entities, which makes the vocabulary too large. To reduce the vocabulary size, we follow the setting by \nameciteZhangEA2017. We recognize the named entities with the Stanford CoreNLP tagger Manning et al. (2014), and replace the named entities with the anonymous symbols NE@N, where NE∈\in{PER, LOC, ORG, MISC} where NN represents the NthN^{th} entity in the sentence. To limit the vocabulary size, we prune the vocabulary to top 50,000 most frequent words, and replace the rest words with the UNK symbols. At test time, we replace the UNK symbols with the highest probability score from the attention alignment matrix following Jean et al. Jean et al. (2015). We filter out sentence pairs whose lengths exceed 100 words in the training set. The encoder is implemented on LSTM, and the decoder is based on LSTM with Luong style attention Luong, Pham, and Manning (2015). We tune our hyper-parameter on the development set. The model has two LSTM layers. The hidden size of LSTM is 256, and the embedding size is 256. We use Adam optimizer Kingma and Ba (2014) to learn the parameters, and the batch size is set to be 64. We set the dropout rate Srivastava et al. (2014) to be 0.4. All of the gradients are clipped when the norm exceeds 5. The evaluation metric is BLEU score.

3 Baseline Systems

Seq2seq. We first compare our model with a basic sequence-to-sequence model Sutskever, Vinyals, and Le (2014). It is a widely used model to generate texts, so it is an important baseline.

Seq2seq-Attention. Seq2seq-Attention Bahdanau, Cho, and Bengio (2014) is a sequence-to-sequence framework with neural attention. Attention mechanism helps capture the context information of source texts. This model is a stronger baseline system.

4 Results

We compare our model with above baseline systems, including Seq2seq and Seq2seq-Attention. We refer to our proposed Semantic Relevance Based neural model as SRB. Table 2 shows the results of our models and baseline systems on LCSTS. As shown in Table 2, the models at the character level achieve better performance than the models at the word level. Therefore, we implement our model at the character level. For fair comparison, we also implement a Seq2seq-Attention model following the details in the previous work Hu, Chen, and Zhu (2015). Our implementation of Seq2seq-Attention has better score, mainly because we tune the hyper-parameters well on the development set. We can see SRB outperforms both Seq2seq and Seq2seq-Attention with the F-score of 33.3 ROUGE-1, 20.0 ROUGE-2 and 30.1 ROUGE-L.

Table 3 shows the results in two text simplification corpus. We compare our model with Seq2seq-Attention and Seq2seq-Attention-w2v. Seq2seq-Attention-w2v is a Seq2seq-Attention with pretrain word embeddings. We also implement a Seq2seq-Attention model, and carefully tune it on the development set. Our implementation get 48.26 BLEU score on PWKP, and 88.97 BLEU score on EW-SEW. Our SRB outperforms all of the baseline systems, with the BLEU score of 50.18 on PWKP and the BLEU score of 89.84 on EW-SEW.

Table 4 summarizes the results of our model and state-of-the-art systems. COPYNET has the highest scores, because it incorporates copying mechanism to deals with out-of-vocabulary word problem. In this paper, we do not implement this mechanism in our model. Our model can also be improved with these additional techniques, which, however, are not the focus of this paper.

We also compare SRB with other models for text simplification, which are not limit to neural models. Table 5 summarizes the results of SRB and the related systems. On PWKP dataset, we compare SRB with NTS, NTS-w2v, DRESS and DRESS-LS. We run the public release code of NTS and NTS-w2v provided by \nameciteNisioiEA2017, and get the BLEU score of 47.52 and 48.10 respectively. As for DRESS and DRESS-LS, we use the scores reported by \nameciteZhangEA2017. The goal of DRESS is not to generate the outputs closer to the references, so BLEU of DRESS and DRESS-LS are relatively lower than NTS and NTS-w2v. SRB achieves a BLEU score of 50.18, outperforming all of the previous systems. On EW-SEW dataset, we compare WEAN-dot with PBMT-R, SBMT-SARI, and the neural models described above. We do not find any public release code of PBMT-R and SBMT-SARI. Fortunately, \nameciteXuEA2016 provides the predictions of PBMT-R and SBMT-SARI on EW-SEW test set, so that we can compare our model with these systems. It shows that the neural models have better performance in BLEU, and WEAN-dot achieves the best BLEU score with 89.84.

5 Case Study

Table 6 is an example to show the semantic relevance between the source text and the summary. It shows that the main idea of the source text is about the reason why Shanghai has few giant company. RNN context produces “Shanghai’s giant companies” which is literally similar to the source text, while SRB generates “Shanghai has few giant companies”, which is closer to the main idea in semantics. It concludes that SRB produces summaries with higher semantic similarity to texts.

Table 7 shows an examples of different text simplification system outputs on EW-SEW. NTS-w2v omits so many words that it lacks a lot of information. PBMT-R generates some irrelevant words, like ’siemens-martin’, ’-rrb-’, and ’-shurba’, which hurts the fluency and adequacy of the generated sentence. SBMT-SARI is able to generate a fluent sentence, but the meaning is different from the source text, and even more difficult to understand. Compared with the statistic model, SRB generates a more fluent sentence. Besides, SRB improves the semantic revelance between the source texts and the generated texts, so the generated sentence is semantically correct, and very close to the original meaning.

Related Work

Abstractive text summarization has achieved successful performance thanks to the sequence-to-sequence model Sutskever, Vinyals, and Le (2014) and attention mechanism Bahdanau, Cho, and Bengio (2014). \nameciteabs first used an attention-based encoder to compress texts and a neural network language decoder to generate summaries. Following this work, recurrent encoder was introduced to text summarization, and gained better performance Lopyrev (2015); Chopra, Auli, and Rush (2016). Towards Chinese texts, \namecitelcsts built a large corpus of Chinese short text summarization. To deal with unknown word problem, \nameciteibmsummarization proposed a generator-pointer model so that the decoder is able to generate words in source texts. \namecitecopynet also solved this issue by incorporating copying mechanism. Besides, \nameciteminimum proposes a minimum risk training method which optimizes the parameters with the target of rouge scores.

ZhuEA2010 constructs a wikipedia dataset, and proposes a tree-based simplification model, which is the first statistical simplification model covering splitting, dropping, reordering and substitution integrally. \nameciteWoodsend2011 introduces a data-driven model based on quasi-synchronous grammar, which captures structural mismatches and complex rewrite operations. \nameciteWubbenEA2012 presents a method for text simplification using phrase based machine translation with re-ranking the outputs. \nameciteKauchak2013 proposes a text simplification corpus, and evaluates language modeling for text simplification on the proposed corpus.

NarayanEA2014 propose a hybrid approach to sentence simplification which combines deep semantics and monolingual machine translation. \nameciteHwangEA2015 introduces a parallel simplification corpus by evaluating the similarity between the source text and the simplified text based on WordNet. \nameciteLigthLS propose an unsupervised approach to lexical simplification that makes use of word vectors and require only regular corpora. \nameciteXuEA2016 design automatic metrics for text simplification, and they introduce a statistic machine translation, and tune with the proposed automatic metrics.

Recently, most works focus on the neural sequence-to-sequence model. \nameciteNisioiEA2017 present a sequence-to-sequence model, and re-ranks the predictions with BLEU and SARI. \nameciteZhangEA2017 propose a deep reinforcement learning model to improve the simplicity, fluency and adequacy of the simplified texts. \nameciteCaoEA2017 introduce a novel sequence-to-sequence model to join copying and restricted generation for text simplification.

Our work is also related to the encoder-decoder framework Cho et al. (2014) and the attention mechanism Bahdanau, Cho, and Bengio (2014). Encoder-decoder framework, like sequence-to-sequence model, has achieved success in machine translation Sutskever, Vinyals, and Le (2014); Jean et al. (2015); Luong, Pham, and Manning (2015), text summarization Rush, Chopra, and Weston (2015); Chopra, Auli, and Rush (2016); Nallapati et al. (2016); Cao et al. (2016), and other natural language processing tasks. Neural attention model is first proposed by \nameciteattention. There are many other methods to improve neural attention model Jean et al. (2015); Luong, Pham, and Manning (2015).

Conclusion

In this work, our goal is to improve the semantic relevance between source texts and generated simplified texts for text summarization and text simplification. To achieve this goal, we propose a Semantic Relevance Based neural network model (SRB). A similarity evaluation component is introduced to measure the relevance of source texts and generated texts. During training, it maximizes the similarity score to encourage high semantic relevance between source texts and simplified texts. In order to better represent a long source text, we introduce a self-gated attention encoder to memory the input text. We conduct the experiments on three corpus, namely LCSTS, PWKP, and EW-SEW. Experiments show that our proposed model has better performance than the state-of-the-art systems on the benchmark corpus.

Acknowledgements

This work was supported in part by National Natural Science Foundation of China (No. 61673028), and an Okawa Research Grant (2016). Xu Sun is the corresponding author of this paper. This work is a substantial extension of the conference version presented at ACL 2017 Ma et al. (2017).

References