Systematically Exploring Redundancy Reduction in Summarizing Long Documents

Wen Xiao, Giuseppe Carenini

Introduction

Summarization is the task of shortening a given document(s) while maintaining the most important information. In general, a good summarizer should generate a summary that is syntactically accurate, semantically correct, coherent, and non-redundant Saggion and Poibeau (2013). While extractive methods tend to have better performance on the first two aspects, they are typically less coherent and more redundant than abstractive ones, where new sentences are often generated by sentence fusion and compression, which helps detecting and removing redundancy Lebanoff et al. (2019). Although eliminating redundancy has been initially and more intensely studied in the field of multi-document summarization Lloret and Sanz (2013), because important sentences selected from multiple documents (about the same topic) are more likely to be redundant than sentences from the same document, generating a non-redundant summary should still be one of the goals for single document summarization Lin et al. (2009).

Generally speaking, there is a trade-off between importance and diversity (non-redundancy) Jung et al. (2019), which is reflected in the two phases, sentence scoring and sentence selection Zhou et al. (2018) in which extractive summarization task can be naturally decomposed. The former typically scores sentences based on importance, while the latter selects sentences based on their scores, but also possibly taking other factors (including redundancy) into account.

Traditionally, in non-neural approaches the trade-off between importance and redundancy has been carefully considered, with sentence selection picking sentences by optimizing an objective function that balances the two aspects Carbonell and Goldstein (1998); Ren et al. (2016). In contrast, more recent works on neural extractive summarization models has so far over-emphasized sentence importance and the corresponding scoring phase, while paying little attention to how to reduce redundancy in the selection phase, where they simply apply a greedy algorithm to select sentences (e.g.,Cheng and Lapata (2016); Xiao and Carenini (2019)). Notice that this is especially problematic for long documents, where redundancy tends to be a more serious problem, as we have observed in key datasets. Improving redundancy reduction in neural extractive summarization for long documents is a major goal of this paper.

Indeed, some recently proposed neural methods aim to reduce redundancy, but they either do that implicitly or inflexibly and only focusing on short documents (e.g., news). For instance, some models learn to reduce redundancy when predicting the scores Nallapati et al. (2016a), or jointly learn to score and select sentences Zhou et al. (2018) in an implicit way. However, whether these strategies actually help reducing redundancy is still an open empirical question. The only neural attempt of explicitly reduce redundancy in the sentence selection phase is the Trigram Blocking technique, used in recent extractive summarization models on news datasets (e.g., Liu and Lapata (2019)). However, the effectiveness of such strategy on the summarization of long documents has not been tested. Finally, a very recent work by Bi et al. (2020) attempts to reduce redundancy in more sophisticated ways, but still focusing on news. Furthermore, since it relies on BERT, such model is unsuitable to deal with long documents (with over 3,000 words).

To address this rather confusing situation, characterized by unclear connections between all the proposed neural models, by their limited focus on short documents, and by spotty evaluations, in this paper we systematically organize existing redundancy reduction methods into three categories, and compare them with respect to the informativeness and redundancy of the generated summary for long documents. In particular, to perform a fair comparison we re-implement all methods by modifying a common basic model Xiao and Carenini (2019), which is a top performer on long documents without considering redundancy. Additionally, we propose three new methods that we argue will reduce redundancy more explicitly and flexibly in the sentence scoring and sentence selection phase by deploying more suitable decoders, loss functions and/or sentence selection algorithms, again building for a fair comparison on the common basic model Xiao and Carenini (2019).

To summarize, our main contributions in this paper are: we first examine popular datasets, and show that redundancy is a more serious problem when summarizing long documents (e.g., scientific papers) than short ones (e.g. news). Secondly, we not only reorganize and re-implement existing neural methods for redundancy reduction, but we also propose three new general and flexible methods. Finally, in a series of experiments, we compare existing and proposed methods on long documents (i.e., the Pubmed and arXiv datasets), with respect to ROUGE scores Lin (2004) and redundancy scores Peyrard et al. (2017); Feigenblat et al. (2017).

As a preview, empirical results reveal that the proposed methods achieve state-of-the-art performance on ROUGE scores, on the two scientific paper datasets, while also reducing the redundancy significantly.

Related Work

In traditional extractive summarization, the process is treated as a discrete optimization problem balancing between importance scores and redundancy scores, with techniques like Maximal Marginal Relevance(MMR)Carbonell and Goldstein (1998), redundancy-aware feature-based sentence classifiers Ren et al. (2016) and graph-based submodular selection Lin et al. (2009).

In recent years, researchers have explored neural extractive summarization solutions, which score sentences by training the neural models on a large corpus, and simply apply a greedy algorithm for sentence selection Cheng and Lapata (2016); Nallapati et al. (2016a). Although a model with a sequence decoder might plausibly encode redundancy information implicitly, Kedzie et al. (2018) empirically show that this is not the case, since non auto-regressive models (the ones scoring each sentence independently), perform on par with models with a sequence decoder. In one of our new methods, to effectively capture redundancy information, we specify a new loss that explicitly consider redundancy when training the neural model.

Beyond a greedy algorithm, the Trigram Blocking is frequently used to explicitly reduce redundancy in the sentence selection phase Liu and Lapata (2019). In essence, a new sentence is not added to the summary if it shares a 3-gram with the previously added one. Paulus et al. (2017) first adopt the strategy for abstractive summarization, which forces the model not to produce the same trigram twice in the generated summaries, as a simplified version of MMR Carbonell and Goldstein (1998). Arguably, this method is too crude for documents with relatively long sentences or specific concentrations (e.g. scientific papers), where some technical terms, possibly longer than 2-grams, are repeated frequently in the ’important sentences’ (even in the reference summaries). To address this limitation, we propose a neural version of MMR to deal with redundancy within the sentence selection phase in a more flexible way, that can be tuned to balance importance and non-redundancy as needed.

The idea of MMR has also inspired Zhou et al. (2018), who propose a model jointly learning to score and select the sentences. Yet, this work not only focuses on summarizing short documents (i.e., news), but also uses MMR implicitly, and arguably sub-optimally, by learning a score that only indirectly captures the trade-off between relevance and redundancy. To improve on this approach, in this paper we propose a third new method, in which importance and redundancy are explicitly weighted, while still making the sentence scoring and selection benefit from each other by fine tuning the trained neural model through a Reinforcement Learning (RL) mechanism.

Finally, Bi et al. (2020) is the most recent (still unpublished) work on reducing redundancy in neural single document summarization. However, their goal is very different form ours, since they focus on relatively short documents in the news domain.

Measuring Redundancy: metrics and comparing long vs. short documents

We use the following two relatively new metrics to measure redundancy in the source documents and in the generated summaries.

Unique n-gram ratioIn this paper, all the unique n-gram ratios are shown in percentage.: proposed in Peyrard et al. (2017), it measures n-grams uniqueness; the lower it is, the more redundant the document is.

Normalized Inverse of Diversity (NID): captures redundancy, as the inverse of a diversity metric with length normalization. Diversity is defined as the entropy of unigrams in the document Feigenblat et al. (2017). Since longer documents are more likely to have a higher entropy, we normalize the diversity with the maximum possible entropy for the document log(∣D∣)log(|D|). Thus, we have:

Note that higher NID indicates more redundancy.

When we compare the redundancy of long vs. short documents with respect to these two metrics on four popular datasets for summarization (CNNDM Nallapati et al. (2016b), Xsum Narayan et al. (2018), Pubmed and arXiv Cohan et al. (2018)), we observe that long documents are substantially more redundant than short ones (as it was already pointed out in the past Stewart and Carbonell (1998)). Table 1 shows the basic statistics of each dataset, along with the average NID scores, while Figure 1 shows the average Unique n-gram Ratio for the same datasets. These observations provide further evidence that redundancy is a more serious problem in long documents. In addition, notice that the sentences in the scientific paper datasets are much longer than in the news datasets, which plausibly makes it even harder to balance between importance and non-redundancy.

Redundancy Reduction Methods

We systematically organize neural redundancy reduction methods into three categories, and compare prototypical methods from each category.

The decoder is designed to implicitly take redundancy into account.

In the sentence scoring phase, explicitly learn to reduce the redundancy.

In the sentence selection phase, select sentences with less redundancy.

In this section, we describe different methods from each category. To compare them in a fair way, we build all of them on a basic ExtSum-LG model (see §4.1), by modifying the decoder and the loss function in the sentence selection phase or the sentence selection algorithm. In Table 2, we summarize the architecture (Encoder, Decoder, Loss Function and sentence selection algorithm) of all the methods we compare.

We consider two baseline models. One is an influential unsupervised method explicitly balancing importance and redundancy (Naive MMR). The other is our basic neural supervised model not dealing with redundancy at all (ExtSum-LG), to which we add different redundancy reduction mechanisms.

MMR Carbonell and Goldstein (1998) is a traditional extractive summarization method, which re-ranks the candidate sentences with a balance between query-relevance(importance) and information novelty(non-redundancy). Given a document DD, at each step, MMR selects one sentence from the candidate set D∖S^D\setminus\hat{S} that is relevant with the query QQ, while containing little redundancy with the current summary S^\hat{S}. Note that if there is no specific query, then the query is the representation of the whole document. The method can be formally specified as:

where Sim1(si,Q)Sim_{1}(s_{i},Q) measures the similarity between the candidate sentence sis_{i} and the query, indicating the importance of sis_{i}, while max⁡sj∈S^Sim2(si,sj)\max_{s_{j}\in\hat{S}}Sim_{2}(s_{i},s_{j}) measures the similarity between the candidate sentence sis_{i} and the current summary S^\hat{S}, representing the redundancy, and λ\lambda is the balancing factor. In this work, all the SimSim are computed as the cosine similarity between the embeddings of the sentences.

ExtSum-LG

For the basic model, we use the current state-of-the-art model Xiao and Carenini (2019) on the summarization of long documents. It is a novel extractive summarization model incorporating local context and global context in the encoder, with an MLP layer as decoder and cross-entropy as the loss function. For the sentence selection phase, it greedily picks the sentences according to the score predicted by the neural model. In this method, redundancy is not considered, so it is a good testbed for adding and comparing redundancy reduction methods.

Specifically, for a document D={s1,s2,...,sn}D=\{s_{1},s_{2},...,s_{n}\}, the output of the encoder is hih_{i} for each sentence sis_{i}, and the decoder gives output P(yi)P(y_{i}) as the confidence score on the importance of sentence sis_{i}. Finally, the model is trained on the Cross Entropy Loss :

2 Implicitly Reduce Redundancy in the neural model (Category A, Table 2)

In this section, we describe two decoders from previous work, in which the redundancy of the summary is considered implicitly.

SummaRuNNer Decoder: Nallapati et al. (2016a) introduce a decoder that computes a sentence score based on its salience, novelty(non-redundancy) and position to decide whether it should be included in the summary. Formally:

where hih_{i} is the hidden state of sentence ii from the encoder, dd is the document representation , summisumm_{i} is the summary representation, updated after each decoding step , and piap_{i}^{a}, pirp_{i}^{r} are absolute and relative position embeddings, respectively. Once P(yi)P(y_{i}) is obtained for each sentence ii, a greedy algorithm selects the sentences to form the final summary. Notice that although SummaRuNNer does contain a component assessing novelty, it would be inappropriate to view this model as explicitly dealing with redunadany because the novelty component is not directly supervised.

NeuSum Decoder: One of the main drawback of SummaRuNNer decoder is that it always score the sentences in order, i.e., the former sentences are not influenced by the latter ones. In addition, it only considers redundancy in the sentence scoring phase, while simply using a greedy algorithm to select sentences according to the resulting scores. To address these problems, Zhou et al. (2018) propose a new decoder to identify the relative gain of sentences, jointly learning to score and select sentences. In such decoder, instead of feeding the sentences and getting the scores in order, they use a mechanism similar to the pointer network Vinyals et al. (2015) to predict the scores of all the sentences at each step, select the sentence with the highest score, and feed it to the next step of sentence selection. As for the loss function, they use the KL divergence between the predicted score distribution and the relative ROUGE F1 gain at each step. To be specific, the loss computed at step tt is:

3 Explicitly Reduce Redundancy in Sentence Scoring (Category B, Table 2)

We propose a new method to explicitly learn to reduce redundancy when scoring the sentences.

RdLoss: Although Zhou et al. (2018) jointly train the decoder to score and select sentences, it still learns to reduce redundancy implicitly, and the method does not allow controlling the degree of redundancy. To address this limitation, we propose a rather simple method to explicitly force the model to reduce redundancy in the sentence scoring phase by adding a redundancy loss term to the original loss function, motivated by the success of a similar strategy of adding a bias loss term in the gender debiasing task Qian et al. (2019). Our new loss term LrdL_{rd} is naturally defined as the expected redundancy contained in the resulting summary, as shown below:

where P(yi),P(yj)P(y_{i}),P(y_{j}) are the confidence scores of sentence ii and jj on whether to select the sentences in the generated summary, and Sim(si,sj)Sim(s_{i},s_{j}) is the similarity, i.e. redundancy between sentence ii and jj. Noting that we define Sim(si,si)Sim(s_{i},s_{i}) as By adding the redundancy loss term, we penalize it more if two sentences are similar to each other and both of them have high confidence scores. β\beta is a balance factor, controlling the degree of redundancy.

4 Explicitly Reduce Redundancy in Sentence Selection (Category C, Table 2)

We first introduce an existing method and then propose two novel methods that explicitly reduce redundancy in the sentence selection phase.

Trigram Blocking is widely used in recent extractive summarization models on the news dataset (e.g. Liu and Lapata (2019)). Intuitively, it borrows the idea of MMR to balance the importance and non-redundancy when selecting sentences. In particular, given the predicted sentence scores, instead of just selecting sentences greedily according to the scores, the current candidate is added to the summary only if it does not have trigram overlap with the previous selected sentences. Otherwise, the current candidate sentence is ignored and the next one is checked, until the length limit is reached.

MMR-Select: Inspired by the existence of a relevance/redundancy trade-off, we propose MMR-Select, a simple method to eliminate redundancy when a neural summarizer selects sentences to form a summary, in a way that is arguably more flexible than Trigram Blocking with a balance factor λ\lambda.

With the confidence score computed by the basic model, P={P(y1),P(y2),...,P(yn)}P=\{P(y_{1}),P(y_{2}),...,P(y_{n})\}, instead of picking sentences greedily, we pick the sentences according to the MMR-score, which is defined based on MMR and updated after each single sentence being selected.

The main difference between the Naive MMR and MMR-Select falls into the computation of the importance score. In the Naive MMR, the importance score is the similarity between each sentence and the query, or the whole document, while in MMR-Select, the importance score is computed by a trained neural model.

MMR-Select+ : The main limitation of MMR-Select is that the sentence scoring phase and the sentence selection phase cannot benefit from each other, because they are totally separate.

To promote synergy between these two phases, we design a new method, MMR-Select+, shown in Figure 2, which synergistically combines three components: the basic model, the original cross-entropy loss LceL_{ce}(in blue), and an RL mechanism (in green) whose loss is LrdL_{rd}. The neural model is then trained on a mixed objective loss LL with γ\gamma as the scaling factor. Zooming on the details of the RL component, it first generates a summary S^\hat{S} by applying the MMR selection described for MMR-Select, which is to greedily pick sentences according to MMR-score, as well as the corresponding label assignment Y^={y1^,y2^,...,yn^}\hat{Y}=\{\hat{y_{1}},\hat{y_{2}},...,\hat{y_{n}}\} (yi^=1\hat{y_{i}}=1 if sis_{i} is selected, yi^=0\hat{y_{i}}=0 otherwise). Then, the expected reward is computed based on the ROUGE score between S^\hat{S} and the gold-standard human abstractive summary SS weighted by the probability of the Y^\hat{Y} labels. Notice that we also adopt the self-critical strategy Paulus et al. (2017) to help accelerating the convergence by adding a baseline summarySˉ\bar{S}, which is generated by greedily picking the sentences according to PP. r(Sˉ)r(\bar{S}) is the reward of this baseline summary and it is subtracted from r(S^)r(\hat{S}) to only positively reward summaries which are better than the baseline. Formally, the whole MMR-Select+ model can be specified as follows:

Experiments

In this section, we describe the settings, results and analysis of the experiments of different methods on the Pubmed and arXiv datasets.

Following previous work, we use GloVe Pennington et al. (2014) as word embedding, and the average word embedding as the distributed representation of sentences. To be comparable with Xiao and Carenini (2019), we set word length limit of the generated summaries as 200200 on both datasets. A document representation in Unsupervised MMR is similarly computed by averaging the embeddings of all the words. We tune the hyperparameter λ\lambda and β\beta in the respective methods on the validation set, and set λ=0.6,β=0.3\lambda=0.6,\beta=0.3 for both datasets. Following previous work (e.g., Li et al. (2019)), γ\gamma was set to 0.990.99. For training MMR-Select+, the learning rate is lr=1e−6lr=1e-6; we start with the pretrained ExtSumm-LG model. As for the evaluation metric, we use ROUGE scores as the measurement of importance while using the Unique N-gram Ratio and NID defined in Section 3 as the measurements of redundancy.

2 Finetuning λ𝜆\lambda

Consistently with previous work Jung et al. (2019), when we finetune λ\lambda of MMR Select on the validation set, we pinpoint the trade off between importance and non-redundancy in the generated summary (see Figure 3). For λ≤0.6\lambda\leq 0.6, as we increase the weight of importance score, the average ROUGE scores continuously increase while the redundancy/diversity increases/drops rapidly. But since extractive methods can only reuse sentences from the input document, there is an upper bound on how much the generated summary can match the ground-truth summary, so when λ>0.6\lambda>0.6, the ROUGE score even drops by a small margin, while the redundancy/diversity still increases/drops. Then the problem to solve for future work is how to increase the peak, which could be done by either applying finer units (e.g., clauses instead of sentences) or further improve the model that predicts the importance score.

3 Overall Results and Analysis

The experimental results for the ROUGE scores are shown in Table 3, whereas results for redundancy scores (Unique N-gram Ratio and NID score) are shown in Table 4. With respect to the balance between importance and non-redundancy, despite the trade-off between the two aspects, all of the three methods we propose can reduce redundancy significantly while also improving the ROUGE score significantly compared with the ExtSum-LG basic neural model. In contrast, the NeuSum Decoder and Trigram Blocking effectively reduce redundancy, but in doing that they hurt the importance aspect considerably. Even worse, the SR Decoder is dominated by the basic model on both aspects.

Focusing on the redundancy aspect (Table 4), Trigram Blocking makes the largest improvement on redundancy reduction, but with a large drop in ROUGE scores. This is in striking contrast with results on news datasets Liu and Lapata (2019), where Trigram Blocking reduced redundancy while also improving the ROUGE score significantly. Plausibly, the difference between the performances across datasets might be the result of the inflexibility of the method. In both Pubmed and arXiv datasets, the sentences are much longer than those in the news dataset (See Table 1), and therefore, simply dropping candidate sentences with 3-gram overlap may lead to incorrectly missing sentences with substantial important information.

Furthermore, another insight revealed in Table 4 is that dealing with redundancy in the sentence selection phase is consistently more effective than doing it in the sentence scoring phase, regardless of whether this happens implicitly (NeuSum >> SR Decoder) or explicitly (Trigram Blocking, MMR-Select/+ >> RdLoss).

Moving to more specific findings about particular systems, we already noted that while the NeuSum Decoder reduces redundancy effectively, it performs poorly on the ROUGE score, something that did not happen with news datasets. A possible explanation is that the number of sentences selected for the scientific paper datasets (on average 6-7 sentences) is almost twice the number of sentences selected for news; and as it was mentioned in the original paper Zhou et al. (2018), the precision of NeuSum drops rapidly after selecting a few sentences.

Other results confirm established beliefs. The considerable difference between Naive MMR and MMR-Select was expected given the recognized power of neural network over unsupervised methods. Secondly, the unimpressive performance of the SR decoder confirms that the in-order sequence scoring is too limited for effectively predicting importance score and reducing redundancy.

4 More Insights of the Experiments

In addition to the main experiment results discussed above, we further explore the performance on informativeness (ROUGE score) and redundancy (Unique N-gram Ratio) of different redundancy reduction methods under two different conditions, namely the degree of redundancy and the length of the source documents. Figure 8 shows the results on the Pubmed dataset, while further results of a similar analysis on the arXiv dataset can be found in the Appendices. With respect to the degree of redundancy, (upper part of Figure 8), the less redundant the document is, the less impact the redundancy reduction methods have. Among all the methods, although Trigram Blocking works the best with respect to reducing redundancy, it hurts the informativeness the most. However, it is still a good choice for a rather less redundant document (e.g. the documents in the last two bins with avg Unique N-gram Ratio over 0.70.7), which is also consistent with the previous works showing the Trigram Blocking works well on the news datasets, which tends to be less redundant (see §3). As for all the other methods, although they have the same trends, MMR-Select+ performs the best on both informativeness and redundancy reduction, especially for the more redundant documents.

Regarding to the length of the source document (bottom part of Figure 8) , as the document become longer, both informativeness and redundancy in the summary generated by all methods increases and then decrease once hitting the peak. MMR-Select+ and MMR-Select are the best choices to balance between the informativeness and redundancy - they are the only two methods having the higher ROUGE scores and higher Unique N-gram ratios across different lengths, even for the short documents with less than 50 sentences.

Besides, we also conduct experiments on generating summaries with different length limit, where we found that our new methods are stable across different summary lengths (Figure. 5).

Conclusion and Future work

Balancing sentence importance and redundancy is a key problem in extractive summarization. By examining large summarization datasets, we find that longer documents tend to be more redundant. Therefore in this paper, we systematically explore and compare existing and newly proposed methods for redundancy reduction in summarizing long documents. Experiments indicate that our novel methods achieve SOTA on the ROUGE scores, while significantly reducing redundancy on two scientific paper datasets (Pubmed and arXiv). Interestingly, we show that redundancy reduction in sentence selection is more effective than in the sentence scoring phase, a finding to be further investigated .

Additional venues for future work include experimenting with generating summaries at finer granularity than sentences, as suggested by our analysis of the λ\lambda parameter. We also intend to explore other ways to assess redundancy, moving from computing the cosine similarity between sentence embeddings, to a pre-trained neural model for sentence similarity. Finally, we plan to run human evaluations to assess the quality of the generated summaries. This is quite challenging for scientific papers, as it requires participants to possess sophisticated domain-specific background knowledge.

Acknowledgments

We thank reviewers and the UBC-NLP group for their insightful comments. This research was supported by the Language & Speech Innovation Lab of Cloud BU, Huawei Technologies Co., Ltd.

References

Appendix A Appendices

In these Appendices, we show more analysis of the experimental results.

Figure 10 shows the performance on informativeness (ROUGE score) and redundancy (Unique N-gram Ratio) of different redundancy reduction methods under different conditions on the arXiv dataset. Comparing with the Pubmed dataset, the documents in the arXiv dataset tend to be longer and more redundant, as the majority of the documents in the Pubmed dataset have less than 100 sentences with average Unique N-gram Ratio in the 0.5−0.60.5-0.6 interval, while the majority of the documents in the arXiv dataset have number of sentences in the range 100 to 300 with average Unique N-gram Ratio in the 0.6−0.70.6-0.7 interval. Consistent with the result on the Pubmed dataset, the Trigram Blocking method is the best choice for rather less redundant documents (with average Unique N-gram Ratio larger than 0.70.7), and the MMR-Select+ is the one better or equivalent to the original model across different degree of redundancy, ignoring the outliers. With respect to the length of the documents, the MMR-Select+ and MMR-Select are consistently the most effective methods for balancing redundancy and informativeness on documents with different length.

A.2 Analysis on Selection Overlap

To explore the difference made by applying different redundancy reduction methods on the original method(ExtSumLG), we compare the selected sentences by all the methods, and show the overlap ratios between every two methods, as well as the total number and the average length of selected sentences in the test set, in Table 5 and Table 6 for Pubmed dataset and arXiv dataset respectively. As we can see from the tables, except for the SR Decoder, all the other methods tend to select more and shorter sentences than the original summarizer. Regarding the overlap between the original method and the others, we observe that among all the three categories, the methods in category A tend to produce large differences, since these methods change the structure of the original model. Comparing the methods in Category C, around 36%36\% of the sentences are regarded as redundant by Trigram Blocking, which means 36%36\% of the sentences have trigram-overlap with other selected sentences, while only around 10%10\% sentences are regarded as redundant by MMR-Select. As the ROUGE scores of MMR-Select are much better than Trigram Blocking on both datasets, this is in line with our analysis in Section 5.3, Triagram Blocking dropping some important sentences incorrectly. Interestingly, we notice that the overlap ratio between Trigram Block and MMR-Select is considerably larger than the overlap ratio of Trigram Block with original method (ExtSumLG) on both datasets. This indicates that there are some sentences, not selected by the original method, which are considered to be important by both the Trigram Blocking and MMR-Select methods.

A.3 Analysis on Recall and Precision of ROUGE Scores

We also provide the Precision and Recall of the ROUGE scores in the main experiment, the results of Pubmed and arXiv datasets are shown in Table 7 and Table 8, respectively. It is interesting to see that the NeuSum Decoder tends to have a high precision but low recall, indicating that the generated summaries tend to be shorter and contain less useful information than the original method.

A.4 Analysis on the Relative Position of Selections

We also show the relative position distribution of the selected sentences on both datasets in Figure 7 to verify if any redundancy reduction method has a particular tendency to select sentences in particular position of the documents. However, as shown in the figure, the trends are all rather similar for all methods.