RefSum: Refactoring Neural Summarization

Yixin Liu, Zi-Yi Dou, Pengfei Liu

Introduction

In neural text summarization, system designers commonly have flexible choices in model architectures Rush et al. 2015; Kedzie et al. 2018, decoding strategies Paulus et al. 2018 (e.g. beam search) and etc. As a result, even on the same dataset, different selection biases of these choices will lead to diverse system outputs Kedzie et al. 2018; Hossain et al. 2020.

To combine complementarity of system’s output under different setups, researchers have made some preliminary efforts on two-stage learning Collins and Koo 2005; Huang 2008; González-Rubio et al. 2011; Mizumoto and Matsumoto 2016, consisting of (i) a base-stage: first generates different outputs under different setups, and (ii) a meta-stage: then aggregates them in diverse ways, exemplified by stacking that uses a high-level model to combine multiple low-level models Ting and Witten 1997, or reranking Collins and Koo 2005, which aims to rerank different outputs of one system. Although these methods each play a role in different scenarios, they suffer from following potential limitations:

(i) Ad-hoc Methods: most existing methods are designed for a specific scenario. For example, Li et al. 2015 and Narayan et al. 2018b resort to reranking techniques to select summary-worthy sentences that are usually generated from one system. By contrast, Hong et al. 2015 focus on summaries generated from different systems and use a non-neural system combination method to make their complementary advantages. Few works explore if the complementarity existing in different scenarios could be utilized in a unified framework.

(ii) Base-Meta Learning Gap: parameterized models between two learning stages are relatively independent. For example, Zhou et al. 2017 and Huang et al. 2020 adapt the seq2seq Sutskever et al. 2014 framework as the meta model for combination, which takes the outputs of multiple base systems as a part of the inputs for machine translation. As a result, there is no parameter sharing between the meta model and base systems as shown in Fig. 1, which prevents the meta model from fully utilizing the knowledge encoded in the base systems.

(iii) Train-Test Distribution Gap: regarding the meta-learning stage, there is a distribution gap between the training and test distributions. Fig. 1 elucidates this phenomenon: the training distribution of Hypo differs from the test distribution of Hypo’. Although both two are outputs from the base stage, Hypo would be more accurate (closer to gold summaries) since it is the output during the training phase.

In this work, we aim to address these limitations by proposing a general framework, named Refactor, which can not only serve as a base system to construct a summary by selecting sentences from the source document but also act as a meta system to select the best system output from multiple candidates. The unification of base and meta systems allows them to share a set of parameters, thereby alleviating the “Base-Meta learning gap”. Besides, we propose a pretrain-then-finetune paradigm for Refactor that mitigates the “Train-Test distribution gap”. In practice, our proposed Refactor can be applied to different scenarios. For example, as a meta system, it can be used for multiple system combination or single system re-ranking.

Our contributions can be briefly summarized as:

(1) We dissect two major factors that influence the performance of two-stage learning when leveraging the complementarity among different systems: (i) Base-Meta Learning Gap (ii) Train-Test Distribution Gap;

(2) We show these two types of gaps can be alleviated by promoting communication between the two stages in §4 , and therefore present a new paradigm where the base and meta learners are parameterized with shared parameters;

(3) We have made comprehensive experiments (twenty-two top-scoring systems, four datasets). In addition to achieving state-of-the-art results on CNN/DailyMail dataset (§5) by a significant margin, the efficacy of the proposed Refactor opens up a thought-provoking direction for performance improvement: instead of pursuing a purely end-to-end system, a promising exploration is to incorporate different types of inductive biases stage-wisely with the same parameterized function. Our experimental results demonstrate that there exists complementarity introduced by decoding algorithms (e.g. beam search) §5.5 or system combination §5.6 among the current state-of-the-art summarization systems, which can be effectively utilized by our model for boosting the system performance.

Preliminaries

Existing works commonly design systems in an end-to-end fashion Sutskever et al. 2014; Sukhbaatar et al. 2015, which, though effective, also proves to be insufficient in some scenarios Glasmachers 2017; Webb et al. 2019. Instead of optimizing a system in an end-to-end fashion, one more flexible paradigm, stage-wise learning, is to break down the holistic process into different stages. The basic idea is to incorporate different types of inductive biases stage-wisely and two typical examples are: Stacking and Reranking.

Stacking (a.k.a, Stacked Generalization) is a general method of using a high-level model to combine lower-level models to achieve greater predictive accuracy Ting and Witten 1997. In NLP research, this method has been widely explored in machine translation (MT) task. Traditionally, it is used to improve the performance of statistical MT systems (González-Rubio et al. 2011; Watanabe and Sumita 2011; Duh et al. 2011; Mizumoto and Matsumoto 2016). Some recent work (Zhou et al. 2017; Huang et al. 2020) also extends this method to neural MT where the meta model and base systems are all neural models. There is a handful of works about system combination for summarization Hong et al. 2015, in which a feature-based meta model is used for combining unsupervised text summarization systems.

Reranking is a technique to improve performance by reranking the output of an existing system, which has been widely used across different NLP tasks, such as constituency parsing (Collins and Koo 2005; Huang 2008), dependency parsing Zhou et al. 2016; Do and Rehbein 2020, semantic parsing (Ge and Mooney 2006; Yin and Neubig 2019), machine translation (Shen et al. 2004; Mizumoto and Matsumoto 2016).

Comparing reranking and stacking, both of them involve two-stage learning and the first stage would provide multiple candidate outputs as the input for the second stage. However, they differ in the way how multiple candidate outputs are generated at the first stage. Specifically, reranking usually decodes kk-most qualified results during inference, using one base system. By contrast, stacking generates multiple outputs that are usually from different base systems.

Summarization as Two-stage Learning

In what follows, we detail how to formulate summarization as a two-stage learning task.

The system in the base stage aims to generate a summary based on the input text. Specifically, given a document D={s1,⋯ ,sn}D=\{s_{1},\cdots,s_{n}\} with nn sentences, we refer to CC as a candidate summary of DD generated by a summarization system, which can be parameterized in diverse forms:

where \textscBase(,Θbase)\textsc{Base}(,\Theta^{\text{base}}) represents a base system that can be instantiated either as an extractive model or abstractive model with a specific experimental setup: training method T\mathcal{T}, decoding strategy S\mathcal{S}.

In practice, different choices of parameterized function \textscBase(⋅)\textsc{Base}(\cdot), training method T\mathcal{T} and decoding strategy S\mathcal{S} commonly lead to different candidate summaries, C={C1,⋯ ,Ck}\mathcal{C}=\{C_{1},\cdots,C_{k}\}, where C\mathcal{C} represents a set of different candidate summaries. The goal of the meta system is to utilize complementarities among C\mathcal{C} by popular techniques, such as reranking and system combination.

Specifically, given a set of candidate summaries C\mathcal{C}, a meta system is used to re-construct a new candidate summary C∗C^{*}

where Θmeta\Theta^{\text{meta}} represents learnable parameters of the meta system.

Refactoring Text Summarization

Despite effectiveness of existing meta systems, they, as briefly mentioned in §1, suffer from two major problems: (i) Base-Meta Learning Gap and (ii) Train-Test Distribution Gap.

In this paper, we propose the model Refactor that unifies the goal of the base and meta systems by the view that a summary can be generated by selecting the best combination of document sentences. Therefore, both base and meta systems aim to select an optimal candidate summary, and they only differ in how the candidate summary set is constructed. For example, Refactor can be a base system when the candidate summary set C\mathcal{C} is formed by directly enumerating different combinations of document sentences and would be a meta system when C\mathcal{C} represents summaries from different systems. This formulation is advantageous in two points:

(1) No matter where a system selects (from document sentences or multiple system outputs), the chosen criteria that define a good summary are shared. Therefore, the learning process of base and meta systems can be parameterized using a set of parameters, maximizing the information-sharing across two stages and mitigating the Base-Meta Learning Gap.

where \textscRefactor(⋅,Θrefactor)\textsc{Refactor}(\cdot,\Theta^{\text{refactor}}) is the Refactor model, and the candidate summaries C\mathcal{C} can be constructed in different ways.

(2) Additionally, learning to select candidate summaries from document sentences enables the system to see more diverse candidates with different distributions. This is effective for solving the Train-Test Distribution Gap, where the distribution of the meta system outputs in training samples deviates from the test one.

Specifically, our proposed Refactor first learns to select candidate summaries from document sentences (pre-trained Refactor) and then learns to select candidate summaries from different system outputs (fine-tuned Refactor).

2 Pre-trained Refactor

Pre-trained Refactor takes as input a document D={s1,⋯ ,sn}D=\{s_{1},\cdots,s_{n}\} as well as a set of candidate summaries C={C1,⋯ ,Cm}\mathcal{C}=\{C_{1},\cdots,C_{m}\}, which can be constructed by enumerating possible combinations of source sentences with heuristic pruning. For example, an extractive system could be used to prune unlikely sentences to control the number of candidates. \textscRefactor(⋅,Θrefactor)\textsc{Refactor}(\cdot,\Theta^{\text{refactor}}) is instantiated as a score function which quantifies the degree to which a candidate summary CiC_{i} is matched with the source document DD.

where D\mathbf{D} and Ci\mathbf{C}_{i} denote document and summary representations respectively, which are calculated by a BERT (Devlin et al. 2019) model. \textscScore(⋅)\textsc{Score}(\cdot) is a function that measures the similarity between a document and candidate summary.

To instantiate \textscScore(⋅)\textsc{Score}(\cdot), we follow the forms as mentioned in Zhang et al. 2019b; Zhao et al. 2019; Gao et al. 2020, which have shown superior performance on measuring semantic similarity between documents and summaries.

Specifically, \textscScore(⋅)\textsc{Score}(\cdot) is defined based on the greedy matching algorithm, which matches every word in one text sequence to the most similar word in another text sequence and vise versa. Given the document embedding matrix D=⟨d1,⋯ ,dk⟩\mathbf{D}=\langle\mathbf{d}_{1},\cdots,\mathbf{d}_{k}\rangle and the candidate embedding matrix C=⟨c1,⋯ ,cl⟩\mathbf{C}=\langle\mathbf{c}_{1},\cdots,\mathbf{c}_{l}\rangle encoded by BERT, \textscScore(⋅)\textsc{Score}(\cdot) can be calculated as:

We use a ranking loss to learn the parameter Θrefactor\Theta^{\text{refactor}}, inspired by the assumption (Zhong et al. 2020) that a good candidate summary should be as close with the source document as possible. Formally,

3 Fine-tuned Refactor

In order to fit the distributions of the specific types of input, we then fine-tune Refactor using the outputs generated by the base systems. Specifically, fine-tuning is also based on Eq. 9 where the candidate summaries CC are generated by the base systems under different application scenarios.

We elaborate on the proposed two-step training using a real case. Fig. 2 depicts the distribution of ROUGE-1 scores regarding the candidate summaries in the pre-training stage training set, fine-tuning stage training set and test set on the XSum dataset, where we sample the same number of {document,candidate summaries}\{\textrm{document},\textrm{candidate summaries}\} pairs. We can observe that:

(i) there is a distribution gap between train and test samples in fine-tuning stage. (ii) in pre-training stage the pre-trained Refactor has seen a large number of candidate summaries with diverse performance (ROUGE value), which improves its generalization ability. In §5 we will show that the Pre-train and Fine-tune paradigm outperforms one-step training where the model is directly trained with data generated from the base systems.

4 Application Scenarios

Our Refactor can be used as different roles in different scenarios as follows.

The pre-trained Refactor can not only be fine-tuned for a better selection of candidate summaries, but also be regarded as a base system, providing one system output. This feature of Refactor maximizes parameter sharing across the two training stages.

4.2 Refactor as Meta Learner

Both pre-trained Refactor and fine-tuned Refactor can be used as a meta system to select the best candidate when we have multiple system summaries. In this work, we explore the following settings:

(1) Single System: It considers re-ranking candidate summaries generated from a single abstractive system using beam search.

(2) Multi-system Summary-level: It is tasked to select the best candidate summary from the results of different systems.

(3) Multi-system Sentence-level: We also take a step towards the fine-grained fusion of summaries from extractive and abstractive systems. Specifically, here candidate summaries are generated by combining the results of different systems at the sentence level.

Experiments

We mainly experiment on four datasets, whose statistics are shown in Tab. 1.

CNNDM https://cs.nyu.edu/~kcho/DMQA/ (Hermann et al. 2015) is a widely used dataset containing news articles and the associated highlights which are used as the reference summaries. We follow the work of Nallapati et al. 2016 for data preprocessing.

XSum https://github.com/EdinburghNLP/XSum (Narayan et al. 2018a) contains online articles collected from BBC with highly abstractive one-sentence summaries.

PubMed https://github.com/acohan/long-summarization (Cohan et al. 2018) contains scientific papers collected from PubMed.com.

WikiHow https://github.com/mahnazkoupaee/WikiHow-Dataset (Koupaee and Wang 2018) is a large-scale dataset constructed from the articles using online WikiHow knowledge base.

2 Base Systems

Below, we mainly use BART, GSum and PEGASUS as the base systems since they have achieved state-of-the-art performance on at least one dataset.

BART (Lewis et al. 2020) is a large pre-trained sequence-to-sequence model that achieves strong performance on the abstractive summarization.

GSum (Dou et al. 2020) enhances the performance of BART using additional guidance information, which achieves the current state-of-the-art performance on the CNNDM dataset.

PEGASUS (Zhang et al. 2020) achieves competitive performance on various summarization datasets and is the current state-of-the-art on the XSum dataset.

To make a comprehensive evaluation of our proposed model, we additionally collect 19 top-scoring systems as base systems on CNNDM. Since CNNDM is the most popular dataset, we can collect more existing systems on it. In details, for §5.7 we use the following systems: pointer-generator+coverage (See et al. 2017), REFRESH (Narayan et al. 2018b), fastAbsRL-rank (Chen and Bansal 2018), CNN-LSTM-BiClassifier (Kedzie et al. 2018), CNN-Transformer-BiClassifier (Zhong et al. 2019), CNN-Transformer-Pointer (Zhong et al. 2019), BERT-Transformer-Pointer (Zhong et al. 2019), Bottom-Up (Gehrmann et al. 2018), NeuSum (Zhou et al. 2018), BanditSum (Dong et al. 2018), twoStageRL (Zhang et al. 2019a), preSummAbs (Liu and Lapata 2019), preSummAbs-ext (Liu and Lapata 2019), HeterGraph (Wang et al. 2020), MatchSum (Zhong et al. 2020), Unilm-v1 (Dong et al. 2019), Unilm-v2 (Dong et al. 2019), T5 (Raffel et al. 2020).

3 Baseline Systems

Neural system combinator: We use BERTScore (Zhang et al. 2019b) as an unsupervised baseline with neural models, which is an automatic evaluation metric computing the similarity of text pairs based on the corresponding BERT-encoded representations. We use it to directly compute the similarity score between the source documents and candidate summaries.

Non-Neural system combinator: We use RankSVM http://www.cs.cornell.edu/people/tj/svm_light/svm_rank.html (Joachims 2002) as a non-neural baseline. We perform cross-validation on the development set for hyper-parameter searching and train the model on the development set. The set of features is listed in Appendix A.

Oracles: We compare our model with sample-wise Min, Max and Random oracles using ROUGE.

4 Training Details

For the following experiments in §5.5, §5.6 and §5.7 on CNNDM, we pre-train the Refactor model with a candidate set generated by enumerating combinations of sentences in the source documents. To reduce the number of candidates, we prune the sentences assigned with lower scores by an extractive model, BERTSum (Liu and Lapata 2019), following Zhong et al. 2020. The maximum number of candidates for one data sample is 20. The pre-trained Refactor is also used a base system in §5.6, whose outputs are used together with other base systems as candidate summaries. For different experiments, we fine-tune pre-trained Refactor on the base system’s output, and name the model as fine-tuned Refactor. To analyze the effectiveness of the proposed two-stage training, we additionally train the model without the pre-training step, which is named as supervised Refactor.

The pre-trained BERT model we used is from Transformers library (Wolf et al. 2020). We use the ‘bert-base-uncased’ version with 110M parameters. We use Adam optimizer (Kingma and Ba 2015) with learning rate scheduling.

5 Exp-I: Single System Reranking

We use BART and GSum for this experiment, and use beam search to generate the candidate summaries where the beam size is set to 4.

The results are listed in Tab. 2, which shows that (1) Refactor can boost the base system’s performance by a significant margin, (2) the fine-tuned Refactor outperforms supervised Refactor directly trained on the base system’s outputs, showing the effectiveness of the two-step training. Notably, we observe the fine-tuned Refactor can boost BART’s performance from 44.26 to 45.15 on ROUGE-1, indicating that the top-11 output selected by beam search is not always the best one, and Refactor can effectively utilize the complementarity introduced by considering all the beam search results.

6 Exp-II: Multiple Systems Stacking

For summary-level combination, we explore two-system combination (BART & pre-trained Refactor) and three-system combination (BART, GSum & pre-trained Refactor). The results are shown in Tab. 3.

For sentence-level combination, we use BART and pre-trained Refactor as the base systems. The sentences of each system’s output are merged together to form the candidate sentence set, and all combinations of three sentences in the candidate set are generated as candidate summaries. To prune the candidates, we use tri-gram blocking to filter out candidates of which there exists an identical tri-gram in two sentences. The average number of candidates in the test set is 15.8. The results are shown in Tab. 4.

We have the following observations: (1) the pre-trained Refactor can already outperform the base systems, and (2) fine-tuning can further improve the performance. Meanwhile, we notice there are two exceptions: (i) For sentence-level combination, supervised Refactor has similar performance as fine-tuned Refactor. We hypothesis that this is because here the number of candidates in the fine-tuning data is relatively large, therefore directly training on the fine-tuning data is sufficient enough. (ii) The pre-trained Refactor cannot outperform GSum model in the three-system combination setting in Tab. 3. The reason might be that GSum has much stronger performance than the other two systems, which intuitively makes the expected gain from system combination lower than other settings.

7 Exp-III: Generalization on 19 Top-performing Systems

To evaluate the Refactor’s generalization ability, we explore another setting where the pre-trained Refactor is directly used to select the outputs of multiple systems without fine-tuning.

To this end, we collect 19 top-performing summarization systems on CNNDM dataset. Here, we investigate if our Refactor can boost the performance of candidate systems with similar performance. In addition, we also aim to investigate how the range width of different systems’ performance affects Refactor’s performance. Therefore, we group the candidate systems into equal-width bins based on their average ROUGE-1 scores, and evaluate our Refactor on each bin separately.

In Tab. 5 we report the average ROUGE-1 scores of the oracles, Refactor, and the best candidate system in each bin whose width is 1. Refactor consistently outperforms the best candidate system, showing its generalization ability.

Next, in Fig. 3 we plot the change of Refactor’s performance with different bin widths. We define the success rate of Refactor with a given bin width to be the number of bins where Refactor outperforms the single best base system normalized by the total number of bins. We observe that Refactor is more likely to improve the performance of base systems when the system-level performance of the base systems is similar. Intuitively, if one base system is significantly better than the other systems, it is more difficult for Refactor to use other systems to complement the best base system.

8 Exp-IV: Effectiveness on More Popular Datasets

Next, we move on to other text summarization datasets to evaluate our proposed method’s strength beyond CNNDM dataset. Some of the datasets used here are not as well-studied as CNNDM dataset, so there are less top-performing systems on these datasets. Therefore, here we focus on the experiments of the single system setting.

Regarding the pre-trained Refactor, we use an extractive oracle to select document sentences and use the combinations of these sentences as candidates. In addition, since on Xsum the abstractive systems outperform extractive systems by a large margin, we use a pre-trained BART model with Diverse Beam Search (Vijayakumar et al. 2018) to generate 16 candidates per sample for pre-training. Regarding system re-ranking, we use BART as the base system to generate the candidate summaries except on Xsum dataset, where we use PEGASUS since it achieves better performance. Similar to §5.5, we use the outputs of beam search as the candidates. We select the first 4 outputs as the candidates.

The results in Tab. 6 show that Refactor is able to bring stable improvement over the base systems. The average summary length of these datasets varies from 23.3 (XSum) to 210.3 (Pubmed). Therefore, the results here demonstrate the Refactor can be applied to datasets with different characteristics. On XSum dataset, the pre-trained Refactor outperforms the fine-tuned Refactor. This may result from the additional pre-training data we introduced using BART, which is effective enough to train the Refactor for reranking PEGASUS output.

9 Fine-grained Analysis

We perform a fine-grained evaluation of Refactor to understand where improvement mainly comes.

We choose the summary-level system combination setting on CNNDM test set in §5.6 as a case study, where the base systems are: BART and pre-trained Refactor, and then we use a fine-tuned Refactor As introduced in §4.4, Refactor could be used as either a base system or a system combinator. to combine them. Specifically, we first

(i) define δ(CBART,CPretrain)\delta(C_{\text{BART}},C_{\text{Pretrain}}) as the performance (i.e., ROUGE) gap on the candidate summary CC.

(ii) then partition test samples into different buckets S1,⋯ ,SnS_{1},\cdots,S_{n} according to the performance gap δ\delta.

(iii) calculate selection accuracy for each bucket, which represents how accurately the Refactor can identify the best one from two candidate summaries.

The results are shown in Fig. 4. We observe that the selection accuracy is increasing as the gap δ\delta becoming larger, indicating that Refactor performs better on the candidate summaries with diverse performance. Combining the results we get in §5.7, we conclude that Refactor has the largest potential gain when the base systems effectively complement each other – They have similar system-level performance but diverse summary-level performance. For example, each base system may perform significantly better than others on a subset of data with different characteristics but could not outperform others across the whole dataset.

Implications and Future Directions

We present a general framework for utilizing the complementarity of modern text summarization systems by formulating text summarization as a two-stage learning problem. Our proposed model, Refactor, can be used either as a base system or a meta system, effectively mitigating the learning gaps introduced in the two-stage learning. Experimental results show that Refactor is able to boost the performance of the base systems, and achieves the state-of-the-art performance on CNNDM and XSum datasets. We believe this work opens up a new direction for improving the performance of text summarization systems apart from an iterative process of searching for better model architectures – The gain of performance could be made by fully investigating and utilizing the complementarity of different systems with various architectures, problem formulations, decoding strategies, etc.

Acknowledgements

We thank Professor Graham Neubig and anonymous reviewers for valuable feedback and helpful suggestions. This work was supported in part by a grant under the Northrop Grumman SOTERIA project and the Air Force Research Laboratory under agreement number FA8750-19-2-0200. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the Air Force Research Laboratory or the U.S. Government.

References

Appendix A Features for RankSVM

We use 18 features as defined below for RankSVM:

rouge-1, rouge-2, rouge-L between source documents and candidates summaries.

copy length: the length of summary’s fragments appeared in the source document.

fragment coverage, fragment density, compression ratio as defined in Grusky et al. 2018.

novelty: the ratio of novel kk-grams (k∈{1,2,3,4}k\in\{1,2,3,4\}) in the candidate summaries.

repetition: the ratio of repeated kk-grams (k∈{1,2,3,4}k\in\{1,2,3,4\}) in the candidate summaries.

sentence fusion ratio: the ratio of sentences in the candidate summaries that combine the content of two source document sentences.