Sequence Level Contrastive Learning for Text Summarization

Shusheng Xu, Xingxing Zhang, Yi Wu, Furu Wei

Introduction

Document summarization is the task of rewriting a long document into a shorter form while still preserving its important content, which requires the model to understand the entire document. Many approaches for summarization has been explored in the literature and the most popular ones are extractive summarization and abstractive summarization (Nenkova and McKeown 2011). Summaries in their nature are abstractive. The summaries generated by extractive summarization methods are usually long and redundant, which bring bad reading experience. Therefore, we focus on abstractive summarization in this paper. Abstractive summarization is usually modeled as a sequence-to-sequence (Seq2Seq) learning problem (Sutskever, Vinyals, and Le 2014), where a document is viewed as a sequence of words and its summary another sequence of words (Nallapati et al. 2016).

Although abstractive models have been more and more powerful due to recent introduction of large pre-trained Transformers (Liu and Lapata 2019; Raffel et al. 2020; Dong et al. 2019; Lewis et al. 2020), the training paradigm for abstractive models is still not changed, which is to minimize the negative log-likelihood (NLL) between the model predicted word distributions and the gold summary. One great property of the summarization task is that a document and its summary should convey the same meaning, which is not modeled explicitly by the NLL loss.

In computer vision, contrastive learning methods for unsupervised image representation learning advanced the state-of-the-art in object detection and image segmentation (He et al. 2020b). The key idea is to minimize distances (or maximize similarities) between feature representations of different views of the same image (positive examples), while to maximize the distances between feature representations of views of different images (negative examples) (He et al. 2020b; Chen et al. 2020). As mentioned earlier, in summarization a document and its summary should convey the same meaning. Therefore, we view a document, its gold summary and its model generated summaries as different views of the same meaning representation and during training, we maximize the similarities between them. To achieve that, we propose SeqCo (as shorthand for Sequence Level Contrastive Learning), which is based on contrastive learning. In addition to the gold summaries, we also use the dynamically generated summaries from our model during training to increase the diversity of inputs to SeqCo. In text summarization, an abstractive summarization model needs to first encode the document and then generate the summary. The contrastive objective in SeqCo tries to map representations of a document and its summary (or generated summary) to the same vector space, which intuitively helps the generation of summaries. Specifically, a document may contain distinct (or unnecessary) information from its summary. During training time, the contrastive objective between the document and summary actually encourages the model to encode important (and necessary) information from the document, otherwise the distance between the representations of document and summary will be large (the objective updates model parameters to make it small). Intuitively, the capability of encoding important information from documents would help to generate better summaries.

In experiments, we find our proposed contrastive learning based model SeqCo consistently improves upon a strong abstractive summarization model based on BART (Lewis et al. 2020) across three different summarization datasets (i.e., CNN/DailyMail (Hermann et al. 2015), New York Times (Sandhaus 2008) and XSum (Narayan, Cohen, and Lapata 2018)). Human evaluation also shows that our model SeqCo achieves better faithfulness ratings compared to its counterpart without contrastive objectives.

Related Work

The most popular paradigms for summarization are extractive and abstractive based approaches. We focus on abstractive summarization. Abstractive summarization may add new words or phrases when generating summaries, which is usually viewed as a sequence to sequence learning problem (Nallapati et al. 2016; See, Liu, and Manning 2017; Paulus, Xiong, and Socher 2018; Gehrmann, Deng, and Rush 2018). Probably because small and shallow LSTM (Hochreiter and Schmidhuber 1997) based attentive seq2seq models (Sutskever, Vinyals, and Le 2014; Bahdanau, Cho, and Bengio 2015) without pre-training are not powerful enough to model documents. Quality of summaries produced by these mdoels are not satisfactory (Liu and Lapata 2019). As the recent introduction of large pre-trained transformer models (Liu and Lapata 2019; Dong et al. 2019; Zou et al. 2020; Lewis et al. 2020; Zhang et al. 2020; Raffel et al. 2020), abstractive models are greatly improved. Best results for summarization are achieved by finetuning large models pre-trained with generation (or summarization) tailored objectives on huge amount of unlabeled text (≥\geq160G). Dong et al. (2019) pre-train jointly designed Transformer encoder and decoder with language model and masked language model objectives. Zhang et al. (2020) predict gapped sentences from a document removing these sentences and Lewis et al. (2020) propose sentence permutation and text infilling tasks to pre-train seq2seq transformers. There is also some work on combining extractive and abstractive summarization models (He et al. 2020a; Dou et al. 2021) or multiple summarization models (Liu, Dou, and Liu 2021). Unfortunately, pre-training transformers from scratch or combining multiple summarization systems are expensive, while our model can be applied to the light-weighted finetuning stage.

Convoluational neural networks pre-trained with contrastive learning methods advance the state-of-the-art in object detection and image segmentation in computer vision (He et al. 2020b). The idea is to minimize the distances between feature representations of different views of the same image (positive examples), while to maximize the distances between feature representations of views of different images (negative examples). To discriminate positive examples from negative examples, He et al. (2020b) maintain a queue of negative sample representations and utilize momentum updates for encoder of the queue to stabilize these representations. Chen et al. (2020) use other examples from the same batch as negative examples and as a result, they need a large batch size. These works above suggest that using a large number of negative examples is crucial to obtain good performance, which also increases the complexity for implementation. There is also an interesting line of work without using negative examples. Caron et al. (2020) employ online clustering to assign codes for two views of the same image and then use representation of one view to predict the cluster codes of the other. During the training of BYOL (Grill et al. 2020), they only minimized the distance between representations of two views of the same image and they use a momentum encoder for the target view to stabilize the training. Chen and He (2020) find that even the momentum encoder can be removed, although there might be a small drop in performance. The contrastive learning method used in our model is most related to BYOL (Grill et al. 2020) in the sense that we do not use negative examples either and we also employ a momentum encoder. In the models above, contrastive learning is applied in the unsupervised pre-training stage, which create different views of the same image by using effective data argumentation methods. In this paper, we take advantage of the nature of the summarization task and use the document, gold summary, and generated summary as different views of the same meaning representation (note that a summary is a shorter form of the original document). To fit sequence-to-sequence learning models for text generation, we handles two sequence of embeddings of discrete words, while the vision models handle two single embeddings of fixed dimensions. In addition, the generated summary are created dynamically during training with a model, which are more diverse than using non-model-based approaches in vision tasks.

In NLP, previously contrastive learning methods are mostly used in pre-training or natural language understanding tasks. For example, word2vec (Mikolov et al. 2013) learns the word embeddings by distinguishing words in a windows (positive examples) w.r.t. the current word and words randomly sampled (negative examples) using negative sampling. (Iter et al. 2020) propose a contrastive learning based method for language model pre-training, which predicts the relative distance between sentences using randomly sampled sentences as negative examples. More recently, MatchSum (Zhong et al. 2020) formulates extractive summarization as a semantic text matching problem using contrastive learning. Wu et al. (2020) measures the summary qualities without reference summaries by contrasting the document with the summaries using a ranking model. GSum (Dou et al. 2021) takes different kinds of external guidance as additional input to the document and advances summarization performance significantly. SimCLS (Liu and Liu 2021) proposes a contrastive based framework for abstractive summarization, which trains a model to rerank the candidate summaries of an abstractive model. We add constrastive learning to the training of an abstractive model by enforcing similarities between document, summary and generated summary, which does not need negative examples.

Model

In this section, we describe our contrastive learning model SeqCo (as shorthand for Sequence Level Contrastive Learning) for abstractive text summarization. We first introduce abstractive text summarization models (i.e., Seq2Seq model), on which our model is based. Then we present SeqCo, which adapts contrastive learning to the sequence-to-sequence learning setting.

For text summarization, we can view the document as a long sequence of tokensWe use tokens instead of words, because the sequence might be a sequence of sub-words. and the summary as a short sequence of tokens. Let X=(x0=<s>,x1,x2,…,x∣X∣=</s>)X=(x_{0}=\text{\tt<s>},x_{1},x_{2},\dots,x_{|X|}=\text{\tt</s>}) denote a document (i.e., the long sequence of tokens) and Y=(y0=<s>,y1,y2,…,y∣Y∣=</s>)Y=(y_{0}=\text{\tt<s>},y_{1},y_{2},\dots,y_{|Y|}=\text{\tt</s>}) its summary (i.e., the short sequence of tokens), where and are begin and end of sequence tokens. We predict YY one token at a time given XX. We adopt the Transformer model (Vaswani et al. 2017), which is composed of an encoder Transformer and a decoder Transformer. Specifically, the encoder Transformer maps XX into a sequence of hidden states E=(e0,e1,…,e∣X∣)\mathbf{E}=(\mathbf{e}_{0},\mathbf{e}_{1},\dots,\mathbf{e}_{|X|}).

Supposing that the first t−1t-1 tokens y1:t−1y_{1:t-1} have been generated and we are generating yty_{t}. The decoder Transformer computes the current hidden state ot\mathbf{o}_{t} by self attending to the encoder hidden states E\mathbf{E} and proceeding tokens y0:t−1y_{0:t-1}.

Note that during training, we can obtain O=(o1,...,o∣Y∣)\mathbf{O}=(\mathbf{o}_{1},...,\mathbf{o}_{|Y|)} in parallel.

The probability of yty_{t} can be estimated using a linear projection and a softmax function

2 SeqCo: Sequence Level Contrastive Learning for Text Summarization

In text summarization, the summary YY is a shorter form of the input document XX and they should convey the same meaning. Therefore, XX and YY should be close in the semantic space at least after certain types of transformations. However, a Seq2Seq model is trained using the negative log-likelihood loss (see Equation (5)) and there is no explicit modeling for the similarity between XX and YY. Further, during the training phase, given XX as input, the model can also generate output sequences from its distribution by either beam search or sampling. Let Y^\hat{Y} denote one sample the model generated from XX. Intuitively, Y^\hat{Y} should also be similar to both XX and YY. As shown in figure 1, we enforce the similarities between XX, YY and Y^\hat{Y} during model training. To do this, we propose SeqCo, which is a contrastive learning based model for text summarization.

Contrastive learning methods are proposed in the context of self-supervised learning for image representations (Wu et al. 2018; He et al. 2020b; Caron et al. 2020; Grill et al. 2020; Chen and He 2020). The training objective tries to make representations of different views of the same image closer (positive examples) while representations of views of different images apart from each other (negative examples). Inspired by Grill et al. (2020) and Chen and He (2020), we propose a model that does not need negative examples. In the following, we first define similarity measures between sequences and then we present how to equip the similarity measures into our training objective.

Suppose that we have two sequences Si=(w0i,w1i,w2i,...,w∣Si∣i)S_{i}=(w_{0}^{i},w_{1}^{i},w_{2}^{i},...,w_{|S_{i}|}^{i}) and Sj=(w0j,w1j,w2j,...,w∣Sj∣j)S_{j}=(w_{0}^{j},w_{1}^{j},w_{2}^{j},...,w_{|S_{j}|}^{j}). SiS_{i} and SjS_{j} are two sequences, which we will maximize their similarity in Eq. 15. For example, SiS_{i} and SjS_{j} can be a document X and its gold summary Y, or document and generated summary, or gold summary and generated summary, just like Fig. 2. Before going to the similarity computation, we first convert them into sequences of hidden representations. We designed two mapping functions here. The first one (fθEf_{\theta}^{\text{E}}) is unconditional, which reuses the encoder of our Seq2Seq model (see Section 3.1):

The second mapping function (fθDf_{\theta}^{\text{D}}) is conditional, which takes of the input sequence into account.Note that in fθDf_{\theta}^{\text{D}} we only consider that SiS_{i} and SiS_{i} as the gold summary and the generated summary Let XX denote the input sequence and SiS_{i} is its gold output sequence or a sequence generated by the Seq2Seq model. In this mapping function, we employ both the encoder and the decoder of the Seq2Seq model (see Section 3.1 for details):

Sequence Similarity

After defining the mapping functions, we are ready to compute sequence similarities. Without losing generality, let fθf_{\theta} denote the mapping function, where θ\theta is the parameter of the function. Note that fθf_{\theta} can be either fθEf_{\theta}^{\text{E}} or fθDf_{\theta}^{\text{D}} (see Eq. (6) and (7) for details). We additionally employ another mapping function fξf_{\xi}, which has the same architecture as fθf_{\theta}, but with parameter ξ\xi. We obtain the representations of SiS_{i} and SjS_{j} by applying fθf_{\theta} and fξf_{\xi} to them:

To fully utilize the word-to-word interactions between the two sequences SiS_{i} and SjS_{j}, we apply a cross attention between Hi\mathbf{H}^{i} and Hj\mathbf{H}^{j}:

We adopt multi-head attention (MHA) for similarity computation for two reasons. 1) The sequences (esp. documents) are long and MHA takes all pairs of tokens across two sequences into account, which is intuitively more powerful than [CLS] pooling based methods (will introduce below). 2) The two sequences we compare may have different lengths (e.g., a document v.s. a summary). MHA can convert the hidden states of one sequence to the same length as the hidden states of another sequence (see Equation 9), which are easier to use for the similarity computation.

Note that we can also define a simpler similarity function using the [CLS] pooling as in BERT (Devlin et al. 2019):

where qq is a feed-forword network to project h0i\mathbf{h}_{0}^{i} following Grill et al. (2020). We obtained worse results using the similarity measure above (see Section 4.4 for details) and the measure also sometimes leads to numerical errors during training.

Training

To make SiS_{i} and SjS_{j} closer, we can minimize the following loss:

As mentioned earlier, fθf_{\theta} (the encoding function for SiS_{i}) and fξf_{\xi} (the encoding function for SjS_{j}) use different set of parameters (i.e., θ\theta and ξ\xi). If we update the parameters in both fθf_{\theta} and fξf_{\xi} simultaneously, the optimization maybe too easy, which may lead to collapsed solutions (Grill et al. 2020). So we use fξf_{\xi} to produce regression targets for fθf_{\theta}. Specifically, we do not update the parameters in fξf_{\xi} during the optimization of the loss above and ξ\xi is a moving average of θ\theta:

where τ∈\tau\in is a hyper-parameter to control the extend of retaining ξ\xi. This contrastive objective is demonstrated in figure 2. Note that Lθ,ξ(Si,Sj)\mathcal{L}_{\theta,\xi}(S_{i},S_{j}) is not symmetric and we make the loss symmetric as follows:

Hence, θ\theta in fθf_{\theta} will have more chances to be updated. As mentioned earlier, the encoding function fθf_{\theta} can be either fθEf_{\theta}^{\text{E}} or fθDf_{\theta}^{\text{D}}. We use LsimE\mathcal{L}_{\text{sim}}^{\text{E}} to denote the loss function using fθEf_{\theta}^{\text{E}} and LsimD\mathcal{L}_{\text{sim}}^{\text{D}} to denote the loss function using fθDf_{\theta}^{\text{D}}.

To enforce the similarities between the document XX, its gold summary YY and one of the model generated summary Y^\hat{Y}, we employ the following loss function as our final training lossWe can also use multiple generated summaries in training, we refrained to do so for efficiency reasons.:

This objective contains five terms. LNLL\mathcal{L}^{\text{NLL}} is the negative log-likelihood; LsimD\mathcal{L}_{sim}^{D} is the similarity loss w.r.t. (Y,Y^)(Y,\hat{Y}) with fθDf_{\theta}^{\text{D}}; LsimE\mathcal{L}_{sim}^{\text{E}} terms are the similarity losses with fθEf_{\theta}^{\text{E}} w.r.t. (X,Y)(X,Y), (X,Y^)(X,\hat{Y}) and (Y,Y^)(Y,\hat{Y}). λx−y\lambda_{x-y}, λx−y^\lambda_{x-\hat{y}}, λy−y^\lambda_{y-\hat{y}} and λy−y^D\lambda^{\text{D}}_{y-\hat{y}} are weight hyper-parameters for the last four terms. We completely train the model end-to-end following this loss function and empirically find that using a single similarity loss works better than using multiple ones (see Section 4.4), which is also more efficient for training. For example, we can set λx−y^=1.0\lambda_{x-\hat{y}}=1.0 and λx−y=λy−y^=λy−y^D=0\lambda_{x-y}=\lambda_{y-\hat{y}}=\lambda^{\text{D}}_{y-\hat{y}}=0. When Y^\hat{Y} is adopted, The model iteratively generates Y^\hat{Y} by using the loss to update parameters and generating new Y^\hat{Y}. Since Y^\hat{Y} can not be perfect, iteratively generating Y^\hat{Y} makes it change toward ground-truth summary and make the positive examples for contrastive learning more accurate and diverse. Since SeqCo is designed for the fine-tuning stage, and the model SeqCo based on (i.e., BART) is pre-trained with a denoising auto-encoding objective, it can naturally generate the sequence with the same meaning as the input even before fine-tuning in a specific dataset. In addition, enforcing the similarity of yy and y^\hat{y} does not equals optimizing NLL, since the similarity loss is on sequence level while the NLL loss is on token level.

Experiments

In this section, we assess the preformance of our contrastive learning model on the task of text summarization. We will first introduce the datasets we used. Then we present our implementation details. Finally, we compare our model with multiple previous models.

We conduct our experiments on three summarization datasets. The CNN/DailyMail dataset (CNNDM; Hermann et al. 2015) contains news articles and their associated highlights (i.e., reference summaries) from the CNN and Daily Mail websites. We follow the standard pre-processing steps in (See, Liu, and Manning 2017)Available at https://github.com/abisee/cnn-dailymail and the resulting dataset contains 287,226 articles for training, 13,368 for validation and 11,490 for test.

NYT

The New York Times dataset (NYT; Sandhaus 2008) is composed of articles published by the New York Times with summaries written by library scientists. Following the pre-processing procedures in (Durrett, Berg-Kirkpatrick, and Klein 2016; Liu and Lapata 2019), we first obtain 110,540 articles with abstractive summaries. The test set is constructed from the 9,706 articles published after January 1, 2007. After removing articles whose summaries are shorter than 50 words, the final test set contains 3,452 articles. The remaining 100,834 articles are filtered and splitted into 38,264 articles for training and 4,000 articles for validation.

XSum

The articles in the XSum dataset (Narayan, Cohen, and Lapata 2018) are from the BBC website with accompanying single sentence summaries, which are professionally written. We use the official splits of (Narayan, Cohen, and Lapata 2018) (i.e., 204,045 articles for training, 11,332 articles for validation and 11,334 articles for test). All datasets are tokenized with the byte-pair encoding of GPT2 (Radford et al. 2019).

2 Implementation Details

Our model is initialized from BARTLarge\text{BART}_{\text{Large}} (Lewis et al. 2020). Therefore, the size is identical with BARTLarge\text{BART}_{\text{Large}} (Lewis et al. 2020). Specifically, the encoder and decoder are all 12-layer transformers with 16 attention heads, hidden size 1,024 and feed-forward filter size 4,096, which amounts to 406M trainable parameters. We also have additional component for contrastive learning. The feedforward network gg (see Equation (6) and (7)) for projecting sequence features contains one hidden layer of 4,096 neurons with ReLU activation function. The multi-head attention module (see Equation (9)) used to compute cross attention between sequences also has 16 heads. These two components above contribute to an extra 13M trainable parameters.

We optimize the model using Adam with β1=0.9,β2=0.999\beta_{1}=0.9,\beta_{2}=0.999. Following (Lewis et al. 2020), we employ a linear schedule for the learning rate. We firstly warmup the model by increasing the learning rate linearly to a peak learning rate and then decrease the learning rate linearly to zero. The peak learning rate, warmup steps, total number of updates and batch size are tuned on validation sets and are different across datasets, which are 10001000, 2000020000, 4e−54e-5, 128128 on CNNDM, 500500, 50005000, 2e−52e-5, 6464 on NYT, 500500, and 1500015000, 6e−56e-5, 6464 on XSum. In all datasets, the number of training epochs are between 5 to 10. During the optimization, parameters ξ\xi in the online encoding function fξf_{\xi} (see Equation (6) and (7)) are not updated. Parameters ξ\xi infξf_{\xi} are updated following Equation (13) with τ=0.99\tau=0.99. We employ label smoothing of 0.1 (Szegedy et al. 2016; Vaswani et al. 2017). The models for CNNDM are trained on 8 Tesla V100 GPUs, and the models for the other datasets are trained on 4 Tesla V100 GPUs. During decoding, we select minimum generated length and length penalty according to ROUGE scores on the validation set. Following (Paulus, Xiong, and Socher 2018), we also blocked repeated trigrams during beam search. Following (Lewis et al. 2020), the articles are truncated to 1024 tokens in both training and decoding.

3 Evaluations

We use ROUGE (Lin 2004) to measure the quality of generated summaries. We reported full-length F1 based ROUGE-1, ROUGE-2 and ROUGE-L scores on CNNDM and XSum datasets. Following (Durrett, Berg-Kirkpatrick, and Klein 2016), we use the limited-length recall based ROUGE-1, ROUGE-2 and ROUGE-L on NYT, where generated summaries are truncated to the length of gold summaries. ROUGE scores are computed with the ROUGE-1.5.5.pl scriptwith -c 95 -r 1000 -n 2 -a -m arguments.

4 Results

We present our main results on the CNNDM dataset in Table 1. We compare our model against both extractive and abstractive systems. The first block summarizes the results for extractive systems. Lead3 is a baseline which simply takes the leading three sentences in a document as its summary. BertExt (Liu and Lapata 2019) employs BERT as encoder and predicts whether a sentence is a summary. MatchSum (Zhong et al. 2020) is the best performing extractive models, which formulates summarization as a semantic text matching problem using contrastive learning. The abstractive models are in the second block. PTGen (See, Liu, and Manning 2017) is a LSTM-based Seq2Seq model augmented with copy and coverage models. Large pre-trained language models mostly dominate summarization. BertSumExtAbs (Liu and Lapata 2019) is an abstractive model with encoder initialized with BERT and decoder randomly initialized. UniLM (Dong et al. 2019) is trained using language modeling and masked language modeling objectives. T5 (Raffel et al. 2020), PEGASUS (Zhang et al. 2020), BART (Lewis et al. 2020) and STEP (Zou et al. 2020) pre-train Seq2Seq transformers using different unsupervised text-to-text tasks. PEGASUS (Zhang et al. 2020) is trained by predicting gapped sentences (selected by some heuristics) in a document given the document with these sentences masked. Similar to BertSumExtAbs, the encoder of STEP is initialized from RoBERTa (Liu et al. 2019). BART + R3F (Aghajanyan et al. 2021) applies a trust region theory based fine-tuning method to BART. Our model is based on BART and therefore we also re-implement BART (BART⋆\star). These models above are single models. We also present the results of recent combination models in the third block. CTRLsum (He et al. 2020a) and GSum (Dou et al. 2021) combine a keywords extraction model (or an extractive model) with an abstractive model by taking the resulting keywords (or sentences) as additional input. SimCLS(Chen et al. 2020) and Refsum (Liu, Dou, and Liu 2021) train re-ranking models to rank multiple candidate summaries.

The fourth block includes results of our model SeqCo. As mentioned in Section 3.2, we can do contrastive learning between document and gold summary (i.e., SeqCo (λx−y\lambda_{x-y})), document and generated summary (i.e., SeqCo (λx−y^\lambda_{x-\hat{y}})) as well as gold summary and generated summary (i.e., SeqCo (λy−y^\lambda_{y-\hat{y}})). Note SeqCo (λ∗−∗\lambda_{*-*}) means that λ∗−∗>0\lambda_{*-*}>0 and all the other λ\lambdas equal to zero in Equation (15)We tune λx−y,λx−y^,λy−y^∈{0.5,1.0}\lambda_{x-y},\lambda_{x-\hat{y}},\lambda_{y-\hat{y}}\in\{0.5,1.0\} on the validation set when >0>0. We can see that SeqCo (λx−y\lambda_{x-y}), SeqCo (λx−y^\lambda_{x-\hat{y}}) and SeqCo (λy−y^\lambda_{y-\hat{y}}) all outperform BART⋆\star significantly (p<0.05p<0.05) measured by the ROUGE script, which demonstrates the effectiveness of our proposed contrastive methods. SeqCo (λy−y^\lambda_{y-\hat{y}}) outperforms all single models in comparison (first two blocks) and differences between them are significant w.r.t. the ROUGE script. We also observe that using generated summaries in contrastive learning leads to better performance (i.e., results of SeqCo (λx−y^\lambda_{x-\hat{y}}) and SeqCo (λy−y^\lambda_{y-\hat{y}}) are better), which is not surprising. Generated summaries are created dynamically during training and they might be more diverse than gold summaries.

It is also possible to employ multiple pairs of text for contrastive learning. Results on validation set with different combinations of text pairs are shown in Table 2. We obtain worse results with more than one pair of text in contrastive learning. Perhaps because the information learned using different pair of text is a bit redundant. We compared the results on the validation and test sets of the other two datasets and observed similar trends. Detailed numbers are shown in Appendix. We find best results are achieved by using a single similarity loss on all datasets except for the validation set of XSum, where SeqCo (λx−y\lambda_{x-y} + λy−y^\lambda_{y-\hat{y}}) and SeqCo (λx−y\lambda_{x-y} + λx−y^\lambda_{x-\hat{y}} + λy−y^\lambda_{y-\hat{y}}) outperform SeqCo(x-y) slightly. Given the fact that adding one more similarity loss increases around 30% training time and the observations above, we recommend using a single similarity loss. We probably need to encourage the “disagreement” between them (we leave this for future work). As mentioned in Section 3.2, we can also use decoder based encoding function fθDf_{\theta}^{\text{D}} (see the SeqCo (λy−y^D\lambda^{D}_{y-\hat{y}}) and SeqCo (λy−y^\lambda_{y-\hat{y}}) rows in Table 2) and we obtain worse results. It may because influencing the decoding during contrastive training is too aggressive. Therefore, we only report results of contrastive models on single pair of text (i.e., SeqCo (λx−y\lambda_{x-y}), SeqCo (λx−y^\lambda_{x-\hat{y}}) and SeqCo (λy−y^\lambda_{y-\hat{y}})) on NYT and XSum. Again in Section 3.2, we propose to employ multi-head attention based similarity modeling (see Equation (9) and (10)) rather than [CLS] based method (see Equation (11)). It also shows attention based similarity, which takes associations across two sequences into account, is better (see SeqCo (λy−y^\lambda_{y-\hat{y}}) and SeqCo (λy−y^\lambda_{y-\hat{y}}) w/ [CLS] rows in Table 2).

Results on NYT are shown in Table 3 and the trend is similar. RoBERTa-S2S is a transformer based Seq2Seq model with encoder initialized from RoBERTa (Liu et al. 2019) and its results are reported in (Zou et al. 2020). SeqCo (λx−y^\lambda_{x-\hat{y}}) outperforms BART⋆\star by +1.0 ROUGE-1, +0.8 ROUGE-2 and +1.0 ROUGE-L and the differences between them are significant measured by the ROUGE script. SeqCo (λx−y^\lambda_{x-\hat{y}}) obtains better results than all models in comparison. We again observe that using generated summaries in SeqCo are better than using gold summaries only.

Table 4 summarizes our results on the XSum dataset. BART⋆\star (our reimplementation) are better at ROUGE-1, but worse at ROUGE-2 and ROUGE-L compared to BART. SeqCo (λx−y\lambda_{x-y}) outperforms BART⋆\star significantly measured with the ROUGE script. Results of SeqCo (λx−y\lambda_{x-y}) are better than all previously published models except for PEGASUS (HugeNews) and Refsum. It is not entirely surprising, because PEGASUS (HugeNews) is trained on 3,800 GB news data (the same genre as the XSum dataset), while PEGASUS(C4) is pre-trained on the C4 dataset consist of text from 350M Web pages (750GB) and performs worse than PEGASUS (HugeNews). Refsum reranks outputs of PEGASUS (HugeNews). Note that the pre-trained transformer (i.e., BART) in SeqCo is trained on only 160 GB data, which also contains data in other domains rather than news data.

We do human evaluations on CNNDM, NYT and XSum with 100 documents each. We asked the participants to rank the outputs of different systems according to their faithfulness and the mean rank scores (lower is better) are shown in table 5. We employed (self-reported) native speakers to annotate our output summaries on Amazon Mechanical Turk. To further guarantee the annotation quality, we filter out the annotated assignments which were done less than two minutes (average time spent per assignment is 6 minutes). After the filtering process, we guarantee each document is annotated by three annotators. In CNNDM and NYT datasets, Seqco outperforms BART significantly. In XSum dataset, there are no significant differences among these systems. It may be because generated summaries in XSum are shorter, which are difficult for annotators to tell the differences. We calculate the ratios of agreement between annotators (i.e., ratio of all three annotators’ agreement and ratios of at least two annotators’ agreement) to measure the agreement for human evaluation. As shown in table 6, there are around 30% of summaries that all of 3 participants give the same annotations, and more than 90% of summaries obtained the same annotations by at least 2 annotators. In addition, the Fleiss’ Kappa scores are 0.329 on CNNDM, 0.313 on NYT and 0.364 on XSum, which demonstrate a fair degree of agreement. We believe the agreement between annotators is reasonable.

Analysis

Different from CNNDM and NYT, why does using generated summaries in contrastive learning perform worse on XSum? As shown in Table 7, it may because XSum is more abstractive (see the novel nngram statistics of Gold on the three datasets) and more difficult. As a result, the generated summaries are easier to have different meanings from their documents and gold summaries (at least in the early stage of training). Maybe that is the reason why the x−y^x-\hat{y} and y−y^y-\hat{y} objective is worse than the x−yx-y objective. CNNDM and NYT are less abstractive and the generated summaries could retain the main meanings more easily and are also more diverse (compared to gold summaries), which leads to the x−y^x-\hat{y} and y−y^y-\hat{y} objectives work better.

We can also see from Table 7 that SeqCo can either be more abstractive than BART or almost as abstractive as BART. To choose the contrastive objective, our suggestion is 1) for the datasets whose summaries are highly abstractive, choose the x−yx-y pair as the contrastive objective; 2) for less abstractive datasets (the case for most datasets), choose either x−y^x-\hat{y} or y−y^y-\hat{y} as the contrastive objective. As far as we observed, the performance of x−y^x-\hat{y} and y−y^y-\hat{y} are similar.

Ablation Study

We list the ablation results on three datasets in the appendix A. We compared single similarity loss v.s. multiple similarity losses on the validation and test sets and observed the similar trends. We find best results are achieved by using a single similarity loss on all datasets except for the validation set of XSum, where SeqCo (λx−y\lambda_{x-y} + λy−y^\lambda_{y-\hat{y}}) and SeqCo (λx−y\lambda_{x-y} + λx−y^\lambda_{x-\hat{y}} + λy−y^\lambda_{y-\hat{y}}) outperform SeqCo(x-y) slightly. Given the fact that adding one more similarity loss increases around 30% training time and the observations above, we recommend using a single similarity loss.

Example Outputs

Some example outputs of SeqCo and BART⋆\star are also listed in appendix B. In conclusion, BART sometimes miss some important points, while SeqCo can do better.

Conclusions

In text summarization, a document, its gold summary and model generated summaries can be viewed as different views of the same meaning representation. We propose SeqCo, a sequence level contrastive learning model for text summarization, which intends to minimize distances between the document, its summary and its generated summaries during training. Experiments on three summarization datasets (CNNDM, NYT and XSum) show that SeqCo consistantly improves a strong Seq2Seq text generation model. In the future, we plan to extend SeqCo in the multi-lingual or cross-lingual text generation tasks. We observed in experiments that using multiple contrastive objectives did not improve the results. We are interested in developing methods for regularizing different contrastive objectives.

References

Appendix A Ablation Results

We list the ablation results on the validation and test set for three datasets in table 8, 9 and 10. We compared single similarity loss v.s. multiple similarity losses on the validation and test sets of the other two datasets and observed similar trends with CNNDM. We find best results are achieved by using a single similarity loss on all datasets except for the validation set of XSum, where SeqCo (λx−y\lambda_{x-y} + λy−y^\lambda_{y-\hat{y}}) and SeqCo (λx−y\lambda_{x-y} + λx−y^\lambda_{x-\hat{y}} + λy−y^\lambda_{y-\hat{y}}) outperform SeqCo(x-y) slightly. Given the fact that adding one more similarity loss increases around 30% training time and the observations above, we recommend using a single similarity loss.

Appendix B Examples

We list some examples of generated summaries and gold summaries in table 11 and 12 on the test set of CNNDM, where we can compare the outputs of BART and SeqCo. In conclusion, BART sometimes miss some important points, while SeqCo can do better.

In the document of Table 11, an important point is that Anne Frank and her older sister died earlier than previously believed. The output of BART doesn’t mention this directly but describes two dates of their death, which is lack of the main idea and confusing. SeqCo points out this emphasis in the first sentence and then further explains, which is quite consistent with the meaning expressed by the gold summary.

For the document in Table 12, the most important thing is that Schuller died. BART focuses on what did do and doesn’t mention the death, while the gold summary and SeqCo both describe his death in the first sentence, and then list some famous deeds in his lifetime. Death is more important than the deeds in this document.