PSP: Pre-trained Soft Prompts for Few-Shot Abstractive Summarization

Xiaochen Liu, Yang Gao, Yu Bai, Jiawei Li, Yinan Hu, Heyan Huang, Boxing Chen

Introduction

Given the high labor-costs of obtaining quality abstractive summaries, few-shot abstractive summarization is very demanding and highly challenging. A widely accepted paradigm for almost all NLP tasks is to fine-tune the entire set of parameters for a large pre-trained language model to suit the target task (Liu and Lapata, 2019; Liu et al., 2020).

However, the fine-tuning with few-shot examples usually leads to disappointing results, especially with generation tasks like abstractive summarization (Fabbri et al., 2020; Yu et al., 2021). The likely outcome is an overfit model. Further, for every specific task, a large number of pre-trained parameters need to be updated and stored, which is not efficient to use.

Pre-trained language models are few-shot learners, i.e., GPT-3 (Brown et al., 2020) that surprisingly perform generation tasks from a few examples without any further gradient updates. Although it lacks a rigorously theoretical proof, prompt learning inherits the few-shot property (Li and Liang, 2021; Schick and Schütze, 2020; Jin et al., 2021; Liu et al., 2021). Commonly, this type of learning is considered to retrieve relevant knowledge from frozen language models, only tuning continuous prompts to quickly adapt to new tasks with very few examples.

More recently, Prompt Tuning (Lester et al., 2021) has received much attention. With large frozen language models (say, >>10 billion parameters), Prompt Tuning simply adds a tunable soft prompt to the input of the encoder, achieving results that are comparable to full-model tuning. Yet, our empirical results, in Section 2, demonstrate that Prompt Tuning for abstractive summarization yields simply abysmal performance. Prefix-Tuning (Li and Liang, 2021) extends the use of prompt learning in the natural language generation area. With this technique, continuous prompts are applied to every layer of the pre-trained model and even shows increase in few-shot generation tasks over fine-tuning. Yet the training process is not stable and updates are required that add to the memory and training costs.See more related work in Section 5.

Given the shortcomings of these two methods, we have developed a soft prompts tuning method that is specifically designed for summarization. The structure is given in Figure 1. The method is capable of performing few-shot language generation task (i.e., abstractive summarization) with an efficient amount of training parameters. Prompt tokens are added before the decoder input tokens to guide the generation process toward the target summary. Moreover, we have designed three inner prompts – interval, sequential, and fixed-length – one of which is placed among the source input tokens. The aim is to capture the structure in the source document and aid in understanding its semantics, so as to better prompt the model to generate document-related content. Each kind of inner prompts focuses on different semantic units (e.g., phrases, sentences, and etc.), differentiating important units from non-informative ones. To bolster the summarization ability of the model and assist the prompts to understand the documents, prompt pre-training is performed before the tuning process, and leveraged by self-supervised pseudo data. As a last step, all the prompts are fine-tuned with few-shot training examples. Experiments conducted on two commonly used datasets - CNNDM (See et al., 2017) and XSum (Narayan et al., 2018) - demonstrate that our method outperforms full-model tuning under few-shot settings only with 0.1% of the parameters. It also surpasses naive Prompt Tuning by a large margin. Our model also yields a performance competitive to Prefix-Tuning with 3% of the trainable parameters. A detailed analysis shows that the designed prompt-pre-training phase and the inner prompts are effective for few-shot text summarization. Thus, the major contributions of this work include : 1) A novel soft prompt architecture for few-shot abstractive summarization. With the well-designed prompts in embedding layer, our model fulfills the task effectively and efficiently; 2) It is necessary to perform prompt pre-training strategy which benefits soft prompts model for few-shot summarization and shows excellent zero-shot capabilities; 3) Experiments that investigate the effect of different prompts by probing the attention weights. The results show our model is able to: extract knowledge from the encoder language model; understand the discourse in the document; and guide the decoder language model to generate fluent summaries.

Pilot Experiments

In a pilot study, we experimented with using Prompt Tuning under 300-shots settings to find reasonable clues as to how to design summary-prompts for the task. Our findings follow.

Consider an encoder-decoder language model pθ(y∣x)p_{\theta}(y|x) based on the Transformer architecture (Vaswani et al., 2017) (e.g., BART (Lewis et al., 2020)) and parameterized by θ\theta. To conduct a few-shot summarization task, we have some few-shot training pairs of a document X={x1,x2,…,x∣X∣}X=\{x_{1},x_{2},\dots,x_{|X|}\} and a corresponding summary Y={y1,y2,…,y∣Y∣}Y=\{y_{1},y_{2},\dots,y_{|Y|}\}. Specifically, we divided XX into different subsets with sentencesNote that, throughout this work, a “sentence” can be an arbitrary span of contiguous text (e.g., fixed length of 10 tokens), or an actual linguistic sentence. as our unit, X={x11,…xji,…,xmn}X=\{x^{1}_{1},\dots x^{i}_{j},\dots,x^{n}_{m}\}, where xjix^{i}_{j} denotes the jthj_{\rm th} token in the ithi_{\rm th} sentence.

First, original Prompt Tuning is applied by concatenating a series of prompt tokens Pen{P}_{en}, parameterized by θpen\theta_{p_{en}}, to the encoder input Xen={e11,…,eji,…emn}X_{en}=\{e^{1}_{1},\dots,e^{i}_{j},\dots e^{n}_{m}\}, where ee represents the embedding of each token (the leftmost structure in Figure 1). The gradients are backpropagated through the prompts and the weights θ\theta of language model are frozen Lester et al. (2021). In this way, the model maximizes the likelihood of the output YY:

The result of original Prompt Tuning is shown on the first line in Table 1, where we see it severely underperforms versus full-model tuning. In further experiments, we added a series of prompts PdeP_{de} to the decoder inputs XdeX_{de} following the generation pθ;θpde(Y∣Xen,Pde)p_{\theta;\theta_{p_{de}}}(Y|X_{en},P_{de}). Here, we found the results to be even worse than the last.

For generation-based tasks, prompts in both the encoder and decoder are equivalently useful. Therefore, our model employs a combination of the two series of prompts mentioned above, and generates YY conditioning on XenX_{en}, PenP_{en} and PdeP_{de}:

The result on the third line in Table 1 again verify our hypothesis. Prompts across the encoder and decoder even achieve comparable results with full-model tuning under few-shot settings. This verifies two things for us. First, prepending simple prompts to only the input embedding layer is effective and efficient for few-shot abstractive summarization. Second, prompts across the encoder and decoder are both necessary for generation tasks.

Lack of Attention on the Document

We further explored the encoder-decoder attention to investigate the effect of the prompts and freezing the language model. From Figure 2, we find the generating output is mainly focused on the soft prompts to come with little attention given to the document itself. This outcome is detrimental to summarization that requires to understand the semantics and inner discourse structure of documents Wang et al. (2019). Without the associations of target summaries and source documents, it is impossible to obtain high-quality summaries using current prompt architectures.

From Figure 2, we can observe that prompts in the encoder and the ones in decoder are consistently and directly associated with each other. We speculate that the mechanism is that encoder prompts retrieve relevant knowledge from the frozen encoder language model as a document representation, and decoder prompts copy the encoder’s behaviour, guiding the decoder language model to generate text.

Method

In light of our findings about the current architectures, we developed a new architecture of pre-trained soft prompts, for few-shot abstractive summarization called PSP. The framework includes continuous prompts across the encoder and decoder inputs, as well as inner-prompts to capture the dependencies between documents and target summaries. To better understand a given document, we add a prompt pre-training process before few-shot tuning. It also brings a good initialization for the prompting. The overall architecture and training scheme are illustrated in Figure 3.

As mentioned in Section 2, in the training phase of current architectures, PenP_{en} is responsible for extracting knowledge from the encoder’s frozen language model as a document representation. Meanwhile, PdeP_{de} mostly copies the behavior of PenP_{en} and guides the frozen decoder’s language model to generate fluent text as a summary.

To strengthen the model’s ability to understand a document, the dependencies and attentions given to the source document need to be embodied in the prompt architecture.

2 Inner-Prompts for Document Understanding

To achieve our goal, we propose the notion of adding inner-prompts within the source document, denoted as Pin={pin1,pin2,…,pinn}P_{in}=\{p_{in}^{1},p_{in}^{2},\dots,p_{in}^{n}\} with the parameters θPin\theta_{P_{in}} to be updated. Each pinip_{in}^{i} corresponds to a single sentence. These inner-prompts are added to the corresponding token embedding, which gives rise to a new Xin′X^{\prime}_{in}:

We believe that by prompting different semantic units (e.g., sentences, phrases, etc.), more attention can be given to understanding the document’s discourse. Furthermore, the inner-prompts help the model to quickly interpret the document by strengthening the associations between outputs and documents. What follows are three different strategies for incorporating the three different inner-prompts. Note that there is more discussion on this point in Section 4.2.

Following Liu and Lapata (2019), the interval inner-prompts comprises two inner-prompt tokens are assigned to each sentence sentisent_{i}, depending on whether ii is odd. Specifically,

In this way, the model can identify important sentences to encode the document at sentence level.

Sequential

To highlight the complex discourse structure of documents, sentence positions need to be considered. Therefore, different tokens are set in sentences by their sequences, formulated as:

Fixed-length

To discover more fine-grained semantic units, a text span with a fixed length kk is manipulated into a new “sentence” and a corresponding sequential token is assigned to it. Further, prompts are assigned to the newly divided sentences [sent1sent_{1}, sent2sent_{2}, …, sentnsent_{n}], as {pin1,pin2,…,pinn}\{p_{in}^{1},p_{in}^{2},\dots,p_{in}^{n}\}. Figure 4 illustrates some examples where the above strategies have been used.

3 Self-supervised Prompt Pre-training

To improve ability of the prompts to understand the documents and to help the model to adapt to the summarization tasks, soft prompts are further pre-trained on the corpus using summarization-oriented self-supervised objectives. Doing this also means that the prompts are well initialized for few-shot tuning.

We tested two strategies for constructing the self-supervised data. Each strategy was designed to suit a particular type of writing bias in the document. These are “lead” and “gap sentences generation”.

Lead bias is common in news articles, which usually follow an inverted pyramid structure where the first few sentences contain the most salient information (See et al., 2017; Yang et al., 2020). With this type of bias, we initially select the first three sentences as our target summary, and treated the rest of the document as the source text. With this type of prompt pre-training process, the model was able to infer the salient information based on the remaining text.

GSG

Gap sentences generation applies to all documents that do not follow the lead bias structure (e.g., XSum (Narayan et al., 2018)). The strategy used here follows Zhang et al. (2020) , where we used ROUGE1-F1 (Lin, 2004) between each sentence xix_{i} and the rest of the document as a proxy for the principal score, si=rouge(xi,D∖{xi}),∀is_{i}=rouge(x_{i},D\setminus\{x_{i}\}),\forall{i}. The top-mm most important sentences were selected according to sis_{i}, and removed from the document. Then these mm sentences are concatenated in the same order as the original text in the form of a pseudo summary. The remainder of the text is treated as a pseudo document.

With the constructed data, our designed prompts can be pre-trained and further tuned with few-shot examples.

4 Training Objective

The model is trained with maximum likelihood estimation (MLE). Given a ground-truth summary Y=[y1,y2,...,y∣Y∣]Y=[y_{1},y_{2},...,y_{|Y|}] for an input passage XX, the objective is to minimize the negative log-likelihood of the target word sequence:

Note that only these prepended-prompts parameters (θpen\theta_{p_{en}}, θpde\theta_{p_{de}}) and the inner-prompts parameters (θpin\theta_{p_{in}}) are optimized, the language model parameters (θ\theta) are all frozen.

Experiments

We experimented with the CNN/DailyMail (CNNDM) dataset (Hermann et al., 2015) and the XSum dataset (Narayan et al., 2018). We chose these datasets because they differ in abstraction level and text length, which helps to show the generalization ability of our results.

We constructed the self-supervised pre-training data for CNNDM with Lead, and for XSum with GSG. We show details in Section A.1 in the appendix. Given that the lead bias structure exists only in some domain-specific datasets, we also conducted experiments to demonstrate the universality of the GSG to construct pseudo-data. The results are shown in Section A.3 in the appendix. Our few-shot training set DtrainD_{train} contained 300 document-summary pairs randomly sampled from the original training data. To tune the hyper-parameters and select the best checkpoint, we composed a validation set DdevD_{dev} from the original validation data. Here, we were careful to ensure that ∣Dtrain∣=∣Ddev∣\lvert D_{train}\rvert=\lvert D_{dev}\rvert so that it fit into a true few-shot learning setting, following Perez et al. (2021). Since few-shot learning may have high variance, we sampled the examples with 5 different random seeds. We used the original test set to report our results, including the mean value and the standard deviation. Table 2 shows the statistics of the pre-processed corpus.

Setup

The base version of BART was used in our work. Following Lester et al. (2021), we used 100 prompt tokens for both the encoder inputs and the decoder inputs. These prompts were randomly initialized from the set of vocabularies. The sequential and fixed-length inner-prompts require a maximum number. Hence, we counted the number of sentences in each document and divided the results into two groups – the 85% with the least sentences (Group A) and the 15% with the most sentences (Group B)We made our division at 85% to ensure all embeddings of inner-prompt tokens could be fully trained, because sentences after the nn-th only exist in 15% of the data.. We then set the number of prompts to the most number of sentences in Group A plus one, i.e., n+1n+1. For CNNDM, that number was 61 and, for XSum, it was 33. In this way, one inner-prompt token was assigned to each sentence up to nn. For the excessively long documents in Group B, the text after nn sentences was assigned an n+1n+1-th token. Further, we drew from a normal distribution N(0,0.05)\mathcal{N}(0,0.05) to initialize the inner-prompt embeddingsMore information about implementation details are shown in Section A.2 in the appendix.. Taking CNNDM as an example, all the tunable parameters that need to be stored amount to only 2×1052\times 10^{5}. This is compared to the (1.4×1081.4\times 10^{8}) parameters of full-model tuning. That equates to around 0.1% of the parameters for each dataset that need to be tuned and stored.

Evaluation Metrics

We adopted ROUGE Lin (2004) to measure the quality of the summaries produced in our experiments. The F1 scores for ROUGE-1, ROUGE-2, and ROUGE-L between the ground-truth and the generated summaries are each reported.

Baseline Models

We compared PSP to: Prompt Tuning (Lester et al., 2021), which only concatenates soft prompts into the encoder input; Prefix Tuning (Li and Liang, 2021), which adds a prefix to all the encoder layers, cross-attention layers, and the decoder layers; and Full-Model Tuning, which does not have any prompts and fine-tunes all the parameters of the pre-trained language model.

1 Experimental Results of Our Method

Table 3 presents the results of all PSP variants and baselines across CNNDM and XSum datasets. With the exception of the ROUGE-2 and ROUGE-L scores for the Prefix-Tuning on the CNNDM dataset, our proposed PSP, outperforms the others. However, PSP delivered a competitive result with only 3% of the parameters, which is an acceptable place to start. To our surprise, we observe that 50% of PSP’s results surpass the full-model tuning, especially on XSum, as underlined in the table. Besides, results on the PPL metric show that PSP can generate more fluent summaries than other models. These results indicate that fine-tuning large language models is not necessarily a good or efficient idea with few-shot generation. It also shows that soft prompts with frozen language models are effective for few-shot abstractive summarization. Moreover, it statistically verifies that PSP with its three inner-prompt strategies is effective.

We gave an overall comparison to baseline models on effectiveness and memory-efficiency, evaluated by ROUGE and the number of parameters, respectively. The results are shown in Table 4. Prompt Tuning has the least number of parameters, while its capacity is limited to this and lacks control over the decoder side, hence it can not perform natural language generation tasks well. We can see that substantial gains are made when going from vanilla Prompt Tuning to PSP. However, even if Prefix-Tuning is nearly thirty times more parameters than ours, there is either a marginal improvement or even performance decrease on some metrics. Besides, Prefix-Tuning relies on reparameterization tricks to stabilize the training, i.e., adds a MLP with large number of parameters to the training stage. Our method provides the best effectiveness-efficiency trade off, and outperforms full-model tuning with only 0.1% parameters, and presents competitive results against Prefix-Tuning with 3% parameters.

Human Evaluation

We conducted a human evaluation study. To this end, we randomly selected 20 instances from the test set of each dataset. Ten graduate students with high levels of fluency in English were asked to assess the generated summaries and golden summaries from independent perspectives Wang et al. (2021): Informativeness (how much useful information does the summary provide?), Relevance (how well does the summary reflect the input document?), and Fluency (how grammatically correct are the summary sentences and how easy are they to read?). Scoring followed the Best-Worst Scaling method (Kiritchenko and Mohammad, 2017). Participants were asked to select the best and worst summaries from each perspective. The scores were computed as the percentage of times a summary was chosen as the best minus the times it was selected as the worst. The scores ranged from -1 (worst) to 1 (best). Results are shown in Table 5. Qualitatively, we show several examples generated by different models and the reference in Table 14 and Table 15 in the appendix. Compared with all baselines, the summaries generated by PSP are always more fluent and relevant to the source document, consistent with the results of human evaluation. Further more, we found summaries generated by PSP and Prefix-Tuning are always similar in sentence patterns and expressions. However, Prefix-Tuning tends to generate texts shorter than PSP, which often leads to lack of information.

Selection of fixed length k𝑘k.

As shown in Table 3, PSPFixed-k performs consistently well on both datasets. So we further explored the influence of different length kk, i.e., k=5,10,15,30k=5,10,15,30, for inner-prompt tokens of the PSPFixed-kThe average number of tokens per sentence in both datasets was about 18, so we did not consider fixed lengths of 20, for its similarity to the PSPSequential.. Table 6 presents the results of the variants on XSum. We observe the segmented spans with 10 tokens achieve the best performance. Interestingly, it can be induced that, to understand a document, it is possible to reorganize the sentence into several semantic units, where the number of the tokens is 10 on average. We also report results of different kk on our validation set in Table 6. The ranking is consistent with the test set. From a practical perspective, when applying PSP to a new dataset, we can choose the best kk based on the validation set.

2 Analyses on Soft Prompts

According to Figure 2, we further present the encoder-decoder attention distribution of the PSP. The comparison visualization is shown in Figure 5. We find the following enhancement of our model by introducing the inner prompts. First, the PSP model strengthens the associations between the encoder prompts and the decoder prompts compared to the original model. Second, the soft prompt PenP_{en} has more opportunities to be related to the output YY, indicating the semantic relations between them. Third, the output YY assigns more attention to the source document XX. This suggests that the hidden structure of the document is emphasized, increasing the capability of understanding its semantics. As such, these prompts can properly elect salient information from the document and prompt the model to generate the output.

Do inner prompts assist the model to understand the content of documents or simply increase the model’s capacity?

Instead of using inner-prompts, we prepended additional tunable tokens (i.e. 150 tokens) in front of the encoder and the decoder inputs. Comparison results are shown in Table 7. Despite the larger capacity, soft prompts with 150 tunable tokens before the input performed the worst, denoted as soft prompts (en.&de., 150). This suggests the inner-prompts with a few parameters do help to understand the document by prompting the structures, rather than simply add more trainable parameters to increase the model’s capacity.

Further insight on soft prompts across the encoder and the decoder.

To verify our hypothesis that the decoder prompts largely copy the behaviour of the encoder prompts, we shared similar embeddings of the soft prompts before the encoder and the decoder. In Table 8, we observe the Soft prompts (en.&de., shared) and (en.&de., separate) almost perform identical results. Although the parameters are only half of the original model, the performance consistently remains competitive. This shows that the shared prompts can extract important information from the document and further guide the language model to generate consistently good summaries more efficiently.

3 Analysis on Few-shot and Zero-shot Summarization

To examine the performance of different methods under few-shots, we further randomly sampled number of {50, 100, 200} as the settings. Figure 6 reports a more detailed overview of all models’ performance across a range of different few-shots. The ROUGE scores of our model generally outperform other baselines and remain steady across different scenarios. Especially, the PSP with only 50 examples receives the most significant improvements, while the Prefix-Tuning doesn’t even work (tuning based on BARTbase) possibly due to its instability of the model. Moreover, we report the results of zero-shot on XSum in Table 9. Benefiting from the knowledge gained in the pre-training phase, our model shows a significant advantage of zero-shot adaptation in generating quality summaries.

4 The Performance of Pre-training on Prefix-Tuning

A crucial strategy for PSP is the pre-training of soft prompts. To give a fairly comparison, we performed prefix pre-training for Prefix-Tuning in the same way with the PSP. The results are shown in Table 10. We can find that the Prefix model obtains improvements on the XSum dataset after adopting the pre-training strategy, but underperforms the original one on the CNNDM dataset. It indicates that Prefix-Tuning shows limited potential compared to our model. We induce that the pre-training for Prefix-Tuning raises over-fitting risk due to its sensitivity to different data or parameter settings.

5 Ablation Study

We conducted experiments to examine the effectiveness of the major components of our model, and Table 11 shows the ablation results across the two datasets. We observed both the prompt pre-training operation and the inner-prompts component contribute to the main model. Notably, with the removal of each component, the model becomes considerably unstable, indicated by the variance shown in the ablation results. Comparably, prompt pre-training in our model accounts for more importance on the XSum dataset whose summaries have a higher abstract level (we assume it’s more “difficult”) than the CNNDM. In sum, these two components support the performance and stability of our model in terms of summarization adaption (by prompt pre-training) and structural documents understanding (by inner-prompts).

Related Work

In practical application scenarios, the lack of manual constructed document-summary pairs or labeled data makes data-driven neural models performs badly Hu et al. (2021, 2020). Fabbri et al. (2020) condense characteristics of the target dataset into Wikipedia data to construct pseudo-summaries. Bražinskas et al. (2020) introduce plug-in networks to reproduce characteristics of the target dataset with only a small set of labeled examples. Bai et al. (2021) conduct cross-lingual summarization in a low-resource setting. Yu et al. (2021) design the second phase of pre-training on large-scale generative models before fine-tuning. In this paper, we construct pseudo-summary corpus with heuristic rules, providing a better parameter initialization for soft prompts under few-shot settings. More importantly, we design summarization-oriented soft prompts to help the model produce few-shot summaries.

Prompt Learning

The emergence of GPT-3 (Brown et al., 2020) introduces the concept of “prompting”. One only needs to assemble a task description and few examples into a prompt, and then prepend it to the task input. With the large-scale frozen parameters, a pre-trained model can generate the output without any task-specific tuning. However, task description is error-prone while there is no unified, explicit, and effective way to build these hard prompts manually (Logan IV et al., 2021). Hence, several works (Gao et al., 2020; Jiang et al., 2020; Shin et al., 2020) are proposed to generate prompts automatically, but they all restrict prompts to discrete spaces. These discrete prompts are less expressive and sub-optimal. To overcome the shortcomings of hard prompts, Li and Liang (2021) propose “Prefix-Tuning”. This method only tunes prefix activation prepended to all transformer layers, and keeps the LM parameters frozen. To further simplify, Prompt Tuning (Lester et al., 2021) only prepends tunable tokens to the encoder input, and keeps all other parameters frozen. Logan IV et al. (2021) and Gu et al. (2021) propose to use pre-training to boost the low performance of Prompt Tuning for few-shot learning. In this work, we fit the structure of Prompt Tuning to text generation models, proposing encoder prompts, decoder prompts, and inner prompts. We successfully apply prompt tuning methods to few-shot abstractive summarization task.

Conclusion

In this paper, we present a novel pre-trained soft prompts architecture (PSP) specifically designed for few-shot abstractive summarization. We design continuous input embeddings across an encoder and a decoder alongside several kinds of inner-prompts placed in the text, assisting the model better to understand documents and guide accurate generation. Empirical results find the necessity of using prompt pre-training for few-shot/zero-shot abstractive summarization. Extensive experiments and analyses show that the proposed PSP provides the best effectiveness-efficiency trade off among all the baseline methods.

Acknowledgments

The research presented in this publication was sponsored by CCF Fund For Young Scholars, and Joint Funds of the National Natural Science Foundation of China (Grant No. U21B2009).

References

Appendix A Appendix

We constructed the pseudo data for CNNDM with Lead. We also conducted a simple data cleaning procedure to the self-supervised pre-train corpus. First, we cleaned away irrelevant information, such as media names, reporter names or dates from the summaries. Second, for those summaries with less than 50 tokens, we iteratively collected the first sentence of the remaining text to the pseudo summary, until the length of summary reaches 70. This procedure was set up to prevent the target text from being too short to form a meaningful summary. Third, for those samples in which the source document is shorter than its summary, we filtered them out.

For XSum, we constructed the pseudo data for pre-training following GSG. The top-1 most important sentence was selected as the pseudo summary. Then we filtered out those pseudo summaries that are not relevant enough to the pseudo passages. In particular, we leveraged hand-written summaries in the few-shot dataset to determine the filtering threshold of pseudo data. We calculated the ROUGE-1 F1 between each ground-truth summary and its corresponding passage, represented as Ri{Ri}. Then we calculated the mean and variance of Ri{Ri}: ϵ=1n∑i=1nRi\epsilon=\frac{1}{n}\sum_{i=1}^{n}Ri, σ2=1n∑i=1n(Ri−ϵ)2\sigma^{2}=\frac{1}{n}\sum_{i=1}^{n}(Ri-\epsilon)^{2}, and ϵ−σ2\epsilon-\sigma^{2} was used as a lower-bound threshold to filter out low quality pseudo data. For those pseudo samples where ROUGE1-F1 between the pseudo summary and the pseudo passage is lower than the threshold ϵ−σ2\epsilon-\sigma^{2}, we filtered them out. Finally, we conducted pre-training on our soft prompts with these filtered pseudo-data. Table 12 shows the statistics for the pre-training data corpus.

A.2 Implementation Details

We first split sentences with the Stanford CoreNLP toolkit (Manning et al., 2014), and the input documents were truncated to 1024 BPE tokens. We adopted BART-base for all the experiments. Our implementation was based on the Hugging Face Transformer models (Wolf et al., 2020). We used a mini-batch size of 8 with a gradient accumulation for 10 iterations. We used Adam optimizer with momentum β1\beta_{1} = 0.9, β2\beta_{2} = 0.998 and noam decay. In the stage of pre-training, the peak value of learning rate was 1e-3, and we set the warm up ratio to 10%. During fine-tuning, the peak value of learning rate was 3e-4, and we set the warm up steps to 100 with 400 epochs. In the decoding stage, we used beam search with a beam size of 4. The decoding process will not stop until an end-of sequence (EOS) token was emitted or the length of the generated summary reached to 256 tokens. All models were trained on 4 TITAN RTX GPUs.

A.3 The Universality of GSG to Construct Pseudo-data

To demonstrate the universality of using the GSG method to construct pseudo-data for prompt pre-training, we conducted a complimentary experiment to testify its effect on the CNNDMWe do not conduct ablation experiments on XSum, as there is no “ lead bias” in this dataset. So it is inappropriate to take the first sentences of the passage as the pseudo summary.. Specifically, we selected m=3m=3 important sentences. Results in Table 13 indicate that the PSP model pre-trained by GSG is equally effective with the original PSPLead, showing that the GSG can be universally employed to pre-train soft prompts for abstractive summarization.