PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, Dacheng Tao
Introduction
Fine-tuning the pretrained language models (PLMs) has become the de facto standard for natural language processing (NLP), and has achieved remarkable success in numerous downstream tasks . The dominant fine-tuning manner, i.e., model-tuning, is tuning the entire pretrained model parameters for each downstream task, which is prohibitively expensive especially for current large-scale PLMs. Hence, various parameter-efficient fine-tuning approaches are further explored , among which prompt-tuning has attached great attention. Specifically, prompt-tuning refers to tuning with the soft prompt that is a set of trainable parameters added to the PLM. Different from the model-tuning that requires characterizing all parameters of PLM for each downstream task, prompt-tuning only trains the task-specific soft prompt while keeping all pretrained parameters of PLM fixed.
Prompt-tuning can achieve competitive performance against model-tuning when the PLM exceeds billions of parameters , but there are still some gaps between prompt-tuning and model-tuning at smaller PLM scales , which can also be observed from our empirical results in Figure 1. Hence, it is crucial to explore how to boost the performance of prompt-tuning across all scales of PLMs. Intuitively, a consensus is that transferring knowledge from an intermediate task to the target task can improve the target performance . Inspired by this, recent works involve leveraging transfer learning in the context of prompt-tuning, i.e., prompt transfer (PoT). Specifically, the vanilla PoT approach first trains a prompt on one or more source tasks, and then employs the trained source prompt to initialize the prompt for a target task. Lastly, this prompt is further tuned on the target task to obtain final task-specific prompt. Although such a vanilla PoT approach can offer some improvements over prompt-tuning, there are several problems that hinder the universality and applicability of PoT. In particular, the first is the performance of PoT is sensitive to the similarity between source and target tasks, i.e., PoT relies on the similarity metric to retrieve the similar source tasks for a given target task, but in our pre-experiments (Section 5), we observe that the prior metrics (e.g., cosine similarity of prompt embeddings ) fall short in distinguishing the task relationships and hardly make principled choices about which source tasks to use. Furthermore, while given a similar source task, the vanilla PoT approach might still perform poorly, as the second problem is directly fine-tuning the prompt initialized with source prompt on target task will lead to catastrophic forgetting of knowledge learned from source task.
To address aforementioned problems, we propose a new metric to better predict the prompt transferability, and improve the Prompt trAnsfer via kNowledge DistillAtion (we name this approach as PANDA) respectively. Specifically, different from the prior metric that simply uses the similarity of prompt parameters as prompt transferability, our proposed metric first maps the source/target tasks into a shared semantic space to obtain the task embeddings based on the source/target soft prompts, and then measures the prompt transferability by the similarity of corresponding task embeddings. On the other hand, regarding the second problem, our PANDA approach introduces the knowledge distillation technique to transfer the knowledge from source prompt to the target prompt in a subtle manner, thus alleviating the problem of catastrophic forgetting effectively. In practice, PANDA first uses the PLM with the source prompt as teacher network and the PLM with randomly initialized prompt as student network. Then, the student network is trained using the supervision signals from both ground truth labels in target task and soft targets predicted by the teacher network. Note that we only update the parameters of student prompt, while keeping others fixed. Furthermore, to adaptively control the knowledge transfer in our PANDA approach, we use the prompt similarity predicted by our metric as the balancing factor between two supervision signals for each source-target pair.
We conduct a large-scale study using the 189 combinations of 21 source and 9 target datasets across 5 scales of PLMs. Qualitative and quantitative analysis shows that our proposed metric not only performs better to distinguish different task relationships, but also aligns with the transfer performance accurately. Moreover, results of PoT demonstrate that our PANDA achieves significant improvements over vanilla PoT across all tasks and model sizes (up to 24.1% average scores in some scenarios), and makes prompt-tuning obtain competitive and even better performances than model-tuning in the full-data scenario. In summary, our contributions are: 1) we recast vanilla PoT with PANDA, a novel prompt transfer approach, which leverages knowledge distillation to alleviate the catastrophic forgetting of knowledge learned from source task; 2) we propose a new metric to predict the prompt transferability of source-target task pairs; 3) extensive and systematic experiments on 189 pairs of source-target tasks across 5 scales of PLMs prove the effectiveness of our proposed methods.
Related Work
In recent years, we have witnessed numerous pretrained lanugage models (PLMs) that achieved tremendous success in the community of NLP . However, with the model size continues increasing, it becomes more costly and impractical to apply these large-scale PLMs to downstream tasks, as the current dominant fine-tuning approach (i.e., model-tuning) needs to tune all pretrained parameters for each task. To address this issue, researchers turn to focus on other parameter-efficient methods for efficiently fine-tuning the PLMs, which aim to only update small parts of model while keeping others fixed , or design and train some task-specific modules (e.g., adapters and low-rank structures ). Among these methods, prompt-tuning, which adds some textual prompt in the input and helps the PLM to produce some desired output text directly, has attracted great attention recently. Earlier works focus on exploring the discrete prompt, such as manually-designed prompt and automatically-searched prompt , which are almost hard tokens and sensitive to the prompt itself, i.e., minor changes in the prompt might lead to significantly different performance. Thus, more recent efforts attempt to investigate soft prompt that is a set of trainable embeddings/parameters and can be updated with task-specific supervision. Our work aims to analyze the effect of soft prompt, thus we refer to the prompt-tuning as that with soft prompts in this paper, instead of discrete hard prompts.
Lester et al. show that prompt-tuning can achieve comparable performances to model-tuning in the full-data scenario when the PLM has billions of parameters . However, the performance of prompt-tuning on smaller PLMs is still sub-optimal and is significantly sensitive to the prompt initialization . Hence, PoT is further proposed, which first learns a soft prompt on source tasks and then uses it to initialize the prompt for a target task. In our study, we argue that such a vanilla PoT approach might lead to the problem of “catastrophic forgetting” during directly tuning the source prompt on target tasks, especially when source and target tasks are significantly dissimilar.
Knowledge Distillation.
Knowledge distillation (KD), which aims to extract the “dark knowledge” from a teacher network to guide the training of a student network, has emerged as an important technique for model compression and transfer learning . Specifically, the student network is trained with the supervision signals from both ground-truth labels and the teacher network. The common way to leverage the supervision information of teacher model is to match the outputs of teacher and student models by minimizing the distance of the output distribution . KD plays an important role in the transfer learning and has the powerful ability to transfer knowledge, which has been proved in many fields of NLP . Thus, an intuitive idea is to employ the KD technique to tackle the problem of “catastrophic forgetting” in PoT. To this end, we introduce the KD into PoT for better knowledge transfer and boost the performance of prompt-tuning.
Method
Prompt Transfer.
Prompt-tuning is a parameter-efficient approach and shows competitive performance against model-tuning in large-scale PLMs, but raises two challenges. The first is the performance drop with the decrease of PLM scales, as shown in Figure 2 (Left), prompt-tuning performs much worse than model-tuning when the PLM is smaller. Another challenge is that prompt-tuning is sensitive to the initialization of prompt parameters . Figure 2 (Right) shows the results of the prompt initialized with different manners on various tasks. Regrading these challenges, Vu et al. propose a new transfer learning approach in the context of prompt-tuning, i.e., prompt transfer (PoT). As illustrated in Figure 3 (Left), vanilla PoT first trains a soft prompt on one source task and then uses the trained prompt to initialize the prompt for a target task. Lastly, the prompt initialized with is further fine-tuned on the target task to obtain the task-specific target prompt .
2 Prompt Transfer meets Knowledge Distillation
As stated in Section 1, in response to the two problems of PoT, we first propose a new metric to predict the prompt transferability, and then introduce the details of our proposed PANDA approach.
Inspired by prior works , we attempt to construct a semantic space of tasks based on soft prompt and map source-target tasks to this vector space. Then, we can measure a cosine similarity between the corresponding task embeddings and use this similarity as prompt transferability. In practice, we randomly select parts of samples (denoted as ), e.g., 100 (the detailed analysis of the number can be found in Appendix A.3), from each task as the representative data for fast computation. Then, the data is respectively fed into the original PLM and the PLM with trained prompt to obtain the hidden vectors of [CLS], i.e., and . To exactly represent the effect of soft prompt, we use the as the prompt-based task embeddings. In this way, we can obtain the corresponding embeddings and for source and target tasks, which are then used to measure the prompt similarity as:
PANDA approach.
As illustrated in Figure 3 (Right), our PANDA approach introduces the knowledge distillation technique to facilitate the knowledge transfer from source task to target task. In particular, PANDA approach first uses the PLM with source prompt as the teacher network , and the PLM with randomly initialized prompt (denoted as ) as the student network . Notably, we can then train on the target task with fewer iterations to obtain the new teacher (denoted as target-like network ), which learns some target information but not forget the source knowledge catastrophically. We state that such a target-like network performs as an intermediate network to bridge the gap between source and target tasksWe empirically find that using the original teacher network (i.e., ) could achieve comparable performance than the target-like network (see in Section 6.2), i.e., we can simply employ the source prompt itself as the teacher prompt, for the novel target tasks required fast computation.. Sequentially, the student network is trained on the target task with the supervision information of both ground-truth labels (same to Equation 1), and soft targets predicted by the teacher network that can be formulated as:
where and denote the distributions predicted by the teacher and student networks respectively. Forcing the student to mimic the prediction of teacher can make the student learn the “dark knowledge” from teacher, thus guiding the knowledge transfer from source task to target prompt. In summary, the complete loss function of student network can be formulated as:
where our metric (i.e., in Equation 2) is used to adaptively control the knowledge transfer, and is a factor to adjust the maximum of transfer ratio, which is empirically set as 0.05. Note that the reason why we use such a loss function in Eq. 3 can be found in Appendix A.4.
Experimental Setup
To investigate the effectiveness of our proposed methods, we conduct contrastive experiments on a total of 189 diverse combinations of 21 source datasets and 9 target datasets, which comprise part of tasks from General Language Understanding Evaluation (GLUE) , SuperGLUE and other NLU benchmarks. Specifically, GLUE is one of the most popular NLP benchmarks, which consists of several NLU tasks covering sentiment analysis, question answering and natural language inference. We use most of GLUE datasets as source datasets, including MNLI, CoLA, SST2, QNLI, MRPC, STSB and QQP. Moreover, as a stickier benchmark, SuperGLUE provides a new set of more difficult language understanding tasks, which additionally include word sense disambiguation and multi-choice question answering. All datasets of SuperGLUE are collected into source datasets, i.e. BoolQ, CB, RTE, WIC, WSC, COPA, Multirc and Record. Furthermore, to enrich the types of source datasets, some other NLU benchmarks are also added, including extractive QA (SQuAD 2.0 ), named entity recognition (NER) (CoNLL03 , CoNLL04 and Ontonotes ) and semantic role labeling (SRL) (CoNLL05 and CoNLL12 ). Following the prior work , we use low-resource datasets as target tasks to simulate a realistic scenario, which covers SuperGLUE (CB, COPA, WSC, RTE, WIC), GLUE (CoLA, MRPC, STSB) and other task (CoNLL04). We measure the development performance of each dataset using its corresponding evaluation metrics. Details of all datasets and evaluation metrics are in Appendix A.1.
2 Implementation
To investigate the universality of our proposed methods, we conduct extensive experiments on 5 different scales of PLMs, i.e., BERT-large (340M), BERT-base (110M), BERT-medium (42M), BERT-small (29M) and BERT-tiny (4.4M). The cutting-edge method P-Tuning-v2 is introduced to perform basic prompt-tuningWe implement our methods based on https://github.com/THUDM/P-tuning-v2. We use AdamW optimizer to tune our models with 8 NVIDIA A100 (40GB) GPUs. For each dataset, we perform grid-search for the learning rate on {5e-3, 7e-3, 1e-2}, batch size on {16, 32, 64} and epoch on {20, 40, 60, 80, 100}. The dropout rate is set to 0.1 and the input sequence length is 128. For the prompt length, we follow prior work and set it to a fixed number of 20 tokens in all cases (The detailed hyper-parameters are listed in Appendix A.1). We compare our PANDA approach with full-parameter model-tuning, basic prompt-tuning without any transfer and vanilla prompt transfer method . Note that the PLMs are fixed during training in all experiments, except those on the model-tuning. Additionally, it is also noteworthy that we evaluate the performance in the full-data scenario for all experiments.
Exploring the Prompt Transferability Metrics
In this section, we explore our proposed metric and other prior metrics in detail. Specifically, to the best of our knowledge, there are two existing prior metrics about prompt similarity : 1) the metric (namely Eavg) in SPoT that computes the cosine similarity of average soft prompt embeddings; 2) the other metric (namely ON) that uses the overlapping rate of activated neurons in the feed-forward network as a similarity of prompts.
Figure 4 (Left) shows heatmaps of the prompt transferability predicted by our metric and others between all 21 tasks. Specifically, these 21 tasks are first clustered into six groups: NLI, Semantic Matching(SM), QA, NER, SRL and others (details are shown in Appendix A.1). Then, we sort these tasks by the groups, i.e. several adjacent tasks are in the same group and have larger task similarities. We can observe that our proposed metric captures most intuitive task relationships , i.e., similar tasks have larger prompt transferability. However, the other metrics fall short in distinguishing the task relationships, except the relationship of same dataset itself. Furthermore, we compare the predicted similarity of these metrics for two trained prompt within the same type of tasks and between different tasks in Table 1. It can also be found that our metric works better to distinguish the prompts of same and different tasks. Takeaway: our proposed metric can better distinguish the different task relationships: similar tasks have larger prompt transferability.
Correlation between these metrics and prompt transfer performance.
Additionally, we conduct a quantitative analysis to examine whether these metrics align with prompt transfer performance. In practice, we follow Su et al. and compute the Spearman’s rank correlation scores between ranks of predicted prompt transferability and prompt transfer performance. Notably, an intuitive manner is using the final PoT performance as reference to measure the rank correlation scores, however, as shown in Section 1, vanilla PoT approach suffers from the problem of catastrophic forgetting when fine-tuning the source prompt on the target task. Thus, we use the performance at earlier epochs, i.e. the first epoch, as the referential prompt transfer performanceDue to space limitation, we show these transfer performance of all source-target pairs in Appendix A.7. Figure 4 (right) shows the results and we can observe that our proposed metric achieves consistent improvements compared to other metrics across all model sizes, indicating that our metric can make better choices to which source tasks for a given target task. Moreover, it can also be found that the Spearman’s rank correlation scores are sensitive to the scales of PLMs. We surmise that the capabilities of prompt-tuning in various PLMs are different, thus affecting the performance of prompt transfer differently. We will explore the in-depth reason in the future work. Takeaway: compared to prior metrics, our metric works better to make appropriate choices to which source tasks for a target task.
Evaluation Results
Table 2 shows contrastive results on 9 target tasks of BERT-large with vanilla PoT and PANDA respectively. Notably, due to the space limitation, we only show the results of some representative source-target pairs (results of all source-target pairs can be found in Appendix A.2). Firstly, it can be found that original prompt-tuning can achieve better results than model-tuning on some target datasets (e.g., COPA and CoLA), but its average performance of all datasets is still sub-optimal (77.5% v.s. 78.5%), showing the limitation of prompt-tuning. Secondly, in the group (a) of vanilla PoT, we can observe that transferring soft prompt from a similar source task can indeed benefit to the target performance, e.g., MNLI (77.5% 78.1%) and QNLI (77.5% 78.2%), which confirms the significance of PoT. Lastly, with our PANDA approach, prompt transfer can be more effective and prompt-tuning can achieve comparable and even better performance than model-tuning (e.g., MNLI: 79.4% v.s. 78.5%).
Knowledge Distillation helps bridge the gap between different types of tasks:
While the vanilla PoT approach can achieve performance improvement compared to the original prompt-tuning on the NLI tasks (many of target datasets belong to NLI task), prompt transfer on the other dissimilar tasks (e.g., NER and SRL) is usually negative and even leads to worse performance. This indicates that the large gap between source and target tasks will hinder the knowledge transfer, confirming our statement. On the contrary, PANDA can consistently boost the performance of prompt-tuning on all types of source tasks. More specially, PANDA significantly surpasses the vanilla PoT in the scenario where source task is dissimilar to target task. For instance, when using CoNLL03 (NER) and Record (QA) as source tasks, the relative improvements of PANDA are up to 24.1% and 15.9%, respectively. We state that this is owing to the knowledge distillation technique, which can effectively bridge the gap between different types of tasks and alleviate the problem of “catastrophic forgetting”.
PANDA shows generality and applicability on various scales of PLMs:
We also conduct experiments on other sizes of PLMs. The average results of these models are listed in Table 3. Compared to model-tuning, the performance drop of prompt-tuning is higher when the PLM is smaller, confirming that prompt-tuning falls short in fine-tuning smaller PLMs. Additionally, it can be seen that our PANDA consistently boosts the effectiveness of prompt transfer and improves the performance of prompt-tuning over the model-tuning approach across all model sizes. This demonstrates the generality and applicability of our PANDA approach on various scales of PLMs.
2 Analysis
As mentioned in Section 3.2, we use the source prompt as the initial teacher prompt and then train it on the target task with fewer iterations to obtain the target-like prompt. To investigate whether the teacher prompt provides useful knowledge to the student prompt, we conduct the contrastive experiments using different teacher prompt, i.e., “source prompt”, “target-like prompt” and “random prompt” (the randomly initialized prompt), and Table 4 lists the results. Obviously, when using “random prompt” as teacher prompt, the average performance is much worse than the others, because the randomly initialized prompt can not provide effective source knowledge and even hinders the learning of student prompt. On the contrary, within the source knowledge, source and target-like prompt performs much better, which proves the effectiveness of our strategy.
Whether PANDA benefits from our metric.
As shown in Equation 4, our metric is used to control the knowledge transfer for each source-target pair is also an important factor and we analyze the influence of different in Appendix A.5. To analyze its effect, we replace our metric with “constant factor (one) ”, and other prior metrics, i.e., “E” and “ON”. Results of different methods are presented in Table 5. Compared to the vanilla PoT approach, all variants of PANDA achieve better performance on both PLMs, continuing confirming the effectiveness of our approach. Additionally, it can be seen that our metric offers most performance improvements over the plain PANDA (more discussions are in Appendix 6.3), i.e., PANDA with constant factor (one), which demonstrates that our metric indeed facilitates the adaptive knowledge transfer and thus boosts the performance of PANDA.
3 Discussion
In this subsection, we explore whether PANDA can be extended to more complex scenarios, e.g, larger PLMs, more baseline methods and cross-model transfer. Moreover, we also discuss the contributions of our metric in detail.
Regarding the question of whether PANDA can be extended to a larger PLM setting, we conduct additional experiments as follows. Specifically, due to the limited amount of computing resources, we have tried our best to extend PANDA to a larger model setting, and lastly choose the DeBERTa-v2-xlarge (900M) for experiments. Similar to the comparing setting in Section 6.1, we compare our PANDA method to the model-tuning, pure prompt-tuning and vanilla PoT. Note that we use the MNLI as the source task in this experiment and list the detailed results in Table 6. It can be seen that PANDA outperforms the other methods on most of the tasks, which can prove that PANDA works well in the larger PLM scenarios as well.
Complementary with other parameter-efficient fine-tuning approaches.
In this paper, our main topic is to improve prompt-tuning with a general KD-based framework, namely PANDA. To further investigate the universality of the framework, we conduct additional experiments on two more parameter-efficient fine-tuning approaches, i.e., adapter and prefix-tuning . Specifically, we show the results of the vanilla approaches and those improved with PANDA in Table 7. The results of prompt-tuning is also listed for reference. We can observe that the original adapter could achieve the comparable performance against the model-tuning on these tasks. Nevertheless, PANDA can still consistently improve the performance of adapter and prefix-tuning by a large margin, i.e., PANDA is complementary with these approaches as well. These results prove the universality of PANDA.
Cross-model transfer to improve the performance on smaller models.
Due to the limited capability of prompt-tuning on smaller models, PANDA performs slightly worse than the model-tuning on some tasks in the smaller model settings. An intuitive method is to transfer the knowledge from a large teacher model to the smaller models. To verify it, we attempt to distill the prompt from a larger model (BERT-medium) to the smaller model (BERT-tiny). Here, we use the MNLI as the source task for PANDA and show the detailed results in Table 8. For reference, we compare the cross-model transfer “medium tiny” to the model-tuning and original PANDA “only tiny” (the teacher and student models have the same model, BERT-tiny). Specifically, when we employ the larger teacher model to distill the student, i.e., “medium→tiny”, PANDA can achieve better performance and even outperform the model-tuning. This proves the effectiveness of cross-model distillation to improve the performance of smaller models.
Main contributions of our proposed metric in this work.
Some readers may point out that the different metrics do not make a big difference in Table 5 of this paper, and show concerns about the effect of our proposed metric. Here, we analyze the results of Table 5 and discuss the contributions of our metric in detail.
Firstly, although the difference between different metrics on the BERT-tiny is slightly small, the difference on the BERT-medium is higher than 0.62, which could prove the effectiveness of our metric to improve the PANDA method. Secondly, we state that the main contribution of our metric is to effectively retrieve similar source tasks for a given target task, and the experiment in Table 5 is used to prove that the metric can further boost the performance of the PANDA method.
To investigate the effect of different metrics intuitively, we report the performance of the target tasks with the most similar source tasks, as measured by different metrics. Table 9 lists the results in the vanilla PoT and PANDA settings across different BERT models respectively. It can be seen that our proposed metric consistently achieves the best performance across all model sizes and prompt transfer settings, proving that our metric works better on choosing the useful similar source task for a given target task. Notably, when using the PANDA method, the difference between multiple metrics is relatively smaller. This is because PANDA can effectively bridge the gap between source and target tasks, and thus alleviate the negative influence of other metrics (E and ON).
Conclusion and Future Work
In this paper, we first introduce a new metric to accurately predict the prompt transferability, and then improve PoT with a knowledge distillation technique, which uses our metric to adaptively transfer the “dark knowledge” from source prompt to the target prompt. Large-scale experiments are conducted to empirically investigate the effectiveness of our methods. We explore the shortcomings of prior metrics and prove that our metric works better to predict the prompt transferability. Additionally, we show that our PANDA consistently improves over vanilla PoT by 2.3% average score across all tasks and models, and makes the prompt-tuning achieve competitive and even better performance than full-parameter model-tuning in various PLM scales scenarios. One limitation of this work is that we only conduct experiments on various sizes of BERT and we will investigate the effectiveness of our proposed methods on more cutting-edge PLMs, e.g., GPT 3 and T5 , in our future work.
References
Checklist
Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]
Did you describe the limitations of your work? [Yes] See Section 7.
Did you discuss any potential negative societal impacts of your work? [N/A]
Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]
If you are including theoretical results…
Did you state the full set of assumptions of all theoretical results? [Yes] See Section 3.
Did you include complete proofs of all theoretical results? [N/A]
Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] We provide them in the supplemental material.
Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] See Table 10.
Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [No] We conduct large-scale experiments (up to 1900 experiments) and it is time-costly for us to run experiments multiple times.
Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] See Section 4.2.
If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…
If your work uses existing assets, did you cite the creators? [Yes]
Did you mention the license of the assets? [Yes]
Did you include any new assets either in the supplemental material or as a URL? [Yes]
Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [Yes]
Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [Yes] See Section 4.1
If you used crowdsourcing or conducted research with human subjects…
Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]
Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]
Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]
Appendix A Appendix
Table 10 shows the details of all 21 tasks used in our study. The tasks are classified into three groups: SuperGLUE, GLUE and others. The training details of prompt-tuning are also listed in the table.
For most datasets, we use accuracy (Acc.) as their evaluation metric in the experiments, except the Matthews correlation (Mcc.) for CoLA, F1 score (F1) for Multirc, Record, SQuAD, CoNLL03, CoNLL04, CoNLL05, CoNLL12 and Ontonotes, the combined score of Pearson-Spearman correlations (Pcor/Scor.) for STS-B. Notably, due to restricted test set access for GLUE and SuperGLUE, we follow Vu et al. and train for a fixed number of steps and report results on the validation set associated with each dataset.
A.2 Full results of our PANDA and vanilla PoT approach
We report all results of our study across model sizes of PLMs. Specifically, Table 14, Table 15, Table 16, Table 17 and Table 18 list results of BERT-large, BERT-base, BERT-medium, BERT-small and BERT-tiny respectively. Our PANDA approach achieves consistent and significant performance improvements compared to the vanilla prompt transfer.
A.3 Sensitivity analysis of our proposed metric on the sampling numbers
As stated in Section 3.2, we sample parts of instances from source and target datasets to calculate the similarity. In practice, we sample the data in the dev set for each task, i.e., the sampling number is determined by the minimum number of data in the dev set among all tasks. Specifically, since the dev set of COPA (one lower-resource task) only contains 100 samples, we set 100 as the sampling number. To examine whether our metric is sensitive to the sampling numbers, we calculate the similarity with different sampling numbers. In practice, for the larger sampling numbers, e.g., 200, we sample the data from the dev set for the higher-resource tasks (whose dev sets contain more than 200 instances), and sample the data from the training set or total dataset for the lower-resource tasks.
Table 11 shows the similarities between the source task (QNLI) and target tasks across different sampling numbers. Note that BERT-base is used in this study. It can be seen that the similarities across different sampling numbers are slightly changed. There are relatively larger changes when we only sample 50 instances to calculate the similarity. One of the possible reasons is that too few samples are difficult to truly reflect the characteristics of the task. This can be further confirmed by the observation that as the sampling number increases (higher than 50), the similarities are almost unchanged.
Additionally, we also list the standard deviation of predicted similarities across different sampling numbers in Table 12. Note that “std” means the standard deviation (average score of all target tasks), and “|100-avg|” denotes the absolute difference (average score) between the predicted similarities in our paper and the average similarities of multiple sampling numbers. It can be seen that the standard deviation of similarities among different sampling numbers is very small, confirming that our metric is insensitive to the number of samples.
A.4 Analysis of different knowledge distillation losses
There are several widely used loss functions, i.e., KL-divergence (KL), cross-entropy (CE) and Mean Squared Error (MSE), which can be used as the knowledge distillation loss. In our preliminary experiments, we investigate the effect of different knowledge distillation losses in our PANDA framework. Specifically, we conducted experiments using three source tasks on two PLMs of different scales. Table 13 shows the detailed results. It can be found the different between multiple loss functions is not large, especially the difference between KL-divergence and cross-entropy. Moreover, we can observe that MSE performs best on most settings, and thus we use it as the knowledge distillation loss in this paper lastly.
A.5 Influence of different weighting factor λ𝜆\lambda
The weighting factor in to Equation 4 for the knowledge distillation auxiliary task is an important parameter. In the subsection, we evaluate the performance of the proposed PANDA approach with different on several pairs of cross-task prompt transfer and show the results in Figure 5.
The weighting factor affects the knowledge transfer between source and target prompts. On the one hand, setting a too small , i.e., = 0.01, would make the student prompt ignore the “dark knowledge” from teach prompt and thus hinder the knowledge transfer. On the other hand, larger would prevent the student prompt from learning good knowledge from the target task, making the student prompt perform worse. More importantly, it can be found that a large is more detrimental than a smaller one. In the case of = 5, the average performance of all pairs of cross-task prompt transfer is much worse than the case of small , e.g., 0.01 or 0.05. Recall that the case of = 0.05 performs best among all settings, thus leaving as the default setting.
A.6 Prompt transferability predicted by our metric and other metrics
In this subsection, we provide more heatmap results of our predicted prompt transferability across all 21 tasks on different PLMs. Specifically, Figure 6, Figure 7 ,Figure 8 and Figure 9 show the results on BERT-base, BERT-medium, BERT-small and BERT-tiny respectively. More interestingly, our predicted similarities tend to drop as the scales of PLMs decrease, while the cosine similarities of prompt embeddings have the opposite tendency. One possible reason is that prompt-tuning works worse on smaller PLMs, but the parameter ratio of prompt itself will increase in smaller PLMs, thus playing an important role in the similarity prediction.
A.7 Transfer performance at the first epoch
As mentioned in Section 5, we calculate the Spearman’s rank correlation scores between ranks of predicted prompt transferability and transfer performance as the first epoch. Here, we report the detailed transfer performance across all model sizes for references. Specifically, besides the cross-task transfer performance, we also list the performance of randomly initialized prompts. Table 19, Table 20, Table 21, Table 22 and Table 23 show the results on BERT-large, BERT-base, BERT-medium, BERT-small and BERT-tiny, respectively.