GPT Understands, Too
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, Jie Tang
Introduction
Pretrained language models (PLMs; Brown et al., 2020) have significantly advanced the performance of natural language understanding (NLU). PLMs are trained with different pretraining objectives, such as masked language modeling Devlin et al. (2018), autoregressive language modeling Radford et al. (2019), seq2seq Raffel et al. (2019), and permutation language modeling Yang et al. (2019). PLMs can be further enhanced with prompting Brown et al. (2020); Schick and Schütze (2020), which employs manually written prompt patterns as additional input to a language model. With prompting while PLMs are either finetuned on a small labeled dataset or frozen for direct inference on downstream tasks. Prompting has significantly improved the performance of many NLU tasks Brown et al. (2020); Schick and Schütze (2020).
However, we observe that manual discrete prompts suffer from a large degree of instability. As shown in Table 1, with a frozen language model, changing a single word in the prompt might result in substantial performance drop. As we will show in Section 3, when the language model is tuned, the instability problem is alleviated but the performance difference between different prompts is still sizeable, especially in the few-shot setting. Such an instability issue of discrete prompts poses a critical challenge in practice. Recent approaches of automatic prompting have attempted to search for a better-performing prompt given a task Shin et al. (2020); Gao et al. (2020); Jiang et al. (2020b), but these methods do not change the unstable nature of discrete prompts.
To reduce the instability of discrete prompts, we propose a novel method P-Tuning that employs trainable continuous prompt embeddings in concatenation with discrete prompts. Specifically, given a discrete prompt as the input, P-Tuning concatenates continuous prompt embeddings with the discrete prompt tokens and feeds them as the input to the language model. The continuous prompts are updated by backpropagation to optimize the task objective. The intuition is that continuous prompts incorporate a certain degree of learnability into the input, which may learn to offset the effects of minor changes in discrete prompts to improve training stability. To further improve performance, we employ a prompt encoder using LSTMs or MLPs to model the dependency between continuous prompt embeddings.
We experiment with two NLU benchmarks: the LAMA Petroni et al. (2019) knowledge probing and SuperGLUE Wang et al. (2019a). On LAMA, with the language model frozen, P-Tuning outperforms manual discrete prompts and searched prompts by 20+ points and 9 points respectively with the same pretrained models. On SuperGLUE, with the language model finetuned, P-Tuning outperforms PET Schick and Schütze (2020) with the best discrete prompts under both the fully-supervised and few-shot settings. In addition to improving performance, our results show that across a wide range of tasks and settings, P-Tuning substantially reduces the performance gap between different discrete prompts, which results in improved stability for language model adaptation.
Method
Prompting employs natural language patterns as additional inputs to pretrained language models for adaptation to downstream tasks Brown et al. (2020); Schick and Schütze (2020). Prior work Zheng et al. (2021) has pointed out that prompting has achieved consistent and substantial improvements on a number of NLP tasks. However, it still remains a challenging problem of how to write high-performing discrete prompts.
We performed preliminary experiments using different manual prompts on the LAMA knowledge probing task Petroni et al. (2019), which aims to extract triplet knowledge from a language model by predicting the tail entities. Results in Table 1 show that manual discrete prompts lead to unstable performance. For example, if we compare the last two prompts in the table, changing a single word in prompt causes a drastic decrease of 20 points in performance.
In light of the challenge, recent works propose to automate the search procedure of discrete prompts by mining the training corpus Jiang et al. (2020b), gradient-based searching Shin et al. (2020), and using pretrained generative models Gao et al. (2020). However, these works aim at searching for better-performing prompts but do not change the nature of instability for discrete prompts. In addition to the instability issue, searching in the discrete space might not be able to fully leverage the gradients from backpropagation, which will potentially result in suboptimal solutions. To this end, we explore the possibility of training continuous prompts to stabilize and improve the performance of language model adaptation.
2 P-Tuning
Formally, let be a pretrained language model with a hidden size of and a vocabulary size of . Let be a labeled dataset for an NLU task, where is an input consisting of a sequence of discrete tokens, and is a label. Our goal is to estimate the conditional probability for classification with parameters of either finetuned or frozen.
Finally, we update the embeddings to optimize a task loss function.
It is noteworthy that we can also concatenate discrete prompts with continuous prompts, which performs better and is adopted throughout our experiments. P-Tuning is applicable to both frozen and finetuned language models.
3 Prompt Encoder
In the aforementioned framework, we employ a mapping function to map trainable embeddings to model inputs . The intuition is that by using a mapping function, it is more convenient to model the dependency between different prompt embeddings, compared to using independent learnable embeddings. In our implementation, we use a lightweight neural network to formulate the function . Specifically, we experiment with using long short-term memory (LSTM) networks, multi-layer perceptrons (MLPs), and the identity mapping function in Section 3.
Experiments
We include two NLU benchmarks: LAMA Petroni et al. (2019) for knowledge probing (§ 3.1) and SuperGLUE Wang et al. (2019a) for general natural language understanding. On SuperGLUE, we consider both the fully-supervised learning (§ 3.2) and few-shot learning (§ 3.3) settings.
On LAMA, following Shin et al. (2020); Jiang et al. (2020b), language models are frozen and only the discrete or continious prompts are tuned. For SuperGLUE, following Schick and Schütze (2020); Zheng et al. (2021), language models are tuned. In our setting, we jointly optimize the language model parameters and the continuous prompts. This setup not only follows the common, standard settings in prior work, but also allows evaluating P-Tuning with both tuned and frozen language models.
The overall task setup and a summary of results are shown in Table 2.
Knowledge probing, or referred to as fact retrieval, evaluates how much real-world knowledge has language models gained from pre-training. The LAMA Petroni et al. (2019) dataset evaluates it with cloze tests created from triples selected in the knowledge bases.
Datasets and vocabulary. LAMA enforces all answers in single-token format. We first adopt the original LAMA-TREx dataset, consisting of 41 Wikidata relations and altogether 34,039 testing triples (namely LAMA-34k, which covers all BERT vocabularies). Since different pretrained models share distinct vocabularies, to allow direct comparison, we follow previous work Shin et al. (2020) to adopt a subset that covers the intersection of GPT’s and BERT’s vocabularies. This is caled LAMA-29k. We again follow Shin et al. (2020) to construct the training, development, and test data to allow for fair comparison.
Setup. LAMA has provided a handcraft prompt for each relation, as shown in Table 1, which are effective but likely sub-optimal. For bidirectional masked language models, we only need to replace “[X]” with the subject entity and “[Y]” with the [MASK] token; for unidirectional language models such as GPT, following LAMA’s original setting on Transformer-XL Dai et al. (2019), we use the network output just before the target position.
The number of prompt tokens and positions are selected based on the development sets, and for simplicity we choose the (3, sub, org_prompt, 3, obj, 3) template for bidirectional models and (3, sub, org_prompt, 3, obj) for unidirectional models as this configuration performs well for most relations (where the number indicates the number of continuous prompt tokens). Continuous prompts are concatenated with original discrete prompts. During the prompt training, we set the learning rate to 1e-5 and use the Adam optimizer.
1.2 Main results
The results are presented in Table 3. P-tuning significantly improves the best results of knowledge probing from 43.3% to 50.6% on LAMA-34k and from 45.2% to 64.2% on LAMA-29k. Moreover, P-tuning outperforms previous discrete prompt searching approaches such as AutoPrompt Shin et al. (2020) and LPAQA Jiang et al. (2020b) on the same-size models. This confirms our intuition in Section 2 that discrete prompts might not be optimal.
2 Fully-supervised Learning
Dataset. To evaluate P-tuning on fully-supervised learning tasks, we adopt the SuperGLUE benchmark Wang et al. (2019b), consisting of 8 challenging natural language understanding (NLU) tasks. We focus on 7 of them since the ReCoRD Zhang et al. (2018) task adopts no discrete prompts, thus P-tuning is not directly applicable. The tasks include question answering (BoolQ Clark et al. (2019a) & MultiRC Khashabi et al. (2018)), textual entailment (CB De Marneffe et al. (2019) & RTE Dagan et al. (2005)), co-reference resolution (WiC Pilehvar and Camacho-Collados (2018)), causal reasoning (COPA Roemmele et al. (2011)), and word sense disambiguation (WSC Levesque et al. (2012)).
Comparison methods. We experiment with P-tuning on both unidirectional and bidirectional pretrained models, i.e., GPT and BERT. We include four variants BERT-Base, BERT-Large, GPT2-Base, and GPT-medium. For each model, we compare standard classification finetuning, PET Schick and Schütze (2020) (a typical finetuning method based on manual discrete prompts) and our P-tuning.
Configuration. We use the same metrics as in Wang et al. (2019b). For fully-supervised learning, we use a large training set to finetune pretrained models and use a development set for hyper-parameter and model selection. Specifically, the AdamW optimizer with a linearly decayed learning rate is used for training. We use a learning rate of , a batch size of , and a warm-up ratio of . For small datasets (i.e., COPA, WSC, CB, RTE), we fine-tune pretrained models for 20 epochs. For larger datasets (i.e., WiC, BoolQ, MultiRC), we reduce the number of training epochs to be 10 as the model converges earlier. Early stopping is used to avoid over-fitting the training data.
2.2 Main Results
The main results of fully-supervised learning are shown in Table 4. We observe that P-tuning can improve fully-supervised learning performance on both BERTs and GPTs. (1) Specifically, on the BERT-Base model, P-tuning achieves best performance on 5/7 tasks, while with BERT-Large, P-tuning outperforms other methods on 4/7 tasks. The exceptions are WiC and MultiRC, both of which have relatively large training sets. We find that P-tuning might not have large gains over CLS-FT on such high-resource tasks, while benefits more on low-resource tasks. On average, P-tuning improves over the considered baselines. (2) On GPT2-Base and GPT2-Medium models, P-tuning consistently achieves the best performance on all tasks.
3 Few-Shot Learning
While GPT-3 has shown decent few-shot learning potential with handcrafted prompts, it still struggles on some of the challenging tasks (e.g., natural language inference) Brown et al. (2020). We are motivated to study whether P-tuning can also improve the few-shot learning performance of pretrained models on challenging tasks.
Few-shot Evaluation. The few-shot performance is sensitive to lots of factors (e.g., the order of training examples, random seed, and prompt patterns), and thus suffers from high variance Zhao et al. (2021a); Lu et al. (2021); Zhang et al. (2020). Therefore, the few-shot evaluation strategy should make sure that the improvements are indeed from an improved method instead of variance. To this end, we follow the FewNLU evaluation procedure Zheng et al. (2021) that has addressed and handled the issue. Specifically, we use random data splits to perform model selection only on a small labeled set to prevent overfitting a large dev set.
Dataset. We use the few-shot SuperGLUE (also known as FewGLUE) benchmark Schick and Schütze (2020) and follow the setting in prior work Zheng et al. (2021) in terms of data split construction.
Baseline and Hyper-parameter. In few-shot learning, we again compare P-tuning with PET Schick and Schütze (2020), which was shown to outperform GPT-3 on some of the tasks. Similar to Schick and Schütze (2020), we use ALBERT-xxLarge as the base model. For hyper-parameters that are shared by PET and P-tuning (e.g., learning rate, maximum training step, evaluation frequency), we use the same search space for fair comparison. Specifically, we search the learning rate in , the maximum training step in , and the evaluation frequency in .
Construction of Prompt Patterns. For PET, we use the same manual prompts reported by Schick and Schütze (2020). When constructing prompt patterns for P-tuning, based on the same manual prompts as PET, we insert different numbers of continuous prompt tokens into different positions, thus formulating a number of pattern candidates. We then select the best pattern for P-tuning using the validation strategy of FewNLU Zheng et al. (2021). We also conduct further analysis of the number and the position of continuous prompt tokens in §8.
3.2 Main Results
Few-Shot Performance. Table 5 shows the main results of few-shot learning. We find that, on ALBERT, P-tuning consistently outperform PET on average by more than 1 points. It outperforms PromptTuning by more than 13 points. It proves that by automatically learning continuous prompt tokens, the pretrained models can achieve better few-shot performance on NLU tasks.
3.3 Ablation Study
Type of Prompt Encoder Prior work Shin et al. (2020) proposes to simply use an MLP as the prompt encoder, we perform further ablation analysis for prompt encoder selection, and results are shown in Table 8. We consider LSTM, MLP, and EMB (i.e., we directly optimize the word embeddings without using additional parameters). From the results, we can see that LSTM, MLP, and EMB all work as a prompt encoder. Results show that both LSTM and MLP generally work well on these tasks, while EMB is unstable and can substantially under-perform the other two on some tasks (e.g,. WiC and CB). To sum up, both LSTM and MLP could be taken into account when working on new tasks.
Location of Prompt Tokens To study at which location to insert continuous prompt tokens, we perform experiments as Table 7 shows. From the results, we have the following findings.
By comparing #1 (or #2) with #3 (or #4), we find that it would be better if we insert continuous prompt tokens at the location where it does not segment the sentences. For example, in case#1, “[P]” breaks the completeness of sentence “[Hypothesis]?” while in case#3, “[P]” is located between sentences.
By comparing #2 (or #3) with #4, we find that there’s no special preference for placing on the edge or in the middle of the inputs.
It is suggested to write a number of pattern candidates and then search over them for the best for each task.
Number of Prompt Tokens We also study the influence of the number of prompt tokens and show the results in Table 7. By comparing #3, #6, #7, and #8, we can conclude that the number of prompt tokens has a great impact on the few-shot performance. However, it is not that a larger number of prompt tokens would always be better. We conjecture that it could be that due to the limited training data, it becomes difficult to learn the parameters when excessively increasing the number of continuous prompt tokens. In practice, it is suggested to search for the best number of prompt tokens through model selection.
3.4 Comparison with Discrete Prompt Search
Prior work Gao et al. (2020) proposed to automatically search discrete prompts and achieved better results than those of manual prompts. We now proceed to compare P-Tuning with auto-searched discrete prompts. For fair comparison, we follow the setting of LM-BFF Gao et al. (2020) to also conduct experiments on some of the GLUE tasks Wang et al. (2018) with RoBERTa-Large model Liu et al. (2019). Since the the evaluation protocols have large impacts on few-shot performance, we use the top-3 discrete prompts searched by LM-BFF and experiment with using only the discrete prompts and additionally applying P-Tuning. For P-Tuning, the prompt patterns are constructed by concatenating the same discrete prompts as well as continuous prompts. Results in Table 9 show that additionally incorporating continuous prompts can further improve few-shot performance. P-Tuning is easy to be combined with existing discrete prompts, while further improving stability as discussed in Section 3.4.
4 Stabilizing Language Model Adaptation
In the above sections, we have shown that P-Tuning improves over performance across multiple settings. Now we present results to demonstrate that P-Tuning also stabilizes language model adaptation; i.e., reducing the differences between different prompts. As we have shown in Table 1, manual prompts have a large impact on the performance. When it comes to few-shot learning, the performance gap of different prompts is prominent due to the sensitivity of few-shot learning Zheng et al. (2021). Results in Table 6 show that P-tuning improves the performance of the worst-performing patterns (e.g., P#5), and achieves a smaller standard deviation over multiple patterns. Compared to PET-FT, P-tuning increases the stability w.r.t. the choice of patterns.
On LAMA, we observe similar a phenomenon that while manual prompts often yield quite volatile results, appending trainable continuous prompts on top of the manual prompts can stabilize their performances, reducing the standard deviation from 10.1 to 0.46.
Related work
Language Model Prompting. GPT-3 Brown et al. (2020) uses in-context examples Liu et al. (2021); Zhao et al. (2021b) as a way of prompting to transfer knowledge from pretraining to downstream tasks. Schick and Schütze (2020) proposed to use cloze patterns, which removes the constraint that the masked token is the last token of the sentence. This further minimizes the gap between pretraining and downstream tasks. To improve prompting for NLU, recent works have proposed methods to automatically search for high-performing prompts by mining the training corpus Jiang et al. (2020b), gradient-based search Shin et al. (2020), or using pretrained generative models Gao et al. (2020). Our approach is different from these prior works in that we resort to using continuous prompt embeddings, which are found to be complementary to discrete prompts in our experiments.
Recently, some concurrent works also proposed the use of continuous prompts. Prefix-tuning Li and Liang (2021) adds continuous prompts at the beginning of the sequence for each layer. In contrast to our work, prefix-tuning targets natural language generation tasks.
In the area of NLU, a few concurrent methods were proposed based on continuous prompts, focusing on improving knowledge probing Qin and Eisner (2021); Zhong et al. (2021). Lester et al. (2021) showed that with large pretrained models, only tuning continuous prompts with a frozen language model achieves comparable performance to full-model tuning.
Compared to these concurrent works on NLU, P-Tuning reaches a unique conclusion that continuous prompts improve performance and stabilize training with either frozen or tuned models under both the few-shot and fully-supervised settings. For example, no concurrent works have shown that continuous prompts can improve performance with a tuned language model. Technically, P-Tuning also has a few unique designs such as using hybrid continuous-discrete prompts and employing a prompt encoder.
Knowledge in Language Models. Self-supervised Liu et al. (2020) pre-trained language models Han et al. (2021) including GPT Radford et al. (2019), BERT Devlin et al. (2018), XLNet Yang et al. (2019), RoBERTa Liu et al. (2019) have been observed to learn not only contextualized text representations but also linguistic and world knowledge. Hewitt and Manning (2019) demonstrates that contextualized representations produced by language models can form a parse tree in the embedding space. Vig (2019); Clark et al. (2019b) look into the multi-head attention patterns within transformers and discover that certain attention heads may correspond to some grammatical functions, including co-reference and noun modifiers. LAMA Petroni et al. (2019, 2020) propose the LAMA task that leverages cloze tests to predict the fact triples of knowledge bases to examine language model’s ability of memorizing facts with answers in the single-token format. In Wang et al. (2020), the authors investigate the attention matrices to find evidence about knowledge triples contained in the context. Jiang et al. (2020a) develops a multi-token fact retrieval dataset based on LAMA.
Conclusions
In this paper, we present a method P-Tuning that uses continuous prompts in concatenation with discrete prompts. P-Tuning improves performance and stabilizes training for pretrained language model adaptation. P-Tuning is effective with both tuned and frozen language models under both the few-shot and fully-supervised setings.