Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Qihuang Zhong, Liang Ding, Yibing Zhan, Yu Qiao, Yonggang Wen, Li Shen, Juhua Liu, Baosheng Yu, Bo Du, Yixin Chen, Xinbo Gao, Chunyan Miao, Xiaoou Tang, Dacheng Tao

Introduction

The last several years have witnessed notable progress across many natural language processing (NLP) tasks, led by pretrained language models (PLMs) such as bidirectional encoder representations from transformers (BERT) [Devlin et al., 2019], OpenAI GPT [Radford et al., 2019] and its most renowned evolution GPT3 [Brown et al., 2020]. The unifying theme of the above methods is that they conduct self-supervised learning with massive easy-to-acquire unlabelled text corpora during the pretraining stage and effectively fine-tune on a few labeled data in target tasks. Such a “pretraining-fine-tuning” paradigm has been widely adopted by academia and industry, and the main research and development direction involves scaling the sizes of foundation models up to extremely large settings, such as Google’s 540B PaLM [Chowdhery et al., 2022b], to determine the upper capacity bounds of foundation models.

In such a context, the SuperGLUE [Wang et al., 2019a] (a more challenging version of the general language understanding evaluation (GLUE) benchmark [Wang et al., 2018]) has become the most influential and prominent evaluation benchmark for the foundation model community. Most high-performing models on the GLUE/ SuperGLUE leaderboard bring new insights and better practices to properly guide future research and applications.

We recently submitted our 6B Vega v2 model to the SuperGLUE leaderboard and, as seen in Figure 1, obtained state-of-the-art records on 4 out of 8 tasks, sitting atop the leaderboard as of Oct. 8, 2022, with an average score of 91.3. Encouragingly, our 6B model with deliberately optimized pretraining and downstream adaptation strategies substantially outperforms 540B PaLM [Chowdhery et al., 2022b], showing the effectiveness and parameter-efficiency of our Vega model. This technical report briefly describes how we build our powerful model under a certain parameter budget, i.e., 6B, from different aspects, including backbone framework (§2.1), the efficient pretraining process (§2.2), and the downstream adaptation approach (§2.3). To fully extract knowledge from the given pretraining data to PLMs, we propose a self-evolution learning (in Figure 2) mechanism to wisely predict the informative tokens that should be masked and supervise the mask language modeling process with rectified smooth labels. To effectively transfer the knowledge to different downstream tasks, especially the low-resource tasks, e.g., CB, COPA, and WSC, we design a knowledge distillation-based prompt transfer method [Zhong et al., 2022b] (in Figure 3) to achieve better performance with improved robustness.

The remainder of this paper is designed as follows. We introduce the major utilized approaches in Section 2. In Section 3, we review the task descriptions and data statistics and present the experimental settings and major results. Conclusions are described in Section 4.

Approaches

In this section, we describe the main techniques in our model, including the backbone framework in §2.1, the efficient pretraining approaches in §2.2, and the downstream adaptation technique in §2.3.

Vanilla transformers [Vaswani et al., 2017] enjoy appealing scalability as large-scale PLM backbones [Devlin et al., 2019, Raffel et al., 2020, Brown et al., 2020, Zan et al., 2022b]; for example, T5 and GPT3 flexibly scale their feedforward dimensions and layers up to 65,534 and 96, respectively. We hereby employ a vanilla transformer, i.e., multihead self-attention followed by a fully connected feedforward network, as our major backbone framework. As encoder-only PLMs have an overwhelming advantage over the existing methods on the SuperGLUE leaderboard, we train our large model in an encoder-only fashion to facilitate downstream language understanding tasks. According to our PLM parameter budget – 6 Billion, we empirically set the model as follows: 24 layers, 4096 as the hidden layer size, an FFN of size 16,384, 32 heads, and 128 as the head size. In addition, ?) empirically demonstrated the necessity of computing self-attention with disentangled matrices based on their contents and relative positions, namely disentangled attentionIn our preliminary ablations, we surprisingly found the Enhanced Mask Decoder [He et al., 2021] technique, which is coupled with disentangled attention strategy in DeBRETa, was useless, therefore we did not this strategy in Vega v2., which is adopted in Vega v2.

2 Efficient Pretraining

Recall that our aim is not to arbitrarily increase the model scales, but to facilitate storing the informative derived knowledge from the pretraining data in PLMs. To approach this goal, we first revisit the representative self-supervision objective – masked language modeling [Devlin et al., 2019] (MLM), and propose a novel self-evolution learning mechanism to enable our PLM to wisely predict the informative tokens that should be masked, and train the model with smooth self-evolution labels.

MLM is a widely used self-supervision objective when conducting large-scale pretraining on large amounts of text to learn contextual word representations. In practice, MLM randomly selects a subset of tokens from a sentence and replaces them with a special mask token, i.e., [MASK]. However, such a random masking procedure is usually suboptimal, as the masked tokens are sometimes too easy to guess with only local cues or shallow patterns. Hence, some prior works focused on more informative masking strategies, such as span-level masking [Joshi et al., 2020], entity-level masking [Sun et al., 2019], and pointwise mutual information (PMI)-based masking [Sadeq et al., 2022]. These efforts achieved better performance than vanilla random masking, which inspires us to explore more approaches for fully extracting knowledge from pretraining data.

Self-Evolution Learning

Based on the above motivations, we propose a novel self-evolution learning mechanism for PLMs, as illustrated in Figure 2. Different from the prior works that designed masking strategies to train language models from scratch, our self-evolution learning approach aims to encourage the given “naive” PLMs to find patterns (tokens) that are not learned well but are more informative, and then fix them. Specifically, there are two stages in our self-evolution learning mechanism, as follows.

Stage 1 is self-questioning. Given a vanilla PLM (e.g., trained with the random masking objective), we first feed the original training samples into the PLM and make it re-predict the output probabilities for each token. As the PLM has seen these samples and learned from them in the pretraining stage, it can make deterministic and correct predictions for most of the tokens, which we denote as learned tokens. However, for some tokens, such as “American” in the sentence “Thomas Edison was an American inventor and businessman”, the PLM tends to predict this token as “excellent” (the probability of “excellent” is 0.44, while the probability of “American” is 0.4). We attribute this phenomenon to the fact that the PLM does not learn this knowledge-intense pattern but only makes its prediction based on the local cues. We refer to these harder and more informative tokens as neglected tokens. After all training samples are fed into the PLM, we can obtain a set of neglected tokens for each training sample. Note that this procedure is conducted offline and does not update the parameters of the original PLM.

Stage 2 is self-evolution training. Given the neglected tokens (obtained in stage 1), we can select them for masking and then encourage the PLM to learn from these informative patterns, thus continuously improving the capability of the PLM. Intuitively, we can make the PLM learn how to predict these tokens, by minimizing the loss between the predicted probabilities and one-hot labels. However, considering the diversity of the neglected token, if we force the PLM to promote one specified ground truth over others, the other reasonable “ground truths” (for a given masking token, there can be more than one reasonable prediction) become false negatives that may plague the training process or cause a performance decrease [Li et al., 2022].

3 Downstream Adaptation

In addition to the above efficient pretraining methods, we also introduce some useful fine-tuning strategies for effectively adapting our Vega v2 to downstream tasks. Specifically, there are two main problems that hinder the adaptation performance of a model. 1) The first concerns the domain gaps between the training and test sets, which lead to poor performance on target test sets. 2) The second is the use of limited training data, e.g., the CB task, which only consists of 250 training samples, as limited data can hardly update the total parameters of PLMs effectively. Note that in addition to the strategies listed below, we have also designed and implemented other methods from different perspectives to improve the generalization and efficiency of models, e.g. the FSAM optimizer for PLMs [Zhong et al., 2022c], SparseAdapter [He et al., 2022], and continued training with downstream data [Zan et al., 2022a]. Although these approaches can help to some extent, they do not provide complementary benefits compared to the listed approaches, so they are not described here.

Regrading the domain or linguistic style gap between the training and test sets (the first problem), we adopt a transductive fine-tuning strategy to improve the target domain performance, which is a common practice in machine translation evaluations [Wu et al., 2020, Ding and Tao, 2021] and some domain adaptation applications [Liu et al., 2020]. The proposed transductive fine-tuning technique is shown in Algorithm 1. Whether we should conduct transductive fine-tuning depends on the practical downstream performance achieved.

Prompt-Tuning

To address the second problem, we replace the vanilla fine-tuning process with a more parameter-efficient method, prompt-tuning [Lester et al., 2021], for low-resource tasks. Despite the success of prompts in many NLU tasks [Wang et al., 2022, Zhong et al., 2022a], directly using prompt-tuning might lead to poor results, as this approach is sensitive to the prompt’s parameter initialization settings [Zhong et al., 2022b]. An intuitive approach, termed as prompt transfer [Vu et al., 2022], is to initialize the prompt on the target task with the trained prompts from similar source tasks. Unfortunately, such a vanilla prompt transfer approach usually achieves suboptimal performance, as (i) the prompt transfer process is highly dependent on the similarity of the source-target pair and (ii) directly tuning a prompt initialized with the source prompt on the target task might lead to forgetting the useful general knowledge learned from the source task.

To this end, we introduce a novel prompt transfer framework [Zhong et al., 2022b] to tackle the above problems. For (i), we propose a new metric to accurately predict prompt transferability. In practice, the metric first maps the source/target tasks into a shared semantic space to obtain their task embeddings based on the source/target soft prompts and then measures the prompt transferability via the similarity of corresponding task embeddings. In our primary experiments, we found that this metric could make appropriately choose which source tasks should be used for a target task. For instance, to perform prompt transfer among the SuperGLUE tasks, WSC is a better source task for the CB task, while COPA benefits more from the RTE task.

Regarding (ii), inspired by the knowledge distillation (KD) paradigm [Hinton et al., 2015, Liu et al., 2021, Rao et al., 2022] that leverages a powerful teacher model to guide the training process of a student model, we propose a KD-based prompt transfer method that leverages the KD technique to transfer the knowledge from the source prompt to the target prompt in a subtle manner, thus effectively alleviating the problem of prior knowledge forgetting. An illustration of our proposed method is shown in Figure 3. More specifically, our KD-based prompt transfer approach first uses the PLM with the source prompt as the teacher network and the PLM with the randomly initialized prompt as the student network. Then, the student network is trained using the supervision signals from both the ground-truth labels in the target task and the soft targets predicted by the teacher network. Note that we only update the parameters of the student prompt, while keeping the other parameters fixed. Furthermore, to adaptively control the knowledge transfer process in our approach, we use the prompt similarity predicted by our metric as the balancing factor between the two supervision signals for each source-target pair.

Adversarial Fine-Tuning

In addition to the above transductive FT and prompt-tuning processed for deliberately solving the training-testing domain gap and low downstream resource problem, respectively, we also adopt the advanced adversarial fine-tuning algorithm [Miyato et al., 2019, Jiang et al., 2020] designed for PLMs, i.e., SiFT [He et al., 2021], to improve the training stability of our approach. In practice, we follow ?) by applying the perturbations to the normalized word embeddings when tuning our Vega foundation model on downstream tasks, where we first normalize the embedding vectors into stochastic vectors and then apply the perturbations to the normalized embedding vectors.

Experiments

For pretraining, we follow many prior works [Liu et al., 2019, He et al., 2021] and use Wikipediahttps://dumps.wikimedia.org/enwiki/ (the English Wikipedia dump, 10 GB), BookCorpus [Zhu et al., 2015]https://github.com/butsugiri/homemade_bookcorpus (6 GB), OpenwebTexthttp://Skylion007.github.io (38 GB), Storieshttps://github.com/tensorflow/models/tree/master/research/lm_commonsense (31 GB) and CC-News [Trinh and Le, 2018] (76 GB) as pretraining datasets, and use 40 NVIDIA DGX nodes with 320 A100 GPUs to train our Vega v2 model. It takes 30 days to finish phase-1 (pretraining, i.e., MLM) with 1 M steps. For phase-2, i.e., self-evolution training, we continuously train Vega v2 for 50K steps. During fine-tuning, we only apply our KD-based prompt transfer strategy to the low-resource tasks, e.g., CB, COPA, and WSC. For the other tasks, the vanilla full-parameter model-tuning method with adversarial and transductive fine-tuning is used. We use AdamW [Loshchilov and Hutter, 2018] as the optimizer for both pretraining and fine-tuning.

2 Tasks

To validate the effectiveness of Vega v2, we use the SuperGLUE benchmark [Wang et al., 2019b] for model evaluation purposes. As one of the most popular NLU benchmarks, SuperGLUE consists of eight challenging NLU tasks, including question answering (BoolQ, ?), MultiRC, ?), ReCoRD, ?)), natural language inference (CB, ?), RTE, ?; ?; ?; ?)), word sense disambiguation (WIC, ?)), coreference resolution (WSC, ?)), and reasoning (COPA, ?)). More detailed data statistics and examples for the above tasks can be found in Appendix Tables 3 and 4.

3 Main Results

Table 1 reports the test results obtained by our Vega v2 and other cutting-edge models on the SuperGLUE benchmarkWe show the detailed ranking results on the SuperGLUE leaderboard in the Appendix. Please refer to Table 2.. Vega v2 significantly surpasses the powerful human baselines in terms of average score (91.3 vs. 89.8) and achieves state-of-the-art performance on four (relatively) low-resource tasks, i.e., CB, COPA, RTE, and WSC. We attribute this success to the novel self-evolution learning mechanism and KD-based prompt transfer method. More specifically, the former enhances Vega v2’s ability to extract informative patterns, while the latter alleviates the problem of overfitting and boosts the model performance on low-resource tasks.

In addition, compared to the other larger PLMs, e.g., PaLM [Chowdhery et al., 2022b] which consists of 540 billion parameters, our 6-billion-parameter Vega v2 can achieve competitive or even better performance on the SuperGLUE benchmark. This inspires us to conclude that scaling PLMs to larger model sizes arbitrarily might not be cost-effective, but would encourage the PLMs to fully extract knowledge from the pretraining data when given a certain parameter budget.

Conclusion

This paper presents the JD Explore Academy large-scale Vega v2 PLM for the SuperGLUE benchmark. Based on an advanced transformer backbone with disentangled attention and a series of advanced fine-tuning strategies, we propose two novel techniques. The first is a self-evolution learning mechanism that fully exploits the knowledge contained in data for a PLM in two steps: 1) the PLM performs self-questioning to determine hard and informative words, and then 2) supervises the MLM process with rectified smooth labels. The second is a prompt transfer strategy for efficiently adapting downstream tasks (especially low-resource tasks) by leveraging the knowledge acquired from the foundation model and related downstream tasks.

We show that these techniques significantly improve the efficiency of model pretraining and the performance achieved on downstream tasks. The Vega v2 model with 6 billion parameters achieves state-of-the-art records on 4 out of 8 tasks and ranks the first in terms of the macro-average score. Our experience with building Vega v2 demonstrates the necessity of 1) fully improving the parameter efficiency of PLMs, and 2) wisely preforming downstream adaptation.

Acknowledgments

The authors wish to thank the leaderboard maintainer of SuperGLUE for their great construction efforts and their prompt responses to our questions. The authors also especially thank Mr. Yukang Zhang (JD Explore Academy), who kindly supports maintaining a stable computing platform.

References