Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE

Qihuang Zhong, Liang Ding, Keqin Peng, Juhua Liu, Bo Du, Li Shen, Yibing Zhan, Dacheng Tao

Introduction

Pretrained language models (PLMs) Devlin et al. (2019); Liu et al. (2019); He et al. (2021); Joshi et al. (2020); Sun et al. (2019); Brown et al. (2020); Raffel et al. (2020) are widely used in the community of natural language processing and have achieved remarkable success in numerous downstream tasks of natural language understanding (NLU), such as sentiment analysis Liu et al. (2021a); Wang et al. (2022a); Zhong et al. (2022a), intent detection Kim et al. (2016); Wu et al. (2020a) and reasoning Qu et al. (2022). These PLMs share a common principle of performing self-supervised learning with massive easy-to-acquire unlabelled text corpora during the pretraining stage and effectively fine-tuning on downstream tasks. In such a context, the general language understanding evaluation (GLUE, Wang et al. 2019) benchmark has emerged as the leading evaluation standard for the pretrained language model community, where most high-performing models (e.g., T5 Raffel et al. (2020)) on its leaderboard provide valuable insights and best practices for future research and applications.

We recently submitted our 1.3B Vega v1 model to the GLUE leaderboard and, as seen in Figure 1, obtained state-of-the-art records on 4 out of 9 tasks, sitting atop the leaderboard as of January 1, 2022, with an average score of 91.3. More encouragingly, our Vega v1 is the first to exceed powerful human performance on the two challenging tasks, i.e., SST-2 Socher et al. (2013) and WNLI Levesque et al. (2012). This technical report briefly describes how we build our powerful model under a certain parameter budget, i.e., 1.3B, from different aspects, including backbone framework (§2.1), efficient pretraining processes (§2.2), and effective downstream adaptation approaches (§2.3). To achieve efficient and sufficient pretraining, we replace the widely-used masked language modeling (MLM, Devlin et al. 2019) with two simple but effective objectives, i.e., denoising and contrastive objectives. The denoising objective Yamaguchi et al. (2021) aims to improve the data efficiency and save training costs, while the contrastive objective involves leveraging contrastive learning Gao et al. (2021) to learn better sentence representations. For downstream adaptation, we focus on two common problems, i.e., domain discrepancy and over-fitting, and adopt several effective fine-tuning methods, such as self-calibrated transductive fine-tuning and adversarial fine-tuning, to achieve better performance and generalization.

The rest of this paper is organized as follows. In Section 2, we introduce the major utilized approaches. Then, Section 3 reports and discusses our evaluation results. Finally, we conclude our study in Section 4.

Approaches

In this section, we describe the main techniques in our Vega v1 model, including the backbone framework in §2.1, the efficient pretraining approaches in §2.2, and the effective downstream adaptation technique in §2.3.

In recent years, the Transformer Vaswani et al. (2017) has become the de-facto standard for neural language modeling, and has achieved great success in the field of large-scale language model pretraining Devlin et al. (2019); Raffel et al. (2020); Brown et al. (2020); Liu et al. (2021b); Zhong et al. (2022b); Wang et al. (2022b); Zan et al. (2022b, c). In terms of network architectures, the Transformer-based PLMs can be classified into three groups: decoder-only models (e.g., GPT-3 Brown et al. (2020)), encoder-only model (e.g., BERT Devlin et al. (2019)) and encoder-decoder models (e.g., T5 Raffel et al. (2020)). As the encoder-only models have an overwhelming advantage over the existing methods on the GLUE leaderboard, we train our large model in an encoder-only manner to facilitate downstream language understanding tasks. In particular, the self-attention mechanism adopted in Transformer can effectively encode the content information, but unfortunately lacks a natural way to encode word position information Dufter et al. (2022); Ding et al. (2020). Thus, following He et al. (2021), we replace the vanilla self-attention mechanism in Transformer with a disentangled attention strategy. According to our Vega v1 parameter budget – 1.3 Billion, we empirically set the model as follows: 48 layers, 1,536 as the hidden layer size, an FFN of size 6,144, 24 heads, and 64 as the head size.

2 Efficient Pretraining

One key component of language model pretraining is the pretraining objective. Most of the prior existing pretrained models are usually based on the masked language modeling (MLM) objective Devlin et al. (2019); Liu et al. (2019) or its variants Sun et al. (2019); Joshi et al. (2020). However, the MLM is usually criticized as being inefficient, as it incurs a substantial compute cost (top-layer vocabulary-dimension embedding) but only produces supervision signals at a small proportion of positions (usually 15%) Clark et al. (2020). To this end, we propose two more efficient pretraining objectives (i.e., denoising and contrastive objectives) for effectively training our Vega v1, which are illustrated in Figure 2. Specifically, the denoising objective and contrastive objective are expected to capture the local token-level and global sentence-level knowledge, respectively. In this part, we first compare our objectives with the vanilla MLM, and then detailed introduce the two-stage pretraining algorithm used in our Vega v1.

MLM randomly selects a subset (usually 15%) of tokens from a sentence and replaces them with a special mask token, i.e., [MASK]. Then, the MLM trains a model to predict a particular token (from the vocabulary space) that has been replaced with a [MASK] placeholder given its surrounding context.

Instead of using the MLM, motivated by the success of ELECTRA Clark et al. (2020) and other similar works Yamaguchi et al. (2021); Alajrami and Aletras (2022), we use a more efficient replaced token detection (RTD) as an alternative. The principle of RTD is to manually corrupt the sentence and encourage the model to denoise the corrupted sentence. In practice, for each sentence, we replace 10% of tokens with shuffled ones from the same sequence and another 5% of tokens with random ones from the vocabulary. Then, the model is forced to learn the linguistic knowledge for detecting the shuffled or replaced tokens among the supervisions of all input tokens. It is noteworthy that, different from the MLM that predicts the token from the vocabulary space, our denoising objective is a ternary classification task, aiming at identifying whether a token in the input sequence has been shuffled or randomly replaced or not, which is more sample-efficient and saves large training costs.

In addition to the above denoising objective that focuses on token-level linguistic information learning, we further introduce a contrastive objective to improve the sentence-level representation learning ability of our Vega v1. Specifically, the motivation of this objective is that many prior studies Li et al. (2020); Gao et al. (2021) have found that Transformer-based models always induce a non-smooth anisotropic semantic space of sentences, which harms the performance of sentence representation.

To alleviate this problem, our contrastive objective adopts the contrastive learning technique on sentence representation. The two key problems of contrastive learning are 1) how to construct positive instances; and 2) how to obtain sentence representations. For problem 1), we simply use the original sentence and its corresponding noising sentence (corrupted with the RTD in the denoising objective) as the positive instance pair. As for problem 2), we can directly use the [CLS] embedding, or simply use basic pooling operations (e.g., mean pooling) to process the hidden representations on the last layer, as the sentence representation. Notably, in the preliminary experiments, we found that there is a slight difference between both methods, whereas the former one is lastly adopted for simplicity.

2.2 Two-stage Pretraining Strategy

Despite the remarkable performance improvement, the contrastive objective leads to external computation overhead, as it requires two forward-propagation processes. Thus, to better trade off the computation costs and performance, we perform the pretraining process in a two-stage manner. Specifically, in phase 1, we only employ the denoising pretraining objective to ensure the model quickly learns the basic token-level linguistic knowledge from pretraining data. In phase 2, we combine both objectives and continue pretraining the model to further encourage it to fully exploit the knowledge and learn better sentence representation. Note that such fine-to-coarse multi-stage training recipe has shown effective performance in many tasks, e.g., classification McDonald et al. (2007) and translation Ding et al. (2021).

3 Effective Downstream Adaptation

In addition to the above efficient and sufficient pretraining methods, we also design some useful fine-tuning strategies for effectively adapting our Vega v1 to downstream tasks. Here, we focus on two main problems that hinder the adaptation performance of our Vega v1. Specifically, (i) Domain gap. The first concerns the domain gaps between the training and test sets, which lead to poor performance on target test sets. (ii) Over-fitting. Due to the limited downstream training data or its hard-to-learn ability, the fine-tuning model usually suffers from the over-fitting problem and shows poor model generalization.

Note that in addition to the strategies listed below, we have also designed and implemented other methods from different perspectives to improve the generalization and efficiency of models, e.g. the FSAM optimizer for PLMs Zhong et al. (2022d), PromptTuning with reusing existing prompts Zhong et al. (2022c), SparseAdapter He et al. (2022), and continued training with downstream data Zan et al. (2022a). Although these approaches can help to some extent, they do not provide complementary benefits compared to the listed approaches, so they are not described here.

Regarding the domain or linguistic style gap between the training and test sets (the problem (i)), we adopt a transductive fine-tuning strategy to improve the target domain performance, which is a common practice in machine translation evaluations Wu et al. (2020b); Ding and Tao (2021) and some domain adaptation applications Liu et al. (2020). Specifically, let the training set (denoted as DsD_{s}) be the source domain and the test set (denoted as DtD_{t}) be the target domain, the key idea of transductive fine-tuning is to transform the target domain into the source domain space with the well-performed model M0M_{0} (trained on the source domain), and obtain the generated synthetic dataset. Then, the model is further tuned on this synthetic dataset. The proposed transductive fine-tuning technique is shown in Algorithm 1. Whether we should conduct transductive fine-tuning depends on the practical downstream performance achieved.

The aforementioned transductive fine-tuning works well when the well-performed finetuned model M0M_{0} is available, but performs poorly when it is hard to train the base model MM. For example, there is usually limited labeled data in the downstream training set (low-resource settings), which hinders the effective training of M0M_{0}. A natural way is to leverage the external (easy-to-obtain but low-quality) labeled data or even unlabeled data (denoted as D∗D^{*}) from a similar source domain for training the M0M_{0}. Obviously, it is sub-optimal or even impractical to directly apply the external noisy data to obtain the satisfactory M0M_{0}.

To this end, we further present a new self-calibrated strategy to boost the effectiveness of transductive fine-tuning. In practice, the self-calibrated strategy contains a three-stage process:

1) Selecting the high-quality sub-dataset Ds∗D^{*}_{s} (similar to DsD_{s}) from D∗D^{*} estimated by a language model (e.g., KenLM Heafield (2011)) trained on DsD_{s};

2) Calibrating Ds∗D^{*}_{s} via adopting the initialized M0M_{0} to re-label the data;

3) Tuning the M0M_{0} on the calibrated external data to obtain the well-performed M0′M^{{}^{\prime}}_{0} and further performing the transductive fine-tuning.

Regarding the problem (ii), we are inspired by many prior adversarial training methods Miyato et al. (2019); Jiang et al. (2020); Zhang et al. (2022) and adopt an adversarial fine-tuning to alleviate the over-fitting problem. In practice, to further improve the training stability, we follow He et al. (2021) and apply the perturbations to the normalized word embeddings when tuning our Vega foundation model on downstream tasks, where we first normalize the embedding vectors into stochastic vectors and then apply the perturbations to the normalized embedding vectors. Similar to the observations of He et al. (2021), we also empirically find that such a normalization improves the generalization and performance of the fine-tuned models on several downstream tasks.

Experiments

For pretraining, we follow many prior works Liu et al. (2019); He et al. (2021) and use Wikipediahttps://dumps.wikimedia.org/enwiki/ (the English Wikipedia dump, 10 GB), BookCorpus Zhu et al. (2015)https://github.com/butsugiri/homemade_bookcorpus (6 GB), OpenwebTexthttp://Skylion007.github.io (38 GB), Storieshttps://github.com/tensorflow/models/tree/master/research/lm_commonsense (31 GB) and CC-News Trinh and Le (2018) (76 GB) as pretraining corpus. For preprocessing, we follow He et al. (2021) and use the same BPE vocabulary. We use 40 NVIDIA DGX nodes (each with 8×\times40GB A100 GPU cards) to train our Vega v1 model with a mixed precision training strategy. It takes about 20 days to finish phase-1 (denoising) pretraining with 1 M steps. For phase-2, i.e., contrastive-augmented training, we continuously train Vega v1 for 100K steps. AdamW Loshchilov and Hutter (2018) is used as the optimizer for the pretraining stage.

2 Downstream Tasks

To validate the effectiveness of Vega v1, we use the widely-used GLUE benchmark Wang et al. (2019) as the test bed. As one of the most popular NLU benchmarks, GLUE consists of nine challenging NLU tasks, including linguistic acceptability (CoLA, Warstadt et al. (2019)), sentiment analysis (SST-2, Socher et al. (2013)), paraphrase (MRPC, Dolan and Brockett (2005)), textual similarity (STS-B, Cer et al. (2017)), question paraphrase (QQP), textual entailment (MNLI, Williams et al. (2018), RTE, Giampiccolo et al. (2007)), question-answer entailment (QNLI, Rajpurkar et al. (2016)), and coreference resolution (WNLI, Levesque et al. (2012)). More detailed data statistics for the above tasks can be found in Appendix (Table 3).

During fine-tuning, we only apply our self-calibrated fine-tuning strategy to CoLA, as it usually suffers from the domain discrepancy between training and test sets, and our strategy can effectively alleviate this problem. Additionally, transductive fine-tuning is used for WNLI, and scale-invariant fine-tuning is used for QNLI, QQP, and MNLI, while vanilla fine-tuning is for the others. Notably, as suggested by Liu et al. (2019), for RTE, STS-B, and MRPC tasks, we first fine-tune our Vega v1 model on the MNLI dataset and then continue fine-tuning on their corresponding single-task corpus for better performance.

3 Main Results

Table 1 reports the final results on the test sets obtained by our Vega v1 and other cutting-edge models on the GLUE benchmarkWe show the detailed ranking results on the GLUE leaderboard in the Appendix. Please refer to Table 2.. As seen, our Vega v1 surpasses the human baselines in terms of average score (91.3 vs. 87.1) by a large margin, and achieves new record state-of-the-art performance among four tasks. More encouragingly, Vega v1 is the first to exceed powerful human performance on the two challenging tasks, i.e., SST-2 (human: 97.8% vs. Vega v1: 97.9%) and WNLI (human: 95.9% vs. Vega v1: 97.9%). We attribute this success to the efficient pretraining objectives, i.e., denoising and contrastive-augmented objectives. Specifically, the sample-efficient denoising objective used in phase 1 ensures sufficient training of Vega v1. On the other hand, with the help of the proposed contrastive-augmented objective, our Vega v1 can learn better sentence representations, which are beneficial to the downstream language understanding tasks.

In general, our Vega v1 outperforms the other cutting-edge counterparts and sits atop the GLUE benchmark ranking as of January 1, 2022. These results show the superiority and effectiveness of our model, indicating the significance of more efficient pretraining objectives.

Conclusion

This paper presents the JD Explore Academy large-scale Vega v1 PLM for the GLUE benchmark. Based on an advanced transformer backbone with disentangled attention, we propose two novel techniques along the pipeline of pretraining and fine-tuning. During the pretraining, we present two efficient pretraining objectives to encourage our Vega v1 model to fully exploit linguistic knowledge and learn better sentence representations. For fine-tuning, a new self-calibrated strategy is proposed to address the domain discrepancy problem though making full use of the model calibration ability.

We show that these techniques significantly improve the efficiency of model pretraining and the performance achieved on downstream tasks. Specifically, results on the GLUE benchmark show that our Vega v1 model with 1.3 billion parameters achieves state-of-the-art records on 4 out of 9 tasks and ranks first in terms of the macro-average score. Our experience with building Vega v1 demonstrates the necessity of 1) exploring more efficient pretraining objectives, and 2) wisely performing downstream adaptation.

We also scale the Vega model to a much larger one – Vega v2 Zhong et al. (2022e) and interestingly find that when scaling up the model size, the current denoising self-supervised objective RTD used in Vega v1 is lightweight but unstable compared to the MLM objective. Therefore, we suggest adopting our training recipes (in Vega v1) for language models smaller than 5 billion, and for larger model scale budgets, the participators could refer to our Vega v2 report Zhong et al. (2022e).

Acknowledgments

The authors wish to thank the leaderboard maintainer of GLUE for their great efforts in the construction, and their prompt responses. The authors also specially thank Mr. Yukang Zhang (JD Explore Academy), who kindly supports maintaining a stable computing platform. Lastly, the R&D of the Vega foundation model series could not have been done without the discussions and support of the JDEA-NLP group.

References