COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
Yu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary, Paul Bennett, Jiawei Han, Xia Song
Introduction
Pretrained language models (PLMs) have reshaped the way AI systems process natural language . Before task-specific training, it is now a common practice to first pretrain the deep neural networks, often Transformers , via a self-supervised token-level language modeling task . Whether it is autoregressive , permutational , or masked language modeling (MLM) , the Transformer networks are pretrained to recover some omitted tokens using the rest of input texts. Then the language semantics captured during pretraining are conveyed to downstream tasks via the pretrained Transformer parameters .
Recent research observed several challenges in this self-supervised learning framework. One challenge is its efficiency. After pretrained for a while with the standard token-level language modeling, the networks have already captured the basic language patterns, making a large fraction of pretraining signals no longer informative. Linear improvement in the model effectiveness often requires exponentially more pretraining compute and parameters , which is unsustainable. Another challenge is the anisotropy of text representations from pretrained models. The sequence representations from many pretrained models are quite irregular and require dedicated fine-tuning approaches to be useful in sequence-level applications .
Clark et al. proposed a new pretraining strategy, ELECTRA, that uses an auxiliary language model (“generator”) to replace tokens in input texts and pretrains the main Transformer (“discriminator”) to detect replaced tokens. This improves the pretraining efficiency and effectiveness, but pretraining via binary classification hinders the model’s usage on applications requiring language modeling capability (e.g., prompt-based learning ). It could further distort the representation space as the Transformers are pretrained to output the same “non-replacement” label for all actual tokens.
In this paper, we present a new self-supervised learning approach, COCO-LM, that pretrains Language Models by COrrecting and COntrasting corrupted text sequences. Following ELECTRA-style pretraining, COCO-LM employs an auxiliary model to corrupt the input texts, upon which it introduces two new pretraining tasks for the main Transformer, one at token level and one at sequence level. The token-level task, corrective language modeling (CLM), pretrains the main Transformer to detect and correct the tokens in the corrupted sequences. It uses a multi-task setup to combine the benefits of replaced token detection and language modeling. The sequence-level task, sequence contrastive learning (SCL), pretrains the model to align text sequences originated from the same source sequence and enforce uniformity of the representation space.
In our experiments on GLUE and SQuAD benchmarks, COCO-LM not only outperforms state-of-the-art pretraining approaches in effectiveness, but also significantly improves the pretraining efficiency. Under the same setting, COCO-LM matches the MNLI accuracy of RoBERTa and ELECTRA with and of their GPU hours in pretraining, respectively. When pretrained with the same number of steps, COCO-LM outperforms the previous best models by GLUE average points under the standard base/large-sized model evaluations. With million parameters, COCO-LM reaches the MNLI accuracy of Megatron , one of the largest BERT-style model with billion parameters. Our analyses provide further insights on the advantage of CLM in learning token representations and its effectiveness in prompted-based fine-tuning, as well as the benefit of SCL in ensuring alignment and uniformity in the representation space for better generalizationCode and pretrained models can be found at https://github.com/microsoft/COCO-LM..
Related Work
Various token-level tasks have been used to pretrain language models. The most classic auto-regressive language modeling is to predict a token given all the previous tokens, or all subsequent ones . BERT uses masked language modeling (MLM) that recovers randomly masked tokens using the rest input. XLNet proposes permutation language modeling that conducts MLM in an autoregressive manner . UniLM uses pseudo MLM which unifies autoregressive and MLM tasks .
Sequence-level tasks are also explored, which often pretrain the model to predict certain co-occurrences of sequence pairs. For example, next sentence prediction , sentence ordering and previous sentence prediction concatenate two sentences (either correlated or random), and train the Transformer to classify the pair.
Empirically, MLM is still among the most effective tasks to pretrain encoders . RoBERTa found the sentence-level task in BERT not benefitial and discarded it. BART and T5 both observed that MLM is often the most effective task. The empirical advantages of other pretraining tasks are more task-specific, for example, entity related masks for knowledge intensive applications , and sequence-level tasks for long form text modeling .
Instead of randomly altering texts, ELECTRA uses a smaller auxiliary Transformer pretrained by MLM to replace some tokens in the text sequences using its language modeling probability, and pretrains the main Transformer to detect the replaced tokens. ELECTRA achieves state-of-the-art accuracy in many language tasks . Later, Clark et el. developed ELECTRIC, which pretrains encoders by contrasting original tokens against negatives sampled from a cloze model. ELECTRIC re-enables the language modeling capability but underperforms ELECTRA in downstream tasks.
Our work is also related to contrastive learning which has shown great success in visual representation learning . Its effectiveness of in language is more observed in the fine-tuning stage, for example, in sentence representation , dense retrieval , and GLUE fine-tuning .
Method
We present the preliminaries of PLMs, their challenges, and the new COCO-LM framework.
In this work we focus on pretraining BERT-style bidirectional Transformer encoders that are widely used in language representation tasks. We first recap the masked language modeling (MLM) task introduced by BERT and then discuss the pretraining framework of ELECTRA .
BERT Pretraining uses the masked language modeling task (MLM) , which is to take an input sequence , with random tokens replaced by [MASK] symbols (e.g., the -th token), and train the model to predict the original tokens at the masked positions:
where the Transformer generates contextualized representations . The MLM Head predicts the masked token from the vocabulary using the hidden representation and token embeddings . The pretraining minimizes the MLM loss on the set of masked positions . Specifically,
ELECTRA Pretraining uses two Transformers, a “generator” pretrained by MLM, and a “discriminator” pretrained using the generator’s outputs. We refer them as auxiliary and main Transformers, as the former is discarded after pretraining and the latter may be trained by “generative” tasks too.
The auxiliary model outputs a corrupted sequence by sampling from its predicted probability:
The masked positions are replaced by sampled tokens considered plausible in context by the auxiliary Transformer, which are more deceiving than random replacements. ELECTRA uses a skinnier auxiliary network (e.g., hidden dimension is of the main model) to control the signal difficulty.
The main Transformer takes and classifies the replaced tokens:
The two Transformers are pretrained jointly. The auxiliary model gradually generates more realistic replacement tokens and the main model learns to better detect them. This forms a natural learning curriculum and significantly improves ELECTRA’s accuracy in downstream tasks .
2 Challenges of ELECTRA-Style Pretraining
Missing Language Modeling Benefits. The classification task in ELECTRA is simpler and more stable , but raises two challenges. The first is the lack of language modeling capability which is a necessity in some tasks . For example, prompt-based learning requires a language model to generate labels . The second is that the binary classification task may not be sufficient to capture certain word-level semantics that are critical for token-level tasks.
Squeezing Representation Space. Another challenge is that the representations from Transformer-based language models often reside in a narrow cone, where two random sentences have high similarity scores (lack of uniformity), and closely related sentences may have more different representations (lack of alignment) . Figure 1 illustrates such behaviors with random sentence pairs (from pretraining corpus) and semantically similar pairs (those annotated with maximum similarity from STS-B ). With RoBERTa, the cosine similarities of most random sentence pairs are near , bigger than many semantically similar pairs. The representation space from ELECTRA is even more squeezed. Nearly all sentence pairs, both random and similar ones, have around cosine similarity. This may not be surprising as ELECTRA is pretrained to predict the same output (“non-replacement”) for all tokens in these sequences. The irregular representation space raises the risk of degeneration and often necessitates sophisticated post-adjustment or fine-tuning to improve the sequence representations .
3 COCO-LM Pretraining
COCO-LM also employs an auxiliary Transformer to construct the corrupted text sequence, as in Eqn. (1), but it introduces two new pretraining tasks upon the corrupted sequences to address the challenges previously described. In the rest of this section, we present these two tasks and then the detailed configurations of COCO-LM. Its framework is illustrated in Figure 2.
Corrective Language Modeling (CLM) trains the main Transformer to recover the original tokens, given the corrupted text sequence :
The CLM Head uses the hidden representations to output a language modeling probability, instead of a binary classification score. The forward pass of the CLM Head is the same as All-Token MLM, a variation of ELECTRA that consists of a language modeling layer and a binary classification layer for the copy mechanism:
where is a learnable weight and is the copy mechanism ( when the input token is original and can be directly copied to the output; when the input token needs to be corrected to another token from the vocabulary).
In ELECTRA, All-Token MLM performs worse than RTD . Language modeling on the corrupted text sequence is hard as the replaced tokens from the auxiliary model are more deceiving than [MASK]. To improve the language model learning, different from All-Token MLM, CLM employs a multi-task setup that combines the RTD task to explicitly train the copy mechanism :
The hyperparameter balances the weights of the two tasks. The binary cross entropy loss in Eqn. (2) explicitly trains the copy probability. We also use stop gradient (sg) to decouple the gradient backpropagation to from the LM task. This way, the main Transformer first learns the easier classification task and then uses it to help learn the harder LM task. The binary classification task is trained on all tokens while the language modeling task is trained only on masked positions.
CLM combines the advantages of MLM and ELECTRA: The main Transformer is trained on all tokens with the help of the binary classification task while also being able to predict words, thus enjoying the efficiency benefits of ELECTRA and preserving the language modeling benefits.
Sequence Contrastive Learning (SCL) forms a contrastive learning objective upon the sequence embeddings to learn more robust representations. Broadly, contrastive learning is to align a positive pair of instances, often different views of the same information , in contrast to unrelated negative instances . The different views are often obtained by applying data augmentations on the same input, for example, rotation, cropping, and blurring on visual representations , so that the neural networks can learn representations robust to these data alterations.
In COCO-LM, the corrupted sequence already provides a form of data augmentation. We pair it with another augmentation, , a randomly cropped contiguous span of (the length of is of so that the major sequence meaning is preserved), to construct the positive pair and to contrast with random negatives.
Specifically, a training batch in SCL includes a random set of corrupted and cropped sequences: , with and originated from . A positive contrastive pair consists of either or (symmetrical contrast). The negative instances are all the remaining sequences in the batch . The contrastive loss is formulated as:
where are the representations of , respectively, from the main Transformer (i.e., ). The similarity metric is cosine similarity (cos) and the temperature is set to .
As shown in Wang et al. , the first term in Eqn. (3) () improves alignment of the space. It encourages representations to be robust to the corruptions and the alterations on the original text. The second term in Eqn. (3) promotes uniformity. It pushes unrelated sequences apart in the representation space and ensures low cosine similarity between random data points. Several studies have observed improved generalization ability from better alignment and uniformity .
Aligning with requires the main Transformer to produce sequence representations robust to both token-level (i.e., MLM replacements) and sequence-level (i.e., cropping) alterations. The model is thus encouraged to reason more using partially altered sequences to recover the original information.
Overall Training. COCO-LM uses the following loss function:
The auxiliary Transformer is pretrained by masked language modeling (MLM) and generates corrupted sequences. The main Transformer is pretrained to correct the corruption (CLM) and to contrast the corrupted sequences with the cropped sequences (SCL). The two Transformers are pretrained jointly with the loss in Eqn. (4). The main Transformer is used in downstream applications.
Network Configurations. Similar to ELECTRA, the auxiliary Transformer is smaller than the main model, but we use different configurations in the auxiliary model: (1) We reduce the number of layers to or (under base or large model setup, respectively) but keep its hidden dimension the same with the main model, instead of shrinking its hidden dimensions; (2) We disable dropout in it when sampling replacement tokens. We find such configurations empirically more effective and use them as the backbone of COCO-LM. The main Transformer follows the standard architecture of BERT/ELECTRA and can be easily adopted by downstream application pipelines with almost no changes.
Experimental Setup
Pretraining Settings. We employ three standard settings, base, base++, and large++. Base is the BERT training configuration : Pretraining on Wikipedia and BookCorpus ( GB of texts) for million samples on token sequences (K batches with batch size). We use the same corpus and uncased BPE vocabulary as with TUPE .
Base++ trains the base size model with larger corpora and/or more training steps. Following recent research , we add in OpenWebText , CC-News , and STORIES , to a total of GB texts, and train for billion (with batch size) samples . We follow the prepossessing of UniLMV2 and use cased BPE vocabulary.
Large++ uses the same training corpora as base++ and pretrains for billion samples ( batch size). Its Transformer configuration is the same with BERT .
Model Architecture. Our base/base++ model uses the BERT architecture : layer Transformer, hidden size, plus T5 relative position encoding . Our large++ model is the same with BERT, layer and hidden size, plus T5 relative position encoding . Our auxiliary network uses the same hidden size but a shallow -layer Transformer in base/base++ and a -layer one in large++. When generating we disable dropout in the auxiliary model.
Downstream Tasks. We use the tasks included in GLUE and SQuAD 2.0 reading compression . Please refer to Appendix A for more details about GLUE tasks. Standard hyperparameter search in fine-tuning is performed, and the search space can be found in Appendix B. The fine-tuning protocols use the open-source implementation of TUPE . The reported results are the median of five random seeds on GLUE and SQuAD.
Baselines. We compare with various pretrained models in each setting. To reduce the variance in data processing/environments, we also pretrain and fine-tune RoBERTa and ELECTRA under exactly the same setting with COCO-LM, marked with “(Ours)”. All numbers unless marked by “(Ours)” are from reported results in recent research (more details in Appendix C).
Implementation Details. Our implementation builds upon the open-source implementation from MC-BERT and fairseq . More implementation details are mentioned in Appendix D.
Evaluation Results
Three groups of experiments are conducted to evaluate COCO-LM and its two new pretraining tasks.
Overall Results are listed in Table 1. Under all three settings, COCO-LM outperforms all recent state-of-the-art pretraining models on GLUE average and SQuAD. It improves the state-of-the-art GLUE score by about one point under all three settings. COCO-LM also enjoys better parameter efficiency. Using less than of Megatron’s parameters, COCO-LM matches the MNLI accuracy of Megatron, one of the largest pretrained BERT-style encoders.
Table 2 shows GLUE test set results which further confirm the advantages of COCO-LM over previous methods.
Efficiency. In downstream tasks, the efficiency of COCO-LM is the same with BERT. In pretraining, the auxiliary model and SCL introduce extra cost. However, as shown in Figure 4, COCO-LM is more efficient in GPU hours. It outperforms RoBERTa & ELECTRA by points on MNLI with the same GPU hours and reaches their accuracy with around & GPU hours, respectively.
Ablation Studies. Table 3 shows the ablations of COCO-LM under the base setting on GLUE DEV.
Pretraining Task. With only RTD, our backbone model with the shallow auxiliary Transformer is quite effective. CLM and SCL both provide additional improvements on MNLI and GLUE average. Their advantages are better observed on different tasks, for example, CLM on MNLI-mm and SCL on RTE and MRPC. Combining the two in COCO-LM provides better overall effectiveness. In later experiments, we further analyze the benefits of these two tasks.
Architecture. Removing relative position encoding (Rel-Pos) leads to better numbers on some tasks but significantly hurts MNLI. Using a shallow auxiliary network and keeping the same hidden dimension () is more effective than ELECTRA’s -layer but -hidden dimension generator.
Pretraining Signal Construction. Using randomly replaced tokens to corrupt text sequence hurts significantly. Using a converged auxiliary network to pretrain the main model also hurts. It is better to pretrain the two Transformers together, as the auxiliary model gradually increases the difficulty of the corrupted sequences and provides a natural learning curriculum for the main Transformer.
CLM Setup. Disabling the multi-task learning and using All-Token MLM reduces model accuracy. The copy mechanism is effective. The benefits of the stop gradient operation are more on stability (preventing training divergence).
2 Analyses of Contrastive Learning with SCL
This group of experiments analyzes the behavior of SCL. All experiments use the base setting.
Ablation on Data Augmentation. Figure 3(b) shows the effects of the cropping operation when forming positive SCL pairs with the corrupted sequence. Using the original sequence results in worse GLUE accuracy. It is less informative as the model no longer needs to learn representations robust to sequence-level alteration. Cropping too much (e.g., only keeping of the original sequence), may hurt as it can alter the semantics too much. Empirically a simple alteration works the best, similar to the observations in recent research .
Alignment and Uniformity. Figure 6 plots the distribution of cosine similarities between random sequence pairs and similar ones using representations pretrained by COCO-LM. The representation space from COCO-LM is drastically different from those in Figure 1. With COCO-LM, similar pairs are more aligned and random pairs are distributed more uniformly. Many similar pairs have near cosine similarity and are clearly separated from random pairs which center around . The t-SNE plot in Figure 6 further demonstrates the benefits of SCL. The similar sentence pairs (marked by same shapes) are aligned closer when pretrained with SCL. Their average cosine similarity is when pretrained with SCL, while is without SCL. This better alignment and uniformity is achieved by COCO-LM with SCL via pretraining, without using task-specific data nor supervised labels.
Regularizing the Representation Learning for Better Few-Shot Ability. One would expect any pretrained Transformers to easily align a pair of corrupted sequence and cropped sequence as the two share about tokens. However, as shown in Figure 7(a), that is not the case: Without SCL, the cosine similarity of the positive pairs is even lower than random negatives. SCL is necessary to regularize the representation space and to reduce the risk of degeneration (Figure 7(b)).
Similar to empirical observations and theoretical analyses in recent research , a more regularized representation space results in better generalization ability in scenarios with limited labels. Figure 7(c) and 7(d) show the results when COCO-LM are trained (via standard fine-tuning) with only a fraction of MNLI labels. The improvements brought by SCL are more significant when fewer fine-tuning labels are available. With MNLI labels, pretraining with SCL improves MNLI-m/mm accuracy by compared to that without SCL. Using only / labels, COCO-LM with SCL reaches similar MNLI accuracy with RoBERTa (Ours)/ELECTRA (Ours) fine-tuned with all labels, respectively.
3 Analyses of Language Modeling with CLM
The last group of experiments studies the effectiveness and benefits of CLM.
Ablations on Training Configurations. Figure 8 illustrates pretraining process with CLM and All-Token MLM. The plots demonstrate the difficulty of language modeling upon corrupted text sequences. It is quite an unbalanced task. For the majority of the tokens (Original) the task is simply to copy its input at the same position. For the replaced tokens ( total), however, the model needs to detect the abnormality brought by the auxiliary model and recover the original token. Implicitly training the copy mechanism as part of the hard LM task is not effective: The copy accuracy of All-Token MLM is much lower, and thus the LM head may confuse original tokens with replaced ones. As shown in Table 3 and ELECTRA , pretraining with All-Token MLM performs worse than using the RTD task, though the latter is equivalent to only training the copy mechanism. The multi-task learning of CLM is necessary for the main Transformer to stably learn the language modeling task upon the corrupted text sequence.
Prompt-Based Fine-Tuning with CLM. Table 4 includes the prompt-based fine-tuning experiments on MNLI for RoBERTa and COCO-LM under base++ and large++ sizes, following the same few-shot manual prompt fine-tuning with demonstration setup in LM-BFF . We use for the learning rate search of COCO-LM base++/large++ model, with everything else kept same as described in LM-BFF. With exactly the same pipeline, COCO-LM outperforms RoBERTa under both base++ and large++ sizes by significant margins on MNLI-m/mm. Such observations are interesting as COCO-LM’s main Transformer does not even see any [MASK] tokens during pretraining but still performs well on predicting masked tokens for prompt-based learning. Note that ELECTRA and COCO-LM variants without the CLM task are not applicable: Their main Transformers are not pretrained by language modeling tasks (thus no language modeling capability is learned to generate prompt label words). This points out the importance, if not necessity, of COCO-LM in the family of ELECTRA-style pretraining models. With the benefits and rapid developments of prompt-based approaches, the lack of language modeling capability is going to limit the potential of ELECTRA’s self-supervised learning framework in many real-world scenarios. COCO-LM not only addresses this limitation but also provides better prompt-based learning results.
Conclusions and Future Work
In this paper, we present COCO-LM, which pretrains language models using Corrective Language Modeling and Sequence Contrastive Learning upon corrupted text sequences. With standard pretraining data and Transformer architectures, COCO-LM improves the accuracy on the GLUE and SQuAD benchmarks, while also being more efficient in utilizing pretraining computing resources and network parameters.
One limitation of this work is that the contrastive pairs are constructed by simple cropping and MLM replacements. Recent studies have shown the effectiveness of advanced data augmentation techniques in fine-tuning language models . A future research direction is to explore better ways to construct contrastive pairs in language model pretraining.
Despite the empirical advantage of this auxiliary-main dual model framework, the auxiliary Transformer training is not influenced by the main Transformer nor learns to generate the optimal pretraining signals for the main model. To better understand and tailor the training of the auxiliary model to the main model is another important future research direction.
Acknowledgments
We sincerely thank Guolin Ke for discussions and advice on model implementation. We also thank anonymous reviewers for valuable and insightful feedback, especially the suggestion of adding prompt-based fine-tuning experiments.
References
Appendix A GLUE Tasks
We provide more details of the tasks included in the GLUE benchmark. Their statistics are listed in Table 5.
MNLI: Multi-genre Natural Language Inference contains K train examples obtained via crowdsourcing. The task is to predict whether a given premise sentence entails, contradicts or neutral with respect to a given hypothesis sentence.
QQP: Question Pairs contains K train examples from the Quora question-answering website. The task is to determine whether a pair of questions asked are semantically equivalent.
QNLI: Question Natural Language Inference contains K train examples derived from the Stanford Question Answering Dataset (SQuAD) . The task is to predict whether a given sentence contains the answer to a given question sentence.
SST-2: Stanford Sentiment Treebank contains K train examples extracted from movie reviews with human-annotated sentiment scores. The tasks is to determine if the sentence has positive or negative sentiment.
CoLA: Corpus of Linguistic Acceptability contains K train examples from books and journal articles on linguistic theory. The task is to determine whether a given sentence is linguistically acceptable or not.
RTE: Recognizing Textual Entailment contains K train examples from textual entailment challenges. The task is to predict whether a given premise sentence entails a given hypothesis sentence or not.
MRPC: Microsoft Research Paraphrase Corpus contains K train examples from online news sources. The task is to predict whether two sentences are semantically equivalent or not.
STS-B: Semantic Textual Similarity contains K train examples drawn from multiple sources with human annotations on sentence pair semantic similarity. The task is to predict how semantically similar two sentences are on a to scoring scale.
Appendix B Hyperparameter Settings
Tuning hyperparameter of pretraining is often too costly and we keep most hyperparameters as default. The auxiliary MLM pretraining uses the standard [MASK] ratio. The crop transformation in the SCL task uses crop ratio, resulting in a sub-sequence that is long of the original sequence. The softmax temperature in the SCL task is . All pretraining tasks in COCO-LM have equal weights except since the loss of the binary classification task is much lower than those of the LM tasks, which are over -way classification tasks. All token embeddings (used in the input embedding layer and the language modeling head) are shared between the auxiliary Transformer and the main Transformer. The detailed hyperparameters used are listed in Table 6 for pretraining, and Tables 7 and 8 for GLUE and SQuAD fine-tuning, respectively.
All reported methods use exactly the same (or equivalent) set of hyperparameters for pretraining and fine-tuning for fair comparison. For COCO-LM and all the baselines implemented under our setting, all fine-tuning hyperparameters are searched per task; the median results of five runs with the same set of five different random seeds are reported on GLUE and SQuAD.
Appendix C The Origins of Reported Baseline Scores
The baseline results listed in Table 1 are obtained from their original papers except the following: BERT from Bao et al. , RoBERTa base/base++ GLUE from and SQuAD from Bao et al. , ELECTRA base/base++ GLUE from Xu et al. , XLNet base++ from Bao et al. , RoBERTa base++ SQuAD from Bao et al. . When multiple papers report different scores for the same method, we use the highest of them in our comparisons.
Appendix D More Implementation Details
Pretraining and Fine-tuning Costs. The pretraining cost of COCO-LM’s CLM task is similar to ELECTRA, which is BERT plus the auxiliary network whose size is of the main network. The addition of SCL task requires one more forward and backward pass on the cropped sequence . With V100 ( GB Memory), one pretraining run takes about hours in base setting, about two-three weeks in base++ setting, and about three-four weeks in large++ setting. The fine-tuning costs are the same with BERT plus relative positive encodings as the same Transformer model is used.
MLM Mode for Corrective Language Modeling. When creating the MLM replaced sequence , we find it slightly improves the downstream task performance to disable dropout (i.e., set the auxiliary MLM in inference mode) for computing the auxiliary network’s output distribution where plausible replacing tokens are sampled. We hypothesize that this leads to more stable generation of challenging replaced tokens to be corrected by the main Transformer and thus improves downstream task results.
Projection Heads. For the auxiliary model trained with MLM, we follow the standard MLM head setup in BERT/RoBERTa that includes a linear layer to project the contextualized embeddings from the encoder to same-dimensional vectors before feeding to the final linear layer that outputs the MLM probability. However, we do not include the projection layer for the main model trained with the CLM task (i.e., only having the final linear layer). We find this improves the training stability.
Masking Special Tokens for Auxiliary Model Training. BERT only masks real tokens (other than artificial symbols like [SEP] and [CLS]) for MLM training, while RoBERTa also masks special tokens. We follow the RoBERTa setting which results in slightly improved performance for some tasks.
Appendix E More Discussions on PLM Research
Currently, the biggest challenge with PLM research is perhaps its prohibitive computation cost. On one hand, PLMs have influenced a wide range of tasks, and any further technical improvement matters a lot for downstream applications. On the other hand, its expensive computing cost and long experimental cycles pose great challenges for careful and thorough studies of the problem space, as any test of new designs comes with a considerable computing cost—pretraining a new language model can easily consume thousands of dollars, or even millions for extra large models.
Such challenges call for more systematic evaluation pipelines that can accurately and reliably judge whether or not a new PLM is really better than previous ones. Currently, the evaluation of PLMs largely relies on GLUE-style benchmark which contains a set of different tasks that are weighed equally for PLM evaluations—usually the average performance over these tasks is treated as a final measure for the effectiveness of a PLM. However, we find that the small tasks in GLUE have very high variances which may provide unreliable indications for a PLM’s performance. For example, on CoLA and RTE, fine-tuning with different random seeds from the same pretrained checkpoint can easily result in a -point difference between the best and the worst seed. In contrast, large tasks like MNLI give relatively stable and consistent results for the same model pretrained/fine-tuned with different random seeds, and thus serve as better indicators for PLMs’ effectiveness.
In this paper, we try to improve the robustness of our observations, for example, by reporting the downstream performance with different training time for future comparisons under limited computing budget, and also by making our code and models publicly available for the reproducibility of our study. We hope our efforts will facilitate more future research to improve the community’s understanding and development of this important problem space.