NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
Xingcheng Yao, Yanan Zheng, Xiaocong Yang, Zhilin Yang
Introduction
Pretrained language models (PLMs) have drawn much attention from the natural language processing (NLP) community. Neural networks based on the Transformer architecture (Vaswani et al., 2017) are trained on large general corpora for self-supervised language modeling tasks such as masked language modeling (Devlin et al., 2019; Liu et al., 2019; Raffel et al., 2019), autoregressive language modeling (Radford et al., 2018; Brown et al., 2020), permutation language modeling (Yang et al., 2019), etc, and then are finetuned on a small amount of labeled data for downstream tasks. This pretraining-finetuning framework has significantly improved the performance of many NLP tasks.
However, while considered effective, large-scale pretraining is usually computationally expensive. For example, RoBERTa-Large (Liu et al., 2019), a widely-used PLM, consumes a computational cost of FLOPsIt was pretrained with 1,000 V100 GPUs each with 32GB memory for approximately one day.. Larger PLMs such as GPT-3 (Brown et al., 2020) consume 50 times more FLOPs for training than RoBERTa-Large. The expensiveness of large-scale pretraining prevents many research groups with limited budgets from pretraining customized language models, exploring new neural architectures, or improving pretraining loss functions. In contrast, a large number of NLP researchers resort to improving the finetuning algorithms, whose performance is largely upper-bounded by the pretraining procedure. This creates a high barrier of NLP research and might not be ideal for the long-term development of the field.
Even though there have been efforts devoted to studying and improving the efficiency of language model pretraining (Clark et al., 2020; So et al., 2021; Tay et al., 2021; Chen et al., 2021), most of them focus on designing sample-efficient self-supervised tasks or discovering efficient Transformer architectures suitable for pretraining. Their improvements are limited, with a reduction of computational costs (in terms of FLOPs) less than one order of magnitude. Another line of works target reducing the sizes of PLMs using distillation (Sanh et al., 2019; Jiao et al., 2020) to improve the efficiency of inference, but these methods rely on pretraining a large PLM before distillation. Moreover, distilled models often do not perform as well as some of the best non-distilled PLMs such as RoBERTa-Large (Sanh et al., 2019; Jiao et al., 2020).
This work explores alternatives to the standard pretraining-finetuning paradigm, aiming at more drastic efficiency improvement without performance drop. We propose a simple, efficient, pretraining-free framework, Task-driven Language Modeling (TLM). Given a large general corpus and some labeled task data, TLM directly trains a model from scratch without relying on PLMs. TLM is motivated by two key ideas. First, humans master a task by using only a small portion of world knowledge (e.g., students only need to review a few chapters, among all books in the world, to cram for an exam). We hypothesize that there is much redundancy in the large corpus for a specific task. Second, training on supervised labeled data is much more data efficient for downstream performance than optimizing the language modeling objective on unlabeled data. Based on these motivations, TLM uses the task data as queries to retrieve a tiny subset of the general corpus. This is followed by jointly optimizing a supervised task objective and a language modeling objective using both the retrieved data and the task data.
We evaluate TLM on eight different tasks covering the domains of news, review, computer science, and biomedical science, following the setting of Gururangan et al. (2020). TLM achieves results better than or similar to BERT (Devlin et al., 2019) and RoBERTa (Liu et al., 2019) while reducing the training FLOPs by two orders of magnitudeThis effectively reduces the cost from training on 1,000 GPUs for one day to training on 8 GPUs for 42 hours..
Related work
Pretrained language models have become the de-facto solution to many of the NLP tasks (Radford et al., 2018; Devlin et al., 2019; Liu et al., 2019; Raffel et al., 2019; Brown et al., 2020; Yang et al., 2019). Those models are usually pretrained on a large-scale corpus in a self-supervised manner to learn a contextualized representation of tokens in natural language, and then are fine-tuned with labeled data for specific tasks. BERT (Devlin et al., 2019), one of the most popular PLMs, is pretrained on a 16GB English corpus using a masked language modeling objective (i.e. predicting randomly masked tokens). RoBERTa (Liu et al., 2019) inherits the training objective of BERT, but is pretrained on a larger corpus consisting of 160GB English texts with larger batch size and dynamic token masking. In this work, we take both BERT and RoBERTa as our major baselines.
There is a line of work dedicated to improving the efficiency of pretraining language models. You et al. (2020) and Shoeybi et al. (2019) utilized the data and model parallelism across different computational devices to accelerate the pretraining process. However, accelerating through parallelism does not actually reduce computational costs in terms of FLOPs for training models at large scale. Chen et al. (2021) and So et al. (2021) tried to identify efficient neural network architectures for language model pretraining, based on the lottery ticket hypothesis and neural architecture search. Such modifications on architecture can bring about reduction in computational costs. Clark et al. (2020) and He et al. (2021) incorporated manually designed mechanisms into language model pretraining, such as adversarial training and disentangled representation of content and position, which brings about reduction in computational costs. Gu et al. (2020) proposed to use task-guided pre-training with selective masking, which reduces the computation cost by around 50%. In this work, orthogonal to the aforementioned works, we investigate improving efficiency by reducing training data redundancy. Our approach also results in more drastic improvements.
Another line of work aims at improving inference efficiency of PLMs. Some works improve inference efficiency by distilling large PLMs into small-sized models and using the distilled models for inference, such as DistilBERT (Sanh et al., 2019), TinyBERT (Jiao et al., 2020), MobileBERT (Sun et al., 2020), FastBERT (Liu et al., 2020), BORT (de Wynter & Perry, 2020), and BERT-of-Theseus (Xu et al., 2020). Other works speed up inference by quantizing PLMs with low-precision representations during inference, such as Q8-BERT (Zafrir et al., 2019), Q-BERT (Shen et al., 2020), and I-BERT (Kim et al., 2021). Another type of works, such as (Michel et al., 2019; Wang et al., 2020; Gordon et al., 2020), adopt pruning by removing parts of PLMs to make it smaller and faster. However, these methods rely on large PLMs, and the performance after distillation, pruning, or quantization often decreases to a certain extent compared with some of the best PLMs (e.g., RoBERTa-Large). In contrast, our approach doesn’t rely on large-scale pre-training and achieves better or at least comparable performance.
Domain-adaptive finetuning is a method that finetunes a pretrained model on in-domain data using a language modeling objective. It has been shown to be effective for domain and task adaptation (Zhang et al., 2019; Gururangan et al., 2020; Li et al., 2020; Lee et al., 2020). There are a few crucial differences between domain-adaptive finetuning and TLM. First, TLM is a general method to improve training efficiency that does not use any additional domain data. It only utilizes the general corpus as in BERT and RoBERTa. In comparison, domain-adaptive finetuning uses domain data to improve domain adaptation. Second, while previous works on domain-adaptive finetuning are built upon a model pretrained on the general corpus, TLM learns from scratch without large-scale pretraining to substantially save computation costs.
Additionally, we observe two techniques related to TLM. They are Co-Training (CT) (Qiao et al., 2018; Yang et al., 2021) and Data-Density-Based Active Learning (DAL) (Zhu et al., 2010; Wang et al., 2017) respectively. Both CT and TLM utilize unlabeled data to aid the learning on a certain task. The difference between TLM and CT is 2-fold: First, CT requires training distinct models from multiple views of unlabeled data, yet TLM only trains a single model through pre-text tasks such as MLM. Second, TLM takes the selection process of unlabeled data into account, which is little discussed in CT. TLM and DAL share the same flavor of finding representative instances in a pool of unlabeled data. However, DAL makes the assumption that every unlabeled sample can be effectively labeled by the definition of the task, which is not required by TLM. Also, DAL tries to find critical instances iteratively from the whole pool of unlabeled data, yet TLM only tries to find relevant instances in a one-shot way with respect to labeled data, which makes TLM more efficient than classic DAL algorithms.
Method
It is an interesting phenomenon that humans are able to quickly master a certain task with limited time and effort by focusing only on pieces of relevant knowledge. For example, when students cram for exams, they review a few chapters instead of going through all books in the world. Following this observation, we conjecture that one of the key aspects of learning a task is to quickly and precisely locate task-relevant information. To this end, we develop TLM that first automatically retrieves relevant training data from a general corpus and then learns on the retrieved data and task data combined.
Formally, given a general corpus where is a document, and labeled task data where is text and is a labelWhile it is straightforward to extend our framework to generation tasks, we focus on classification tasks in this work., our goal is to train a model to estimate the conditional probability for classification .
TLM consists of two steps as shown in Figure 2.
Retrieve data from a general corpus using task data as queries.
Train a model from scratch by jointly optimizing the task objective and the language modeling objective on the retrieved data and task data.
We use BM25 (Robertson & Zaragoza, 2009) for retrieval due to its efficiency. While using embedding-based dense retrievers (Karpukhin et al., 2020) might lead to better retrieval results, we do not consider these methods to keep our approach as simple as possible. Moreover, dense retrievers rely on pretraining, which might bring additional computational costs. The exploration of achieving a better tradeoff between efficiency and retrieval performance is left to future work. Moreover, for tasks with extremely long texts (e.g., Helpfulness (McAuley et al., 2015)), we find it more efficient to extract keywords (e.g., using the RAKE algorithm (Rose et al., 2010)) to form the queries for retrieval instead of using the entire input sequence. We call the retrieved data external data and the task data internal data.
Note that our data retrieval method is task-agnostic—it only depends on text without dependency on . Moreover, the retrieval procedure does not assume the availability of domain-specific data. It operates on a general corpus and has the same input as the pretraining-finetuning paradigm.
Given both the internal and external data, we train a language model from scratch. Let be the masked language modeling loss as in BERT (Devlin et al., 2019), and let be the task loss function (e.g., cross entropy for classification). TLM optimizes the following loss function:
where and are hyperparameters. The network architecture we employ is identical to BERT, where we use a CLS head for classification and an LM head for masked language modeling. TLM can also be extended to other architectures for non-classification tasks. Our implementation involves a two-stage training procedure. In the first stage, we interleave one batch of internal data with batches of external data for mini-batch stochastic gradient descent, where is set as an integer. In the second stage, we set both and as zero to only finetune the model on internal data with the task objective.
2 Comparison Between TLM and PLMs
Both TLM and pretraining-finetuning have two stages. In fact, the second stage of TLM equals the traditional finetuning stage. The main difference between the first stage of TLM and pretraining (PLMs) is shown in Table 1. Unlike PLMs which learn as much task-agnostic knowledge as possible at an extremely high cost, TLM learns task-related knowledge for each task with very low costs.
Given the above difference between TLM and PLMs, we will discuss the pros and cons of TLM in detail.
In pretraining-finetuning paradigm, the finetuning performance is largely upper bounded by the pretrained model. However, due to the constraints of computational resources, the majority of NLP researchers cannot afford training large-scale language models and resort to studying the finetuning algorithms. Since only a small portion of researchers are working on the architectures, loss functions, and other design choices of PLMs, there is a risk that the development of the field might be slowing down. On the other hand, TLM is efficient and highly performant. As a result, TLM has the potential of democratizing NLP and expediting its development by allowing most researchers to freely explore the architectures, loss functions, algorithms, and other design choices in the neighborhood of a state-of-the-art solution.
TLM improves over PLMs in terms of per-task FLOPs. In many cases when there are only a few target tasks, TLM is favorable. For example, a researcher might be interested in solving four textual entailment datasets, or an industrial team might want to improve a recommender system which can be viewed as one task. However, if the goal is to solve 1,000 tasks at once (e.g., building an NLP platform to serve multiple business units within a corporate), PLMs might still be preferred.
Since TLM is task-driven, there is a larger degree of flexibility. Researchers can use custom strategies for tokenization, sequence length, data representations, hyperparameter tuning, etc, which might improve performance and/or efficiency.
PLMs learn task-agnostic general representations and can be used for few-shot and zero-shot learning (Brown et al., 2020). In comparison, TLM trades generality for efficiency by learning only task-specific representations. How to further improve TLM in terms of learning more general representations poses a challenge for future work. We believe multi-task learning might alleviate this issue given recent observations (Wei et al., 2021; Zhong et al., 2021), especially for in-domain zero-shot generalization. It might also be possible to combine pretraining with TLM, e.g., using a small PLM with TLM to match a larger PLM, to achieve a better tradeoff between generality and efficiency.
Experiments
Following (Gururangan et al., 2020), we conduct experiments on eight tasks over four domains, including biomedical science, computer science, news, and reviews (two tasks in each domain). The tasks can be categorized into high-resource and low-resource tasks. High-resource tasks has more than 5K task data, including AGNews (Zhang et al., 2015), IMDB (Maas et al., 2011), RCT (Dernoncourt & Lee, 2017), and Helpfulness (McAuley et al., 2015), while low-resource tasks include ChemProt (Kringelum et al., 2016), ACL-ARC (Jurgens et al., 2018), SciERC (Luan et al., 2018), and HyperPartisan (Kiesel et al., 2019). For the general training corpus, we collected two corpora that respectively match the original training corpora of BERT and RoBERTa. We name them respectively Corpus-BERT () and Corpus-RoBERTa (). The size of is 10 times larger than .
Our experiments focus on comparison with general PLMs. We finetuned both BERT (Devlin et al., 2019) and RoBERTa (Liu et al., 2019) of base and large scales as the baselines. Although TLM is a general method without using addition in-domain data, it even performs close to domain-adaptive finetuning methods (Gururangan et al., 2020) (see Appendix A for detailed comparison).
We report the average performance across three random seeds, together with the standard deviation. We follow Beltagy et al. (2019) and Gururangan et al. (2020) to report the test micro-F1 for ChemProt and RCT, and macro-F1 for the rest of the datasets.
For fair comparison, we evaluate TLM of different training scales. The training scale is defined by three factors, including the number of parameters, the size of the general corpus, and the number of total training tokens. The number of total training tokens is calculated as the product of training steps, batch size, and sequence length. We report TLM at three training scales as shown in Table B.1, namely small, medium, and large scales. Each scale of TLM is accordingly compared to the PLM baselines with an increasing computational cost.
For each experiment of TLM, while fixing the training scale hyper-parameters (i.e., training steps, batch size and sequence length), we perform a grid search over and . We listed the hyper-parameters used in Table B.1 in Appendix.
2 Main Results
Table 2 shows the main results that compare TLM of three different scales and the according PLM baselines. In conclusion, TLM can achieve results that are better than or comparable to the baselines with substantial reduction in FLOPs and the size of training data. Specifically, at a small scale, TLM achieves comparable results to BERT-Large with an average of 1/33 of FLOPs and 1/16 of the training corpus. At the medium and large scales, TLM improves the performance by and points on average respectively, while significantly reducing both FLOPs and the training data size by two orders of magnitude or more. These results confirm that TLM is highly accurate and much more efficient than PLMs. Moreover, TLM gains more advantages in efficiency at a larger scale. This indicates that larger-scale PLMs might have been trained to store more general knowledge that is not useful for a specific task.
3 Ablation Study
Table 3 shows the comparison between different retrieval methods (i.e., BM25 and random retrieval) and different sizes of the general corpus. We find that given the same general corpus, the results of BM25 significantly outperform those of random retrieval by a large margin on all tasks, showing that using task-relevant data for joint training is crucial for the best performance. Specifically, BM25 shows an advantage of almost 1 point against random retrieval on high-resource tasks such as IMDB, and more significant advantages on low-resource tasks such as SciERC and ChemProt by around 3-4 points. This is aligned with our intuition that low-resource tasks rely more on external data.
By comparing the results of and with BM25, we observe that increasing the size of the general corpus improves performance (by 0.5, 1.34, and 1.35 points on IMDB, SciREC, and ChemProt respectively). The gains of using 10 times more data are similar to the ones observed in PLMs (Yang et al., 2019; Liu et al., 2019). This indicates that although TLM only uses a small amount of data, it is able to scale when a larger general corpus is available while maintaining efficiency. On the other hand, the gains of using a larger corpus diminish with random retrieval, showing that random retrieval, as a task-agnostic method, is not very sensitive to the general corpus size.
Data retrieval selects the top- similar documents from the general corpus. Table 4 shows the results of different values. We observe that high-resource tasks such as AGNews only need a small value, while low-resource tasks such as SciREC and ChemProt require a large to obtain the best performance. The observation is consistent with the above analysis that low-resource tasks rely more on external data to improve from joint training.
The hyperparameters and are the weights for the LM loss on external and internal data respectively. We conduct sensitivity analysis over and . Results are shown in Table 5 and Table 6.
For , we find that high-resource tasks such as Helpfulness perform better with a smaller (i.e., Helpfulness achieves best when ) while low-resource tasks such as SciERC and ChemProt achieve their best when is large (i.e., both tasks use ). This is in line with conclusions in Section 4.3.1 that low-resource tasks rely more on external data. In addition, removing task data and only using external data for training (i.e., =#), it performs worse than when incorporating the task data, proving the indispensability of small task data.
Results in Table 6 show that language modeling on internal data is necessary: consistently better results are achieved when is non-zero. Based on our observations, competitive performance can be achieved when is set to a proper value between 20 and 1000.
3.3 Second Stage of Training
TLM contains two training stages—first training on all three terms combined and then finetuning using only the task objective. To validate the effectiveness of the second stage of TLM, we compare the performance of two-stage training against using only stage one. Results are shown in Table 7. We find that removing the second stage hurts the ultimate performance consistently, proving its indispensability. Particularly, the second stage has much more influence on low-resource tasks (with a huge decrease of 19.37 points on ACL-ARC and 14.34 points on ChemProt) than on high-resource tasks (with a performance decrease of 0.53 points on AGNews and 2.17 points on IMDB).
3.4 MLM Loss on Task Data
During the first training stage, TLM uses masked language loss on task data. To examine whether the trick attains the main improvements, we compare results on PLM, PLM with additional MLM loss on task data (PLM+MLM) and TLM. Results in Table 8 show that adding MLM loss on task data into PLM has only marginal gains and does not affect the main conclusion of the paper. In addition, results in Table 3 and Table 4 show that retrieving appropriate relevant data is also essential for the performance of TLM.
4 Analysis
We also study the difference between the model behaviors of TLM and pretraining-finetuning by visualizing their attention weights. Voita et al. (2019) found that a specific kind of heads, referred to as ”positional head” in which at least 90% of the maximum attention weights are assigned to adjacent tokens, have vital contributions to final predictions of the model. Another sort of heads we are interested in are those in which most maximum attention weights are assigned to [CLS],[SEP] or the period token(”.”), which potentially encode less semantic or syntactic information (Kovaleva et al., 2019). In our experiments, if more than 90% maximum weights are assigned to [CLS], [SEP] or the period token, we categorize this head as a “vertical head”. Results in Figure 3 show that on the task ChemProt, more positional heads and less vertical heads are observed in TLM than in PLMs. We also observe similar patterns across various tasks (see Appendix C). These phenomena suggest that TLM learns different (probably more informative) attention patterns compared to PLMs.
4.2 Case Study of Retrieved Data
We have shown several casess of retrieved data in Table 9. TLM retrieves relevant data from a general corpus using BM25 (Robertson & Zaragoza, 2009). Since BM25 is based on sparse features, it focuses more on lexical similarity instead of semantic similarity. This might be specifically beneficial for professional domains, e.g., SciERC for computer science and ChemProt for biomedical science), since there are a large number of proper nouns in these domains. For other domains, it seems BM25 also performs reasonably well for retrieving related documents.
5 Results on More Datasets
So far we have followed the setting of Gururangan et al. (2020) and adopted the datasets therein. In this section, we additionally experiment with the GLUE benchmark (Wang et al., 2018) following the setting of BERT (Devlin et al., 2019) to examine the performance of TLM on a more diverse set of tasks including natural language understanding. We follow the small-scale setting in Section 4.2 in terms of model size, data, and FLOPs. Results in Table 10 show that given the advantages in efficiency, the average performance of TLM is comparable to BERT across 8 tasks, which is consistent with our previous findings and demonstrates the effectiveness of TLM.
Conclusions
In this paper, we have proposed a simple, efficient, pretraining-free framework, TLM. The core idea is to only use a tiny, task-relevant subset of the general corpus for language model training. Our experiments show that TLM achieves results similar to or even better than PLMs, with a reduction of training FLOPs by two orders of magnitude. TLM opens the possibility of reducing the heavy reliance on large-scale PLMs and training a model from scratch in an efficient manner, while not hurting the overall performance. We hope TLM will contribute to democratizing NLP and expediting its development by allowing most researchers to freely explore the architectures, loss functions, algorithms, and other design choices in the neighborhood of a state-of-the-art solution.
As discussed in Section 3.2, there are several potential directions for future work. It will be interesting to study how to use TLM to match the performance even larger-scale PLMs. Moreover, further extending and improving TLM for few-shot and zero-shot learning is a crucial problem.
References
Appendix A Comparison to Domain Adaptation
Our work is different from domain adaptation such as Gururangan et al. (2020). While domain adaptation aims to address how to effectively adapt a pretrained LM into one domain-specific task with sufficient domain data, this work targets to provide a method that is general enough to solve any task without domain data. Nevertheless, we still compare TLM with (Gururangan et al., 2020) as Table A.2 shows. We hope to figure out that, under the harsh but practical condition that no domain data is accessible, whether our proposed framework TLM can still match or even outperform the traditional domain adaptation methods with large pretrained language models as well as domain data.
From results in Table A.2, we have observations:
We reproduced the RoBERTa-Base results using the hyper-parameters reported by Gururangan et al. (2020) as well as our own hyper-parameters. Results show that the baseline RoBERTa-Base results are underestimated in the paper with a gap of around 3 points. We list our hyper-parameters for fine-tuning RoBERTa in Table A.1.
We also reproduced the DAPT+TAPT results using our own hyper-paraemters. Results show that DAPT+TAPT with new hyper-parameters also performs slightly better than it was reported by Gururangan et al. (2020).
From the perspective of total training computes (FLOPs), DAPT+TAPT consumes a comparable FLOPs with TLM (large-scale), and TLM (large-scale) achieved comparable results with DAPT+TAPT (i.e., 85.70 vs 85.57). However, from the perspective of data usage, DAPT+TAPT uses large amounts of domain data, the amount of which for each domain almost equals the amount of BERT total training corpus. TLM does not rely on it.
Appendix B Detailed Experiment Settings
Table B.1 lists the detailed hyperparameters for TLM at stage 1 of different scales for each task. At small and medium scales, for tasks with less than 5K training examples (HyperPartisan, ChemProt, SciERC, ACL-ARC), we set ; for tasks with more than 100K training examples (RCT, AGNews, Helpfulness), we set , for the rest of the tasks (IMDB), we set . At the large scale, is doubled for each task. At each scale on every task, we conduct grid search for and , and adjust training steps, batch size and sequence length to minimize the training cost while preserving competitive performance. We observe that for almost all the tasks, the larger the training scale, the more reliance on external data, indicated by the increasing trend of and as the total training tokens goes up.
Appendix C Attention visualization on other tasks
Besides ChemProt (Figure 3), we also experimented on RCT (Figure C.1) and SciERC (Figure C.2) to get attention visualizations. We find TLM consistently contains more positional heads (in red box) and less vertical heads (in gray mask). These results reveal that the aforementioned pattern generally holds for TLM.