Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora
Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, Xiang Ren
Introduction
Pretrained language models (PTLMs) have achieved remarkable performance on benchmark datasets for a range of NLP tasks Liu et al. (2019b); Brown et al. (2020). However, when deployed in the wild, NLP systems must deal with emerging data that have constantly shifting data distribution, different from the text corpora they were initially pretrained on — for example, when new data domains are introduced (upper part of Fig. 1) Gururangan et al. (2020), or when the language uses and vocabulary change over time (lower part of Fig. 1) Lazaridou et al. (2021). Fine-tuning from a static and possibly “outdated" PTLM may limit the model performance on downstream tasks, as the PTLM may no longer provide an effective model initialization Beltagy et al. (2019); Müller et al. (2020). Here we look to understand whether continuously adapting a PTLM to emerging data can yield gains on various downstream tasks, and how to achieve better downstream performance for such lifelong PTLM adaptation.
A number of recent works make attempts on adapting PTLMs to a new data domain. Gururangan et al. (2020); Yao et al. (2021) adapt language models to corpora of different genres and topics and observe performance improvement in domain-specific downstream tasks. Arumae et al. (2020) further show that by regularizing the parameters of PTLMs, the downstream tasks performance on the general domain can be preserved. Another line of works focuses on temporal domain shift Hombaiah et al. (2021), which analyzes the effect of pretraining over up-to-date data to the downstream tasks. Röttger and Pierrehumbert (2021) further study vocabulary composition approaches for improving adaptation to up-to-date corpora. However, these work focus their study on adapting PTLM to a single new domain; while in practice, corpora from distinct domains and time stamps may emerge sequentially. Whether one can maintain a single, up-to-date PTLM remains an open problem. Related to this, Lazaridou et al. (2021) study adaptation of PTLMs over temporal data streams, but solely focus on language modeling instead of fine-tuning performance. It is also important to understand multiple aspects of the utility of lifelong PTLM pretraining, such as knowledge retention over all the seen data, and study what methods can improve the utility of PTLMs in such a continual pretraining process.
In this paper, we formulate a Lifelong Language Model Pretraining task to simulate practical scenarios of maintaining and adapting a PTLM over emerging corpora, create a testbed (along with pretraining data streams and downstream tasks) for studying continual pretraining algorithms, and present a systematic evaluation protocol for measuring the progress made on this challenging problem (see Figure 2 for an illustration). We consider two types of text corpus sequences when constructing pretraining data streams, each of which simulates a representative use case and that has slightly different focuses on the evaluation: continuously learning a single model that is applicable to both old and new domains; and improving the model’s ability to handle latest data. Specifically, we construct 1) a domain-incremental text stream that consists of academic papers published in four research fields, and 2) a temporal tweet stream that consists of tweets collected from four different years. By conducting systematic experiments on these two data streams, we look to answer a series of analysis questions: 1) whether continual pretraining retains fine-tuning performance over earlier corpora compared to traditional offline pretraining, 2) whether pretraining improves downstream performance on the latest data, and 3) whether pretraining improves temporal generalization where training and evaluation have distribution gaps because of time.
To address the research questions above, we conduct a systematic evaluation of existing continual learning (CL) algorithms, spanning over model-expansion based, memory-based, and distillation-based approaches. Our results show distillation-based approaches are most effective in knowledge retention in the research paper stream, while simultaneously improve adaptation to latest data and temporal generalization in the tweet stream. We believe our problem formulation, evaluation setup, methods and analysis can inspire more future work on continual pretraining of language models.
Problem Formulation
Here we present the problem formulation for lifelong pretraining of PTLM, provide details about the data stream construction process and downstream tasks, and introduce the evaluation protocol.
We consider the scenario where one needs to deploy and/or maintain NLP models over a sequence of data domains. At each time step the model visits an unlabeled text corpus from a domain with a data distribution . The data distribution evolves as the time step , forming a data stream . In practice, the data domain shift can refer to the topic change of the text content (from computer science research papers to biomedical papers), or temporal evolution of the text (from past to recent tweets). The task of lifelong pretraining of PTLM looks to continuously adapt a language model as the model visits (unlabeled) text corpus from the data stream , in order to provide a good model initialization for fine-tuning on downstream tasks from the same domain. With slight abuse in notations, we also use to directly refer to a data domain.
Here, we assume a language model is updated sequentially over each pretraining corpora , without accessing the full earlier corpora in the data stream . This aims to capture practical constraints such as privacy restriction for storing earlier data, or computation budget for training over all the text corpora in . We use to denote the language model right after updating on the domain . In our study, is a RoBERTa-base transformer Liu et al. (2019b) and the model () is initialized with pretrained RoBERTa weights.
The utility of the PTLMs is evaluated based on their fine-tuned model performance on various downstream tasks. After updating on a domain , the model can be fine-tuned over downstream tasks from visited domains where . We note the set of downstream tasks related to domain as , assuming the number of downstream tasks is . Note that in the fine-tuning stage, model has no access to any of the pretraining corpus .
2 Data Streams & Downstream Datasets
We construct data streams to simulate two representative scenarios of data domain shifts in practice (also see Fig. 1): one domain-incremental stream to simulate the sequential changes of research paper areas; and one chronologically-ordered stream to simulate tweets emerging over time.
This paper stream consists of the full text of research papers published in four research areas: biomedical, computer science, material science, and physics, filtered from the S2ORC datasetWe use the 20200705v1 version of the S2ORC dataset at https://github.com/allenai/s2orc, which are presented sequentially to the model. For each domain, we evaluate downstream performance over two datasets. The downstream tasks span over various tasks such as relation extraction and named entity recognition, and are summarized in Table 1. We detail these datasets in Appendix D.
Chronologically-ordered Tweet Stream.
This tweet data stream consists of tweets from the year 2014, 2016, 2018 and 2020, collected by the Archive Teamhttps://archive.org/details/twitterstream and preprocessed following Nguyen et al. (2020). These four tweet corpora are presented sequentially to the language model following the chronological order of the tweet year. For downstream tasks, we hold out 1M tweets from each year’s corpus to construct multi-label hashtag prediction datasets Gong and Zhang (2016) and single-label emoji prediction datasets Barbieri et al. (2018). On two datasets, we report label ranking average precision scores (a multi-label version of MRR) of models Azeemi and Waheed (2021) and Macro-F1 respectively. The detailed dataset construction process is included in Appendix D.
3 Evaluation Protocol
We consider three key aspects for evaluating the utility of the language models that are continuously updated over the data stream , also illustrated in Figure 2: 1) knowledge retention and transfer over the pretraining corpora seen earlier; 2) adaptation to the latest data domain, and 3) temporal generalization when training and evaluation data are from different time steps.
A key utility of continual language model pretraining is to obtain a single model applicable to all domains. We focus on the evaluation of the ability with the domain-incremental paper stream, because for the tweet stream, the practical need of performance over outdated data is limited. Knowledge retention is measured with the downstream task performance from earlier or the current domains that the pretrained model has visited. More formally, for each pretrained model checkpoint in , we fine-tune over downstream tasks where and evaluate the corresponding test set performance. It is important that the models do not suffer from catastrophic forgetting Robins (1995), i.e., significantly reduced helpfulness when is fine-tuned for downstream tasks from earlier domains with .
Adaption to Latest Data Domain.
In certain scenarios, performance of downstream models over the latest data domain should be emphasized. For example, classifiers in the tweet domain are usually trained and evaluated with up-to-date data for practical deployment. Formally, we focus on the downstream task performance of models fine-tuned from the final pretrained model checkpoint , where the downstream tasks are also from the latest domain. To succeed in these metrics, it is crucial for the model to transfer knowledge from earlier domains to the latest domain.
Temporal Generalization Ability.
We consider another practical fine-tuning scenario in the tweet stream where the model is trained on outdated data and evaluated on the latest data Rijhwani and Preotiuc-Pietro (2020); Huang and Paul (2018), referred to as the temporal generalization ability. Formally, we fine-tune the final pretrained model checkpoint over the training set of downstream tasks from an earlier time step (), and evaluate on the test set of the downstream tasks from the latest time step .
Methods
Lifelong language model pretraining introduces novel challenges because of the large training sets and more comprehensive evaluation protocols compared to classification tasks. We establish several strong baselines, and evaluate the performance of continual learning algorithms from different categories spanning over model-expansion, memory-based, and distillation-based approaches, We illustrate the approaches in Figure 3.
We consider several simple baselines which continual learning algorithms will be compared against. RoBERTa-base () corresponds to not pretraining on any of the domain-specific corpora. By separately pretraining on each corpus , we obtain Task-Specific pretrained models. We also pretrain sequentially over , which we refer to as sequential pretraining. While it allows knowledge transfer between domains compared to domain-specific models, without any continual learning algorithms, sequential pretraining is prone to catastrophic forgetting Robins (1995). Finally, we randomly shuffle corpora from all domains before pretraining, noted as Multi-Task Learning (MTL). MTL corresponds to an offline training paradigm that models new corpora by re-training over all corpora seen before. The drawback is that it requires storing full data from earlier domains, and that it can be extremely costly to repetitively retrain over earlier data if new data keeps emerging.
2 Model-expansion and Regularization-based Methods
We first introduce model-expansion based approaches, which add small trainable modules (e.g., multi-layer perceptron) to the model per new domain while keeping other parts of the model frozen. The Adapter approach is a representative approach that learns a set of “adapter” layers for each domain and each of the transformer layers Houlsby et al. (2019). We also experiment with a simple Layer Expansion approach, which learns separate top two layers of the transformer and the prediction head for each domain. We also involve a regularization-based continual learning baseline, online EWC Schwarz et al. (2018), which directly penalize change of model parameters.
3 Memory Replay Methods
We also experiment with Experience Replay (ER) Chaudhry et al. (2019), which alleviates forgetting by storing a subset of earlier examples and periodically re-training (replaying) over them. We maintain a fixed-size memory ( examples by default) and populate the memory each time pretraining on a domain finishes with examples in the current domain. We ensure always contains a balanced sample of examples from all seen domains . We replay a mini-batch of examples from the memory every 10 training steps.
4 Distillation-based CL Methods
While knowledge distillation (KD) Hinton et al. (2015) techniques have been studied intensively for pretrained language models Sun et al. (2019), applying them to continual learning has been under-explored outside image classification tasks Li and Hoiem (2018); Rebuffi et al. (2017); Hou et al. (2018). Distillation based CL approaches store one previous model checkpoint of the model (noted as ) and regularize the differences between and the current model . We adapt several existing knowledge distillation techniques to PTLMs and utilize them for continual learning. We note, while individual distillation techniques are not original, their adaptation to CL algorithms can be novel.
In logit distillation Hinton et al. (2015), we collect the output logits of and , noted as and respectively. The distillation loss is computed as , where is the Kullback–Leibler divergence function.
Representation Distillation.
We also consider minimizing the representational deviation of sentences between previous and current models Sun et al. (2019); Jiao et al. (2020). We extract the representation of each word of two models, noted as and , before the masked language modeling prediction head, where is the length of the sentence. Then, we compute MSE loss as the distillation loss.
Contrastive Distillation.
Self-Supervised Distillation (SEED).
SEED distillation proposed by Fang et al. (2021) has a similar spirit as the contrastive distillation. The only difference is that it distills representational similarity between the batch and a large set of other examples. We leave the details of the algorithm in Appendix E. We further combine SEED Distillationwith logit distillation and refer to the approach as SEED-Logit Distillation.
Results
We summarize our findings over the created data streams. We ask whether lifelong pretraining and continual learning algorthms are effective base on our evaluation protocol proposed in Sec. 2.3.
We use the RoBERTa-base model Liu et al. (2019b), initialized with RoBERTa-base weights throughout the experiments. We set the maximal sequence length to 128 and an effective training batch size of 2,048. On the research paper stream, models are trained for 8 steps in the first domain and 4 steps in the subsequent domains. On the Tweet stream, we train the models for 4 steps in each domain. These correspond to less than a single pass of data in each domain. See Appendix A for detailed setups.
2 Domain Incremental Data Stream
As we introduced in Sec. 2.2, in the domain incremental research paper stream, we expect a model to perform well on all downstream tasks from domains . In Table 2, we report the performance of models on all downstream tasks fine-tuned from the final pretraining checkpoint, . We visualize more complete change of downstream task performance over different time steps of pretraining (i.e.,, ) in Fig. 4. We also report the log perplexity of masked language modeling (MLM) in Table 2 as additional information. With these results, we address the research questions below.
We first examine whether task-specific or lifelong pretraining improves performance over domain-specific downstream tasks. Comparing Task-Specific LMs with RoBERTa-base in Table 2, we notice consistent performance improvements, especially on Biomedical and Computer Science domains (). We also see Sequential Pretraining could consistently outperform RoBERTa-base. However, the comparison between Sequential Pretraining and Task Specific LMs are mixed: on , Sequential Pretraining could outperform Task-Specific LMs only except MNER; while on the earliest biomedical domain (), Sequential Pretraining achieves substantially lower performance. From Figure 4, we see the performance of Sequential Pretraining on Chemprot and RCT (from ) drops significantly from to . The results imply lifelong pretraining allows later domains to benefit from knowledge transfer from earlier domains, but the performance on earlier domains is limited because of forgetting.
Does continual learning algorithms help retain knowledge in sequential pretraining?
Next, we compare different kinds of CL algorithms and investigate the effect of CL algorithms in alleviating forgetting and improving knowledge transfer. Table 2 shows that Online-EWC slightly improves MLM perplexity compared to Sequential PT, but brings no improvement to the fine-tuning performance. We hypothesize that regularization directly in the parameter space as in Online-EWC is not effective when the parameter space is very high dimensional. Adapter improves downstream task F1 scores on the bio-medical domain () by 1.2% and 0.8%, but does not outperform Sequential Pretraining in other domains (similarly for Simple Layer Expansion approach), likely because a great portion of the model is kept frozen.
In contrast, the memory-replay based approach (ER) allows training the full parameters of the model and has been shown to be highly effective in continual learning of classification tasks Wang et al. (2019); Chaudhry et al. (2019). However, we surprisingly find that ER could hardly improve over Sequential Pretraining except . A similar pattern can be found in the MLM perplexity. We hypothesize that the positive effect of example replay has diminished because of the overfitting to the memory examples. Table 3 summarizes the effect of tuning hyperpameters in ER. When we reduce the frequency of replay (from every 10 steps to 100 steps), the MLM performance improves, which implies reduced overfitting; however, the performance of downstream task performance does not improve. When we increase the size of the memory from to , the MLM perplexity also improves; still, there are still no improvements in downstream tasks. It may imply ER itself is not an effective approach for continual pretraining.
Unlike ER, distillation approaches utilize richer information such as output logits or representation similarity to preserve past knowledge. We find either Logit KD or SEED-Logit KD to be most effective depending on the task, while Rep-KD and Contrastive-KD are less effective. The best performing distillation approach improves F1 over Sequential Pretraining on downstream tasks from , at least by 1.0%. However, performance on , which come later in the data stream, does not improve over Sequential Pretraining, possibly because the distillation loss term makes the model rigid in obtaining new knowledge.
What is the gap between lifelong pretraining and multi-task learning across all the domains?
Multi-Task Learning refers to the offline training paradigm, which retrain PTLMs over all corpora () each time a new corpus becomes available. We examine whether lifelong pretraining is comparable to multi-task pretraining in terms of performance. From Table 2 and Figure 4, we see Sequential Pretraining in general underperforms MTL except for the final domain. However, certain CL approaches, such as Logit-Distillation, could improve over MTL on all downstream tasks from the first and the second domain. We speculate the reason is that continual learning naturally provides a curriculum Xu et al. (2020); Shi et al. (2015) to models where each individual task is easier to learn. The results have a positive implication that lifelong pretraining is not only more computationally efficient and requires less storage of past data, but may also improve the performance of pretraining.
Does lifelong pretraining make models more data efficient?
In Table 5, we further examine the performance of final pretrained models under different amounts of training examples. We include full results in Appendix B. We find in general, performance improvements are more significant in the low-resource setup.
Computational Costs.
We quantify computational costs of different CL algorithms with the number of forward and backward passes in Table 4 and present additional experiments with controlled computational costs in Appendix F. We find additional computational cost is necessary for performance improvement of distillation-based CL. However, it is not possible to trade performance simply by investing more computation budget with arbitrary CL algorithms. We leave detailed discussions in Appendix F.
3 Temporal Data Stream
We conduct analysis on pretraining PTLM on chronologically-ordered tweet corpora, to understand whether lifelong pretraining helps adaptation to the latest data and improves temporal generalization ability. The results are summarized in Table 5.
We compare the performance of Task-Specific (2014) to the Task-Specific models pretrained on the year of downstream datasets (noted as Task-Specific (Latest)) and notice consistent improvements in downstream tasks in 2018 and 2020 (first two columns in Table 5). Sequential Pretraining could also outperform the Task-Specific (2014) model. It verifies that language models may get outdated over time, but the issue can be addressed by task-specific or lifelong pretraining over the latest corpora.
Does lifelong pretraining help improve the downstream model’s performance on latest data?
We show that downstream model’s performance over later data () can be improved over Task-Specific models when continual learning algorithms are applied. From the first two columns of Table 5, we see Logit-KD and SEED-KD improve Hashtag prediction score over data of years 2018 and 2020. SEED-Logit KD further improves prediction F1 on Emoji prediction. Note that these findings are in contrast to the research paper stream, where CL algorithms do not improve performance in the latest domain . The reason can be the higher similarity between domains in the tweet corpora making the knowledge transfer easier, which is further discussed in Appendix I.
Does lifelong pretraining improve temporal generalization?
Temporal generalization evaluates downstream performance over latest test data when fine-tuned over outdated training data. We show lifelong pretraining brings clear improvement to temporal generalization. From Table 5, we see even Sequential Pretraining could improve over the model pretrained merely on the year 2020 data (Task-Specific (2020)) consistently. We find performance further improves with CL algorithms applied. SEED-Logit-KD performs best in general on crossyear hashtag prediction tasks. In crossyear emoji prediction, we find Contrast-KD and SEED-KD perform best. We also find that SEED-Logit-KD could slightly outperform Logit-KD.
Related Works
Gururangan et al. (2020) study adaptation of PTLMs to domain-specific corpora. Arumae et al. (2020) study algorithms to mitigate forgetting in original PTLMs, but does not investigate forgetting that happens over a sequence of domains. Maronikolakis and Schütze (2021); Röttger and Pierrehumbert (2021); Luu et al. (2021) proposes sequential pretraining over domains or emerging data, but did not investigate CL algorithms. Several recent studies have demonstrated the necessity of adapting LMs over time Lazaridou et al. (2021) while specifically focusing on factual knowledge Dhingra et al. (2021); Jang et al. (2021).
Continual Learning Algorithms in NLP.
Continual learning in NLP has mainly been studied for classification tasks. An effective approach is to utilize a number of stored past examples de Masson d’Autume et al. (2019); Wang et al. (2020), or pseudo examples (e.g., the ones generated with a PTLM Sun et al. (2020); Kanwatchara et al. (2021)). Recent extensions of the algorithm Chuang et al. (2020) perform knowledge distillation with generated pseudo examples. Other lines of works focus on regularization over the sentence representations Wang et al. (2019); Huang et al. (2021); Liu et al. (2019a) or directly merging models in the parameter space Matena and Raffel (2021). Model expansion-based approaches Liu et al. (2019a); Pfeiffer et al. (2021), including learning domain specific expert models Gururangan et al. (2021), are also actively studied. Wu et al. (2022) present a comparative study of algorithms in the context of continual fine-tuning over NLP tasks.
Conclusion
In this paper, we formulated the lifelong language model pretraining problem and constructed two data streams associated with downstream datasets. We evaluated knowledge retention, adaptation to the latest data, and temporal generalization ability of continually pretrained language models. Our experiments show distillation-based approaches being most effective in these evaluation setups. A limitation of the work is that it has not been fully addressed whether there exists a variant of distillation-based CL approach that consistently outperforms Logit-KD. Based on the current observation, we conclude the performance of different KD approaches for CL is highly task-dependent. It asks for more future works into continual learning algorithms within the proposed problem setup.
References
Appendix A Detailed Experiment Settings
We use a linearly decreasing learning rate initialized with 5e-4 on the research paper stream and 3e-4 on the tweet stream. On the research paper stream, we train the model for 8,000 steps in the first task, and 4,000 steps in the subsequent tasks. On the tweet stream, we train the model for 8,000 steps in all tasks. We hold out 128,000 sentences from each corpus to evaluate MLM performance. As the size of pretraining corpora is large, during training, each training example is visited only once. We use the masked language modeling perplexity over held-out validation sets of the pretraining corpora as the metrics for hyperparameter tuning. Common hyperparameters such as learning rate and batch sizes are tuned with Task-specific models with the first task. Hyperparameters that are specific to continual learning algorithms, such as the scale of the distillation loss, is tuned using the first two domains in the stream according to the MLM performance over validation sets. The weight of the distillation term is set as 1.0 for logit distillation and 0.1 for other distillation algorithms. By default, we replay or perform distillation with a mini-batch of examples from the replay memory every 10 training steps in ER and Distillation-based CL approaches. We use the huggingface transformers library https://github.com/huggingface/transformers for implementation.
Appendix B Low-Resource Fine-Tuning
Figure 6 summarizes the performance of fine-tuned models from the final model checkpoint () using different amount of downstream training examples. We see on Chemprot and SciERC, the benefit of Sequential Pretraining over RoBERTa-base is more significant in low-resource fine-tuning setups. Whenever Seqential Pretraining outperforms RoBERTa-base, we notice Logit-KD could further improve over Sequential Pretraining.
Appendix C Full Results over the Tweet Stream
Tables 6 and 7 summarize full results over the Tweet stream. Compared to the table 5 in the main text, we add downstream performance over data from years 2014 and 2016 (, ), and temporal generalization from year 2014 to 2020 ().
Appendix D Dataset Details
The research paper stream consists of full text of 6.6M, 12.1M, 7.8M, and 7.5M research papers from the S2ORC Lo et al. (2020) dataset. We evaluate downstream fine-tuning performance on two in-domain datasets for each research area: Chemprot relation exaction dataset Vindahl (2016) and RCT abstract sentence role labeling dataset Dernoncourt and Lee (2017) for the bio-medical domain; ACL-ARC citation intent classification dataset Jurgens et al. (2018) and SciERC relation extraction dataset Luan et al. (2018) for the computer science domain; relation extraction over Synthesis procedures Mysore et al. (2019) and named entity recognition over material science papers (MNER) Olivetti et al. (2020) for material science domain; keyphrase classification and hyponym classification after filtering out physics papers for the physics domain Augenstein et al. (2017). We report micro-averaged F1 on Chemprot, RCT, MNER datasets following the evaluation metrics in the original work, and report macro-averaged F1 on all other datasets. We use the official data splits for all datasets except for RCT, where we employ a low-resource training setup following Gururangan et al. (2020).
The pretraining corpora for the tweet stream consist of 25M tweets in each year. For downstream tasks, we use a separate set of 1M tweets from each year to construct multi-label hashtag prediction Gong and Zhang (2016) datasets and single-label emoji prediction datasets Barbieri et al. (2018). We replace user names to special tokens. For Hashtag prediction, the label space consists of tweets containing 200 most frequent hashtags in each year. We independently sample 500 tweets per label (hashtag) as training, validation and test sets, which results 10 examples in each of the data splits. For emoji prediction, we construct 20-way single-label emoji prediction datasets for each year following Barbieri et al. (2018) with the 1M held out tweets. We sample 5,000 tweets per emoji in each split, resulting in balanced datasets of the same size as the hashtag prediction datasets.
Appendix E Details of Continual Learning Algorithms
During continual pretraining, in addition to the language model pretraining objective, we add a unsupervised contrastive learning objective, namely the SimCSE Gao et al. (2021) objective, so that the similarity in the sentence representation better reflects the semantic similarity in the sentence. We use the -normalized representation of the start-of-sequence token at the final layer as the sentence representation, noted as . Then, we distill the intra-batch representational similarity from the previous model to the current model . Given a mini-batch of examples , we compute the representational dot-product similarity matrix between normalized sentence representations between each pair of examples with and , noted as and , where each element is,
where is a temperature hyperparameter. We specify a temperature for the teacher model and a temperature for the student model . We compute the cross-entropy between and as the distillation loss,
E.2 SEED Distillation
Appendix F Analysis and Controlled Experiments of Computational Costs
Computational cost is a crucial matter for online continual learning systems. In this section, we analyze the computational costs of continual learning algorithms and perform controlled experiments of computational costs.
We quantify computational costs with the total number of forward () and backward () computations () over the PTLMs, which is easy to control; in practice, we find the wall clock time of training was approximately linear to . We summarize the number of forward and backward passes and the wall clock time of training in Table 4. In the visit of batches from the training stream, Sequential PT performs forward and backward passes respectively over the PTLM, resulting in . Experience replay further replays 1 batch of examples every steps over the training stream, which results in . In our main experiments, is set to 10 (Sec. 3.3). Logit-Distill and Rep-Distill require one additional forward pass over a frozen PTLM to compute the target of distillation, resulting in . Distillation algorithms that perform contrastive learning with SimCSE (i.e. SEED-Distill and SEED-Logit-Distill) additionally require one forward and backward pass using the same batch of examples with different dropout masks. Therefore, for SEED-Logit-Distill, .
To control the number of forward and backward passes, we present approaches to compensate the lower computation costs compared to Distillation algorithms and one approach to shrink the computational cost of distillation algorithms: (1) for Sequential PT, we train the models for 1.2 times more steps so that , noted as Sequential PT; (2) for ER, we increase the replay frequency to 5 from the default setup 10, so that . We also decrease the cost of Logit-KD and SEED-Logit-KD by reducing the frequency of distillation from every 1 batch to every 10 steps, while still replaying and distilling knowledge over 1 batch of memory examples every 10 training steps. This results in and , where when both and are 10. The approach is referred to as Sparse Logit-KD. Finally, for SEED-Logit-KD, we remove the SimCSE loss from training and perform sparse distillation similar to Sparse-Logit-KD, which also results in .
The performance of the models is presented in Table 9. We notice that at the end of pretraining, increasing the number of training steps in Sequential PT by 1.2 times does not lead to performance boost on the latest domain (), while the performance over tasks from earlier domains (Chemprot, ACL-ARC, SciERC) slightly dropped, possibly due to increased forgetting. For ER, we notice replaying only slightly more frequently (ERk=5) than the default setup (=10) greatly increased the perplexity of MLM, implying significantly increased overfitting to the memory; while the performance differences of downstream tasks compared to the default ER is mixed. When we decrease the replay frequency of distillation, the performance on Logit-KD and SEED-KD also decreased and does not outperform ER.
The results show additional computation costs can be necessary for continual learning algorithms such as Logit-KD and SEED-Logit-KD. However, the results also show that there is no simple trade-off between computational cost and performance. We have seen that it is not always beneficial to increase the number of training steps over the emerging data, as it increases forgetting in earlier domains. Similarly, increasing the frequency of replay may lead to significant overfitting to the replay memory. Investigating into more effective continual learning algorithms, despite increased computation costs, allows us to obtain performance improvement that cannot be simply traded with more computation with arbitrary continual learning algorithms. We leave more thorough studies into this topic as future work.
Appendix G Experiments with RoBERTa-large
We present additional experiments on RoBERTa-large. Figure 7 and Table 8 summarizes the results of selected continual learning algorithms and baselines. On Chemprot, RCT-Sample, ACL-ARC and SciERC, either SEED-Logit-KD or Logit-KD achieves best performance with the final pretrained model checkpoint. We notice that sometimes certain continual learning algorithms (ER, Logit-KD) achieves lower F1 at the initial time step (e.g., =2 in Figure 7(c)). In these cases, we hypothesize continual learning algorithms may hurt model’s performance in capturing new knowledge, despite its potential to reduce forgetting.
Appendix H Experiments with BERT on Tweet Stream After 2019
In this section, we present an additional set of experiments on BERT-base Devlin et al. (2019) model, which is originally pretrained with Wikipedia articles before 2019, with Tweets only after 2019. The training corpora consist of tweets from the first half of 2019, the second half of 2019, the first half of 2020, and the second half of 2020 respectively. We accordingly construct hashtag prediction and cross-year hashtag prediction datasets. The performance of downstream tasks fine-tuned from the final pretrained model is presented in Table 10. We see Sequential PT clearly outperforms BERT-base which is not continually pretrained, and that Logit-KD generally improves hashtag prediction performance compared to Sequential PT except on the first half of 2019. We hypothesize the small temporal gap between makes improvements less significant than our main experiment setup. We present temporal generalization performance in cross-year hashtag prediction tasks in Table 11. Similarly, Logit-KD improves over Sequential PT in two out of three cross-year hashtag prediction setups.
Appendix I Analysis of Data Streams
In this section, we provide further analysis about the created research paper stream and the tweet stream. We measure cosine distances of vocabulary distributions between each pair of different domains and summarize the results in Figure 8. The results indicate that the Tweet stream has a magnitude smaller vocabulary distribution gap between domains, which is in the scale of , compared to the research paper stream, which is in the scale of . On the Tweet stream, we see the differences of vocabulary distributions align with the temporal gap between domains. On the research paper stream, we find some domains to be more similar than others. For example, Bio-medical () and Material Science domains have larger similarity in their vocabulary distributions, which explains general downstream performance increase on after the model is pretrained on (Fig. 4 (a,b)).
The differences in vocabulary distribution explain inconsistency in results between two data streams, specifically, whether lifelong pretraining improves downstream model performance on the latest domain, as we mentioned in Sec. 4.3. Other than this, our main findings, such as the effect of distillation-based CL algorithms on reducing forgetting, are consistent over two datasets with such significant differences in their changes of vocabulary distribution. We believe it implies the conclusions in this paper should be reliable in diverse data streams.
Appendix J Ethic Risks
We would like to note that, in practice, continually pretrained models over real-world data streams would require identification and removal of biased contents from pretraining corpora, which may affect the prediction of downstream models. As PTLMs are continuously updated, the bias in earlier pretraining may have a profound negative impact. In future works, it is preferable to develop algorithms to “forget” certain biased knowledge from language models. We further note that any data released in this paper, especially the tweet stream, should only be used for research purposes.