Re-evaluating Continual Learning Scenarios: A Categorization and Case for Strong Baselines

Yen-Chang Hsu, Yen-Cheng Liu, Anita Ramasamy, Zsolt Kira

Introduction

While current learning-based methods can achieve high performance on tasks, they only perform well when the testing data is similarly distributed as the training data. In other words, they cannot adapt continuously in dynamic environments where situations can significantly change. Such adaptation is desirable for any intelligent system, and is the hallmark of learning in biological systems. One approach to this problem is continual learning, where models are updated incrementally as data streams in. However, deep neural networks, which are currently state of art for many applications, are known to suffer catastrophic interference or forgetting when incrementally updated through gradient-based methods. This leads to the model forgetting how to solve old tasks after being exposed to a new one due to interference caused by parameter updates.

To address this problem, several approaches have been proposed and a number of experimental methodologies (i.e. datasets, learning curricula, and architectures) have been used for evaluation. In this paper, we argue that the current set of evaluations have significant limitations, including lack of uniformity across the different experimental methodologies, lack of hyper-parameter tuning of reasonable baselines under similar tuning budgets as the proposed methods, and simplicity of the tasks (e.g. short task queues). Towards this end, we make several contributions in this paper: 1) A categorization of a large number of experimental methodologies into a few canonical settings along with a comparison of their difficulty, 2) A uniform but flexible framework for generating scenarios under this categorization and systematic evaluation of current state of art methods, and 3) Demonstration that very simple baselines can be surprisingly effective if used properly and result in comparable or better performance against the current state of art. We have released our framework (written in PyTorch) to enable fair and uniform evaluation to aid the community, and conclude with some suggested modifications to the scenarios to increase realism of the evaluation.

Generating Task Sequences for Evaluating Continual Learning

In order to evaluate continual learning methods, scenarios are commonly generated from datasets using two operations: permutation and splitting. The typical source dataset is MNIST (1), an image dataset of hand-written digits. The Permuted MNIST experiment (2) involves ten-digit classification, where each task consists of different permutations of the pixels in the images. The number of different permutations represents the length of the task sequence. This evaluation scenario is widely adopted (3; 4; 5; 6; 7; 8; 9), despite criticism that it is less challenging in terms of forgetting (10). Another typical scenario, the Split MNIST experiment, was initially introduced in a multi-headed form where the ten digits are split into five two-class classification tasks (the model has five output heads, one for each task) (9; 8; 6), and the task identity (1 to 5) is given at testing time. This scenario is argued to be easier since the selection of output head is given by the task identity (10). Farquhar and Gal (10) propose a single-headed variant which does not require task identity, where it always requires the model to make a prediction over all classes (digits 0 to 9). Such single-headed Split MNIST is known as incremental class learning (11; 12; 13). Van and Tolias (14) propose another variant of single-headed Split MNIST, where the model always predicts over two classes instead of ten classes. Furthermore, the similar multi-headed/single-headed strategies in Split MNIST can apply to Permuted MNIST (14) resulting in many combinations. These scenarios are used in different works, and therefore there is a lack of coherent comparison. This paper addresses this problem by providing a systematic interpretation of the differences between an old task and a new one (see Section 2.2).

2 A Categorization of Current Scenarios

Incremental Domain Learning: First, we discuss change in the marginal probability distribution of inputs P(X)P(X), specifically P(X1)≠P(X2)P(X_{1})\neq P(X_{2}). The input distribution (domain) difference has been discussed extensively in the setting of transfer learning (15), primarily as a domain adaptation problem (16). Unlike domain adaptation, which aims to transfer knowledge from an old task to a new task where only the performance of the new task is considered, the continual learning setting aims to keep performance on old tasks while achieving reasonable performance on the new one as well using a single model. This domain difference can be caused by the permutation and splitting strategies mentioned previously. In fact, the most widely adopted Permuted MNIST experiment (3; 4; 5; 6; 7; 8; 9) generates such a domain difference. However, the inputs generated by the random permutation protocol is highly uncorrelated (9; 10); thus they can not represent all possible scenarios. To generate a scenario with better-correlated tasks, one should avoid random permutation to keep the original spatial correlation between image pixels, allowing the possibility of shared features among tasks. This requirement can be fulfilled by a variant of the Split MNIST experiment, where ten digits are split into five binary classification tasks and the model has only a binary classifier (single-headed). Such a scenario is illustrated in the middle column of Figure 1. The requirement of using a single-headed model is essential to control for other types of differences. Specifically, the single-headed model ensures the same output space {Y1}={Y2}\{Y_{1}\}=\{Y_{2}\} (task identity becomes unnecessary), and the equal amount of MNIST digits ensures the output distributions of the binary classification are the same (P(Y1)=P(Y2)P(Y_{1})=P(Y_{2})). As a result, the only difference in the middle column of Figure 1 is that the input images for label {0,1}\{0,1\} shift from digit 0/1 to digit 2/3.

Incremental Class Learning: The second scenario relates to multiclass classification. Each task in the sequence contains an exclusive subset of classes in a dataset. P(X1)≠P(X2)P(X_{1})\neq P(X_{2}) is true by the nature of this setting. All class labels are in the same naming space (single-headed) and the number of output nodes equals the number of total classes in the task sequence. Due to the multiclass property, P(Y1)≠P(Y2)P(Y_{1})\neq P(Y_{2}) is a natural consequence. The right-most column in Figure 1 demonstrates how split-dataset strategy (11; 10) generates this scenario. The permutation strategy can also generate the task sequence (14). In the latter case, each permuted digit represents a new class, so the total number of classes is multiplied by the number of permutations, e.g., total 10×1010\times 10 classes in a ten Permuted MNIST experiment.

Incremental Task Learning: In the last scenario, the output spaces are disjoint between tasks, denoted as {Y1}≠{Y2}\{Y_{1}\}\neq\{Y_{2}\}. This definition makes P(Y1)≠P(Y2)P(Y_{1})\neq P(Y_{2}) true as well, while P(X1)≠P(X2)P(X_{1})\neq P(X_{2}) is generally true since the semantic classes differ. The differences between the output spaces are the output dimension and their associated semantic meaning. For example, the old task can be a classification problem of 5 classes while the new task is a regression task of a single value. To allow a model to produce an output for a specific task, a model requires task-specific output components which are selected by additional information, the task identifier tt. A typical neural network for this scenario has a multi-headed output layer (one head for each task) (9). At testing time, only the head matching the tt will be activated to make predictions. One common approach to generate this scenario is having multiple datasets (ex: MNIST, CIFAR10, SVHN, AudioSet, CUB-200, etc.) and use one of them in one task (17; 18; 19; 20). The splitting and permutation strategies can also generate task sequences for this scenario, illustrated in Figure 1 and Appendix Figure 2. The prior works mentioned in this paragraph commonly have the task identity given during testing; thus the experiments in this work follow the same setting.

Experiments

We now describe the experimental configuration. Here, we use the MNIST dataset with the splitting strategy (Figure 1) to generate the three continual learning scenarios in Table 1. The standard train/test split was used, with 60k training images (∼\sim6k per digit) and 10k test images (∼\sim1k per digit). The preprocessing of images includes zero padding to 32x32 pixels and a standard normalization to zero mean with unit variance. No other data augmentation, (e.g. random translation) is applied.

For a fair comparison, all methods use the same neural network architecture, which is a multi-layered perceptron with two hidden layers of 400 nodes each, followed by a softmax output layer. Both hidden layers use ReLU for the activation function. The loss function is a standard cross-entropy for classification in all methods and scenarios. All models are trained for 4 epochs per task with mini-batch size 128 using the Adam optimizer (β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, learning rate=0.001=0.001) as the default unless explicitly described. In all experiments, the optimizer is never reset.

We use three baseline strategies. The most common baseline used in prior work is a neural network sequentially trained on all tasks in the standard way, in that the parameters learned from old tasks are fine-tuned to the new task. Such a model is usually optimized with Adam (21), but here we use different optimizers including Adam, SGD, and Adagrad (22). In fact, we show that Adam is in general a poor choice for this task. The latter two optimizers use 0.01 for the learning rate without momentum in all scenarios. Another baseline, L2 regularization, prevents the parameters from deviating too much from those previously learned. Note that this is similar to EWC in that the identity matrix replaces the Fisher information matrix. In other work this is only evaluated in one specific scenario (permuted MNIST) with limited length (3) of task sequence (3). Similar to other regularization methods, the L2 regularization requires tuning of the single-valued regularization coefficient, which is done by a grid search (3; 9).

The third baseline is a naive rehearsal strategy, which is sometimes called experience replay. The model has a small replay buffer to store a fraction of previous data randomly. While training a new task, each mini-batch is constructed by an equal amount (64/64) of new data and the rehearsal data. The buffer size is predefined and fixed to match the space overhead introduced by online EWC and SI (#parameter ≈(1024×400+400×400)×2=1,139,200\approx(1024\times 400+400\times 400)\times 2=1,139,200, ignoring small overhead), which converts to 1.1k images when the pixel value is saved in a 32-bit floating number (named Naive rehearsal). Additionally, with the same memory space, it can store more images when compression is used. One naive compression is using an 8-bit integer to represent the pixel value; thus 4.4k images can be stored (named Naive rehearsal-C). For the buffer management, all tasks seen so far have an equal amount of images in the buffer while keeping the total number the same. This management is similar to iCaRL (11), except that we randomly pick the images for staying in the buffer.

For comparison, we pick several popular methods (EWC(3), online EWC(23), SI(9), MAS(24), LwF(25), GEM(5)) and state-of-the-art rehearsal-based methods (DGR(8), RtF(14)) with generative model. The hyperparameter is tuned by a grid search, and the results with the best setting are reported. The total static memory overhead is controlled to be the same among Naive rehearsal, Naive rehearsal-C, L2, online EWC, SI, MAS, GEM, and DGR. RtF has only half the overhead of DGR since its classification model is shared with its generative model. The hyperparameter of EWC, online EWC, SI, and MAS is tuned with grid search. We use the results from Ven and Tolias (14), which provides an analysis and comparison developed concurrently with our work, for LwF, DGR, and RtF since the same model and training procedures are used.

Results and Discussion

Three interesting points can be seen in the results in Table 2. First, Adagrad and L2 achieve better performance than online EWC and similar to SI. This shows that while Adam is popularly used for this task, Adagrad is more appropriate. This is possibly due to the fact that it results in small updates for parameters frequently used for past tasks. Second, naive rehearsal achieves performance similar to state-of-the-art methods with the same space overhead, and performs much better than online EWC and SI, especially in the incremental class scenario. This highlights the limitation of regularization-based methods and raises a question about the benefit of using a generative model, which is more difficult to train. The third is the obvious trend of difficulty among the three scenarios. The easiest one is incremental task learning, while the incremental class learning is harder than the incremental domain learning.

A similar trend also happens in the scenarios generated by the permutation strategy, which can be seen in Appendix Table 3. In that case, SI and online EWC are significantly better than Adagrad in only one of the three scenarios (incremental class), although all three methods present poor performance in that scenario. One aspect that is not apparent from these results, but is illustrated in Appendix Figure 4, is that EWC and SI variants require significant hyper-parameter tuning, with a wide gap between their worst and best performance (no such tuning was done for Adagrad). This would be difficult to tune automatically in real-world scenarios, however. One interesting cross-table comparison is that the performance of our six baselines in the scenarios generated by permutation (Appendix Table 3) is generally comparable or higher than the same scenarios generated by splitting (Table 2), although the Permuted MNIST scenarios have a larger number of classes and tasks. Such a result indicates that the permutation strategy creates simpler scenarios.

The last highlight is the comparison between Adagrad and EWC. We note that the difference between their performance is not significant in most of the experimental settings, yet EWC requires knowing the boundaries of a task to calculate Fisher information and store the parameters before switching to the next task. This creates a requirement that makes EWC less applicable than Adagrad in a real-world scenario, where the task boundaries are usually not available. Other regularization-based methods, such as SI and MAS, also suffer from the same limitation.

The strong baseline performance, especially in the incremental task learning, does not mean a scenario is solved. One can easily increase the difficulty by using a harder dataset or using a set of datasets to extend the length of a task sequence, as demonstrated in Appendix Section B. Under longer task sequences, regularization-based methods continue to degrade over time leading to questions about whether they fundamentally address catastrophic forgetting. Indeed, biological systems continue to learn new tasks with very little degradation in performance, even when that task has not been seen for a while. One avenue of research in this respect is to take a closer look at continual learning for tasks where feature sharing is possible. How such shared features can be learned when possible, or augmented with new features when tasks distributions differ significantly, is an open question. We also argue that more future effort should be put into scenarios that do not require knowing the task identity (incremental domain/class). Such scenarios are not only harder but also closer to a real scenario where the prior information about task selection is usually weak.

This research is supported by DARPA’s Lifelong Learning Machines (L2M) program, under Cooperative Agreement HR0011-18-2-001.

References

Appendices

Appendix A Permuted MNIST Experiments

In the Permuted MNIST experiments, the pre-processing of an image is similar to Split MNIST except that the order of pixels are permuted. The neural network architecture is also similar to the one in Split MNIST, except that the number of nodes in both hidden layers is 1000. Since the network size is larger, the space overhead introduced in online EWC and SI is larger (#parameter ≈(1024×1000+1000×1000)×2=4,048,000\approx(1024\times 1000+1000\times 1000)\times 2=4,048,000) and converts to a buffer of 4k images in Naive rehearsal, or 16k compressed images in Naive rehearsal-C. The generative model in DGR (implemented by ) uses a variational autoencoder whose encoder has the same architecture as the classification model; therefore DGR has a similar space overhead as online EWC and SI. In all scenarios, the standard cross-entropy for classification is optimized for 10 epochs with a learning rate ten-times smaller than Split MNIST experiments. Here we use 0.0001 for Adam, and 0.001 for SGD and Adagrad.

The best regularization coefficient for L2, EWC, online EWC, SI, and MAS is obtained through a grid search in each scenario. Our results in Table 3 is averaged from ten runs with random neural network initialization (include randomly initialized parameters in network heads which leads to much better baseline performance, compared to the zero initialization used in ). Note that the results of LwF, DGR, and RtF are from , which uses the same neural network architecture and training procedure.

Figure 2 illustrated how to use the permutation strategy to generate the three learning scenarios. Note that the number of classes here is larger than the scenarios generated by the splitting strategy (incremental task/domain: 10 versus 2; incremental class: 100 versus 10).

Table 4 compares two different initialization strategies in the hardest scenario, the incremental class learning. The two initialization strategies have the output nodes of all classes pre-allocated or not. In the pre-allocated setting, which has all output nodes be subject to the classification loss right from the beginning, an output node firstly sees negative samples since the first task, followed by some positive samples (in its corresponding task that contains the class), then again sees negative samples in the remaining tasks till the end. In the setting without pre-allocating the output nodes (the default used in our Table 2 and 3 as well as prior works), an output node is created while a new class arrived; thus the output node firstly sees some positive samples (in the newly arrived task), followed by negative samples in the remaining tasks. The results show that the scenario becomes easier with pre-allocating which enables the output nodes to learn from the beginning of the scenario.

Appendix B Lengthened Task Queue Experiments

Most continual learning methods examine their models with less than 2020 tasks in the queue. A more challenging yet practical case of continual learning is to have a dynamic environment where varied levels of domain shifting are encountered; thus we augment the incremental task learning in Section 3 by extending the length of task queue from 55 to 7878 with five datasets, including MNIST, Fashion MNIST, EMNIST letter, SVHN, and CIFAR100. Each task only contains two classes, while each class presents only once in the task queue.

The evaluation uses two different neural network architectures, CNN and MLP, to enrich the comparison, and we list the detailed architecture in Table 7. For a fair comparison, the number of parameters is similar (∼\sim300K parameters) between CNN and MLP. In the training stage, we adopt Adam as optimizer with a learning rate of 0.001 to train 10 epochs for each task with a batch size of 128. The learning rate for Adagrad is 0.001.

To examine the robustness of the regularization-based methods, we include SI and online EWC in the scenario of longer task queue. The results are presented in Table 5 (MLP) and 6 (CNN), demonstrating that the regularization-based methods can be worse than the baseline (Adam) if the regularization coefficient is not in tuned well. In contrast, the Adagrad achieves a similar level of performance without any hyperparameter search. The sensitivity analysis is provided in Figure 4 and 5, which shows that regularization-based methods are prone to the choice of the regularization coefficient and are sensitive to how the parameters (of the heads) been initialized.