Variational Continual Learning

Cuong V. Nguyen, Yingzhen Li, Thang D. Bui, Richard E. Turner

Introduction

Continual learning (also called life-long learning and incremental learning) is a very general form of online learning in which data continuously arrive in a possibly non i.i.d. way, tasks may change over time (e.g. new classes may be discovered), and entirely new tasks can emerge (Schlimmer & Fisher, 1986; Sutton & Whitehead, 1993; Ring, 1997). What is more, continual learning systems must adapt to perform well on the entire set of tasks in an incremental way that avoids revisiting all previous data at each stage. This is a key problem in machine learning since real world tasks continually evolve over time (e.g. they suffer from covariate and dataset shift) and the size of datasets often prohibits frequent batch updating. Moreover, practitioners are often interested in solving a set of related tasks that benefit from being handled jointly in order to leverage multi-task transfer. Continual learning is also of interest to cognitive science, being an intrinsic human ability.

The ubiquity of deep learning means that it is important to develop deep continual learning methods. However, it is challenging to strike a balance between adapting to recent data and retaining knowledge from old data. Too much plasticity leads to the infamous catastrophic forgetting problem (McCloskey & Cohen, 1989; Ratcliff, 1990; Goodfellow et al., 2014a) and too much stability leads to an inability to adapt. Recently there has been a resurgence of interest in this area. One approach trains individual models on each task and then carries out a second stage of training to combine them (Lee et al., 2017). A more elegant and more flexible approach maintains a single model and uses a single type of regularized training that prevents drastic changes in the parameters which have a large influence on prediction, but allows other parameters to change more freely (Li & Hoiem, 2016; Kirkpatrick et al., 2017; Zenke et al., 2017). The approach developed here follows this venerable work, but is arguably more principled, extensible and automatic.

This paper is built on the observation that there already exists an extremely general framework for continual learning: Bayesian inference. Critically, Bayesian inference retains a distribution over model parameters that indicates the plausibility of any setting given the observed data. When new data arrive, we combine what previous data have told us about the model parameters (the previous posterior) with what the current data are telling us (the likelihood). Multiplying and renormalizing yields the new posterior, from which point we can recurse. Critically, the previous posterior constrains parameters that strongly influence prediction, preventing them from changing drastically, but it allows other parameters to change. The wrinkle is that exact Bayesian inference is typically intractable and so approximations are required. Fortunately, there is an extensive literature on approximate inference for neural networks. We merge online variational inference (VI) (Ghahramani & Attias, 2000; Sato, 2001; Broderick et al., 2013) with Monte Carlo VI for neural networks (Blundell et al., 2015) to yield variational continual learning (VCL). In addition, we extend VCL to include a small episodic memory by combining VI with the coreset data summarization method (Bachem et al., 2015; Huggins et al., 2016). We demonstrate that the framework is general, applicable to both deep discriminative models and deep generative models, and that it yields excellent performance.

Continual Learning by Approximate Bayesian Inference

Consider a discriminative model that returns a probability distribution over an output yy given an input x\bm{x} and parameters θ\bm{\theta}, that is p(y∣θ,x)p(y|\bm{\theta},\bm{x}). Below we consider the specific case of a softmax distribution returned by a neural network with weight and bias parameters, but we keep the development general for now. In the continual learning setting, the goal is to learn the parameters of the model from a set of sequentially arriving datasets {xt(n),yt(n)}n=1Nt\{\bm{x}_{t}^{(n)},y_{t}^{(n)}\}_{n=1}^{N_{t}} where, in principle, each might contain a single datum, Nt=1N_{t}=1. Following a Bayesian approach, a prior distribution p(θ)p(\bm{\theta}) is placed over θ\bm{\theta}. The posterior distribution after seeing TT datasets is recovered by applying Bayes’ rule:

Here the input dependence has been suppressed on the right hand side to lighten notation. We have used the shorthand Dt={yt(n)}n=1Nt\mathcal{D}_{t}=\{y_{t}^{(n)}\}_{n=1}^{N_{t}}. Importantly, a recursion has been identified whereby the posterior after seeing the TT-th dataset is produced by taking the posterior after seeing the (T−1)(T-1)-th dataset, multiplying by the likelihood and renormalizing. In other words, online updating emerges naturally from Bayes’ rule.

Variational continual learning employs a projection operator defined through a KL divergence minimization over the set of allowed approximate posteriors Q\mathcal{Q},

The zeroth approximate distribution is defined to be the prior, q0(θ)=p(θ)q_{0}(\bm{\theta})=p(\bm{\theta}). ZtZ_{t} is the intractable normalizing constant of pt∗(θ)=qt−1(θ) p(Dt∣θ)p_{t}^{*}(\bm{\theta})=q_{t-1}(\bm{\theta})~{}p(\mathcal{D}_{t}|\bm{\theta}) and is not required to compute the optimum.

VCL will perform exact Bayesian inference if the true posterior is a member of the approximating family, p(θ∣D1,D2,…,Dt)∈Qp(\bm{\theta}|\mathcal{D}_{1},\mathcal{D}_{2},\ldots,\mathcal{D}_{t})\in\mathcal{Q} at every step tt. Typically this will not be the case and we might worry that performing repeated approximations may accumulate errors causing the algorithm to forget old tasks, for example. Furthermore, the minimization at each step may also be approximate (e.g. due to employing an additional Monte Carlo approximation) and so additional information may be lost. In order to mitigate this potential problem, we extend VCL to include a small representative set of data from previously observed tasks that we call the coreset. The coreset is analogous to an episodic memory that retains key information (in our case, important training data points) from previous tasks which the algorithm can revisit in order to refresh its memory of them. The use of an episodic memory for continual learning has also been explored by Lopez-Paz & Ranzato (2017).

Algorithm 1 describes coreset VCL. For each task, the new coreset CtC_{t} is produced by selecting new data points from the current task and a selection from the old coreset Ct−1C_{t-1}. Any heuristic can be used to make these selections, e.g. KK data points can be selected at random from Dt\mathcal{D}_{t} and added to Ct−1C_{t-1} to form an unbiased new coreset CtC_{t}. Alternatively, the greedy KK-center algorithm (Gonzalez, 1985) can be used to return KK data points that are guaranteed to be spread throughout the input space. Next, a variational recursion is developed. Bayes’ rule can be used to decompose the true posterior taking care to break out contributions from the coreset,

Variational Continual Learning in Deep Discriminative Models

The VCL framework is general and can be applied to many discriminative probabilistic models. Here we apply it to continual learning of deep fully-connected neural network classifiers. Before turning to the application of VCL, we first consider the architecture of neural networks suitable for performing continual learning. In simple instances of discriminative continual learning, where data are arriving in an i.i.d. way or where only the input distribution p(x1:T)p(\bm{x}_{1:T}) changes over time, a standard single-head discriminative neural network suffices. In many cases the tasks, although related, might involve different output variables. Standard practice in multi-task learning (Bakker & Heskes, 2003) uses networks that share parameters close to the inputs but with separate heads for each output, hence multi-head networks. Graphical models depicting the network architecture for deep discriminative and deep generative models are shown in fig. 1.

Recent work has explored more advanced structures for continual learning (Rusu et al., 2016) and multi-task learning more generally (Swietojanski & Renals, 2014; Rebuffi et al., 2017). These architectural advances are complementary to the new learning schemes developed here and a synthesis of the two would be potentially more powerful. Moreover, a general solution to continual learning would perform automatic continual model building adding new bespoke structure to the existing model as new tasks are encountered. Although this is a very interesting research direction, here we make the simplifying assumption that the model structure is known a priori.

VCL requires specification of q(θ)q(\bm{\theta}) where θ\bm{\theta} in the current case is a DD dimensional vector formed by stacking the network’s biases and weights. For simplicity we use a Gaussian mean-field approximate posterior qt(θ)=∏d=1DN(θt,d;μt,d,σt,d2)q_{t}(\bm{\theta})=\prod_{d=1}^{D}\mathcal{N}(\theta_{t,d};\mu_{t,d},\sigma_{t,d}^{2}). Taking the most general case of a multi-head network, before task kk is encountered the posterior distribution over the associated head parameters will remain at the prior and so q(θkH)=p(θkH)q(\bm{\theta}_{k}^{H})=p(\bm{\theta}_{k}^{H}). This is convenient as it means the variational approximation can be grown incrementally, starting from the prior, as each task emerges. Moreover, only tasks present in the current dataset Dt\mathcal{D}_{t} need to have their posterior distributions over head parameters updated. The shared parameters, on the other hand, will be constantly updated. Training the network using the VFE approach in eq. 1 is equivalent to maximizing the negative online variational free energy or the variational lower bound to the online marginal likelihood

Variational Continual Learning in Deep Generative Models

Deep generative models (DGMs) have garnered much recent attention. By passing a simple noise variable (e.g. Gaussian noise) through a deep neural network, these models have been shown to be able to generate realistic images, sounds and videos sequences (Chung et al., 2015; Kingma et al., 2016; Vondrick et al., 2016). Standard approaches for learning DGMs have focused on batch learning, i.e. the observed instances are assumed to be i.i.d. and are all available at the same time. In this section we extend the VCL framework to encompass variational auto-encoders (VAEs) (Kingma & Welling, 2014; Rezende et al., 2014), a form of DGM. The approach could be extended to generative adversarial networks (GANs) (Goodfellow et al., 2014b) for which continual learning is an open problem (see Seff et al. (2017) for an initial attempt).

where ϕ\bm{\phi} are the variational parameters of the approximate posterior or “encoder”.

The approximate MLE approach is unsuitable for the continual learning setting as it does not return parameter uncertainty estimates that are critical for weighting the information learned from old data. So, instead the VCL approach will approximate the full posterior distribution over parameters, qt(θ)≈p(θ∣D1:t)q_{t}(\bm{\theta})\approx p(\bm{\theta}|\mathcal{D}_{1:t}), after observing the tt-th dataset. Specifically, the approximate posterior qtq_{t} is obtained by maximizing the full variational lower bound with respect to qtq_{t} and ϕ\bm{\phi}:

where the encoder network qϕ(zt(n)∣xt(n))q_{\bm{\phi}}(\bm{z}_{t}^{(n)}|\bm{x}_{t}^{(n)}) is parameterized by ϕ\bm{\phi} which is task-specific. It is likely to be beneficial to share (parts of) these encoder networks, but this is not investigated in this paper.

As was the case for multi-head discriminative models, we can split the generative model into shared and task-specific parts. There are two options: (i) the generative models share across tasks the network that generates observations x\mathbf{x} from the intermediate-level representations h\mathbf{h}, but have private “head networks” for generating h\mathbf{h} from the latent variables z\mathbf{z} (see fig. 1(b)), and (ii) the other way around. Architecture (i) is arguably more appropriate when data are composed of a common set of structural primitives (such as strokes in handwritten digits) that are selected by high level variables (character identities). Moreover, initial experiments on architecture (ii) indicated that information about the current task tended to be encoded entirely in the task-specific lower-level network negating multi-task transfer. For these reasons, we focus on architecture (i) in the experiments.

Related Work

Continual Learning for Deep Discriminative Models: Many neural network continual learning approaches employ regularized maximum likelihood estimation, optimizing objectives of the form:

Lt(θ)=∑n=1Ntlog⁡p(yt(n)∣θ,xt(n))−12λt(θ−θt−1)⊺Σt−1−1(θ−θt−1).{\hskip 56.9055pt}\mathcal{L}^{t}(\bm{\theta})=\sum_{n=1}^{N_{t}}\log p(y_{t}^{(n)}|\bm{\theta},\mathbf{x}_{t}^{(n)})-\frac{1}{2}\lambda_{t}(\bm{\theta}-\bm{\theta}_{t-1})^{\intercal}\Sigma^{-1}_{t-1}(\bm{\theta}-\bm{\theta}_{t-1}).

Here the regularization biases the new parameter estimates towards those estimated at the previous step θt−1\bm{\theta}_{t-1}. λt\lambda_{t} is a user-selected hyper-parameter that controls the overall contribution from previous data and Σt−1\Sigma_{t-1} is a matrix (normally diagonal in form) that encodes the relative strength of the regularization on each element of θ\bm{\theta}. We now discuss specific instances of this scheme:

Laplace Propagation (LP) (Smola et al., 2004): applying Laplace’s approximation at each step leads to a recursion for Σt−1\Sigma_{t}^{-1}, which is initialized using the covariance of the Gaussian prior,

Σt−1=Φt+Σt−1−1{\hskip 14.22636pt}\Sigma_{t}^{-1}=\Phi_{t}+\Sigma_{t-1}^{-1} where \Phi_{t}=-\nabla\nabla_{\bm{\theta}}\sum_{n=1}^{N_{t}}\log p(y_{t}^{(n)}|\bm{\theta},\mathbf{x}_{t}^{(n)})\Big{|}_{\bm{\theta}=\bm{\theta}_{t}} and λt=1\lambda_{t}=1.

To avoid computing the full Hessian of the likelihood, diagonal Laplace propagation retains only the diagonal terms of Σt−1\Sigma_{t}^{-1}.

Elastic Weight Consolidation (EWC) (Kirkpatrick et al., 2017) builds on diagonal Laplace propagation by approximating the average Hessian of the likelihoods using well-known identities for the Fisher information:

{\hskip 71.13188pt}\Phi_{t}\approx\text{diag}\left(\sum_{n=1}^{N_{t}}\left(\nabla_{\bm{\theta}}\log p(y_{t}^{(n)}|\bm{\theta},\mathbf{x}_{t}^{(n)})\right)^{2}\Big{|}_{\bm{\theta}=\bm{\theta}_{t}}\right).

EWC also modifies the Laplace regularization, 12(θ−θt−1)⊺(Σ0−1+∑t′=1t−1Φt′)(θ−θt−1)\frac{1}{2}(\bm{\theta}-\bm{\theta}_{t-1})^{\intercal}(\Sigma_{0}^{-1}+\sum_{t^{\prime}=1}^{t-1}\Phi_{t^{\prime}})(\bm{\theta}-\bm{\theta}_{t-1}), introducing hyper-parameters, removing the prior and regularizing to intermediate parameter estimates, rather than just those derived from the last task, 12∑t′=1t−1λt′(θ−θt′−1)⊺Φt′(θ−θt′−1){\frac{1}{2}\sum_{t^{\prime}=1}^{t-1}\lambda_{t^{\prime}}(\bm{\theta}-\bm{\theta}_{t^{\prime}-1})^{\intercal}\Phi_{t^{\prime}}(\bm{\theta}-\bm{\theta}_{t^{\prime}-1})}. These changes may be unnecessary (Huszár, 2017; 2018) and require storing θ1:t−1\bm{\theta}_{1:t-1}, but may slightly improve performance (see Kirkpatrick et al. (2018) and our experiments).

Synaptic Intelligence (SI) (Zenke et al., 2017): SI computes Σt−1\Sigma_{t}^{-1} using a measure of the importance of each parameter to each task. Practically, this is achieved by comparing the changing rate of the gradients of the objective and the changing rate of the parameters.

VCL differs from the above methods in several ways. First, unlike MAP, EWC and SI, it does not have free parameters that need to be tuned on a validation set. This can be especially awkward in the online setting. Second, although the KL regularization penalizes the mean of the approximate posterior through a quadratic cost, a full distribution is retained and averaged over at training time and at test time. Third, VI is generally thought to return better uncertainty estimates than approaches like Laplace’s method and MAP estimation, and we have argued this is critical for continual learning.

There is a long history of research on approximate Bayesian training of neural networks, including extended Kalman filtering (Singhal & Wu, 1989), Laplace’s approximation (MacKay, 1992), variational inference (Hinton & Van Camp, 1993; Barber & Bishop, 1998; Graves, 2011; Blundell et al., 2015; Gal & Ghahramani, 2016), sequential Monte Carlo (de Freitas et al., 2000), expectation propagation (EP) (Hernández-Lobato & Adams, 2015), and approximate power EP (Hernández-Lobato et al., 2016). These approaches have focused on batch learning, but the framework described in section 2 enables them to be applied to continual learning. On the other hand, online variational inference has been previously explored (Ghahramani & Attias, 2000; Broderick et al., 2013; Bui et al., 2017), but not for neural networks or in the context of sets of related complex tasks.

Continual Learning for Deep Generative Models: A naïve continual learning approach for deep generative models would directly apply the VAE algorithm to the new dataset Dt\mathcal{D}_{t} with the model parameters initialized at the previous parameter values θt−1\bm{\theta}_{t-1}. The experiments show that this approach leads to catastrophic forgetting, in the sense that the generator can only generate instances that are similar to the data points from the most recently observed task.

Alternatively, EWC regularization can be added to the VAE objective:

However computing Φt\Phi_{t} requires the gradient of the intractable marginal likelihood ∇θlog⁡p(x∣θ)\nabla_{\bm{\theta}}\log p(\mathbf{x}|\bm{\theta}). Instead, we can approximate the marginal likelihood by the variational lower bound, i.e.

Similar variational lower-bound approximations apply when computing the Hessian matrices for LP and Σt−1\Sigma_{t}^{-1} for SI. An importance sampling estimate could also be used (Burda et al., 2016).

Experiments

The experiments evaluate the performance and flexibility of VCL through three discriminative tasks and two generative tasks. Standard continual learning benchmarks are used where possible. Comparisons are made to EWC, diagonal LP and SI that employ tuned hyper-parameters λ\lambda whereas VCL’s objective is hyper-parameter free. More details of the experiment settings and an additional experiment are available in the appendix.An implementation of the methods proposed in this paper can be found at: https://github.com/nvcuong/variational-continual-learning.

We consider the following three continual learning experiments for deep discriminative models.

Permuted MNIST: This is a popular continual learning benchmark (Goodfellow et al., 2014a; Kirkpatrick et al., 2017; Zenke et al., 2017). The dataset received at each time step Dt\mathcal{D}_{t} consists of labeled MNIST images whose pixels have undergone a fixed random permutation. We compare VCL to EWC, SI, and diagonal LP. For all algorithms, we use fully connected single-head networks with two hidden layers, where each layer contains 100 hidden units with ReLU activations. We evaluate three versions of VCL: VCL with no coreset, VCL with a random coreset, and VCL with a coreset selected by the K-center method. For the coresets, we select 200 data points from each task.

Figure 2 compares the average test set accuracy on all observed tasks. From this figure, VCL outperforms EWC, SI, and LP by large margins, even though they benefited from an extensive hyper-parameter search for λ\lambda. Diagonal LP performs slightly worse than EWC both when λ=1\lambda=1 and when the values of λ\lambda are tuned. After 10 tasks, VCL achieves 90% average accuracy, while EWC, SI, and LP only achieve 84%, 86%, and 82% respectively. The results also show that the coresets perform poorly by themselves, but combining them with VCL leads to a modest improvement: both random coresets and K-center coresets achieve 93% accuracy.

We also investigate the effect of the coreset size. In fig. 3, we plot the average test set accuracy of VCL with random coresets of different sizes. At the coreset size of 5,000 examples per task, VCL achieves 95.5% accuracy after 10 tasks, which is significantly better than the 90% of vanilla VCL. Performance improves with the coreset size although it asymptotes for large coresets as expected: if a sufficiently large coreset is employed, it will be fully representative of the task and thus training on the coreset alone can achieve a good performance. However, the experiments show that the combination of VCL and coresets is advantageous even for large coresets.

Split MNIST: This experiment was used by Zenke et al. (2017) to assess the SI method. Five binary classification tasks from the MNIST dataset arrive in sequence: 0/1, 2/3, 4/5, 6/7, and 8/9. We use fully connected multi-head networks with two hidden layers comprising 256 hidden units with ReLU activations. We compare VCL (with and without coresets) to EWC, SI, and diagonal LP. For the coresets, 40 data points from each task are selected through random sampling or the K-center method.

Figure 4 compares the test set accuracy on individual tasks (averaged over 10 runs) as well as the accumulated accuracy averaged over tasks (right). As an upper bound on the algorithms’ performance, we compare to batch VI trained on the full dataset. From this figure, VCL significantly outperforms EWC and LP although it is slightly worse than SI. Again, unlike VCL, EWC and SI benefited from a hyper-parameter search for λ\lambda, but a value close to 1 performs well in both cases. After 5 tasks, VCL achieves 97.0% average accuracy on all tasks, while EWC, SI, and LP attain 63.1%, 98.9%, and 61.2% respectively. Adding the coreset improves VCL to around 98.4% accuracy.

Split notMNIST: This experiment is similar to the previous one, but it uses the more challenging notMNIST dataset and deeper networks. The notMNIST datasetThe notMNIST dataset is available at: http://yaroslavvb.blogspot.com/2011/09/notmnist-dataset.html. here contains 400,000 images of the characters from A to J with different font styles. We consider five binary classification tasks: A/F, B/G, C/H, D/I, and E/J using deeper networks comprising four hidden layers of 150 hidden units with ReLU activations. The other settings are kept the same as the previous experiment. VCL is competitive with SI and significantly outperforms EWC and LP (see fig. 5), although the SI and EWC baselines benefited from a hyper-parameter search. VCL achieves 92.0% average accuracy after 5 tasks, while EWC, SI, and LP attain 71%, 94%, and 63% respectively. Adding the random coreset improves the performance of VCL to 96% accuracy.

2 Experiments with Deep Generative Models

We consider two continual learning experiments for deep generative models: MNIST digit generation and notMNIST (small) character generation. In both cases, ten datasets are received in sequence. For MNIST, the first dataset comprises exclusively of images of the digit zero, the second dataset ones and so on. For notMNIST, the datasets contain the characters A to J in sequence. The generative model consists of shared and task-specific components, each represented by a one hidden layer neural network with 500 hidden units (see fig. 1(b)). The dimensionality of the latent variable z\mathbf{z} and the intermediate representation h\mathbf{h} are 50 and 500, respectively. We use task-specific encoders that are neural networks with symmetric architectures to the generator.

We compare VCL to naïve online learning using the standard VAE objective, LP, EWC and SI (with hyper-parameters λ=1,10,100\lambda=1,10,100). For full details of the experimental settings see Appendix E. Samples from the generative models attained at different time steps are shown in fig. 6. The naïve online learning method fails catastrophically and so numerical results are omitted. LP, EWC, SI and VCL remember previous tasks, with SI and VCL achieving the best visual quality on both datasets.

The algorithms are quantitatively evaluated using two metrics in fig. 7: an importance sampling estimate of the test log-likelihood (test-LL) using 5,0005,000 samples and a measure of quality we term “classifier uncertainty”. For the latter, we train a discriminative classifier for the digits/alphabets to achieve high accuracy. The quality of generated samples can then be assessed by the KL-divergence from the one-hot vector indicating the task, to the output classification probability vector computed on the generated images. A well-trained generator will produce images that are correctly classified in high confidence resulting in zero KL. We only report the best performance for LP, EWC and SI.

We observe that LP and EWC perform similarly, most likely due to the fact that both LP and EWC use the same Σt\Sigma_{t} matrices. EWC achieves significantly worse performance than SI. VCL is on par with or slightly better than SI. VCL has a superior long-term memory of previous tasks which leads to better overall performance on both metrics even though it does not have tuned hyper-parameters in its objective function. For MNIST, the performance of LP and EWC deteriorate markedly when moving from task “digit 0” to “digit 1” possibly due to the large task differences. Also for all experimental settings we tried, SI fails to produce high test-LL results after task “digit 7”. Future work will investigate continual learning on a sequence of tasks that follows “adversarial ordering”, i.e. the ordering that makes the next task maximally different from the current task.

Conclusion

Approximate Bayesian inference provides a natural framework for continual learning. Variational Continual Learning (VCL), developed in this paper, is an approach in this vein that extends online variational inference to handle more general continual learning tasks and complex neural network models. VCL can be enhanced by including a small episodic memory that leverages coreset algorithms from statistics and connects to message-scheduling in variational message passing. We demonstrated how the VCL framework can be applied to both discriminative and generative models. Experimental results showed state-of-the-art performance when compared to previous continual learning approaches, even though VCL has no free parameters in its objective function. Future work should explore alternative approximate inference methods using the same framework and also develop more sophisticated episodic memories. Finally, we note that VCL is ideally suited for efficient model refinement in sequential decision making problems, such as reinforcement learning and active learning.

The authors would like to thank Brian Trippe, Siddharth Swaroop, and Matej Balog for insightful comments and discussion. Cuong V. Nguyen is supported by EPSRC grant EP/M0269571. Yingzhen Li is supported by the Schlumberger FFTF Fellowship. Thang D. Bui is supported by the Google European Doctoral Fellowship. Richard E. Turner is supported by Google as well as EPSRC grants EP/M0269571 and EP/L000776/1.

References

Appendix

In this experiment, we use fully connected single-head networks with two hidden layers, where each layer contains 100 hidden units with ReLU activations. The metric used for comparison is the test set accuracy on all observed tasks. We train all the models using the Adam optimizer (Kingma & Ba, 2015) with learning rate 10−310^{-3} since we found that it works best for all models. All the VCL algorithms are trained with batch size 256 and 100 epochs. For all the algorithms with coresets, we choose 200 examples from each task to include into the coresets. The algorithms that use only the coresets are trained using the VFE method with batch size equal to the coreset size and 100 epochs. We use the prior N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}) and initialize our optimizer for the first task at the mean of the maximum likelihood model and a very small initial variance (10−610^{-6}).

We compare the performance of SI with hyper-parameters λ=0.01,0.1,0.5,1,2\lambda=0.01,0.1,0.5,1,2 and select the best one (λ=0.5\lambda=0.5) as our baseline (see fig. 8). Following Zenke et al. (2017), we train these models with batch size 256 and 20 epochs. We also compare the performance of EWC with λ=1,10,102,103,104\lambda=1,10,10^{2},10^{3},10^{4} and select the best value λ=102\lambda=10^{2} as our baseline (see fig. 9). The models are trained without dropout and with batch size 200 and 20 epochs. We approximate the Fisher information matrices in EWC using 600 random samples drawn from the current dataset. For diagonal LP, we compare the performance of λ=0.01,0.1,1,10,100\lambda=0.01,0.1,1,10,100 and use the best value λ=0.1\lambda=0.1 as our baseline (see fig. 10). The models are also trained with prior N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}), batch size 200, and 20 epochs. The Hessians of LP are approximated using the Fisher information matrices with 200 samples.

B Further Details for Split MNIST Experiment

In this experiment, we use fully connected multi-head networks with two hidden layers, each of which contains 256 hidden units with ReLU activations. At each time step, we compare the test set accuracy of the current model on all observed tasks separately. We also plot the average accuracy over all tasks in the last column of fig. 4. All the results for this experiment are the averages over 10 runs of the algorithms with different random seeds.

We use the Adam optimizer with learning rate 10−310^{-3} for all models. All the VCL algorithms are trained with batch size equal to the size of the training set and 120 epochs. We use the prior N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}) and initialize our optimizer for the first task at the mean of the maximum likelihood model and a very small initial variance (10−610^{-6}). For the coresets, we choose 40 examples from each task to include into the coresets. In this experiment, the final approximate posterior used for prediction in eq. 3 is computed for each task separately using the coreset points corresponding to the task. The algorithms that use only the coresets are trained using the VFE method with batch size equal to the coreset size and 120 epochs.

We compare the performance of SI with λ=0.01,0.1,1,2,3\lambda=0.01,0.1,1,2,3 and use the best value λ=1\lambda=1 as our baseline (see fig. 14). We also compare EWC with both single-head and multi-head models and λ=1,10,102,103,104\lambda=1,10,10^{2},10^{3},10^{4} (see fig. 14). We approximate the Fisher information matrices using 200 random samples drawn from the current dataset. The figure shows that the multi-head models work better than the single-head models for EWC, and the performance is insensitive to the choice of λ\lambda. Thus, we use the multi-head model with λ=1\lambda=1 as the EWC baseline for our experiment. For diagonal LP, we also use the multi-head model with λ=1\lambda=1, prior N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}), and approximate the Hessians using the Fisher information matrices with 200 samples.

C Further Details for Split notMNIST Experiment

The settings for this experiment are the same as those in the Split MNIST experiment above, except that we use deeper networks with 4 hidden layers, each of which contains 150 hidden units. Figures 14 and 14 show the performance of SI and EWC with different hyper-parameter values respectively. In the experiment, we choose λ=104\lambda=10^{4} for multi-head EWC, λ=1\lambda=1 for multi-head LP, and λ=0.1\lambda=0.1 for SI.

D Additional Experiment on a Toy 2D Dataset

Here we consider a small experiment on a toy 2D dataset to understand some of the properties of EWC with λ=1\lambda=1 and VCL. The experiment comprises two sequential binary classification tasks. The first task contains two classes generated from two Gaussian distributions. The data points (with green and black classes) for this task are shown in the first column of fig. 15. The second task contains two classes also generated from two Gaussian distributions. The green class for this task has the same input distribution as the first task, while the input distribution for the black class is different. The data points for this task are shown in the third columns of fig. 15. Each task contains 200 data points with 100 data points in each class.

We compare the multi-head models trained by VCL and EWC on these two tasks. In this experiment, we use fully connected networks with one hidden layer containing 20 hidden units with ReLU activations. The first column of fig. 15 shows the contours of the prediction probabilities after observing the first task, and both methods perform reasonably well for this task. However, after observing the second task, the EWC method fails to learn the classifiers for both tasks, while the VCL method are still able to learn good classifiers for them.

E Further Details on Deep Generative Model Experiments

In the experiments on Deep Generative Models, the learning rates and numbers of optimization epochs are tuned on separate training of each tasks. This gives a learning rate of 10−410^{-4} and the number of epochs 200 for MNIST (except for SI) and 400 for notMNIST. For SI we optimize for 400 epochs on MNIST. For the VCL approach, the parameters of qt(θ)q_{t}(\bm{\theta}) are initialized to have the same mean as qt−1(θ)q_{t-1}(\bm{\theta}) and the log standard deviation is set to 10−610^{-6}.

The generative model consists of shared and task-specific components, each represented by a one hidden layer neural network with 500 hidden units (see fig. 1(b)). The dimensionality of the latent variable z\mathbf{z} and the intermediate representation h\mathbf{h} are 50 and 500, respectively. We use task-specific encoders that are neural networks with symmetric architectures to the generator.

F Memory of a Linear Regression Model Trained with VCL on Random Patterns

In many probabilistic models with conjugate priors, the exact posterior of the parameters/latent variables can be obtained. For example, a Bayesian linear regression model with a Gaussian prior over the parameters and a Gaussian observation model has a Gaussian posterior. If we insist on using a diagonal Gaussian approximation to this posterior and use either the variational free energy method or the Laplace’s approximation, we will end up at the same solution – a Gaussian distribution with the same mean as that of the exact posterior and the diagonal precisions being the diagonal precisions of the exact posterior. Consequently, the online variational Gaussian approximation will give the same result to that given by the online Laplace’s approximation. However, when a diagonal Gaussian approximation is used, the batch and sequential solutions are different.

In the following, we will explicitly detail the sequential variational updates for a Bayesian linear regression model to associate random binary patterns to binary outcomes (Kirkpatrick et al., 2017), and show its relationship to the online Laplace’s approximation and the EWC approach of Kirkpatrick et al. (2017). The task consists of associating a random DD-dimensional binary vector xt\mathbf{x}_{t} to a random binary output yty_{t} by learning a weight vector WW. Note that the possible values of the features and outputs are 11 and −1-1, and not 0 and 1. We also assume that the model sees only one input-output pair {xt,yt}\{\mathbf{x}_{t},y_{t}\} at the tt-th time step and the previous approximate posterior qt−1(W)=∏d=1DN(wd;mt−1,d,vt−1,d)q_{t-1}(W)=\prod_{d=1}^{D}\mathcal{N}(w_{d};m_{t-1,d},v_{t-1,d}), and that the observation noise variance σy2\sigma_{y}^{2} is fixed. The variational updates for the mean and precisions are available in closed-form as follows:

By further assuming that σy=1\sigma_{y}=1 and ∥xt∥2=1\lVert\mathbf{x}_{t}\rVert_{2}=1, the equations above become:

where xˉt=ytxt\bar{\mathbf{x}}_{t}=y_{t}\mathbf{x}_{t}. When v0,d−1v_{0,d}^{-1} = 0, i.e. the prior is ignored, the update for the mean above is exactly equation S4 in the supplementary material of Kirkpatrick et al. (2017). Therefore, in this case, the memory of the network trained by online variational inference is identical to that of the online Laplace’s method, provided in (Kirkpatrick et al., 2017). These methods differ, in practice, when the prior is not ignored or when the parameter regularization constraints are accumulated as discussed in the main text. This equivalence also does not hold in the general case, as discussed by Opper & Archambeau (2009) where the Gaussian variational approximation can be interpreted as an averaged and smoothed Laplace’s approximation.