Adaptive Aggregation Networks for Class-Incremental Learning

Yaoyao Liu, Bernt Schiele, Qianru Sun

Introduction

AI systems are expected to work in an incremental manner when the amount of knowledge increases over time. They should be capable of learning new concepts while maintaining the ability to recognize previous ones. However, deep-neural-network-based systems often suffer from serious forgetting problems (called “catastrophic forgetting”) when they are continuously updated using new coming data. This is due to two facts: (i) the updates can override the knowledge acquired from the previous data , and (ii) the model can not replay the entire previous data to regain the old knowledge.

To encourage solving these problems, defined a class-incremental learning (CIL) protocol for image classification where the training data of different classes gradually come phase-by-phase. In each phase, the classifier is re-trained on new class data, and then evaluated on the test data of both old and new classes. To prevent trivial algorithms such as storing all old data for replaying, there is a strict memory budget due to which a tiny set of exemplars of old classes can be saved in the memory. This memory constraint causes a serious data imbalance problem between old and new classes, and indirectly causes the main problem of CIL – the stability-plasticity dilemma . In particular, higher plasticity results in the forgetting of old classes , while higher stability weakens the model from learning the data of new classes (that contain a large number of samples). Existing CIL works try to balance stability and plasticity using data strategies. For example, as illustrated in Figure 1 (a) and (b), some early methods train their models on the imbalanced dataset where there is only a small set of exemplars for old classes , and recent methods include a fine-tuning step using a balanced subset of exemplars sampled from all classes . However, these data strategies are still limited in terms of effectiveness. For example, when using the models trained after 2525 phases, LUCIR and Mnemonics “forget” the initial 5050 classes by 30%30\% and 20%20\%, respectively, on the ImageNet dataset .

In this paper, we address the stability-plasticity dilemma by introducing a novel network architecture called Adaptive Aggregation Networks (AANets). Taking the ResNet as an example of baseline architectures, we explicitly build two residual blocks (at each residual level) in AANets: one for maintaining the knowledge of old classes (i.e., the stability) and the other for learning new classes (i.e., the plasticity), as shown in Figure 1 (c). We achieve these by allowing these two blocks to have different levels of learnability, i.e., less learnable parameters in the stable block but more in the plastic one. We apply aggregation weights to the output feature maps of these blocks, sum them up, and pass the result maps to the next residual level. In this way, we are able to dynamically balance the usage of these blocks by updating their aggregation weights. To achieve auto-updating, we take the weights as hyperparameters and optimize them in an end-to-end manner .

Technically, the overall optimization of AANets is bilevel. Level-1 is to learn the network parameters for two types of residual blocks, and level-2 is to adapt their aggregation weights. More specifically, level-1 is the standard optimization of network parameters, for which we use all the data available in the phase. Level-2 aims to balance the usage of the two types of blocks, for which we optimize the aggregation weights using a balanced subset (by downsampling the data of new classes), as illustrated in Figure 1 (c). We formulate these two levels in a bilevel optimization program (BOP) that solves two optimization problems alternatively, i.e., update network parameters with aggregation weights fixed, and then switch. For evaluation, we conduct CIL experiments on three widely-used benchmarks, CIFAR-100, ImageNet-Subset, and ImageNet. We find that many existing CIL methods, e.g., iCaRL , LUCIR , Mnemonics Training , and PODNet , can be directly incorporated in the architecture of AANets, yielding consistent performance improvements. We observe that a straightforward plug-in causes memory overheads, e.g., 26%26\% and 15%15\% respectively for CIFAR-100 and ImageNet-Subset. For a fair comparison, we conduct additional experiments under the settings of zero overhead (e.g., by reducing the number of old exemplars for training AANets), and validate that our approach still achieves top performance across all datasets.

Our contribution is three-fold: 1) a novel and generic network architecture called AANets specially designed for tackling the stability-plasticity dilemma in CIL tasks; 2) a BOP-based formulation and an end-to-end training solution for optimizing AANets; and 3) extensive experiments on three CIL benchmarks by incorporating four baseline methods in the architecture of AANets.

Related Work

Incremental learning aims to learn efficient machine models from the data that gradually come in a sequence of training phases. Closely related topics are referred to as continual learning and lifelong learning . Recent incremental learning approaches are either task-based, i.e., all-class data come but are from a different dataset for each new phase , or class-based i.e., each phase has the data of a new set of classes coming from the identical dataset . The latter one is typically called class-incremental learning (CIL), and our work is based on this setting. Related methods mainly focus on how to solve the problems of forgetting old data. Based on their specific methods, they can be categorized into three classes: regularization-based, replay-based, and parameter-isolation-based . Regularization-based methods introduce regularization terms in the loss function to consolidate previous knowledge when learning new data. Li et al. proposed the regularization term of knowledge distillation . Hou et al. introduced a series of new regularization terms such as for less-forgetting constraint and inter-class separation to mitigate the negative effects caused by the data imbalance between old and new classes. Douillard et al. proposed an effective spatial- based distillation loss applied throughout the model and also a representation comprising multiple proxy vectors for each object class. Tao et al. built the framework with a topology-preserving loss to maintain the topology in the feature space. Yu et al. estimated the drift of previous classes during the training of new classes. Replay-based methods store a tiny subset of old data, and replay the model on them (together with new class data) to reduce the forgetting. Rebuffi et al. picked the nearest neighbors to the average sample per class to build this subset. Liu et al. parameterized the samples in the subset, and then meta-optimized them automatically in an end-to-end manner taking the representation ability of the whole set as the meta-learning objective. Belouadah et al. proposed to leverage a second memory to store statistics of old classes in rather compact formats. Parameter-isolation-based methods are used in task-based incremental learning (but not CIL). Related methods dedicate different model parameters for different incremental phases, to prevent model forgetting (caused by parameter overwritten). If no constraints on the size of the neural network is given, one can grow new branches for new tasks while freezing old branches. Rusu et al. proposed “progressive networks” to integrate the desiderata of different tasks directly into the networks. Abati et al. equipped each convolution layer with task-specific gating modules that select specific filters to learn each new task. Rajasegaran et al. progressively chose the optimal paths for the new task while encouraging to share parameters across tasks. Xu et al. searched for the best neural network architecture for each coming task by leveraging reinforcement learning strategies. Our differences with these methods include the following aspects. We focus on class-incremental learning, and more importantly, our approach does not continuously increase the network size. We validate in the experiments that under a strict memory budget, our approach can surpass many related methods and its plug-in versions on these related methods can bring consistent performance improvements.

Bilevel Optimization Program can be used to optimize hyperparameters of deep models. Technically, the network parameters are updated at one level and the key hyperparameters are updated at another level . Recently, a few bilevel-optimization-based approaches have emerged for tackling incremental learning tasks. Wu et al. learned a bias correction layer for incremental learning models using a bilevel optimization framework. Rajasegaran et al. incrementally learned new tasks while learning a generic model to retain the knowledge from all tasks. Riemer et al. learned network updates that are well-aligned with previous phases, such as to avoid learning towards any distracting directions. In our work, we apply the bilevel optimization program to update the aggregation weights in our AANets.

Adaptive Aggregation Networks (AANets)

Class-Incremental Learning (CIL) usually assumes (N+1)(N+1) learning phases in total, i.e., one initial phase and NN incremental phases during which the number of classes gradually increases . In the initial phase, data D0\mathcal{D}_{0} is available to train the first model Θ0\Theta_{0}. There is a strict memory budget in CIL systems, so after the phase, only a small subset of D0\mathcal{D}_{0} (exemplars denoted as E0\mathcal{E}_{0}) can be stored in the memory and used as replay samples in later phases. Specifically in the ii-th (i≥1i\geq 1) phase, we load the exemplars of old classes E0:i−1={E0,…,Ei−1}\mathcal{E}_{0:i-1}=\{\mathcal{E}_{0},\dots,\mathcal{E}_{i-1}\} to train model Θi\Theta_{i} together with new class data Di\mathcal{D}_{i}. Then, we evaluate the trained model on the test data containing both old and new classes. We repeat such training and evaluation through all phases.

The key issue of CIL is that the models trained at new phases easily “forget” old classes. To tackle this, we introduce a novel architecture called AANets. AANets is based on a ResNet-type architecture, and each of its residual levels is composed of two types of residual blocks: a plastic one to adapt to new class data and a stable one to maintain the knowledge learned from old classes. The details of this architecture are elaborated in Section 3.1. The steps for optimizing AANets are given in Section 3.2.

In Figure 2, we provide an illustrative example of our AANets with three residual levels. The inputs xx^{} are the images and the outputs xx^{} are the features used to train classifiers. Each of our residual “levels” consists of two parallel residual “blocks” (of the original ResNet ): the orange one (called plastic block) will have its parameters fully adapted to new class data, while the blue one (called stable block) has its parameters partially fixed in order to maintain the knowledge learned from old classes. After feeding the inputs to Level 1, we obtain two sets of feature maps respectively from two blocks, and aggregate them after applying the aggregation weights α\alpha^{}. Then, we feed the resulted maps to Level 2 and repeat the aggregation. We apply the same steps for Level 3. Finally, we pool the resulted maps obtained from Level 3 to train classifiers. Below we elaborate the details of this dual-branch design as well as the steps for feature extraction and aggregation.

where ⊙\odot donates the element-wise multiplication. Assuming there are QQ layers in total, the overall scaling weights can be denoted as ϕ={ϕq}q=1Q\phi=\{\phi_{q}\}_{q=1}^{Q}.

Feature Extraction and Aggregation. We elaborate on the process of feature extraction and aggregation across all residual levels in the AANets, as illustrated in Figure 2. Let Fμ[k](⋅)\mathcal{F}^{[k]}_{\mu}{(\cdot)} denote the transformation function of the residual block parameterized as μ\mu at the Level kk. Given a batch of training images xx^{}, we feed them to AANets to compute the feature maps at the kk-th level (through the stable and plastic blocks respectively) as follows,

The transferability (of the knowledge learned from old classes) is different at different levels of neural networks . Therefore, it makes more sense to apply different aggregation weights for different levels of residual blocks. Let αϕ[k]\alpha^{[k]}_{\phi} and αη[k]\alpha^{[k]}_{\eta} denote the aggregation weights of the stable and plastic blocks, respectively, at the kk-th level. Then, the weighted sum of xϕ[k]x_{\phi}^{[k]} and xη[k]x_{\eta}^{[k]} can be derived as follows,

In our illustrative example in Figure 2, there are three pairs of weights to learn at each phase. Hence, it becomes increasingly challenging to choose these weights manually if multiple phases are involved. In this paper, we propose an learning strategy to automatically adapt these weights, i.e., optimizing the weights for different blocks in different phases, see details in Section 3.2.

2 Optimization Steps

In each incremental phase, we optimize two groups of learnable parameters in AANets: (a) the neuron-level scaling weights ϕ\phi for the stable blocks and the convolutional weights η\eta on the plastic blocks; (b) the feature aggregation weights α\alpha. The former is for network parameters and the latter is for hyperparameters. In this paper, we formulate the overall optimization process as a bilevel optimization program (BOP) .

The Formulation of BOP. In AANets, the network parameters [ϕ,η][\phi,\eta] are trained using the aggregation weights α\alpha as hyperparameters. In turn, α\alpha can be updated when temporarily fixing network parameters [ϕ,η][\phi,\eta]. In this way, the optimality of [ϕ,η][\phi,\eta] imposes a constraint on α\alpha and vise versa. Ideally, in the ii-th phase, the CIL system aims to learn the optimal αi\alpha_{i} and [ϕi,ηi][\phi_{i},\eta_{i}] that minimize the classification loss on all training samples seen so far, i.e., Di∪D0:i−1\mathcal{D}_{i}\cup\mathcal{D}_{0:i-1}, so the ideal BOP can be formulated as,

Data Strategy. To solve Problem 4, we need to use D0:i−1\mathcal{D}_{0:i-1}. However, in the setting of CIL , we cannot access D0:i−1\mathcal{D}_{0:i-1} but only a small set of exemplars E0:i−1\mathcal{E}_{0:i-1}, e.g., 2020 samples of each old class. Directly replacing D0:i−1∪Di\mathcal{D}_{0:i-1}\cup\mathcal{D}_{i} with E0:i−1∪Di\mathcal{E}_{0:i-1}\cup\mathcal{D}_{i} in Problem 4 will lead to the forgetting problem for the old classes. To alleviate this issue, we propose a new data strategy in which we use different training data splits to learn different groups of parameters: 1) in the upper-level problem, αi\alpha_{i} is used to balance the stable and the plastic blocks, so we use the balanced subset to update it, i.e., learning αi\alpha_{i} on E0:i−1∪Ei\mathcal{E}_{0:i-1}\cup\mathcal{E}_{i} adaptively; 2) in the lower-level problem, [ϕi,ηi][\phi_{i},\eta_{i}] are the network parameters used for feature extraction, so we leverage all the available data to train them, i.e., base-training [ϕi,ηi][\phi_{i},\eta_{i}] on E0:i−1∪Di\mathcal{E}_{0:i-1}\cup\mathcal{D}_{i}. Based on these, we can reformulate the ideal BOP in Problem 4 as a solvable BOP as follows,

where Problem 5a is the upper-level problem and Problem 5b is the lower-level problem we are going to solve.

Then, we use a balanced exemplar set to solve the upper-level problem, i.e., training αi{\alpha}_{i} as follows,

where γ1\gamma_{1} and γ2\gamma_{2} are the lower-level and upper-level learning rates, respectively.

3 Algorithm

In Algorithm 1, we summarize the overall training steps of the proposed AANets in the ii-th incremental learning phase (where i∈[1,...,N]i\in[1,...,N]). Lines 1-4 show the preprocessing including loading new data and old exemplars (Line 1), initializing the two groups of learnable parameters (Lines 2-3), and selecting the exemplars for new classes (Line 4). Lines 5-12 optimize alternatively between the network parameters and the Adaptive Aggregation weights. In specific, Lines 6-8 and Lines 9-11 execute the training for solving the upper-level and lower-level problems, respectively. Lines 13-14 update the exemplars and save them to the memory.

Experiments

We evaluate the proposed AANets on three CIL benchmarks, i.e., CIFAR-100 , ImageNet-Subset and ImageNet . We incorporate AANets into four baseline methods and boost their model performances consistently for all settings. Below we describe the datasets and implementation details (Section 4.1), followed by the results and analyses (Section 4.2) which include a detailed ablation study, extensive comparisons to related methods, and some visualization of the results.

Datasets. We conduct CIL experiments on two datasets, CIFAR-100 and ImageNet , following closely related work . CIFAR-100 contains 60,00060,000 samples of 32×3232\times 32 color images for 100100 classes. There are 500500 training and 100100 test samples for each class. ImageNet contains around 1.31.3 million samples of 224×224224\times 224 color images for 10001000 classes. There are approximately 1,3001,300 training and 5050 test samples for each class. ImageNet is used in two CIL settings: one based on a subset of 100100 classes (ImageNet-Subset) and the other based on the full set of 1,0001,000 classes. The 100100-class data for ImageNet-Subset are sampled from ImageNet in the same way as .

Architectures. Following the exact settings in , we deploy a 3232-layer ResNet as the baseline architecture (based on which we build the AANets) for CIFAR-100. This ResNet consists of 11 initial convolution layer and 33 residual blocks (in a single branch). Each block has 1010 convolution layers with 3×33\times 3 kernels. The number of filters starts from 1616 and is doubled every next block. After these 33 blocks, there is an average-pooling layer to compress the output feature maps to a feature embedding. To build AANets, we convert these 33 blocks into three levels of blocks and each level consists of a stable block and a plastic block, referring to Section 3.1. Similarly, we build AANets for ImageNet benchmarks but taking an 1818-layer ResNet as the baseline architecture . Please note that there is no architecture change applied to the classifiers, i.e., using the same FC layers as in .

Hyperparameters and Configuration. The learning rates γ1\gamma_{1} and γ2\gamma_{2} are initialized as 0.10.1 and 1×10−81\times 10^{-8}, respectively. We impose a constraint on each pair of αη\alpha_{\eta} and αϕ\alpha_{\phi} to make sure αη+αϕ=1\alpha_{\eta}+\alpha_{\phi}=1. For fair comparison, our training hyperparamters are almost the same as in . Specifically, on the CIFAR-100 (ImageNet), we train the model for 160160 (9090) epochs in each phase, and the learning rates are divided by 1010 after 8080 (3030) and then after 120120 (6060) epochs. We use an SGD optimizer with the momentum 0.90.9 and the batch size 128128 to train the models in all settings.

Benchmark Protocol. We follow the common protocol used in . Given a dataset, the model is trained on half of the classes in the -th phase. Then, it learns the remaining classes evenly in the subsequent NN phases. For NN, there are three options as 55, 1010, and 2525, and the corresponding settings are called “NN-phase”. In each phase, the model is evaluated on the test data for all seen classes. The average accuracy (over all phases) is reported. For each setting, we run the experiment three times and report averages and 95%95\% confidence intervals.

2 Results and Analyses

Table 1 summarizes the statistics and results in 88 ablative settings. Table 2 presents the results of 44 state-of-the-art methods w/ and w/o AANets as a plug-in architecture, and the reported results from some other comparable work. Figure 3 compares the activation maps (by Grad-CAM ) produced by different types of residual blocks and for the classes seen in different phases. Figure 4 shows the changes of values of αη\alpha_{\eta} and αϕ\alpha_{\phi} across 1010 incremental phases.

Comparing to the State-of-the-Art. Table 2 shows that taking our AANets as a plug-in architecture for 44 state-of-the-art methods consistently improves their model performances. E.g., for CIFAR-100, LUCIR w/ AANets and Mnemonics w/ AANets respectively gains 4.9%4.9\% and 3.3%3.3\% improvements on average. From Table 2, we can see that our approach of using AANets achieves top performances in all settings. Interestingly, we find that AANets can boost more performance for simpler baseline methods, e.g., iCaRL. iCaRL w/ AANets achieves mostly better results than those of LUCIR on three datasets, even though the latter method deploys various regularization techniques.

Visualizing Activation Maps. Figure 3 demonstrates the activation maps visualized by Grad-CAM for the final model (obtained after 55 phases) on ImageNet-Subset (NN=5). The visualized samples from left to right are picked from the classes coming in the -th, 33-rd and 55-th phases, respectively. For the -th phase samples, the model makes the prediction according to foreground regions (right) detected by the stable block and background regions (wrong) by the plastic block. This is because, through multiple phases of full updates, the plastic block forgets the knowledge of these old samples while the stable block successfully retains it. This situation is reversed when using that model to recognize the 55-th phase samples. The reason is that the stable block is far less learnable than the plastic block, and may fail to adapt to new data. For all shown samples, the model extracts features as informative as possible in two blocks. Then, it aggregates these features using the weights adapted from the balanced dataset, and thus can make a good balance of the features to achieve the best prediction.

Aggregation Weights (αη\alpha_{\eta} and αϕ\alpha_{\phi}). Figure 4 shows the values of αη\alpha_{\eta} and αϕ\alpha_{\phi} learned during training 1010-phase models. Each row displays three plots for three residual levels of AANets respectively. Comparing among columns, we can see that Level 1 tends to get larger values of αϕ\alpha_{\phi}, while Level 3 tends to get larger values of αη\alpha_{\eta}. This can be interpreted as lower-level residual blocks learn to stay stabler which is intuitively correct in deep models. With respect to the learning activity of CIL models, it is to continuously transfer the learned knowledge to subsequent phases. The features at different resolutions (levels in our case) have different transferabilities . Level 1 encodes low-level features that are more stable and shareable among all classes. Level 3 nears the classifiers, and tends to be more plastic such as to fast to adapt to new classes.

Conclusions

We introduce a novel network architecture AANets specially for CIL. Our main contribution lies in addressing the issue of stability-plasticity dilemma in CIL by a simple modification on plain ResNets — applying two types of residual blocks to respectively and specifically learn stability and plasticity at each residual level, and then aggregating them as a final representation. To achieve efficient aggregation, we adapt the level-specific and phase-specific weights in an end-to-end manner. Our overall approach is generic and can be easily incorporated into existing CIL methods to boost their performance.

Acknowledgments. This research is supported by A*STAR under its AME YIRG Grant (Project No. A20E6c0101), the Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 1, Alibaba Innovative Research (AIR) programme, Major Scientific Research Project of Zhejiang Lab (No. 2019DB0ZX01), and Max Planck Institute for Informatics.

References

A Results for Different CIL Settings.

B Strict Memory Budget Experiments

C More Ablation Results

In Table S3, we supplement the ablation results obtained in more settings. “4×4\times” denotes that we use 44 same-type blocks at each residual level. Comparing Row 7 to Row 2 (Row 5) shows the efficiency of using different types of blocks for representing stability and plasticity.

D Additional Plots

In Figures S2, we present the phase-wise accuracies obtained on CIFAR-100, ImageNet-Subset and ImageNet, respectively. “Upper Bound” shows the results of joint training with all previous data accessible in every phase. We can observe that our method achieves the highest accuracies in almost every phase of different settings. In Figures S3 and S4, we supplement the plots for the values of αη\alpha_{\eta} and αϕ\alpha_{\phi} learned on the CIFAR-100 and ImageNet-Subset (NN=55, 2525). All curves are smoothed with a rate of 0.80.8 for a better visualization.

E More Visualization Results

Figure S1 below shows the activation maps of a “goldfinch” sample (seen in Phase 0) in different-phase models (ImageNet-Subset, NN=55). Notice that the plastic block gradually loses its attention on this sample (i.e., forgets it), while the stable block retains it. AANets benefit from its stable blocks.

F Source Code in PyTorch

We provide our PyTorch code on https://class-il.mpi-inf.mpg.de/. To run this repository, we kindly advise you to install Python 3.6 and PyTorch 1.2.0 with Anaconda.