Class-incremental Learning via Deep Model Consolidation
Junting Zhang, Jie Zhang, Shalini Ghosh, Dawei Li, Serafettin Tasci, Larry Heck, Heming Zhang, C. -C. Jay Kuo
Introduction
Despite the recent success of deep learning in computer vision for a broad range of tasks , classical training paradigm of deep models is ill-equipped for incremental learning (IL). Most deep neural networks can only be trained when the complete dataset is given and all classes are known prior to training. However, the real world is dynamic and new categories of interest can emerge over time. Re-training a model from scratch whenever a new class is encountered can be prohibitively expensive due to training data storage requirements and the computational cost of full retrain. Directly fine-tuning the existing model on only the data of new classes using stochastic gradient descent (SGD) optimization is not a better approach either, as this might lead to the notorious catastrophic forgetting problem , which can result in severe performance degradation on old tasks.
We consider a realistic, albeit strict and challenging, setting of class-incremental learning, where the system must satisfy the following constraints: 1) the original training data for old classes are no longer accessible when learning new classes — this could be due to a variety of reasons, e.g., legacy data may be unrecorded, proprietary, too large to store, or subject to privacy constraint when training the model for a new task; this is a practical concern in various academic and industrial applications, where the model can be transferred from one party to another but data should be kept private, and a practitioner wants to augment the model to learn new classes; 2) the system should provide a competitive multi-class classifier for the classes observed so far, i.e. single-headed classification should be supported, which does not require any prior information of the test data; 3) the model size should remain relatively unchanged after learning new classes.
Several attempts have been made to enable IL for DNNs, but none of them satisfies all of these constraints. Some recent works that rely on the storage of partial old data have made impressive progress. They are arguably not memory efficient and storing data for the life time involves violate some practical constraints such as copyright or privacy issues, which is common in the domains like bio-informatics . The performance of the existing methods that do not store any past data is yet unsatisfactory. Some of these methods rely on incrementally training generative models , which is a harder problem to solve; while others fine-tune the old model on the new data with certain regularization techniques to prevent forgetting . We argue that the ineffectiveness of these regularization-based methods is mainly due to the asymmetric information between old classes and new classes in the fine-tuning step. New classes have explicit and strong supervisory signal from the available labeled data, whereas the information for old classes is implicitly given in the form of a noisy regularization term. Moreover, if we over-regularize the model, the model will fail to adapt to the new task, which is referred to as intransigence in the IL context. As a result, these methods have intrinsic bias towards either the old or the new classes in the final model, and it is extremely difficult to find a sweet spot considering that in practice we do not have a validation dataset for the old classes during incremental learning.
As depicted in Fig. 1, we propose a novel paradigm for class-incremental learning called deep model consolidation (DMC), which first trains a separate model for the new classes using labeled data, and then combines the new and old models using publicly available unlabeled auxiliary data via a novel double distillation training objective. DMC eliminates the intrinsic bias caused by the information asymmetry or over-regularization in the training, as the proposed double distillation objective allows the final student model to learn from two teacher models (the old and new models) simultaneously. DMC overcomes the difficulty introduced by loss of access to legacy data by leveraging unlabeled auxiliary data, where the abundant transferable representations are mined to facilitate IL. Furthermore, using the auxiliary data rather than the training data of the new classes ensures the student model absorbs the knowledge transferred from the both teacher models in an unbiased way.
Crucially, we do not require the auxiliary data share the class labels or generative distribution of the target data. The only requirement is that they are generic, diversified, and generally related to the target data. Usage of such unlabeled data incurs no additional dataset construction and maintenance cost since they can be crawled from the web effortlessly when needed and discarded once the IL of new classes is complete. Furthermore, note that the symmetric role of the two teacher models in DMC has a valuable extra benefit in the generalization of our method; it can be directly applied to combine any two arbitrary pre-trained models that can be downloaded from the Internet for easy deployment (i.e., only one model needs to be deployed instead of two), without access to the original training data.
To summarize, our main contributions include:
A novel paradigm for incremental learning which exploits external unlabeled data, which can be obtained at negligible cost. This is an illuminating perspective for IL, which bypasses the constraint of having old data stored by finding some cheap substitute that does not need to be stored.
A new training objective function to combine two deep models into one single compact model to promote symmetric knowledge transfer. The two models can have different architectures, and they can be trained on data of distinct set of classes.
An approach to extend the proposed paradigm to incrementally train modern one-stage object detectors, to which the existing methods are not applicable.
Extensive experiments that demonstrate the substantial performance improvement of our method over existing approaches on large-scale image classification and object detection benchmarks in the IL setting.
Related work
McCloskey et al. first identified the catastrophic forgetting effect in the connectionist models, where the memory about the old data is overwritten when retraining a neural network with new data. Recently, researchers have been actively developing methods to alleviate this effect.
Regularization methods. Regularization methods enforce additional constraints on the weight update, so that the new concepts are learned while retaining the prior memories. Goodfellow et al. found that dropout could reduce forgetting for multi-layer perceptrons sometimes. One line of work constrains the network parameters that are important to the old tasks to stay close to their old values, while looking for a solution to a new task in the neighborhood of the old one. EWC and its variants use Fisher information matrix to estimate the weight importance; MAS uses the gradients of the network output; SI uses the path integral over the optimization trajectory instead. RWalk combines EWC and SI . Information about the old task and new task is not symmetric during learning in these methods; besides, the network may become stiffer and stiffer to adapt to the new task as it learns more tasks over time. Li and Hoiem pursued another direction by proposing the Learning without Forgetting (LwF) method, which finetunes the network using the images of new classes with knowledge distillation loss, to encourage the output probabilities of old classes for each image to be close to the original network outputs. However, information asymmetry between old classes and new classes still exists. Image samples from new data may severely deviate from the true distribution of the old data, which further aggravates the information asymmetry. Instead, we assign two teacher models to one student network to guarantee the symmetric information flow from old- and new-class models into the final model. IMM first finetunes the network on the new task with regularization, and then blends the obtained model with the original model through moment matching. Though conceptually similar, our work is different from IMM in the following ways: 1) we do not use regularized-finetuning from old-class model when training the new model for the new classes, so we can avoid intrinsic bias towards the old classes and suboptimal solution for the new task; 2) we do not assume the final posterior distribution for all the tasks is Gaussian, which is a strong assumption for DNNs.
Dynamic network methods. Dynamic network methods dedicate a part of the network or a unique feed-forward pathway through neurons for each task. At test time, they require the task label to be specified to switch to the correct state of the network, which is not applicable in the class-IL where task labels are not available.
Rehearsal and pseudo-rehearsal methods. In rehearsal methods , past information is periodically replayed to the model to strengthen memories it has already learned, which is done by interleaving data from earlier sessions with the current session data . However, storage of past data is not resource efficient and may violate some practical constraints such as copyright or privacy issues. Pseudo-rehearsal methods attempt to alleviate this issue by using generative models to generate pseudopatterns that are combined with the current samples. However, this requires training a generative model in the class-incremental fashion, which is an even harder problem to solve. Existing such methods do not produce competitive results unless supported by real exemplars .
Incremental learning of object detectors. Shmelkov et al. adapted LwF for the object detection task. However, their framework can only be applied to object detectors in which proposals are computed externally, e.g., Fast R-CNN . In our experiments, we show that our method is applicable to more efficient modern single-shot object detection architectures, e.g., RetinaNet .
Exploiting external data. In computer vision, the idea of employing external data to improve performance of a target task has been explored in many contexts. Inductive transfer learning aims to transfer and reuse knowledge in labeled out-of-domain instances. Semi-supervised learning attempts to exploit the usefulness of unlabeled in-domain instances. Our work shares a similar spirit with self-taught learning , where we use unlabeled auxiliary data but do not require the auxiliary data to have the same class labels or generative distribution as the target data. Such unlabeled data is significantly easier to obtain compared to typical semi-supervised or transfer learning settings.
Method
We perform IL in two steps: the first step is to train a -class classifier using training data , which we refer as the new model ; the second step is to consolidate the old model and the new model.
The new class learning step is a regular supervised learning problem and it can be solved by standard back-propagation. The model consolidation step is the major contribution of our work, where we propose a method called Deep Model Consolidation (DMC) for image classification which we further extend to another classical computer vision task, object detection.
We start by training a new CNN model on new classes using the available training data with standard softmax cross-entropy loss. Once the new model is trained, we have two CNN models specialized in classifying either the old classes or the new classes. After that, the goal of the consolidation is to have a single compact model that can perform the tasks of both the old model and the new model simultaneously. Ideally, we have the following objective:
where denotes the index of the classification score associated with -th class, and denotes the joint distribution from which samples of class are drawn. We want the output of the consolidated model to approximate the combination of the network outputs of the old model and the new model. To achieve this, the network response of the old model and the new model is employed as supervisory signals in joint training of the consolidated model.
Due the absence of the legacy data, we cannot consolidate the two models using the old data. Thus some auxiliary data has to be used. If we assume that natural images lie on an ideal low-dimensional manifold, we can approximate the distribution of our target data via sampling from readily available unlabeled data from a similar domain. Note that the auxiliary data do not have to be stored persistently; they can be crawled and fed in mini-batches on-the-fly in this stage, and discarded thereafter.
Specifically, the training objective for consolidation is
where denotes the unlabeled auxiliary training data, and the double distillation loss is defined as:
in which is the logit produced by the consolidated model for the -th class, and
where is the concatenation of and .
The regression target is the concatenation of normalized logits of the two specialist models. We normalize by subtracting its mean over the class dimension (Eq. 4). This serves as a step of bias calibration for the two set of classes. It unifies the scale of logits produced by the two models, but retains the relative magnitude among the classes, so that the symmetric information flow can be enforced.
Notably, to avoid the intrinsic bias toward either old or new classes, should not be initialized from or ; we should also avoid the usage of training data for the new classes in the model consolidation stage.
2 DMC for object detection
We extend the IL approach given in Section 3.1 for modern one-stage object detectors, which are nearly as accurate as two-stage detectors but run much faster than the later ones. A single-stage object detector divides the input image into a fixed-resolution 2D grid (the resolution of the grid can be multi-level), where higher resolution means that the area corresponding to the image region (i.e., receptive field) of each cell in the grid is smaller. There are a set of bounding-box templates with fixed sizes and aspect ratios, called anchor boxes, which are associated with each spatial cell in the grid. Anchor boxes serve as references for the subsequent prediction. The class label and the bounding box location offset relative to the anchor boxes are predicted by the classification subnet and bounding boxes regression subnet, respectively, which are shared across all the feature pyramid levels .
In order to apply DMC to incrementally train an object detector, we have to consolidate the classification subnet and bounding boxes regression subnet simultaneously. Similar to the image classification task, we instantiate a new detector whenever we have training data for new object classes. After the new detector is properly trained, we then use the outputs of the two specialist models to supervise the training of the final model.
Anchor boxes selection. In one-stage object detectors, a huge number of anchor boxes have to be used to achieve decent performance. For example, in RetinaNet , 100k anchor boxes are used for an image of resolution . Selecting a smaller number of anchor boxes speeds up forward-backward pass in training significantly. The naive approach of randomly sampling some anchor boxes doesn’t consider the fact that the ratio of positive anchor boxes and negative ones is highly imbalanced, and negative boxes that correspond to background carry little information for knowledge distillation. In order to efficiently and effectively distill the knowledge of the two teacher detectors in the DMC stage, we propose a novel anchor boxes selection method to selectively enforce the constraint for a small set of anchor boxes. For each image sampled from the auxiliary data, we first rank the anchor boxes by the objectness scores. The objectness score () for an anchor box is defined as:
where are classification probabilities produced by the old-class model, and are from the new-class model. Intuitively, a high objectness score for a box implies a higher probability of containing a foreground object. The predicted classification probabilities of the old classes are produced by the old model, and new classes by the new model. We use the subset of anchor boxes that have the highest objectness scores and ignore the others.
DMC for classification subnet. Similar to the image classification case in Sec. 3.1, for each selected anchor box, we calculate the double distillation loss between the logits produced by the classification subnet of the consolidated model and the normalized logits generated by the two existing specialist models . The loss term of DMC for the classification subnet is identical to Eq. 3.
DMC for bounding box regression subnet. The output of the bounding box regression subnet is a tuple of spatial offsets , which specifies a scale-invariant translation and log-space height/width shift relative to an anchor box. For each anchor box selected, we need to set its regression target properly. If the class that has the highest predicted class probability is one of the old classes, we choose the old model’s output as the regression target, otherwise, the new model’s output is chosen. In this way, we encourage the predicted bounding box of the consolidated model to be closer to the predicted bounding box of the most probable object category. Smooth loss is used to measure the closeness of the parameterized bounding box locations. The loss term of DMC for the bounding box regression subnet is as follows:
Overall training objective. The overall DMC objective function for the object detection is defined as
where is a hyper-parameter to balance the two loss terms.
Experiments
There are two evaluation protocols for incremental learning. In one setting, the network has different classification layers (multiple “heads”) for each task, where each head can differentiate the classes learned only in this task; it relies on an oracle to decide on the task at test time, which would result in a misleading high test accuracy . In this paper, we adopt a practical yet challenging setting, namely “single-head” evaluation, where the output space consists of all the classes learned so far, and the model has to learn to resolve the confusion among the classes from different tasks, when task identities are not available at test time.
2 Incremental learning of image classifiers
We evaluate our method on iCIFAR-100 benchmark as done in iCaRL , which uses CIFAR-100 data and learn all 100 classes in groups of or classes at a time. The evaluation metric is the standard top-1 multi-class classification accuracy on the test set. For each experiment, we run this benchmark 5 times with different class orderings and then report the averages and standard deviations of the results. We use ImageNet3232 dataset as the source of auxiliary data in the model consolidation stage. The images are down-sampled versions of images from ImageNet ILSVRC training set. We exclude the images that belong to the CIFAR-100 classes, which results in 1,082,340 images. Following iCaRL , we use a 32-layer ResNet for all experiments and the model weights are randomly initialized.
2.2 Experimental results and discussions
We compare our method against the state-of-the-art exemplar-free incremental learning methods EWC++ , LwF , SI , MAS , RWalk and some baselines with . Finetuning denotes the case where we directly fine-tune the model trained on the old classes with the labeled images of new classes, without any special treatment for catastrophic forgetting. Fixed Representation denotes the approach where we freeze the network weights except for the classification layer (the last fully connected layer) after the first group of classes has been learned, and we freeze the classification weight vector after the corresponding classes have been learned, and only fine-tune the classification weight vectors of new classes using the new data. This approach usually underfits for the new classes due to the limited degree of freedom and incompatible feature representations of the frozen base network. Oracle denotes the upper bound results via joint training with all the training data of the classes learned so far.
1𝑥log(1+x) for better visibility. Fig. 4(b), 4(c) and 4(d) are from . (Best viewed in color.) The results are shown in Fig. 2. Our method outperforms all the methods by a significant margin across all the settings consistently. We used the official codehttps://github.com/facebookresearch/agem for to get the results for EWC++ , SI , MAS and RWalk . We found they are highly sensitive to the hyperparameter that controls the strength of regularization due to the asymmetric information between old classes and new classes, so we tune the hyperparameter using a held-out validation set for each setting separately, and report the best result for each case. The results of LwF are from iCaRL and they are the second-best in all the settings.
It can be also observed that DMC demonstrates a stable performance across different , in contrast to other regularization-based methods, where the disadvantages of inherent asymmetric information flow reveal more, as we incrementally learn more sessions. They struggle in finding the good trade-off between forgetting and intransigence.
Fig. 3 illustrates how the accuracy on the first group of classes changes as we learn more and more classes over time. While the previous methods all suffer from catastrophic forgetting on the first task, DMC shows considerably more gentle slop of the forgetting curve. Though the standard deviations seems high, which is due to the random class ordering in each run, the relative standard deviations (RSD) are at a reasonable scale for all methods.
We visualize the confusion matrices of some of the methods in Fig. 4. Finetuning forgets old classes and makes predictions based only on the last learned group. Fixed Representation is strongly inclined to predict the classes learned in the first group, on which its feature representation is optimized. The previous best performing method LwF does a better job, but still has many more non-zero entries on the recently learned classes, which shows strong evidence of information asymmetric between old classes and new classes. On the contrary, the proposed DMC shows a more homogeneous confusion matrix pattern and thus has visibly less intrinsic bias towards or against the classes that it encounters early or late during learning.
Impact of the distribution of auxiliary data. Fig. 5 shows our empirical study on the impact of the distribution of the auxiliary data by using images from datasets of handwritten digits (MNIST ), house number digits (SVHN ), texture (DTD ), and scenes (Places365 ) as the sources of the auxiliary data. Intuitively, the more diversified and more similar to the target data the auxiliary data is, the better performance we can achieve. Experiments show that usage of overparticular datasets like MNIST and SVHN fails to produce competitive results, but using generic and easily accessible datasets like DTD and Places365 can already outperform the previous state-of-the-art methods. In the applied scenario, one may use the prior knowledge about the target data to obtain the desired auxiliary data from a related domain to boost the performance.
Choices of loss function. We compare some common distance metrics used in knowledge distillation in Table 1 . We observe DMC is generally not sensitive to the loss function chosen, while loss and KD loss with performs slightly better than others. As stated in , both formulations should be equivalent in the limit of a high temperature , so we use loss throughout this paper for its simplicity and stability over various training schedules.
Effect of the amount of auxiliary data. Fig. 6 illustrates the effect of the amount of auxiliary data used in consolidation stage. We randomly subsampled images for from ImageNet3232 . We report the average of the classification accuracies over all steps of the IL (as in , the accuracy of the first group is not considered in this average). Overall, our method is robust against the reduction of auxiliary data to a large extent. We can outperform the previous state-of-the-art by just using 8,000, 16,000 and 32,000 unlabeled images ( of full auxiliary data) for , respectively. Note that it also takes less training time for the consolidated model to converge when we use less auxiliary data.
Experiments with larger images. We additionally evaluate our method on CUB-200 dataset in IL setting with . The network architecture (VGG-16 ) and data preprocessing are identical with REWC . We use BirdSnap as the auxiliary data source where we excluded the CUB categories. As shown in Table 2, DMC outperforms the previous state-of-the-art by a considerable margin. This demonstrates that DMC generalizes well to various image resolutions and domains.
3 Incremental learning of object detectors
Following , we evaluate DMC for incremental object detection on PASCAL VOC 2007 in the IL setting: there are 20 object categories in the dataset, and we incrementally learn classes and classes. The evaluation metric is the standard mean average precision (mAP) on the test set. We use training images from Microsoft COCO dataset as the source of auxiliary data for the model consolidation stage. Out of 80 object categories in the COCO dataset, we use the 98,495 images that contain objects from the 60 non-PASCAL categories.
We perform all experiments using RetinaNet , but the proposed method is applicable to other one-stage detectors with minor modifications. In the experiment, we use ResNet-50 as the backbone network for both 10-class models and the final consolidated 20-class model. In experiment, we use ResNet-50 as the backbone network for the 19-class model as well as the final consolidated 20-class model, and ResNet-34 for the 1-class new model. In all experiments, the backbone networks were pretrained on ImageNet dataset .
3.2 Experimental results and discussions
We compare our method with a baseline method and with the state-of-the-art IL method for object detection by Shmelkov et al. . In the baseline method, denoted by Inference twice, we directly run inference for each test image using two specialist models separately and then aggregate the predictions by taking the class that has the highest classification probability among all classes, and use the bounding box prediction of the associated model. The method proposed by Shmelkov et al. is compatible only with object detectors that use pre-computed class-agnostic object proposals (e.g., Fast R-CNN ), so we adapt their method for RetinaNet by using our novel anchor boxes selection scheme to determine where to apply the distillation, denoted by Adapted Shmelkov et al. .
Learning classes. The results are given in Table 3. Compared to Inference twice, our method is more time- and space-efficient since Inference twice scales badly with respect to the number of IL sessions, as we need to store all the individual models and run inference using each one at test time. The accuracy gain of our method over the Inference twice method might seem surprising, but we believe this can be attributed to the better representations that were inductively learned with the help of the unlabeled auxiliary data, which is exploited also by many semi-supervised learning algorithms. Compared to Adapted Shmelkov et al. , our method exhibits remarkable performance improvement in detecting all classes.
Learning classes. The results are given in Table 10. We observe an mAP pattern similar to the experiment. Adapted Shmelkov et al. suffers from degraded accuracy on old classes. Moreover, it cannot achieve good AP on the “tvmonitor” class. Heavily regularized on 19 old classes, the model may have difficulty learning a single new class with insufficient training data. Our DMC achieves state-of-the-art mAP of all the classes learned, with only half of the model complexity and inference time of Inference twice. We also performed the addition of one class experiment with each of the VOC categories being the new class. The behavior for each class is very similar to the “tvmonitor” case described above. The mAP varies from 64.88% (for new class “aeroplane”) to 71.47% (for new class “person”) with mean 68.47% and standard deviation of 1.75%. Detailed results are in the supplemental material.
Impact of the distribution of auxiliary data. The auxiliary data selection strategy that was described in Sec. 4.3.1 would potentially include images that contain objects from target categories. To see the effect of data distribution, we also experimented with a more strict data in which we exclude all the MS COCO images that contain any object instance of 20 PASCAL categories, denoted by DMC- exclusive aux. data in Table 3 and 10. This setting can be considered as the lower bound of our method regarding the distribution of auxiliary data. We see that even in such a strict setting, our method outperforms the previous state-of-the-art . This study also implies that our method can benefit from auxiliary data from a similar domain.
Consolidating models with different base networks. As mentioned in Sec. 4.3.1, originally we used different base network architectures for the two specialist models in classes experiment. As shown in Table 5, we also compare the case when using ResNet-50 backbone for both the 19-class model and the 1-class model. We observed that ResNet-50 backbone does not work as well as ResNet-34 backbone, which could result from overfitting of the deeper model to the training data of the new class and thus it fails to produce meaningful distillation targets in the model consolidation stage. However, since our method is architecture-independent, it offers the flexibility to use any network architecture that fits best to the current training data.
Conclusion
In this paper, we present a novel class-incremental learning paradigm called DMC. With the help of a novel double distillation training objective, DMC does not require storage of any legacy data; it exploits readily available unlabeled auxiliary data to consolidate two independently trained models instead. DMC outperforms existing non-exemplar-based methods for incremental learning on large-scale image classification and object detection benchmarks by a significant margin. DMC is independent of network architectures and thus it is applicable in many tasks.
Future directions worth exploring include: theoretically characterize how the “similarity” between the unlabeled auxiliary data and target data affects the IL performance; 2) continue the study on using of exemplars of old data with DMC (presented in supp. material), in terms of exemplar selection scheme and rehearsal strategies; 3) generalize DMC to consolidate multiple models at one time; 4) extend DMC to other applications where consolidation of deep models is beneficial, e.g., taking ensemble of models trained with the same or partially overlapped sets of classes.
Acknowledgments. This work was started during internship at Samsung Research America and later continued at USC. We also acknowledge the support of NVIDIA Corporation with the donation of a Titan X Pascal GPU.
References
Appendix Overview
In this supplemental document, we provide additional detailed experimental results and analyses of the proposed method, Deep Model Consolidation (DMC), for class-incremental learning.
Appendix A Detailed experimental results of DMC for object detection
In the experiments of DMC for incremental learning of object detectors, we incrementally learn classes using RetinaNet . In the main paper, we presented the results of adding “tvmonitor” class as the new class. Here, we show the results of addition of one class experiment with each of the VOC categories being the new class in Table 10, where Old Model denotes the 19-class detector trained on the old 19 classes, New Model denotes the 1-class detector trained on the new class and DMC denotes the final consolidated model that is capable of detecting all the 20 classes. Per-class average precisions on the entire test set of PASCAL VOC 2007 are reported.
Appendix B Effect of the amount of auxiliary data for object detection
We studied the effect of the amount of auxiliary data for DMC for image classification task. To see how the amount of auxiliary data affects the final performance in the incremental learning of object detection, we performed additional experiments on PASCAL VOC 2007 with the classes setting. We randomly sampled , and of the full auxiliary data from Microsoft COCO dataset for consolidation. As shown in Table 6, with just of full data, i.e., 12.3k images, DMC can still outperform the state-of-the-art, which demonstrates its robustness and efficiency in the detection task as well.
Appendix C Implementation and training details
Training details for the image classification experiments. Following iCaRL , we use a 32-layers ResNet for all experiments and the model weights are randomly initialized. When training the individual specialist models, we use SGD optimizer with momentum for 200 epochs. In the consolidation stage, we train the network for 50 epochs. The learning rate schedule for all the experiments is same, i.e., it starts with 0.1 and reduced by 0.1 at 7/10 and 9/10 of all epochs. For all experiments, we train the network using mini-batches of size 128 and a weight decay factor of and momentum of 0.9. We apply the simple data augmentation for training: 4 pixels are padded on each side, and a crop is randomly sampled from the padded image or its horizontal flip.
Training details for the object detection experiments. We resize each image so that the smaller side has 640 pixels, keeping the aspect ratio unchanged. We train each model for 100 epochs and use Adam optimizer with learning rate on two NVIDIA Tesla M40 GPUs simultaneously, with batch size of 12. Random horizontal flipping is used for data augmentation. Standard non-maximum suppression (NMS) with threshold 0.5 is applied for post-processing at test time to remove the duplicate predictions. For each image, we select 64 anchor boxes for DMC training. Empirically we found selecting more anchor boxes (128, 256 etc.) did not provide further performance gain. The is set to 1.0 for all experiments.
Hyperparameters used for the baseline methods. We report results of EWC++ , SI , MAS and RWalk on iCIFAR-100 benchmark in the main paper. Table 7 summarizes the hyperparamter that controls the strength of regularization used in each experiment, and they are picked based on a held-out validation set.
Appendix D Preliminary experiments of adding exemplars
While DMC is realistic in applied scenarios due to its scalablity and immunity to copyright and privacy issues, we additionally tested our method in the scenario where we are allowed to store some exemplars from the old data with a fixed budget when learning the new classes. Suppose we are incrementally learning a group of classes at a time, With the same total memory budget as in iCaRL , we fill the exemplar set by randomly sampling training images from each class when we learn the first group of classes; then every time we learn more classes with training data in the -th incremental learning session, we augment the exemplar set by randomly sampled training images of the new classes, and we fine-tune the consolidated model using these exemplars for 15 epochs with a small learning rate of . After fine-tuning, we reduce the size of the exemplar set by keeping exemplars for each class. We refer to this variant of our method as DMC+. We validate the effectiveness of DMC+ on the iCIFAR-100 benchmark, and Table 8 summarizes the results as the average of the classification accuracies over all steps of the incremental training (as in , the accuracy of the first group is not considered in this average). We can get comparable performance to iCaRL in all settings. Note that we also tried the herding algorithm to select the exemplars as in iCaRL, but we did not observe any notable improvement.
The confusion matrices comparison between DMC+ and iCaRL is shown in Fig. 7, and we find: 1) fine-tuning with exemplars can indeed further reduce the intrinsic bias in the training; 2) our DMC+ is on a par with iCaRL, even though we use naive random sampling rather than the more expensive herding approach to select exemplars.
These preliminary results demonstrate that DMC may also hold promise for exemplar-based incremental learning, and we would like to further study the potential improvement of DMC+, e.g. in terms of exemplar selection scheme and rehearsal strategies.
Appendix E Preliminary experiments of consolidating models with common classes
The original DMC assumes the two models to be consolidated are trained with distinct sets of classes, but it can be easily extended to the case where we have two models that are trained with partially overlapped set of classes. We first normalize the logits produced by the two models as Eq. 4 in the main paper. We then set the double distillation regression target as the follows: for the common classes, we take the mean of normalized logits from the two model; for each of the other classes, we take the normalized logit from the corresponding specialist model that was trained with this class.
Below we present a preliminary experiment on CIFAR-100 dataset in this setting, where we have separately trained two 55-class classifiers for Class 1-55 and Class 46-100, respectively, where 10 classes (Class 46-55) are in common. The results are shown in Table 9. For the common classes, DMC can be considered as an ensemble learning method, where at least the accuracy of the weaker model is maintained; for learning the rest of classes, it does not exhibit catastrophic forgetting or intransigence. This shows that DMC is promisingly extensible to the special case of incremental learning with partially overlapped categories.
Appendix F Enlarged plots
We provide enlarged plots of accuracy curves for iCIFAR-100 () in Fig. 8 for better visibility.