Federated Class-Incremental Learning

Jiahua Dong, Lixu Wang, Zhen Fang, Gan Sun, Shichao Xu, Xiao Wang, Qi Zhu

Introduction

Federated learning (FL) enables multiple local clients to collaboratively learn a global model while providing secure privacy protection for local clients. It successfully addresses the data island challenge without completely compromising clients’ privacy . Recently, it has attracted significant interests in academia and achieved remarkable successes in various industrial applications, e.g., autonomous driving , wearable devices , medical diagnosis and mobile phones .

Generally, most existing FL methods are modeled in a static application scenario, where data classes of the overall FL framework are fixed and known in advance. However, real-world applications are often dynamic, where local clients receive the data of new classes in an online manner. To handle such a setting, existing FL methods typically require storing all training data of old classes at the local clients’ side so that a global model can be obtained via FL, however the high storage and computation overhead may render the FL unrealistic when new classes arrive dynamically . And if these methods are required to learn new classes continuously with very limited storage memory, they may suffer from significant performance degradation (i.e., catastrophic forgetting ) on old classes. Moreover, in real-world scenarios, new local clients that collect the data of new classes in a streaming manner may want to participate in the FL training, which could further exacerbate the catastrophic forgetting on old classes in the global model training.

To address these practical scenarios, we consider a challenging FL problem named Federated Class-Incremental Learning (FCIL) in this work. In the FCIL setting, each local client collects the training data continuously with its own preference, while new clients with unseen new classes could join in the FL training at any time. More specifically, the data distributions of the collected classes across the current and newly-added clients are non-independent and identically distributed (non-i.i.d.). FCIL requires these local clients to collaboratively train a global model to learn new classes continuously, with constraints on privacy preservation and limited memory storage . To better comprehend the FCIL problem, we here use COVID-19 diagnosis among different hospitals as a possible example . Imagine that before the pandemic, there could be hundreds of hospitals working collaboratively to train a global infectious disease diagnosis model via FL. Due to the sudden emergence of COVID-19, these hospitals will collect a large amount of new data related to COVID-19 and add them into the FL training as new classes. Moreover, new hospitals whose main focus is not infectious diseases may join the fight against COVID-19, where they have little data of the old infectious diseases, and all hospitals should learn to diagnose the old diseases and new COVID-19 variants. In such scenarios, most existing FL methods will likely suffer from catastrophic forgetting on old diseases diagnosis under the sudden emergence of new COVID-19 variants data.

An intuitive way to address new classes (e.g., learning new COVID-19 variants) continuously in the FCIL setting is to simply integrate FL and class-incremental learning (CIL) together. However, such strategy needs the central server to know when and where the data of new classes arrives (privacy-sensitive information), which violates the requirement of privacy preservation in FL. In addition, although local clients can utilize conventional CIL to address their local catastrophic forgetting, the non-i.i.d. class imbalance across clients may still cause heterogeneous forgetting on different clients, and this simple integration strategy could further exacerbate local catastrophic forgetting due to the heterogeneous global catastrophic forgetting on old classes across clients.

To tackle these challenges in FCIL, we propose a novel Global-Local Forgetting Compensation (GLFC) model in this paper, which effectively addresses local catastrophic forgetting occurred on local clients and global catastrophic forgetting across clients. Specifically, we design a class-aware gradient compensation loss to alleviate the local forgetting brought by the class imbalance at the local clients via balancing the forgetting of different old classes, and propose a class-semantic relation distillation loss to distill consistent inter-class relations across different incremental tasks. To overcome the global catastrophic forgetting caused by the non-i.i.d. class imbalance across clients, we design a proxy server to select the best old global model for the class-semantic relation distillation at the local side. Considering the privacy preservation, the proxy server collects perturbed prototype samples of new classes from local clients via a prototype gradient-based communication mechanism, and then utilizes them to monitor the performance of the global model for selecting the best one. Our model achieves 4.4%∼\sim15.1% improvement in terms of average accuracy on several benchmark datasets, when compared with a variety of baseline methods. The major contributions of this paper are summarized as follows:

We address a practical FL problem, namely Federated Class-Incremental Learning (FCIL), in which the main challenges are to alleviate the catastrophic forgetting on old classes brought by the class imbalance at the local clients and the non-i.i.d class imbalance across clients.

We develop a novel Global-Local Forgetting Compensation (GLFC) model to tackle the FCIL problem, alleviating both local and global catastrophic forgetting. To our best knowledge, this is the first attempt to learn a global class-incremental model in the FL settings.

We design a class-aware gradient compensation loss and a class-semantic relation distillation loss to address local forgetting, by balancing the forgetting of old classes and capturing consistent inter-class relations across tasks.

We design a proxy server to select the best old model for class-semantic relation distillation on the local clients to compensate global forgetting, and we protect the communication between this proxy server and clients with a prototype gradient-based mechanism for privacy.

Related Work

Federated Learning (FL) is a decentralized learning framework that can train a global model by aggregating local model parameters . To collaboratively learn a global model, proposes to aggregate local models via a weight-based mechanism. introduces a proximal term to help local model approximate the global ones. focuses on minimizing the model discrepancies across clients via an improved EWC. Moreover, designs a layer-wise aggregation strategy to reduce computation overhead . sacrifices the local optimality for rapid convergence, while aim to improve the performance of local models. integrates unsupervised domain adaptation into federated learning framework . However, these existing FL methods cannot effectively learn new classes continuously, due to the limited memory to store old classes at the local clients’ side.

Class-Incremental Learning (CIL) aims to learn new classes continuously while tackling forgetting on old classes . Without access to the data of old classes, designs new regulators for balancing the biased model optimization caused by new classes, and use the knowledge distillation to surmount catastrophic forgetting. introduce generative adversarial networks to produce synthetic data of old classes. As claimed in , the class imbalance between old and new classes is a crucial challenge for exemplar replay methods. design a self-adaptive network to balance biased predictions. uses the causal effect on knowledge distillation to rectify class imbalance. introduces geodesic path to traditional knowledge distillation. combines task-wise knowledge distillation and separated softmax for bias compensation. These CIL methods however cannot be applied to tackle our FCIL problem, due to their strong assumptions on when and where the data of new classes arrive.

Problem Definition

In the standard class-incremental learning , there are a sequence of streaming tasks T={Tt}t=1T\mathcal{T}=\{\mathcal{T}^{t}\}_{t=1}^{T}, where TT denotes the task number, and the tt-th task Tt={xit,yit}i=1Nt\mathcal{T}^{t}=\{\mathbf{x}_{i}^{t},\mathbf{y}_{i}^{t}\}_{i=1}^{N^{t}} consists of NtN^{t} pairs of samples xit\mathbf{x}_{i}^{t} and their one-hot encoding labels yit∈Yt\mathbf{y}_{i}^{t}\in\mathcal{Y}^{t}. Yt\mathcal{Y}^{t} represents the label space of the tt-th task including CtC^{t} new classes that are different from Cp=∑i=1t−1Ci⊂∪j=1t−1YjC^{p}=\sum_{i=1}^{t-1}C^{i}\subset\cup_{j=1}^{t-1}\mathcal{Y}^{j} old classes in previous t−1t-1 tasks. Inspired by , we construct an exemplar memory M\mathcal{M} to select ∣M∣Cp\frac{|\mathcal{M}|}{C^{p}} exemplars for each old class in the tt-th incremental task, and it satisfies ∣M∣Cp≪NtCt\frac{|\mathcal{M}|}{C^{p}}\ll\frac{N^{t}}{C^{t}}.

We then extend conventional class-incremental learning to Federated Class-Incremental Learning (FCIL). Given KK local clients {Sl}l=1K\{\mathcal{S}_{l}\}_{l=1}^{K} and a global central server SG\mathcal{S}_{G}, for the rr-th global round (r=1,⋯ ,Rr=1,\cdots,R), a set of local clients are randomly selected to participate in the gradient aggregation. Specifically, once the ll-th client Sl\mathcal{S}_{l} is selected at each global round for the tt-th incremental task, it will receive the latest global model Θr,t\Theta^{r,t}, and train Θr,t\Theta^{r,t} on its privately accessible tt-th incremental task Tlt∪Ml∼Pl∣Tlt∣+∣Ml∣\mathcal{T}_{l}^{t}\cup\mathcal{M}_{l}\sim\mathcal{P}_{l}^{|\mathcal{T}_{l}^{t}|+|\mathcal{M}_{l}|}, where Tlt={xlit,ylit}i=1Nlt⊂Tt\mathcal{T}_{l}^{t}=\{\mathbf{x}_{li}^{t},\mathbf{y}_{li}^{t}\}_{i=1}^{N_{l}^{t}}\subset\mathcal{T}^{t} is the training data of new classes, Ml\mathcal{M}_{l} denotes its exemplar memory, and Pl\mathcal{P}_{l} is the class distribution of the ll-th client. {Pl}l=1K\{\mathcal{P}_{l}\}_{l=1}^{K} are non-independent and identically distributed (i.e., non-i.i.d.) from each other. At the tt-th incremental task, the label space Ylt⊂Yt\mathcal{Y}_{l}^{t}\subset\mathcal{Y}^{t} of the ll-th local client is a subset of Yt=∪l=1KYlt\mathcal{Y}^{t}=\cup_{l=1}^{K}\mathcal{Y}_{l}^{t}, and it includes CltC_{l}^{t} new classes (Clt≤CtC_{l}^{t}\leq C^{t}), different from Clp=∑i=1t−1Cli⊂∪j=1t−1YljC_{l}^{p}=\sum_{i=1}^{t-1}C_{l}^{i}\subset\cup_{j=1}^{t-1}\mathcal{Y}_{l}^{j} old classes. After loading Θr,t\Theta^{r,t} and conducting the local training at the tt-th incremental task, SlS_{l} can get a locally updated model Θlr,t\Theta_{l}^{r,t}. All locally updated models of selected clients are then uploaded to the global server SGS_{G} to be aggregated as the global model Θr ⁣+ ⁣1,t\Theta^{r\!+\!1,t} of next round. The global server SG\mathcal{S}_{G} then distributes parameters Θr ⁣+ ⁣1,t\Theta^{r\!+\!1,t} to local clients for the next global round.

In the FCIL setting, we divide local clients {Sl}l=1K\{\mathcal{S}_{l}\}_{l=1}^{K} into three categories (i.e., {Sl}l=1K=So∪Sb∪Sn\mathcal{S}_{l}\}_{l=1}^{K}={\mathbf{S}_{o}}\cup\mathbf{S}_{b}\cup\mathbf{S}_{n}) in each incremental task. Specifically, So\mathbf{S}_{o} consists of KoK_{o} local clients that cannot receive the new data of current task but have the exemplar memory stored via previous learned tasks; Sb\mathbf{S}_{b} includes KbK_{b} clients collecting the new data of current task and the exemplar memory of previous tasks; and Sn\mathbf{S}_{n} consists of KnK_{n} newly-added clients that receive the new data of current task, but have no any exemplar memory of old classes. These clients are dynamically changing as the incremental tasks arriving. Namely, we randomly determine {So,Sb,Sn}\{\mathbf{S}_{o},\mathbf{S}_{b},\mathbf{S}_{n}\} at each global round, and Sn\mathbf{S}_{n} are irregularly added at any global round in the FCIL. It causes the gradual increase of K=Ko+Kb+KnK=K_{o}+K_{b}+K_{n} in streaming tasks.

Moreover, we have no any prior knowledge about the number of streaming tasks TT, data distributions {Pl}l=1K\{\mathcal{P}_{l}\}_{l=1}^{K}, when to collect new classes or add new local clients. The goal of FCIL is to effectively train a global model ΘR,T\Theta^{R,T} to learn new classes consecutively while alleviating the catastrophic forgetting on old classes with the requirement of privacy preservation, via communicating the local model parameters with the global central server SG\mathcal{S}_{G}.

The Proposed GLFC Model

The overview of our model is depicted in Figure 1. To address the FCIL requirements, our model solves local forgetting via a class-aware gradient compensation loss and a class-semantic relation distillation loss (Section 4.1), while tackling global forgetting via a proxy server to select the best old model for local clients (Section 4.2).

As aforementioned, the class imbalance between old and new classes (Tlt\mathcal{T}_{l}^{t} and Ml\mathcal{M}_{l}) at the local side enforces the local training to suffer from significant performance degradation (i.e., local catastrophic forgetting) on old classes. To prevent local forgetting, as shown in Figure 1, we develop a class-aware gradient compensation loss and a class-semantic relation distillation loss for local clients, which can correct imbalanced gradient propagation and ensure inter-class semantic consistency across incremental tasks.

∙\bullet Class-Aware Gradient Compensation Loss: After SG\mathcal{S}_{G} distributes Θr,t\Theta^{r,t} to local clients, the class-imbalanced distributions at local side cause imbalanced gradient back-propagation of the last output layer in Θr,t\Theta^{r,t}. It forces the update of local model Θlr,t\Theta_{l}^{r,t} to perform different learning paces within new classes and different forgetting paces within old classes after local training. This phenomenon heavily worsens the local forgetting on old classes, when new streaming data becomes part of old classes continuously.

where Plt(xlit,Θlr,t)ylitP_{l}^{t}(\mathbf{x}_{li}^{t},\Theta_{l}^{r,t})_{y_{li}^{t}} is the ylity_{li}^{t}-th softmax probability of the ii-th sample.

To this end, we propose a task transition detection mechanism to accurately identify when local clients receive new classes. Specifically, at the rr-th global round, every client computes the average entropy Hlr,t\mathcal{H}_{l}^{r,t} via the received global model Θr,t\Theta^{r,t} on its current training data Tlt\mathcal{T}_{l}^{t}:

where I(⋅)=∑ipilog⁡pi\mathcal{I}(\cdot)=\sum_{i}p_{i}\log p_{i} is the entropy function. When Hlr,t\mathcal{H}_{l}^{r,t} encounters a sudden rise and satisfies Hlr,t−Hlr−1,t≥rh\mathcal{H}_{l}^{r,t}-\mathcal{H}_{l}^{r-1,t}\geq r_{h}, we argue that the local clients are receiving new classes, and update tt by t←t+1t\leftarrow t+1. Then they can update memory Ml\mathcal{M}_{l} and store old model Θlt−1\Theta_{l}^{t-1}. We empirically set rh=1.2r_{h}=1.2.

2 Global Catastrophic Forgetting Compensation

where Pt(xˉnt,Γ)P^{t}(\bar{\mathbf{x}}_{n}^{t},\Gamma) is the probability predicted via Γ\Gamma. η\eta denotes the learning rate to update xˉnt\bar{\mathbf{x}}_{n}^{t}.

∙\bullet Perturbed Prototype Samples Construction: Although the network Γ\Gamma is only privately accessible to SP\mathcal{S}_{P} and local clients, malicious attackers may steal Γ\Gamma and these gradients to reconstruct raw prototype sample {xlc∗t,ylc∗t}∈Tlt\{\mathbf{x}_{lc^{*}}^{t},\mathbf{y}_{lc^{*}}^{t}\}\in\mathcal{T}_{l}^{t} of the ll-th local client. To achieve privacy preservation, we propose to add perturbations to these prototype samples. The attackers can get little useful information from the perturbed prototype samples even if they can reconstruct them.

To be specific, given a prototype sample {xlc∗t,ylc∗t}∈Tlt\{\mathbf{x}_{lc^{*}}^{t},\mathbf{y}_{lc^{*}}^{t}\}\in\mathcal{T}_{l}^{t}, we forward it into the local model Θlr,t\Theta_{l}^{r,t} that has been trained via Eq. (6), and apply back-propagation to update this sample. In order to produce perturbed prototype sample, we introduce a Gaussian noise into the latent feature of prototype sample, and then update xlc∗t\mathbf{x}_{lc^{*}}^{t} via Eq. (11):

where Φ(xlc∗t)\Phi(\mathbf{x}_{lc^{*}}^{t}) denotes the latent feature of xlc∗t\mathbf{x}_{lc^{*}}^{t}, and Plt(Φ(xlc∗t)+γN(0,σ2),Θlr,t)P_{l}^{t}(\Phi(\mathbf{x}_{lc^{*}}^{t})+\gamma\mathcal{N}(0,\sigma^{2}),\Theta_{l}^{r,t}) is the probability predicted via Θlr,t\Theta_{l}^{r,t} when adding Gaussian noise N(0,σ2)\mathcal{N}(0,\sigma^{2}) to Φ(xlc∗t)\Phi(\mathbf{x}_{lc^{*}}^{t}). σ2\sigma^{2} represents the variance of features of all samples belonging to ylc∗t\mathbf{y}_{lc^{*}}^{t}, and we empirically set γ=0.1\gamma=0.1 to control the effect of Gaussian noise in this paper. Some reconstructed prototype samples are visualized in Figure 2.

3 Optimization Pipeline of Our GLFC Model

Starting from the first incremental task, all clients are required to compute the average entropy of their private training data via Eq. (7) at the beginning of each global round, and follow iCaRL to update their exemplar memory Ml\mathcal{M}_{l}. For each global training round, the central server SG\mathcal{S}_{G} randomly selects a set of local clients to conduct local training. After that, when the selected clients identify new classes via the task transition detection strategy, they will construct perturbed prototype samples of these new classes and share the corresponding gradients to the proxy server SP\mathcal{S}_{P} via the prototype gradient-based communication mechanism. After receiving these gradients, SP\mathcal{S}_{P} reconstructs these prototype samples, and utilizes them to select the best global model Θt\Theta^{t} until collecting gradients next time. Starting from the second task (t ⁣= ⁣2t\!=\!2), SP\mathcal{S}_{P} will distribute best models of the last and current task (i.e., Θt−1\Theta^{t-1}, and Θt\Theta^{t}) to selected clients. Then the ll-th client uses Θt−1\Theta^{t-1} as its Θlt−1\Theta_{l}^{t-1} to update the current local model Θlr,t\Theta_{l}^{r,t} via optimizing Eq. (6), when it doesn’t detect new classes via task transition detection. Otherwise, it uses Θt\Theta^{t} to train the current local model Θlr,t\Theta_{l}^{r,t}. Finally, SG\mathcal{S}_{G} aggregates the updated local models Θlr,t\Theta_{l}^{r,t} to get the global model Θr+1,t\Theta^{r+1,t} of next ground. The optimization pipeline is provided in supplementary material.

Experiments

We use three datasets: CIFAR-100 , ImageNet-Subset , and TinyImageNet in our experiments. For a fair comparison with baseline class-incremental learning methods in the FCIL setting, we follow the same protocols proposed by to set incremental tasks, utilize the identical class order generated from iCaRL , and employ the same backbone (i.e., ResNet-18 ) as the classification model . The SGD optimizer whose learning rate is 2.0 is used to train all models. The exemplar memory Ml\mathcal{M}_{l} of each client is set as 2,000 for all streaming tasks. A shallow LeNet with only 4 layers is used as the network Γ\Gamma. We employ a SGD optimizer with the learning rate as 0.1 to construct perturbed samples for local clients, while utilizing a L-BFGS optimizer with the learning rate as 1.0 to reconstruct prototype samples for proxy server. We initialize the number of local clients as 30 in the first incremental task, introduce 10 additional new local clients as the learning tasks arrive consecutively. At each global round, we randomly select 10 clients to conduct 20-epoch local training. Each client randomly receives 60% classes from the label space of its seen task. We run our experiments for 3 times with 3 random seeds (2021, 2022, 2023) and report the averaged results. Please refer to more details in supplementary material.

2 Performance Comparison

This section shows comparison experiments to illustrate the effectiveness of our GLFC model, as shown in Tables 1, 2, 3, where ‘△\triangle’ denotes the improvements of our model compared with other comparison methods. We observe that our model outperforms existing class-incremental methods in the FCIL setting by a margin of 4.4%∼\sim15.1% in terms of average accuracy. It validates that our model could enable local clients to collaboratively train a global class-incremental model. Moreover, our model has stable performance improvement in comparison with other methods for all incremental tasks, which verifies the effectiveness of our model to address the forgetting in FCIL.

3 Ablation Studies

4 Qualitative Analysis of Incremental Tasks

In this section, we conduct qualitative analysis of various incremental tasks (T ⁣= ⁣5,10,20T\!=\!5,10,20) on benchmark datasets to validate the superior performance of GLFC, as shown in Figures 3, 4. According to these curves, we can easily observe that our model performs better than other competing baseline methods for all incremental tasks, in the settings with different number of tasks (T ⁣= ⁣5,10,20T\!=\!5,10,20). It demonstrates that the GLFC model enables multiple local clients to learn new classes in a streaming manner while addressing local and global forgetting.

5 Qualitative Analysis of Exemplar Memory

As shown in Table 4, we study the effects of different exemplar memories on the performance of our GLFC model, by respectively setting Ml\mathcal{M}_{l} as {500,1000,1500,2000}\{500,1000,1500,2000\} on CIFAR-100 . From the results in Table 4, we can conclude that the better performance of GLFC model on learning new classes in a streaming manner will be, the larger size of exemplar memory Ml\mathcal{M}_{l} is. It validates that storing more training data of old classes at the local side could strengthen the memory ability of our proposed GLFC model for old classes. Moreover, the presented results also illustrate the effectiveness of our GLFC model about identifying new classes via the task transition detection mechanism and updating the exemplar memory.

Conclusion

In this paper, we propose a practical Federated Class-Incremental Learning (FCIL) problem, and develop a novel Global-Local Forgetting Compensation (GLFC) model to address the local and global catastrophic forgetting in FCIL. Specifically, a class-aware gradient compensation loss and a class-semantic relation distillation loss are designed to locally address the catastrophic forgetting, by correcting the imbalanced gradient propagation and ensuring consistent inter-class relations across tasks. We also employ a proxy server to tackle global forgetting by selecting the best old model for preserving the memory of old classes. Extensive experiments on representative benchmark datasets demonstrate the effectiveness of our proposed GLFC model.

Acknowledgement

This work was partially supported by National Nature Science Foundation of China under Grant 62003336; National Science Foundation of US under Grants 1834701, 2016240; and research awards from Facebook, Google, PlatON Network, and General Motors.

References

Appendix A Implementation Details

CIFAR-100 is composed of 60,000 color images from 100 classes. Each class has 500 training and 100 evaluation samples with the size 32×3232\times 32.

ImageNet-Subset is a subset of ImageNet , and it includes 100 classes sampled in the same way as . We split each class into 500 training and 100 test samples with the size 224×224224\times 224.

TinyImageNet includes 100,000 samples for 200 classes, and each sample is downsized to 64×6464\times 64. Each class has 500 training and 50 test samples.

A.2 Experimental Settings

For the federated learning setting, in the first task, there are total 30 local clients, and for each global round, we randomly select 10 clients to conduct 20-epoch local training. After the local training, these clients will share their updated models to participate in the global aggregation of this round. When the number of streaming tasks is T=10T=10, for CIFAR-100 and ImageNet-Subset, each task includes 10 new classes for 10 global rounds, and each task transition will introduce 10 additional new clients. For TinyImageNet, each task includes 20 new classes for the same 10 global rounds, and each task transition also includes 10 new clients. Therefore, in the last task of T=10T=10, the total number of local clients is 120, and we also randomly select 10 clients to perform global aggregation. For the case of T=20T=20, each task has 5 classes with 10 global rounds of training for CIFAR-100 and ImageNet-Subset, and the number of classes will be 10 for TinyImageNet. Note that the number of newly introduced clients is 5 now at each task transition. We also conduct experiments of T=5T=5, CIFAR-100 and ImageNet-Subset contain 20 classes for each task, while TinyImageNet has 40 classes per task. For all three datasets, each task covers 20 global rounds and there will be 20 new clients joining in the framework at each task transition. For building the non-i.i.d. setting, every client can only own 60% classes of the label space in the current task, and these classes are randomly selected. During the task transition global round, we assume 90%90\% existing clients are from Sb\mathcal{S}_{b}, while the resting 10%10\% clients are from So\mathcal{S}_{o}.

For fair comparisons with other class-incremental learning methods in the FCIL setting, we follow the same protocols proposed by to split classes into incremental tasks, utilize the identical class order generated from iCaRL . Moreover, we use ResNet-18 as the classification model. As for the gradient encoding network, we use a shallow LeNet with only 4 layers. We use a SGD optimizer whose initial learning rate is 2.0 to train all classification models, and the learning rate is divided by 5, 25 and 125 when the accumulated local epochs of a task hit 100, 150 and 180, respectively. What’s more, the optimizer of back-propagation perturbation generation is SGD with learning rate as 0.1, and the sample reconstruction optimization happened on the proxy server utilizes L-BFGS with learning rate as 1.0 for saving memory storage. The batch size is 128, and the exemplar memory Ml\mathcal{M}_{l} of each client has the sample size of 2,000 during all streaming tasks. As for the optimization iterations, the prototype perturbation generation has 100 iterations, and prototype sample reconstruction conducts 200 iterations for each gradient. We repeatedly run our experiments for three times with three random seeds (2021, 2022, 2023) and report the average results in our comparison experiments.

A.3 Comparison Methods

This paper is the first exploration to address the federated class-incremental learning (FCIL) problem, and there is not any baseline method that is built on similar settings. Therefore, for fair comparisons, we compare our GLFC model with several state-of-the-art class-incremental methods (i.e., iCaRL , BiC , PODNet , DDE , GeoDL and SS-IL under the federated learning (FL) settings, to validate the effectiveness of our proposed GLFC model. Besides, top-1 accuracy metric is employed to evaluate the performance of other comparison methods and our proposed GLFC model.

Appendix B Optimization Pipeline of Our GLFC Model

Starting from the first incremental task, all clients are required to compute the average entropy of their private training data via Eq. (7) at the beginning of each global round, and follow iCaRL to update their exemplar memory Ml\mathcal{M}_{l}. For each global training round, the central server SG\mathcal{S}_{G} randomly selects a set of local clients to conduct local training. After that, when the selected clients identify new classes via the task transition detection strategy, they will construct perturbed prototype samples of these new classes and share the corresponding gradients to the proxy server SP\mathcal{S}_{P} via the prototype gradient-based communication mechanism. After receiving these gradients, SP\mathcal{S}_{P} reconstructs these prototype samples, and utilizes them to select the best global model Θt\Theta^{t} until collecting gradients next time. Starting from the second task (t ⁣= ⁣2t\!=\!2), SP\mathcal{S}_{P} will distribute best models of the last and current task (i.e., Θt−1\Theta^{t-1}, and Θt\Theta^{t}) to selected clients. Then the ll-th client uses Θt−1\Theta^{t-1} as its Θlt−1\Theta_{l}^{t-1} to update the current local model Θlr,t\Theta_{l}^{r,t} via optimizing Eq. (6), when it doesn’t detect new classes via task transition detection. Otherwise, it can use Θt\Theta^{t} to train the current local model Θlr,t\Theta_{l}^{r,t}. Finally, SG\mathcal{S}_{G} aggregates the updated local models Θlr,t\Theta_{l}^{r,t} to get the global model Θr+1,t\Theta^{r+1,t} of next ground. The detailed optimization pipeline is provided in Algorithm 1.

Appendix C Experiments on TinyImageNet Dataset

As shown in Tables 5, 6, we present comparison experiments between our model and other baseline class-incremental learning methods on TinyImageNet dataset. The presented results show that our GLFC model significantly outperforms other state-of-the-art comparison methods by 4.7%∼\sim11.0% in terms of average accuracy. It illustrates the effectiveness of our model to address both local and global catastrophic forgetting in the FCIL setting. Moreover, the performance of our model is the best among all incremental tasks, and there is a large performance improvement for each incremental task. This phenomenon validates that the proposed proxy server is effective to address global catastrophic forgetting brought by non-i.i.d. class imbalance across clients via prototype sample construction mechanism. Meanwhile, the proposed class-aware gradient compensation loss and class-semantic relation distillation loss guarantee that our model could effectively alleviate local catastrophic forgetting at local client side.

C.2 Ablation Studies

This subsection investigates the effectiveness of different variants of our model on TinyImageNet dataset, as presented in Tables 5, 6. When compared with Ours, Ours-w/oCGC degrades the performance of 2.7%∼\sim2.8% in terms of average accuracy, which validates the effectiveness of the class-aware gradient compensation loss to compensate imbalanced gradient propagation. We observe that Ours performs better than Ours-w/oCRD by 10.1%∼\sim10.2% in terms of average accuracy. The class-semantic relation distillation loss ensures inter-class semantic consistency across different incremental tasks to address local catastrophic forgetting. Moreover, the performance of Ours-w/oPRS is worse than Ours by 3.2%∼\sim4.6% in terms of average accuracy. This performance degradation verifies that global catastrophic forgetting brought by non-i.i.d. class imbalance across clients could be effectively addressed via the proxy server. All proposed modules in our GLFC model could cooperate well to address the FCIL problem. When any one of proposed components is removed, as shown in Tables 5, 6, Ours-w/oCGC, Ours-w/oCRD and Ours-w/oPRS achieve significant performance degradation.

C.3 Effects of Incremental Tasks

As presented in Tables 7, 8, in this subsection, we introduce the qualitative analysis of various incremental tasks (T=5,10T=5,10) on TinyImageNet dataset to validate the effectiveness of the proposed GLFC model. From the results in Tables 7, 8, we observe that the performance of our proposed model has a large improvement (3.2%∼\sim10.0% in terms of average accuracy) over other state-of-the-art comparison methods for all incremental tasks. Even though there are different settings with different number of tasks (T=5,10T=5,10), our proposed GLFC model still has the best performance, which verifies that our model could effectively tackle both local and global catastrophic forgetting in the FCIL setting. Moreover, the significant performance improvement illustrates that our model enables multiple local clients to learn new classes consecutively, while addressing catastrophic forgetting on old learned classes under the privacy preservation and limited memory of local clients.

Appendix D Qualitative Analysis of Exemplar Memory

In this subsection, as shown in Table 9, we further conduct extensive experiments (T=10T=10) on CIFAR-100 dataset to investigate the effects of different exemplar memories on the performance of our proposed GLFC model when setting Ml\mathcal{M}_{l} as {500,1000,1500,2000}\{500,1000,1500,2000\}. From the presented results in Table 9, we easily observe that our model achieves the better performance for all incremental tasks, when local clients have large memory storage to store the exemplar samples of old classes. Moreover, storing more training data of old classes at the local side could promote the memory replay on old classes, which further addresses catastrophic forgetting at local clients’ side for old classes. Besides, it validates that our proposed model is efficient to distinguish new classes via the task transition detection strategy and update the corresponding exemplar memory Ml\mathcal{M}_{l} at local side. The updated exemplar memory plays an essential role in tackling local catastrophic forgetting on old classes.

Appendix E Limitation and Societal Impact

This section discusses the limitation for our proposed model and the potential societal impact of this paper.

This paper mainly focuses on addressing Federated Class-Incremental Learning (FCIL) problem from the algorithm perspective. In the future, it is necessary to develop mathematical theoretical supports for understanding the FCIL problem and the proposed GLFC model. A possible way to develop mathematical theories for our model is considering existing mathematical explanations of federated learning (FL) and class-incremental learning (CIL) simultaneously. However, as we know, there is rare theoretical analysis about CIL. Therefore, it might be tough to propose a brand-new theoretical support to analyze regular CIL problem. Instead, we will try to establish a theoretical analysis for the FCIL problem from a FL perspective in the future work.

E.2 Potential Societal Impact

The FCIL problem discussed in our paper doesn’t have any negative societal impact. On the contrary, we believe our work can solve real-world problems and bring about extensive benefits. In comparison with standard federated learning (FL), the proposed Federated Class-Incremental Learning (FCIL) is more practical as we assume the data of new classes as well as new clients will indiscriminately and continuously participate in FCIL. To solve the FCIL problem, our proposed GLFC model can enable a global class-incremental learning model to be trained on decentralized devices without data sharing (uploading decentralized data to a central server or data exchange between participated devices). Compared to regular class-incremental learning methods that always need access to the training data, FCIL can protect the private information of participants by remaining the local data where it is collected.

We have faith that the proposed GLFC model can bring beneficial gains to a number of information-sensitive scenarios, such as medical diagnosis, smartphone applications, pharmaceutical companies, and high-technology enterprises, etc. In summary, this work is the first attempt to learn a global class-incremental model in the setting of FL, which expedites the development of FL-based applications with the requirement of privacy preservation.