Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning

Lin Zhang, Li Shen, Liang Ding, Dacheng Tao, Ling-Yu Duan

Introduction

With the explosive growth of data and the strict privacy-protection policy, reckless data transmission and aggregation gradually become unacceptable due to the high bandwidth cost and risk of privacy leakage. Recently, Federated Learning (FL) has been proposed to replace the traditional heavily centralized learning paradigm and protect data privacy. It has been successfully applied in real-world tasks, such as smart city , health care , and recommender system , etc.

One of the main challenges in FL is the data heterogeneity, i.e., the data in clients are non-identically and independently distributed (Non-IID). It has been verified that the vanilla FL algorithm, FedAvg , leads to drifted local models and forgets the global knowledge catastrophically in this scenario, which further induces degraded performance and slow convergence . This is because the local model is updated merely with local data, i.e., minimizing the local empirical loss. However, minimizing the local empirical loss is fundamentally inconsistent with minimizing the global empirical loss in Non-IID FL.

To tackle the data heterogeneity challenge, most existing methods, e.g., FedProx , SCAFFOLD , FedDyn , MOON constrain the direction of local model update to align the local and global optimization objectives. Recently, FedGen learns a lightweight generator to generate pseudo feature and broadcasts it to clients to regulate local training. However, all these methods merely conduct simple model aggregation to get the global model in server, which ignores local knowledge incompatibility and induces knowledge forgetting in the global model. In addition, shows that directly aggregating models will largely degrade the performance while fine-tuning can greatly boost the accuracy. These motivate us to fine-tune the aggregated global model in the server with the knowledge in local models. On the other hand, merely aggregating local models in server ignores the server’s rich computing resources that could be potentially utilized to improve the performance of FL, such as the computing source in cross-silo FL .

Motivated by these observations, we propose a novel approach that boosts the performance of standard FL by on-the-fly fine-tuning the global model via data-free knowledge distillation (FedFTG), which simultaneously refines the model aggregation procedure and exploits the rich computing power of the sever. Concretely, FedFTG models the input space of local models through an auxiliary generator in the server, then generates pseudo data to transfer the knowledge in local models to the global model to improve the performance. To facilitate effective knowledge distillation throughout the training, FedFTG iteratively explores the hard samples in data distribution, which will induce prediction disagreement between local models and global model. Figure 1 compares FedFTG with FedAvg. FedFTG fine-tunes the global model with the hard samples to correct the model shift after model aggregation. The generator and global model are adversarially trained in a data-free manner, thus the whole procedure will not violate the privacy policy in FL. Considering the label distribution shift in data heterogeneity scenario, we further propose customized label sampling and class-level ensemble techniques, which explore the distribution correlation of clients and exploit maximum utilization of knowledge.

FedFTG is orthogonal to several existing local optimizers, such as FedAvg, FedProx, FedDyn, SCAFFOLD and MOON, as it only modifies the procedure of global model aggregation in the server. Consequently, FedFTG can be seamlessly embedded into these local FL optimizers, taking their advantages to further improve the performance of FedFTG. Extensive experiments on various settings verify that FedFTG achieves superior performance compared with state-of-the-art (SOTA) methods.

The main contributions of this work are four-fold:

We propose FedFTG to fine-tune the global model in server via data-free distillation, which simultaneously enhances the model aggregation step and utilizes the computing power of the server.

We develop hard sample mining to effectively transfer knowledge to global model. Besides, we propose customized label sampling and class-level ensemble to facilitate maximum utilization of knowledge.

We demonstrate that FedFTG is orthogonal to exiting local optimizers and can serve as a strong and versatile plugin to enhance the performance of FedAvg, FedProx, FedDyn, SCAFFOLD and MOON.

We verify the superiority of FedFTG against several SOTA methods for FL, including FedAvg, FedProx, FedDyn, SCAFFOLD, MOON, FedGen and FedDF, with extensive experiments on five benchmarks.

Related Work

There exist extensive works on improving the global performance of FL via client selection , split learning , domain adaptation , etc. The readers may refer to monographs and the reference therein to follow up its recent advances. Below, we mainly summarize the most relevant techniques to our work.

Federated Optimizer. The vanilla FL algorithm, i.e. FedAvg periodically aggregates the local models in server and updates the local model with its individual data. FedProx adds a proximal term to the local subproblem to restrict the local update closer to the initial (global) model. SCAFFOLD uses a variance reduction technique to correct the drifted local update. FedDyn modifies the objective of client with linear and quadratic penalty terms to align global and local objectives. In summary, all these methods focus on aligning the local and global model to narrow the distribution drift during the local training without enhancing the global model directly as in FedFTG.

Knowledge Distillation in Federated Learning. With the help of an unlabeled dataset, FedDF proposes an ensemble distillation for model fusion, trains the global model using the averaged logits from local models. FedAUX finds a model initialization for the local models, and weights the logits from local models using (ε,δ)(\varepsilon,\delta)-differentially private certainty scoring. FedBE generates a series of global models from Bayesian perspective using the local models, then summarizes these models into one global model by ensemble knowledge distillation. All these methods rely on an unlabeled auxiliary dataset in the server, while it is unclear to which extent should the auxiliary dataset be related to training data to guarantee effective knowledge distillation. Though FedDF maintains the auxiliary dataset can be replaced with a pretrained generator, it does not instantiate how to acquire the generator.

Data-Free Knowledge Distillation (DFKD). DFKD methods generate pseudo data from a pretrained teacher model, and use them to transfer knowledge of teacher model to another student model. The data is generated by maximizing the response of fake data on teacher model. DeepImpression models the output space of teacher model and recovers the real data by fitting the output space. DeepInversion further optimizes the pseudo data by regularizing the distribution of intermediate feature maps. DAFL and DFAD use a generator to generate data efficiently, where DAFL optimizes the generator by maximizing the response on prediction and feature level, and DFAD uses an adversarial training scheme to exploit the knowledge in teacher model effectively.

FedGen also learns a lightweight generator to ensemble knowledge of local models in a data-free manner, but uses the generator to regularize the local training. Besides, we design hard samples mining scheme, customized label sampling, and class-level ensemble to effectively transfer the knowledge from local models to global model in data heterogeneity scenario.

Methodology

In this section, we describe the proposed novel federated learning method: FedFTG. In each communication round, FedFTG randomly selects a set of clients and broadcasts the global model to them. Each client initializes the local model using the global model and trains it with a local optimizer. The server collects the local models and aggregates them as a preliminary global model. Instead of broadcasting the aggregated model back to each client directly, FedFTG fine-tunes this preliminary global model in server using the knowledge extracted from local models. Concretely, we develop a data-free knowledge distillation method with hard sample mining to effectively explore and transfer the knowledge to global model. Considering the label distribution shift in clients, we propose customized label sampling and class-level ensemble to facilitate more effective knowledge utilization. Figure 2 visualizes the training procedure on the server, and the corresponding algorithm is summarized in Algorithms 1&2. Note that FedFTG is orthogonal to efforts on optimizing local model training, such as SCAFFOLD, FedAvg, FedProx, and FedDyn.

Let ω\omega be the model parameter in the server and clients. In this work, we consider there exist KK clients, where Dk={(xk,i,yk,i)}i=1Nk\mathcal{D}_{k}=\{(x_{k,i},y_{k,i})\}_{i=1}^{N_{k}} is the dataset individually stored in kk-th client, NkN_{k} is the corresponding number of samples. Generally speaking, federated learning can be formulated as the following problem:

where L\mathcal{L} is the loss function to measure training error, and dataset Dk\mathcal{D}_{k} for each k∈{1,2,...,K}k\in\{1,2,...,K\} could be distributed heterogeneously. Due to the privacy protection constraint in FL, the server can not directly access local data of clients. To solve Eq. (1), for each communication round tt, existing methods send the global model ω\omega to a random set of clients StS_{t} and optimize it by min⁡ωfk(ω),k∈St\min_{\omega}f_{k}(\omega),k\in S_{t}. The server collects the local models {ωk}k∈St\{\omega_{k}\}_{k\in S_{t}} and aggregates them by averaging the gradients to update the global model ω\omega.

However, the local models are greatly drifted from each other in data heterogeneity scenario. Thus, traditional gradient averaging could lose the knowledge in local models, and the performance of updated global model is much lower than local models . To address this issue, we propose a data-free knowledge distillation method to fine-tune the global model, so that the global model can preserve the knowledge in local models and maintain their performance as much as possible. Concretely, the server maintains a conditional generator GG that generates pseudo data to capture the data distribution of clients as follows,

where θ\theta is the parameter of GG, z∼N(0,1)z\sim\mathcal{N}(\mathbf{0},\mathbf{1}) is a standard Gaussian noise, and yy is the class label of x~\widetilde{x} sampled from predefined distribution pt(y)p_{t}(y).

As shown in Figure 2, we then input the pseudo data x~\widetilde{x} to the global model to solve the following problem,

where Lmdk\mathcal{L}_{md}^{k} is the model discrepancy between global model ω\omega and local model ωk\omega_{k},

where DD is the classifier. σ\sigma is the softmax function, which will output the prediction score of x~\widetilde{x}. DKLD_{KL} denotes the Kullback-Leibler divergence. αtk,y\alpha_{t}^{k,y} controls the weight of knowledge from different local models during ensemble. By minimizing Lmd\mathcal{L}_{md}, we transfer the knowledge in local models to the global model. In Section 3.2 we will introduce how to acquire pt(y)p_{t}(y) and αtk,y\alpha_{t}^{k,y} to adapt label distribution shift in data heterogeneity scenario.

Data Fidelity and Diversity Constraints. To better extract knowledge from local models, the pseudo data x~\widetilde{x} should fit the input space of local models. Therefore, we use semantic loss Lcls\mathcal{L}_{cls} to train the generator GG, which will facilitate the fidelity of pseudo data,

where Lclsk\mathcal{L}_{cls}^{k} is the cross-entropy loss between the prediction of local model on pseudo data x~\widetilde{x} and the class label yy,

where LCE\mathcal{L}_{CE} is the cross-entropy loss. By minimizing Lcls\mathcal{L}_{cls}, x~\widetilde{x} is enforced to yield higher prediction on class yy, thus it fits the data distribution of class yy.

Simply using Lcls\mathcal{L}_{cls} will lead to model collapse of the generator: GG outputs the same data for every class. To address this issue, we use diversity loss Ldis\mathcal{L}_{dis} in to improve the diversity of the generated data,

where x~i\widetilde{x}_{i} is generated using ziz_{i}. By minimizing Ldis\mathcal{L}_{dis}, the pseudo data will be diverse and scattered in the data space.

Hard Sample Mining. Training the generator GG using Lcls\mathcal{L}_{cls} will generate pseudo data x~\widetilde{x} with low classification error, which means x~\widetilde{x} contains the most discriminative feature of class yy and is easy to be classified. However, these naive samples will not cause the prediction disagreement between global model and local models, i.e., Lmd=0\mathcal{L}_{md}=0, thus the global model is not optimized during training. As illustrated in Figure 3, the naive samples are already correctly classified by the global model. To effectively exploit the knowledge in local models and transfer them to the global model, we explore the hard samples in data distribution that cause prediction disagreement between local models and global model. Concretely, we adversarially train the generator and the global model with Lmd\mathcal{L}_{md}: (1) the generator is enforced to generate hard samples that maximize Lmd\mathcal{L}_{md}, and (2) the global model is trained to minimize Lmd\mathcal{L}_{md} using the hard samples. As a result, the global model can be gradually fine-tuned to fit the data distribution as in Figure 3.

To the end, the overall objective of FedFTG in the server is formulated as an adversarial learning scheme,

2 Adaptation to Label Distribution Shift for Effective Knowledge Distillation

In data heterogeneity scenario, label distributions are different among clients, i.e., pi(y)≠pj(y)p^{i}(y)\neq p^{j}(y) for different clients ii and jj. This indicates that: (1) the local dataset Dk\mathcal{D}_{k} of client is class-imbalanced, and the local model trained by Dk\mathcal{D}_{k} contains imbalanced data information; (2) for one class, the importance of knowledge are different among local models of clients. To facilitate more effective knowledge distillation, we propose customized label sampling and class-level ensemble to adapt these two problems respectively.

Customized Label Sampling. Typically, dataset in local client is class-imbalanced in data heterogeneity scenario, even having no data for some classes. It has been proved that deep neural networks tend to learn the majority classes and ignore the minority classes . Hence, the data information of minority classes in local models could be wrong and misleading, and the generated pseudo data are invalid to measure the model discrepancy. If uniformly sample the class label yy, these invalid data will influence the global model training and induce performance decrease. To mitigate this issue, we customize the sampling probability pt(y)p_{t}(y) according to the distribution of whole training data in each round, so that more pseudo data with effective information can be generated,

Class-Level Ensemble. The widely used ensemble method in knowledge distillation assigns the same weight to the knowledge from different teacher models , i.e., αtk,y=1∣St∣\alpha_{t}^{k,y}=\frac{1}{|S_{t}|} in Eq. (3) and Eq. (5). Due to the label distribution shift, for one class the importance of knowledge are different among local models. If assigning the same weight to clients, the important knowledge can not be figured out and utilized properly. Therefore, we propose class-level ensemble, which assigns the ensemble weight according to the individual data distribution of clients. Specifically, we compute the weight αtk,y\alpha_{t}^{k,y} via the data proportion of class yy in client kk against the total data in StS_{t},

As a result, the knowledge from clients can be flexibly integrated according to their importance on classes, so that FedFTG can facilitate maximum utilization of knowledge from local models.

Experiments

In this section, we empirically verify the effectiveness of FedFTGCode is available at https://github.com/ZhangLin-PKU/FedFTG.. We summarize the implementation details in Section 4.1, and compare FedFTG with several SOTA FL algorithms in Section 4.2. Ablation studies are conducted to verify the necessity of each component of FedFTG in Section 4.3. To further validate the effect of FedFTG on real-world FL applications, we evaluate the performance of FedFTG on three real-world datasets in Section 4.4.

Baselines. We compare FedFTG against FedAvg , FedProx , SCAFFOLD , FedDyn , MOON , FedGen and FedDF . Since FedDF does not explain how to obtain the generator, we train it in the same way as FedGen.

Datasets. CIFAR10 and CIFAR100 datasets with heterogeneous dataset partition are used to test the efficacy of FedFTG, which are two difficult tasks in FL scenario and are widely adopted in FL research. Similar to existing works , we use Dirichlet distribution Dir(β)\mathbf{Dir}(\beta) on label radios to simulate the non-iid data distribution among clients, where a smaller β\beta indicates higher data heterogeneity. During the implementation, we set β=0.3\beta=0.3 and β=0.6\beta=0.6.

Network Architecture. For both CIFAR10 and CIFAR100, we employ ResNet18 as the basic backbone. We borrow the generator network architecture from DFAD for FedFTG and FedDF. For FedGen, the network of generator is composed of two embedding layers (for inputs zz and yy, respectively) and two fully-connected (FC) layers with LeakyReLU and BatchNorm layers between them.

Hyperparameters. For all methods, we set the number of local training epoch E=5E=5, communication round T=1000T=1000, the client number K=100K=100 with the active fraction C=0.1C=0.1 (i.e., ∣St∣=10|S_{t}|=10). For local training, the batchsize is 5050 and the weight decay is 1e−31e-3. The learning rates for classifier and generator are initialized to be 0.10.1 and 0.010.01 respectively, and they are decayed quadratically with weight 0.9980.998. The dimension of zz is 100100 for CIFAR10 and 256256 for CIFAR100. II, IgI_{g}, IdI_{d} in Algorithm 2 are 1010, 11 and 55, respectively. If not specifically declared, we adopt λcls=1.0\lambda_{cls}=1.0 and λdis=1.0\lambda_{dis}=1.0, and adopt SCAFFOLD as the FL optimizer in FedFTG.

We further provide detailed implementations and extra experiment results in the supplementary material.

2 Performance Comparison

Test Accuracy. Table 1 reports the test accuracy of all compared algorithms on CIFAR10 and CIFAR100 datasets. We provide the performance of centralized learning in the first line. All experiments are repeated over 3 random seeds. In Table 1, FedFTG achieves the best performance in all scenarios, surpassing the second one (i.e., SCAFFOLD) by at least 1.5%1.5\%. FedDF also employs data-free knowledge distillation to improve the global model in server. It outperforms FedAvg and FedProx, and outperforms FedDyn and MOON in some cases, which further validates the superiority of the scheme “fine-tuning the global model in server”. However, it is worse than SCAFFOLD and FedFTG. FedGen yields lower accuracy compared with FedDF and FedFTG, and shows marginal performance gains than FedAvg in some cases. The performance of FedDF and FedGen further verifies the effectiveness of the proposed modules in FedFTG.

Communication Rounds. Table 2 evaluates different FL methods in term of the number of communication rounds to reach target test accuracy (75%75\% and 80%80\% for CIFAR10, 40%40\% and 50%50\% for CIFAR100, respectively). In Table 2, FedFTG achieves the second best and the best results on CIFAR10 and CIFAR100, respectively. Besides, FedFTG reduces the round number required by its FL optimizer (SCAFFOLD) in all scenarios. For CIFAR10, although FedDyn uses fewer rounds to achieve the target accuracy, its final accuracy is much worse than FedFTG, as displayed in Table 1. Below, we provide the results of using FedDyn as the optimizer of FedFTG, and the derived method FedDyn+FedFTG requires fewer rounds to reach target accuracy than FedDyn. Figure 4 displays the learning curve of different methods in 1000 communication rounds, where FedFTG achieves distinct performance gain after 1000 rounds. Though FedDyn has a faster increase rate in the beginning, the increase trend is gradually slowing down as training goes, and its accuracy is falling behind FedFTG after 150 and 50 rounds for CIFAR10 and CIFAR100 respectively.

Data heterogeneity and Partial Client Participant. Figure 5(a) displays the test accuracy on different β\beta values. In this figure, FedFTG achieves the best accuracy on all settings, which validates that FedFTG is effective in various data heterogeneity scenarios. Besides, FedFTG gains more accuracy improvement in extreme data heterogeneity scenario β=0.2\beta=0.2. In addition, as the degree of data heterogeneity decreases i.e., β\beta increases, the accuracy of each method is ascending. Figure 5(b) displays the test accuracy of FL methods with different fractions of active clients in each communication round. FedFTG also yields the best performance in this figure. Besides, the more clients involved in communication, the higher accuracy will be achieved.

Orthogonality of FedFTG with existing FL optimizers. Table 3 provides the performance of FedFTG using FedAvg, FedProx, FedDyn, SCAFFOLD and MOON optimizers. In Table 3, SCAFFOLD+FedFTG yields the best test accuracy among all the optimizers. FedDyn+FedFTG performs better than SCAFFOLD+FedFTG in terms of the round number to reach the target accuracy. This is consistent with the results in Table 2, where FedDyn requires fewer rounds than SCAFFOLD. Comparing Table 3 with Tables 1 and 2, we notice that for any FL optimizer, its performance can be largely boosted by using FedFTG. This validates the effectiveness and the orthogonality of FedFTG. Besides, simply using FedAVG+FedFTG as local optimizer already exceeds the other methods in Table 1 except SCAFFOLD.

3 Ablation Study

Necessity of each component in FedFTG. Table 4 displays the test accuracy of FedFTG after discarding some modules and losses, trained with 500 communication rounds on CIFAR10, β=0.3\beta=0.3. Here hsm, cls and abe represent the hard sampling mining, customized label sampling and class-level ensemble, respectively. We can see that removing any module leads to worse and unstable performance, i.e., lower accuracy and larger confidence interval. In addition, their joint absence can cause a further decrease on accuracy. On the other hand, a similar tendency is observed for the losses: the absence of single loss will lead to performance decrease, and removing multiple losses will enlarge the decrease. It should be noticed that, if replacing the KL divergence with Mean Average Square (Lmse\mathcal{L}_{mse}) to measure the model discrepancy, the model will collapse, which leads to severe performance degradation.

Robustness of FedFTG on hyperparameters. To measure the influence of hyperparameter selection, we select λcls\lambda_{cls} and λdis\lambda_{dis} from [0.5,0.75,1.0,1.25,1.5]\left[0.5,0.75,1.0,1.25,1.5\right] and select the dimension dd of noise data zz in [50,100,150,200,250]\left[50,100,150,200,250\right]. Figure 6 illustrates the test accuracy in term of the box plot, where FedFTG achieves similar performance among all the choices. Besides, the worst accuracy in Figure 6(a)-(b) is better than the best of previous works in Table 1. This indicates that FedFTG is not sensitive to the selection of hyperparameter in a large range.

Comments on feature-level pseudo data. FedGen advises that the data in feature space are more compact than in input space, thus it generates pseudo data in feature-level and fine-tunes the last few FC layers of federated model. Motivated by this, we compare the performance of FedFTG using feature-level generation (F) and input-level generation (I) in Table 5. Though FedFTG(F) still exceeds the other methods in Table 1, it suffers significant performance drop compared with FedFTG(I), which indicates the input-level generation is more effective for FedFTG. This is because FedFTG(F) only fine-tunes the last few layers of the global model, so the effect of knowledge transfer is limited.

4 Experiments on Real-World Datasets

In this section, we test the performance of FedFTG on more challenging real-world datasets - vehicle classification datasets MIO-TCD and CompCar , and large-scale image classification dataset Tiny-ImageNethttps://www.kaggle.com/c/tiny-imagenet. To better validate the effectiveness of FedFTG, we use the surveillance subset of CompCar, of which the images are collected by surveillance cameras. For MIO-TCD and Tiny-ImageNet, we assign the training data to 100 clients, while for CompCar the client number is 50. β\beta of Dirichlet distribution is 0.60.6 for all these datasets. The images of MIO-TCD and CompCar are resized to 112∗112112*112 before training, and we adopt a deeper generator for them. The communication round is 50, 100 and 1000 for MIO-TCD, CompCar and Tiny-ImageNet respectively. The other settings are the same as in Section 4.1. Experiment results are presented in Table 6.

From this table, we find that FedFTG consistently outperforms the other methods in all scenarios, which verifies the effectiveness of FedFTG in real-world FL applications. FedDF and FedGen also adopt data-free knowledge generation to improve the federated model. Though they yield higher performance than FedAvg and FedProx, FedFTG exceeds them by 1%∼6%1\%\sim 6\%. This further validates the effectiveness of the proposed modules in FedFTG.

Discussion

Privacy issue. Since FedFTG recovers the training data of clients in server, it may violate the privacy regulation in FL. However, according to our observation, the pseudo data only captures the high-level feature pattern of real data, which cannot be understood by human beings (see Figure 2). Besides, as the generator is trained by all local models, the pseudo data tend to show shared features of data in clients, which means the attribute of individual data will not be revealed. Uploading label statistics of data in clients may also leak privacy. One optional solution is adding noise to label statistics. According to our experiments, when the noise ratio is less than 10%, its influence on performance is less than 0.1% on CIFAR10, β=0.3\beta=0.3 setting.

Communication cost. Compared with other methods, FedFTG only need to additionally transmit the label statistics of training data (i.e., {ntk,y}k∈St,y∈[1,..,M]\{n_{t}^{k,y}\}_{k\in S_{t},y\in\left[1,..,M\right]}, MM is the class number), which induces negligibly extra transmission cost. If the training data keep the same during training, the label statistics can be reported to server before training, thus no extra transmission cost will be introduced.

Limitations. The main limitation of this work mainly exists in computation efficiency. As FedFTG additionally trains the global model apart from local training, it will make the whole training time longer than the other methods. In our experiment, FedFTG requires about double the time of FedAVG in each communication round. Besides, as the global model training is conducted in the server, FedFTG is more applicable to the cross-silo FL applications as defined in , where the server can be organizations that own sufficient computation source.

Conclusion

In this paper we propose a new data-free knowledge distillation method FedFTG to fine-tune the global model and to boost the performance of federated learning. A hard sample mining scheme is proposed to effectively explore the knowledge in local models and transfer it to global model. Facing the label distribution shift in data heterogeneity scenario, we propose customized label sampling and class-level ensemble to derive maximum utilization of knowledge. Extensive experiments on five benchmarks validate the effectiveness of the proposed FedFTG.

This work was supported by the National Natural Science Foundation of China under Grant 62088102, and in part by the PKU-NTU Joint Research Institute (JRI) sponsored by a donation from the Ng Teng Fong Charitable Foundation, and in part by Science and Technology Innovation 2030 –“Brain Science and Brain-like Research” Major Project (No. 2021ZD0201402 and No. 2021ZD0201405).

References

Supplementary

To explore the influence of long-tailed data on model performance, we train models using multiple subsets of CIFAR10, which have different degrees of imbalance. Here the subset is generated by Dirichlet distribution Dir(β)\mathbf{Dir}(\beta), where a smaller β\beta indicates more imbalanced data. The data number of each subset is 5000, and the architecture of the model is ResNet34 . The results are illustrated in Figure 7. Here, the curves in green and blue are the test accuracy on total test data and partitive test data respectively, where the distribution of partitive test data is the same as the distribution of training data. We can see there is a performance gap between two curves, and the gap becomes larger when the degree of imbalance is increased. This is because the model only learns the majority classes, and these classes also dominate the partitive test data, thus the model achieves high accuracy on partitive test data; whereas for the total test data that contains balanced data for every class, the model can not correctly predict the data of minority classes, thus the model yields lower test accuracy on total test data. The results in Figure 7 verifies that the model tends to learn majority data from imbalanced training data and ignore the minority classes. In the following, we term the model trained using long-tailed data as “the biased model”.

To further explore the influence of the biased model on pseudo data generation, we evaluate the accuracy of a biased model trained by a class-imbalanced CIFAR10 subset, and the quality of pseudo data generated via the biased model. The data quality is displayed in terms of the percentage of pseudo data that are correctly classified by a well-trained classifier, which is trained on all data of CIFAR10 and achieves 81.38% test accuracy. The results are illustrated in Figure 8. We can see that the model tends to learn majority classes and yields extremely low even zero accuracies for minority classes 7,10 and 9. Moreover, the quality of pseudo data is highly related to original data distribution. For the minority classes, the test accuracy of the pseudo data is less than 10%, i.e., the quality of the pseudo data is even worse than random noise. This indicates that the pseudo data generated via biased model could be invalid to conduct knowledge transfer, which motivates as to customize the sample probability of label during data generation to facilitate effective knowledge transfer.

2 Visualization of Data Heterogeneity

In Figure 9, we figure out the data distributions of clients that generated by Dirichlet distribution Dir(β)\mathbf{Dir}(\beta) with different β\beta as well as IID data distributions. For each β\beta value, we display the data distributions of 10 clients. In Figure 9, the data distributions of clients are significantly different when β\beta is small, and the client even has no data for some classes. When β\beta grows, the data is distributed more evenly in each client, and the discrepancy of data distributions among clients becomes smaller.

3 Detailed Hyperparameters

Here we introduce the setting of hyperparameters for baselines during experiments. For FedProx, the proximal regularization parameter μ\mu is 1e−41e-4. α\alpha in FedDyn is 1e−21e-2. We set the local update round in SCAFFOLD following , which is 5050 according to our experiment setting. Following , we set τ=0.5\tau=0.5, tune μ\mu from {0.1,1,5}\{0.1,1,5\} and report the best result. For FedGen and FedDF, the learning rate for the generator is the same as FedFTG, i.e., it is initialized as 0.010.01 and is decayed quadratically with weight 0.9980.998. As Resnet18 only has one fully-connected layer, ll in FedGen is L−1L-1, where LL is the total layer number.

4 Detailed Architecture of Generator

Table 7 lists the architectures of generators for FedFTG, FedDF and FedGen used in Section 4.1 ∼\sim Section 4.3. Here, dd is the dimension of noise data zz, and it is 100100 and 256256 for CIFAR10 and CIFAR100, respectively. MM is the class number of datasets, and it is 1010 and 100100 for CIFAR10 and CIFAR100 respectively. The inplace of LeakReLU is 0.20.2 here. Note that in Table 7(b) the output of generator is 512512-dimensional, as the input of the last FC layer in ResNet18 is 512512-dimensional. If using the other classifiers, the dimension of the generator’s output should be adjusted accordingly.

Table 8 lists the architectures of generators used in Section 4.4. Here d=256d=256 for all the datasets MIO-TCD, CompCar and Tiny-ImageNet. ss is the image size, and s=112,112,64s=112,112,64 for MIO-TCD, CompCar and Tiny-ImageNet respectively. Note that for the experiments of VGG11 in Table 3 in the main paper, we also adopt these two generators for FedDF, FedGen and FedFTG.

5 Supplementary Experiment Results

Table 9 illustrates the communication rounds of different methods to reach the target test accuracy (75% and 80% for CIFAR10, 40% and 50% for CIFAR100) when β=0.6\beta=0.6, which is a supplement to Table 2 in the main paper. Same as Table 2, FedFTG achieves the second best and the best convergence for CIFAR10 and CIFAR100 respectively. Besides, it greatly reduces the round numbers required by its FL optimizer SCAFFOLD.

Table 10 displays the test accuracy when adopting VGG11 and ResNet34 as the classifier. In this table, FedFTG yields the best performance in all scenarios, which validates the effectiveness of FedFTG on various architectures of deep neural network.