Delving into the Adversarial Robustness of Federated Learning

Jie Zhang, Bo Li, Chen Chen, Lingjuan Lyu, Shuang Wu, Shouhong Ding, Chao Wu

Introduction

Nowadays, end devices are generating massive amounts of potentially sensitive user data, raising practical concerns over security and privacy. Federated Learning (FL) (McMahan et al. 2017) emerges as a privacy-aware learning paradigm that allows multiple clients to collaboratively train neural networks without revealing their raw data. Recently, FL has attracted increasing attention from different areas, including medical image analysis (Liu et al. 2021a; Chen et al. 2021b), recommender systems (Liang, Pan, and Ming 2021; Liu et al. 2021b), natural language processing (Zhu et al. 2020; Wang et al. 2021), etc.

Prior studies have demonstrated that neural networks are vulnerable to evasion attacks by adversarial examples (Goodfellow, Shlens, and Szegedy 2014) during inference time. The goal of inference-time adversarial attack (Li et al. 2021a; Chen et al. 2022c; Zhang et al. 2022b; Chen et al. 2022b) is to damage the global model by adding a carefully generated imperceptible perturbation on the test examples. As shown in Table 1, federated models are as fragile to adversarial examples as centrally trained models (i.e. zero accuracy under PGD-40 attack (Madry et al. 2017)). Hence, it is also important to consider how to defend against adversarial attacks in federated learning.

There are several works that aim to deal with adversarial attacks in FL (Zhang et al. 2022c, a), i.e, federated adversarial training (FAT) (Zizzo et al. 2020; Hong et al. 2021; Shah et al. 2021; Chen, Zhang, and Lyu 2022; Chen et al. 2022a). (Zizzo et al. 2020) and (Hong et al. 2021) proposed to conduct adversarial training (AT) on a proportion of clients but conduct plain training on other clients. (Shah et al. 2021) investigated the impact of local training rounds in FAT. Nevertheless, these methods all ignore the issue that the clean accuracy of federated adversarial training is very low.

To further show the problems of federated adversarial training, we first begin with the comparison between the plainly-trained models and AT-trained (Madry et al. 2017) models in both the IID (Independent and Identically Distributed) and non-IID FL settings, measured by clean accuracy AclnA_{cln} and robust accuracy ArobA_{rob}, respectively. We show the test accuracy of plain training and adversarial training (AT) on CIFAR10 dataset under both IID and non-IID FL settings in Fig. 1 (left sub-figure). We summarize some valuable observations as follows: 1) Compared with the plainly-trained models, AT-trained models achieve a lower accuracy, which indicates that directly adopting adversarial training in FL can hurt AclnA_{cln}; 2) AclnA_{cln} drops heavily for both the plainly-trained models and AT-trained models under non-IID distribution, which is exactly the challenge that typical federated learning with heterogeneous data encountered (Zhao et al. 2018); 3) The performance of AT-trained models with non-IID data distribution decrease significantly compared with IID data distribution. Motivated by these observations, we focus on improving both adversarial robustness and clean accuracy of adversarial training in FL, i.e., we aim to increase AclnA_{cln} while keeping ArobA_{rob} as high as possible.

To achieve this goal, in this paper, we investigate the impact of decision boundary, which can greatly influence the performance of the model in FAT. Specifically, 1) we apply adversarial training with a re-weighting strategy in local update to get a better ArobA_{rob}. Our method takes the limited data of each client into account, those samples that are close to/far from the decision boundary are assigned larger/smaller weight. 2) Moreover, since the global model in FL has a more accurate decision boundary through model aggregation, we take advantage of the logits from the global model and introduce a new regularization term to increase AclnA_{cln}. This regularization term aims to alleviate the accuracy reduction across distributed clients.

We conclude our major contributions as follows:

We conduct systematic studies on the adversarial robustness of FL, and provide valuable observations from extensive experiments.

We reveal the negative impacts of adopting adversarial training in FL, and then propose an effective algorithm called Decision Boundary based Federated Adversarial Training (DBFAT), which utilized local re-weighting and global regularization to improve both the accuracy and robustness of FL systems.

Extensive experiments on multiple datasets demonstrate that our proposed DBFAT consistently outperforms other baselines under both IID and non-IID settings. We present the performance of our method in Fig. 1 (right sub-figure), which indicates the improvement in both robustness and accuracy of adversarial training in FL.

Related Works

Following the success of DNNs in various tasks (Li et al. 2019; Li, Sun, and Guo 2019; Huang et al. 2022b, a; Dong et al. 2021), FL has attracted increasing attention. A recent survey has pointed out that existing FL systems are vulnerable to various attacks that aim to either compromise data privacy or system robustness (Lyu et al. 2022). In particular, robustness attacks can be broadly classified into training-time attacks (data poisoning and model poisoning) and inference-time attacks (evasion attacks, i.e., using adversarial examples to attack the global model during inference phase). In FL, the architectural design, distributed nature, and data constraints can bring new threats and failures (Kairouz 2021).

Adversarial Attacks.

The white-box attacks have access to the whole details of threat models, including parameters and architectures. Goodfellow et al. (Goodfellow, Shlens, and Szegedy 2014) introduced the Fast Gradient Sign Method (FGSM) to generate adversarial examples, which uses a single-step first-order approximation to perform gradient ascent. Kurakin et al. (Kurakin, Goodfellow, and Bengio 2017) iteratively applied FGSM with a small step-size to develop a significantly stronger multi-step variant, called Iterative FGSM (I-FGSM). Based on these findings, more powerful attacks have been proposed in recent years including MIM (Dong et al. 2018), PGD (Madry et al. 2017), CW (Carlini and Wagner 2017), and AA (Croce and Hein 2020).

Adversarial Training.

Adversarial training has been one of the most effective defense strategies against adversarial attacks. Madry et al. (Madry et al. 2017) regarded adversarial training as a min-max formulation using empirical risk minimization under PGD attack. Kannan et al. (Kannan, Kurakin, and Goodfellow 2018) presented adversarial logit pairing (ALP), a method that encourages logits for pairs of examples to be similar, to improve robust accuracy. To quantify the trade-off between accuracy and robustness, Zhang et al. (Zhang et al. 2019) introduced a TRADES loss to achieve a tight upper bound on the gap between clean and robust error. Based on the margin theory and soft-labeled data augmentation, Ding et al. (Ding et al. 2020) proposed Max-Margin Adversarial (MMA) training and Lee et al. (Lee, Lee, and Yoon 2020) introduced Adversarial Vertex mixup (AVmixup).

Federated Adversarial Training.

In terms of the adversarial robustness, Zizzo et al. (Zizzo et al. 2020) investigated the effectiveness of the federated adversarial training protocol for idealized federated settings, and showed the performance of their models in a traditional centralized setting and a distributed FL scenario. Zhou et al. (Zhou et al. 2022) decomposed the aggregation error of the central server into bias and variance. However, all these methods sacrificed clean accuracy (compared to plainly trained models) to gain robustness. In addition, certified defense (Chen et al. 2021a) against adversarial examples in FL is another interesting direction, which will be discussed in the future.

Adversarial Robustness of FL

In this section, we briefly define the goal of federated adversarial training. Then we conduct a systematic study on some popular federated learning algorithms with the combination of various adversarial training methods and evaluate their robustness under several attacks. Besides, we further reveal the challenges of adversarial training in non-IID FL.

In typical federated learning, training data are distributed across all the KK clients, and there is a central server managing model aggregations and communications with clients. In general, federated learning attempts to minimize the following optimization:

Here, we denote that the global approximate optimal is a sum of local objectives weighted by the local data size nkn_{k}, and nn is the total data size of all clients that participate in a communication round. Moreover, each local objective measures the empirical risk over possibly different data distributions DkD_{k}, which can be expressed as:

Let xx denote the original image, xadvx^{adv} denote the corresponding adversarial example, and δ\delta denote the perturbation added on the original image, then xadv=x+δx^{adv}=x+\delta. To generate powerful adversarial examples, we attempt to maximize the loss L(x+δ;w)L(x+\delta;w), where LL is the loss function for local update.

To improve the robustness of the neural networks, many adversarial defense methods have been proposed. Among them, adversarial training (Carlini and Wagner 2017) is one of the most prevailing and effective algorithms. Combined with adversarial training, the local objective becomes solving the following min-max optimization problem:

The inner maximization problem aims to find effective adversarial examples that achieve a high loss, while the outer optimization updates local models to minimize training loss.

In this work, we conduct a systematic study on several state-of-the-art FL algorithms including FedAvg (McMahan et al. 2017), FedProx (Li et al. 2018), FedNova (Wang et al. 2020) and Scaffold (Karimireddy et al. 2020), and explore their combinations with AT methods to defend against adversarial attacks. We report detailed results in Table 2, here robustness is averaged over four popular attacks (FGSM (Kurakin, Goodfellow, and Bengio 2017), MIM (Dong et al. 2018), PGD (Madry et al. 2017), and CW (Carlini and Wagner 2017)). Besides, we implement some prevailing adversarial training methods including PGD_AT (Madry et al. 2017) , TRADES (Zhang et al. 2019), ALP (Kannan, Kurakin, and Goodfellow 2018), MMA (Ding et al. 2020) and AVMixup (Lee, Lee, and Yoon 2020). We observe that there is no federated adversarial learning algorithm that can outperform all the others in all cases. Moreover, the clean accuracy drops heavily under non-IID distribution. As such, we are motivated to develop a more effective method. Due to the similar performance of these FL methods observed from Table 2, we design our method based on FedAvg – a representative algorithm in FL.

Adversarial Traning with non-IID Data

Federated learning faces the statistical challenge in real-world scenarios. The IID data makes the stochastic gradient as an unbiased estimate of the full gradient (McMahan et al. 2017). However, the clients are typically highly heterogeneous with various kinds of non-IID settings, such as label skewness and feature skewness (Li et al. 2021b). According to previous studies (Wang et al. 2020; Karimireddy et al. 2020), the non-IID data settings can degrade the effectiveness of the deployed model.

Similarly, due to the non-IID data, the performance of AT may vary widely across clients. To better understand the challenge of adversarial training with non-IID data, we examine the performance of both clean accuracy and robustness on a randomly selected client and report the results in Fig. 2. Observed from Fig. 2, we can find that: 1) AclnA_{cln} on the plainly trained model drops from majority classes to minority classes, which is exactly what traditional imbalanced learning attempts to solve; 2) A similar decreasing tendency reasonably occurs in ArobA_{rob}. It is obvious that adopting adversarial training in federated learning with non-IID data is more challenging.

According to above observations, we conjecture that AT-trained local models with imbalanced data lead to a more biased decision boundary than plainly trained ones. Since adversarial examples need a larger number of epochs to achieve near-zero error (Zhang et al. 2021), it becomes harder to fit adversarial examples than clean data. However, for the local client itself, imbalanced clean data generates imbalanced adversarial examples, making it more difficult for training and enlarging the accuracy gap, which can reduce the performance both in accuracy and robustness. In Fig. 3, we also show the differences between plain training and adversarial training in federated settings. Compared with the plainly trained models, the aggregation of adversarially trained models can enlarge the accuracy gap, which results in poor consistency between different clients. To overcome this problem, we propose a novel method to utilize local re-weighting and global regularization to improve both the accuracy and robustness of FL systems.

Methodology

The generalization performance of a neural network is closely related to its decision boundary. However, models trained in the federated setting are biased compared with the centrally trained models. This is mainly caused by heterogeneous data and objective inconsistency between clients (Kairouz 2021). Moreover, a highly skewed data distribution can lead to an extremely biased boundary (Wang et al. 2020). We tackle this problem in two ways: 1) locally, we take full advantage of the limited data on the distributed client; 2) globally, we utilize the information obtained from the global model to alleviate the biases between clients.

Subsequently, we propose a simple yet effective approach called Decision Boundary based Federated Adversarial Training (DBFAT), which consists of two components. For local training, we re-weight adversarial examples to improve robustness; while for global aggregation, we utilize the global model to regularize the accuracy for a lower boundary error AbdyA_{bdy}. We show the training process of DBFAT in the supplementary and illustrate an example of the decision boundary of our approach in Fig. 4.

Adversarial examples have the ability to approximately measure the distances from original inputs to a classifier’s decision boundary (Heo et al. 2018), which can be calculated by the least number of steps that iterative attack (e.g. PGD attack (Madry et al. 2017)) needs in order to find its misclassified adversarial variant. To better utilize limited adversarial examples, we attempt to re-weight the adversarial examples to guide adversarial training. For clean examples that are close to the decision boundary, we assign larger weights; while those examples that are far from the boundary are assigned with smaller weights.

In this paper, we use PGD-SS to approximately measure the geometric distance to the decision boundary, SS denotes the number of maximum iteration. We generate adversarial examples as follows (Madry et al. 2017):

Here ΠB[x,ϵ]\Pi_{\mathcal{B}[x,\epsilon]} is the projection function that projects the adversarial data back into the ϵ\epsilon-ball centered at natural data, α\alpha is the steps size, ϵ\epsilon is perturbation bound.

We find the minimum step dd, such that after dd step of PGD, the adversarial variant can be misclassified by the network, i.e., arg  maxcf(c)(xadv)≠yarg\;max_{c}f^{(c)}(x^{adv})\neq y, where f(c)(xadv)f^{(c)}(x^{adv}) is the logits of the cc-th label.

In this way, given a mini-batch samples {(xi,yi)}i=1m\{(x_{i},y_{i})\}_{i=1}^{m}, then the weight list ρ\rho can be formulated as :

Regularization with Global Model

Early work (Zhang et al. 2019; Cui et al. 2021) claims that there exists a trade-off between accuracy and robustness, standard adversarial training can hurt accuracy. To achieve a lower boundary error AbdyA_{bdy}, we take advantage of logits from the global model fglof^{glo}, which is trained after aggregation. Particularly, in federated learning, the model owns the information obtained from the averaged parameters on distributed clients.

Let flocf^{loc} denote the adversarially trained model at each local client, fglof^{glo} has the most desirable classifier boundary for natural data. Then we can modify the local objective mentioned in Equation 3 as below:

To show the difference between our DBFAT and existing defense methods, we list the loss functions of different adversarial training methods in Table 3.

Experimental Results

Following the previous work of FL (McMahan et al. 2017), we distribute training data among 100 clients in both IID and non-IID fashion. For each communication round, we randomly select 10 clients to average the model parameters. All experiments are conducted with 8 Tesla V100 GPUs. More details can be referred to the supplemental material.

In this section, we show that DBFAT improves the robust generalization and meanwhile maintains a high accuracy with extensive experiments on benchmark CV datasets, including MNIST (Lecun et al. 1998), FashionMNIST (Xiao, Rasul, and Vollgraf 2017) (FMNIST), CIFAR10 (Krizhevsky and Hinton 2009), CIFAR100 (Krizhevsky and Hinton 2009), Tiny-ImageNet (Le and Yang 2015), and ImageNet-12 (Deng et al. 2009). The ImageNet-12 is generated via (Li et al. 2021c), which consists of 12 classes. We resize the original image with size 224*224*3 to 64*64*3 for fast training.

Data partitioning

In the federated learning setup, we evaluate all algorithms on two types of non-IID data partitioning: Dirichlet sampled data and Sharding. For Dirichlet sampled data, each local client is allocated with a proportion of the samples of each label according to Dirichlet distribution (Li et al. 2020). Specifically, we follow the setting in (Yurochkin et al. 2019), for each label cc, we sample pc∼Dir⁡J(0.5)p_{c}\sim\operatorname{Dir}_{J}(0.5) and allocate pc,jp_{c,j} proportion of the whole dataset of label cc to client jj. In this setting, some clients may entirely have no examples of a subset of classes. For Sharding (McMahan et al. 2017), each client owns data samples of a fixed number of labels. Let KK be the number of total clients, and qq is the number of labels we assign to each client. We divide the dataset by label into K∗qK*q shards, and the amount of samples in each shard is nK⋅q\frac{n}{K\cdot q}. We denote this distribution as shards_qq, where qq controls the level of difficulty. If qq is set to a smaller value, then the partition is more unbalanced. An example of these partitioning strategies is shown in Fig. 5, in which we visualize IID and non-IID distribution (Dirichlet sampled with pc∼Dir⁡J(0.5)p_{c}\sim\operatorname{Dir}_{J}(0.5) and Sharding with shards_55) on five randomly selected clients.

MNIST and FMNIST setup

We use a simple CNN with two convolutional layers, followed by two fully connected layers. Following the setting used in (Goodfellow, Shlens, and Szegedy 2014), for MNIST, we set perturbation bound ϵ=0.3\epsilon=0.3, and step size α=0.01\alpha=0.01, and apply adversarial attacks for 20 iterations. For FMNIST, we set perturbation bound ϵ=32/255\epsilon=32/255, and step size α=0.031\alpha=0.031, we adversarially train the network for 10 steps and apply adversarial attacks for 20 iterations. Due to the simplicity of MNIST and FMNIST, we mainly use non-IID data (Sharding), which is hard to train.

CIFAR10, CIFAR100, Tiny-ImageNet and ImageNet-12 setup

We apply a larger CNN architecture, and follow the setting used in (Madry et al. 2017), i.e., we set the perturbation bound ϵ=0.031\epsilon=0.031, step size α=0.007\alpha=0.007. To evaluate the robustness, we conduct extensive experiments with various data partitioning.

Baselines

For attack methods, we perform five popular attacks including FGSM (Kurakin, Goodfellow, and Bengio 2017), MIM (Dong et al. 2018), PGD (Madry et al. 2017), CW (Carlini and Wagner 2017) and AA (Croce and Hein 2020). We further use Square (Andriushchenko et al. 2020) for black-box attack. To investigate the effectiveness of existing FL algorithms, we implement FedAvg(McMahan et al. 2017), FedProx(Li et al. 2018), FedNova(Wang et al. 2020) and Scaffold(Karimireddy et al. 2020). To defend against adversarial attacks, we implement four most prevailing methods including PGD_AT(Madry et al. 2017), TRADES (Zhang et al. 2019), ALP (Kannan, Kurakin, and Goodfellow 2018), MMA (Ding et al. 2020) and AVMixup (Lee, Lee, and Yoon 2020). We compare the performance of our DBFAT with various kinds of defense methods combined with FL methods.

Convergence For Local Training

To show the convergence rate of DBFAT, we use the Dirichlet sampled CIFAR10 dataset, where each client owns 500 samples from 5 classes. Fig. 6 (left sub-figure) shows the impact of local epoch EE during adversarial training. Indeed, for a very small epoch (e.g., E=2E=2), it has an extremely slow convergence rate, which may incur more communications. Besides, a large epoch (e.g., E=20E=20) also leads to a slow convergence, as model may overfit to the local data. Considering both the communication cost and convergence issues, we set E=5E=5 in our experiments, which can maintain a proper communication efficiency and fast convergence.

Effectiveness of Our Method

We verify the effectiveness of our method compared with several adversarial training techniques on Dirichlet sampled CIFAR10. Evaluation of model robustness is averaged under four attacks using the the same setting for a fair comparison and all defense methods are combined with FedAvg.

To show the differences between DBFAT and above mentioned defense methods, we report the training curves on non-IID CIFAR10 dataset in the right sub-figure of Fig. 6. Fig. 6 confirms that our DBFAT achieves the highest clean accuracy. We speculate that this benefit is due to the regularization term and re-weighting strategy introduced in Equation 6. It is worth mentioning that in the training curves, the model trained with PGD_AT performs very poorly. It indicates that standard AT may not be a suitable choice for adversarial robustness in FL, as it only uses cross-entropy loss with adversarial examples, but ignores the negative impact on clean accuracy. We further report the results on various datasets under both IID and non-IID settings in Table 4, which indicates that DBFAT significantly outperforms other methods in terms of both accuracy and robustness.

In Table 5, we show the accuracy and robustness of each method on large datasets (e.g., CIFAR100, Tiny-ImageNet, and ImageNet-12). All results are tested under PGD-20 attack (Madry et al. 2017), AutoAttack (Croce and Hein 2020), and Square attack (Andriushchenko et al. 2020) in non-IID settings. From the results reported in Table 5, we can find that our method still outperforms other baselines in terms of both clean accuracy and robustness. Note that our method can achieve the highest accuracy and robustness of 61.38% and 22.08% under AutoAttack, respectively. It thus proves that our method can also be used to improve the accuracy and robustness of the model on large datasets. We think that the higher clean accuracy is a result of the regularization term introduced in Equation 6, while maintaining a high robustness.

Ablation Study

As part of our ablation study, we first investigate the contributions of different modules introduced in DBFAT. As shown in Table 6, turning off both the re-weighting strategy and regularization term will lead to poor performance, which demonstrates the importance of both modules. Moreover, cut-offing the re-weighting strategy can lead to a more severe degradation. We conjecture this is a reasonable phenomenon. As mentioned in Fig. 1, non-IID data can cause a serious accuracy reduction. Our re-weighting strategy can alleviate the bias by taking the limited data on each client into account.

Effects of Regularization

The regularization parameter β\beta is an important hyperparameter in our proposed method. We show how the regularization parameter affects the performance of our robust classifiers by numerical experiments on two datasets, MNIST and FMNIST. In Equation 6, β\beta controls the accuracy obtained from the global model, which contains information from distributed clients. Since directly training on adversarial examples could hurt the clean accuracy, here we explore the effects of β\beta on both accuracy and robustness. As shown in Table 7, we report the clean accuracy and robustness by varying the value of β\beta. We empirically choose the best β\beta for different datasets. For example, for MNIST, β=1.5\beta=1.5 can achieve better accuracy and robustness. For FMNIST, we let β=2\beta=2 for a proper trade-off in accuracy and robustness.

Conclusion

In this paper, we investigate an interesting yet not well explored problem in FL: the robustness against adversarial attacks. We first find that directly adopting adversarial training in federated learning can hurt accuracy significantly especially in non-IID setting. We then propose a novel and effective adversarial training method called DBFAT, which is based on the decision boundary of federated learning, and utilizes local re-weighting and global regularization to improve both accuracy and robustness of FL systems. Comprehensive experiments on various datasets and detailed comparisons with the state-of-the-art adversarial training methods demonstrate that our proposed DBFAT consistently outperforms other baselines under both IID and non-IID settings. This work would potentially benefit researchers who are interested in adversarial robustness of FL.

References