When the Curious Abandon Honesty: Federated Learning Is Not Private
Franziska Boenisch, Adam Dziedzic, Roei Schuster, Ali Shahin Shamsabadi, Ilia Shumailov, Nicolas Papernot
Introduction
With machine learning (ML) being increasingly applied to sensitive data in critical use-cases such as health care , smart metering , or the internet of things , there is a growing need for privacy-preserving training schemes that do not leak sensitive information. Federated learning (FL) is a widely popular distributed learning protocol where user data can be utilized for jointly training an ML model without the data ever leaving the users’ device. Instead, the device computes and sends model updates to a central party which aggregates them to produce a shared model. Assuming the model updates do not reveal the user data, FL would, thereby, preserve a notion of privacy.
This assumption has been repeatedly contested by prior work. It has been shown how the model updates sent to the central party not only leak training data membership (i.e. allow the attacker to tell if a given data point was used in training) but also properties of the training data . Inspecting model updates allows attackers to even (partially) reconstruct users’ training data. Ultimately, FL in its naive implementation offers little to no guarantees regarding potential leakage of user data to other users or to the central party.
Yet, existing data reconstruction attacks either are computationally expensive and yield low-fidelity extraction , are limited to small mini-batch sizes , or require modifications of the model architecture that are trivially detected . Another other line of concurrent work proposes modifications to the model parameters that are still easily noticeable: The attack introduced by Pasquini et al. sets a noticeable portion of parameters to zero or negative values. Similarly, the attacks by Wen et al. zero out many parameters of the last fully connected classification layer. In this work, we perform data extraction from large mini-batches of local data based on inconspicuous manipulations of the shared model weights. We start by showing scenarios where the gradients sent to the central party include full, memorized training data points. We then proceed to show that a malicious central party can significantly amplify this leakage by simply adversarially setting the model’s weights with our trap weights method, prior to dispatching the weights to users.
Our trap weights mainly rely on re-scaling components in the model’s weights matrix and can be applied to unmodified model architectures, which makes the attack more stealthy. By adversarially initializing the shared model with our trap weights, the central party can ensure they are able to perfectly extract a significant portion of the users’ training data, as depicted in Figure 1. This even holds when the gradients are computed over large training data mini-batches containing only data from the same class, a scenario in which previous optimization-based attacks usually fail to obtain high-fidelity reconstructions . Since in FL, the central party holds full control over the shared model weights that are sent out to users, our attack integrates naturally in the FL protocol. Furthermore, our attack is highly computationally efficient since it extracts individual training inputs by simply projecting the appropriate portions of the users’ gradients onto the input domain. Finally, we show both in theory and in practice that our attack is equally successful when users perform multiple rounds of local training (Fed-Avg ) and send the model updates instead of the gradients to the central party.
In summary, we make the following contributions:
We observe that in neural networks starting with a fully-connected layer, even gradients of large training data mini-batches contain individual training data points. In other words, in FL, mini-batch training data points are often directly sent from users to the central party, such that even an honest-but-curious central party has access to them.
We show that a dishonest and active central party can amplify the leakage of individual training data points and extend it to other model architectures by adversarially initializing the weight of the shared model.
In this setting, we perform data reconstruction on image and text data. Our attack is able to perform an extremely computationally efficient extraction of individual training data points in only a single-step computation over the received model updates with the attack setup depicted in Figure 2. For complex image datasets such as ImageNet , the attack yields perfect reconstruction of more than % of the training data points, even for large training data mini-batches that contain as many as data points. For textual tasks such as IMDB sentiment analysis , it perfectly extracts more than 65% of the data points for mini-batches with data points.
Background: Neural Networks, Federated Learning, and Differential Privacy
The goal of the model is to map an input to its desired ground-truth . Therefore, the model weights are adapted in a training process, most commonly with the mini-batch Stochastic Gradient Descent (SGD). To adjust the initial , mini-batch SGD repeats the following sequence of steps: (1) sample a mini-batch of size from the training data , (2) take a forward pass through the model to obtain its predictions on the mini-batch, (3) compute the difference between predictions and ground-truth labels, called the loss , (4) compute the gradient of w.r.t. the weights, called the weight gradient , and update the weights accordingly.
To bootstrap mini-batch SGD, weights need to be initialized by sampling from a random distribution; popular distributions include the zero-mean Gaussian , Xavier or He distributions. The choice of distribution has a large effect on learning success . In fact, when weights are maliciously initialized, the final model’s utility might be degraded .
Federated Learning. FL is a communication protocol for training a shared ML model on decentralized data owned by different users . Since collecting and managing all the data centrally might be costly, time consuming, and stand in conflict with the confidentiality of these respective users’ data, FL enables each user to keep their data locally. A central party coordinates the training of the shared model by iteratively aggregating gradients computed locally by users.
More formally, let be the current iteration of the FL protocol. At iteration , the model is initialized (at random) by the central party denoted as . Let be the model with its weights at iteration . At every iteration , out of the () users are selected to contribute to the learning. Then, each of the selected users obtains from and calculates the gradients for based on one mini-batch sampled from their local dataset . In other words, the user computes the gradient . Each uploads their gradients to , who then averages all of these gradients to update the shared model’s parameters:
FL, thereby, represents a decentralization of the mini-batch SGD (i.e. distributed training from mini-batches of user data).
Existing Data Reconstruction Attacks
This section introduces prior work on passive and active data reconstruction attacks in FL and discusses the limitations of attacks based on iterative optimization.
Passive attackers performing data reconstruction attacks in FL can simply observe the received gradients but not maliciously manipulate the protocol. Phong et al. were the first to show how gradients leak information that can be used to recover training data from single neurons or linear layers. Recent work proposed that the central party or users involved in FL training launch data reconstruction attacks based on either training a Generative Adversarial Network (GAN) or solving a second order optimization problem.
Optimization-based Instance Reconstruction Attacks. Several attacks aim to reconstruct individual user data points while also relaxing the assumption that data labels are available to the attacker. Zhu et al. proposed Deep Leakage from Gradients (DLG), where a data reconstruction attack is formulated as a joint optimization problem on the labels and input data, see Algorithm 2 in Appendix A. iDLG sped up the convergence rate of DLG by analytically computing the labels based on the users’ gradients of the last layer. These works, and other optimization-based ones , are limited to a setting where mini-batches only contain a single example, i.e., . GradInversion regularizes DLG’s objective to improve the extraction fidelity, attaining some success in extraction for mini-batches of size . In Section 7.4 we compare performance of our approach against a state-of-the-art optimization-based attack. Our attack is superior in extracting individual training data even for large mini-batch sizes of , and being far more computationally efficient (even for passive adversaries in the honest-but-curious model). We present a more thorough overview on passive data reconstruction attacks in Appendix A.
Limitations of Optimization-Based Attacks. We hereby provide a brief exposition to Zhu et al. ’s DLG, as a representative case study of an optimization-based attack. Their approach, characteristic of optimization-based data reconstruction attacks, is given in Algorithm 2 in Appendix A. It firstly randomly initializes a “dummy data point and corresponding label” and computes the resulting “dummy gradients” as . Then, they iteratively optimize the dummy data to produce gradients that are close to the original gradients by solving:
DLG often fails to reconstruct high-fidelity data points and discover the ground-truth labels consistently because of a lack of convergence in the optimization. While other methods offer improvements (e.g. iDLG sped up the convergence by simplifying the objectives in Equations 2 and 3 from both data and label reconstruction to only data reconstruction; and GradInversion adds useful regularization), they suffer from the same pathology.
We identify several reasons for this. First, the gradient of the loss is non-injective i.e. is not invertible everywhere: different mini-batches may yield nearly identical gradients . This holds whether the user samples mini-batches that contain multiple data points or a single data point only, i.e. . Second, optimization-based attacks converge to different minima due to the underlying randomness (see step 1 in Algorithm 2 in the Appendix). These minima correspond to different possible reconstructions of the input that often differ from the original training points . Third, optimization-based attacks are computationally expensive: they either need to train a GAN or solve a second-order gradient optimization problem. Instead, our attack extracts exact data points from the gradients without any optimization or GAN training.
2 Active Attackers
In the work most similar to ours, considers a threat model with an active and dishonest central party, similar to our setup. This attack relies on the existence of a fully-connected layer early within the network (otherwise, the attack adds it). Since this layer’s weights have to contain many weight rows with the exact same weight values, this layer is inherently detectable.This is inherent to the attack because the method relies on each row computing the exact same function on the data and binning its result by varying only the bias term, such that it becomes likely that a bin contains only one input. Conversely, our trap weights are initialized such that it is likely that an output neuron is only activated for a single input in a mini-batch while avoiding imposing a highly regular structure on the weight matrix. Additionally, they do not discuss passive analytical-extraction attacks. Finally, our work generalizes their setup and performs successful extraction also for textual data.
In follow-up work, proposes an attack that requires modifications to the model parameters (specifically, to the last fully-connected classification layer) sent to a user but without changing the model architecture. The attack extracts single data points by increasing the gradient contribution of a target data point and decreasing the gradient contribution of other data points. The final goal of an attacker is to reduce an aggregated gradient to an update calculated on a single sample. The attack is easily detectable since it requires many parameters in the last layer to be zeroed out. Moreover, our attack extracts individual data points in a single training round while their approach requires a collection of many updates from an individual user. In the cross-device FL setting where participants get randomly sampled from millions of users, it is possible that single users participate fewer times than required by the attack.
In another active attack proposed by , a server sends distinct malicious parameters to individual users. The main purpose of the attack is to circumvent the protection of Secure Aggregation (SA) in FL and enable the central party to learn individual model updates from a target user. However, the work does not propose individual user-data point extraction, as enabled by out trap weights. We argue that by including our trap weights into their attack and sending our trap weights to the target user, they could efficiently extract this target user’s private data.
Threat Model and Assumptions
This section presents our threat model in terms of the assumed attacker, the FL deployment, and the assumptions required for our attack to succeed.
Our attacker aims at extracting individual training data points from a chosen subset of the participating users. Therefore, the attacker’s primary vantage point is the central party who is in charge of orchestrating the FL protocol. The assumption here is that, for example, the company orchestrating the FL protocol or potential rogue employees, are untrusted. This is the same attacker that FL is meant to defend against by leaving data on the users’ devices. For brevity, in the following, we will refer to the central party as the attacker, even though the attacker can be a third party controlling the central party to deploy our trap weights-attack.
In FL, the central party initiates the FL protocol and chooses the task to train the shared ML model for. Therefore, the central party is aware of the type, domain, and dimensionality of data held by the users. It instantiates the shared model appropriately to learn from this data. Furthermore, in the standard FL scenario considered in this work, the central party holds full control over the shared model weights and can read users’ gradient updates that are sent back. Finally, our central party is in charge of sampling the users who contribute their gradients in a given round—following standard deployments of the protocol . This allows the central party to even run targeted attacks against specific users.
2 Assumptions and FL Setup
Following prior work , we consider an FL protocol where users calculate the model gradients locally on one (potentially large) mini-batch of their training data and share the resulting gradients directly with the central party. We assume that the data features are scaled in the range $$, which is a standard pre-processing step in ML. When users have abundant amounts of data, they can perform local gradient calculation and averaging over more than one mini-batch, see evaluation in Section 7.3.
We, furthermore, assume that the attacker is in possession of a small amount (e.g., one mini-batch) of data from the users’ private data domain. This is no strong assumption given that the central party chooses the ML task and has to instantiate the ML model appropriately.
The weight-manipulation attacks we study in this paper are not agnostic of the model architecture. In designing the attack, we focus on victim models that contain a ReLU-based fully-connected layer, and we experiment with several different types of such networks. This is not a material limitation: the approach of manipulating shared model weights to promote leakage is very flexible, and can be extended to cover many more architectures as needed using similar techniques. For example in Section 6.3 and Appendix B, we show how to extend the attack to work on networks that contain convolutional layers and in Section 7, also experimentally evaluate extraction under the presence of a token-embedding layer.
3 Course of Attack
The course of our attack is illustrated in Figure 2. In a given round of the protocol, the central party maliciously manipulates the shared model with our trap weights. Note that the central party does not necessarily attack all users in every round of the protocol. Instead, it can target one or several specific users in one or more chosen round(s). To target a subset of the users at iteration , the central party can send out different models to users under attack and other users : while the targeted users receive a model initialized with our trap weights, all other users receive the shared model used to train the ML task. Attacking only a few users in a few rounds makes the attack more stealthy and allows the central party to train a performant shared model based on the gradient updates received in the benign rounds or from non-targeted users.
After receiving the gradients from users under attack, the central party simply projects the appropriate portions these gradients onto the input domain. In the following section, we will show how this approach can yield perfect extraction of the users’ data points.
Passive Analytical Extraction for FC-NNs
Here, we show how the gradients of an FC-NN directly leak the individual training data points they are computed on, even to a passive attacker who just observes said gradients. In Section 5.1, we formally show that for a single training data point, i.e. a mini-batch size of , perfect extraction from the network gradients is possible. Then, in Section 5.2, we motivate why it is also possible to perfectly extract a small number of individual data points from gradients, even when working with larger mini-batches of size . However, the success of this passive extraction attack drops as the mini-batch sizes increase. This limitation motivates our active adversarial weight initialization attack, which we introduce in Section 6.
It has been shown by Geiping et al. that a single input data point can be reconstructed from the gradients of any fully-connected layer which is preceded only by fully-connected layers and contains a bias . This holds if the gradient of the loss w.r.t. the layer’s output contains at least one non-zero entry. For detailed proof of the above see Proposition D.1 in . In particular, when considering the first model layer, reconstructing its input data directly corresponds to obtaining the original input data point . Let denote the output of the neuron of the first and fully-connected layer of a model, and let be the corresponding row in the weight matrix and the corresponding component in the bias vector. Assume , and therefore, . The reconstruction of the input is done by calculating the gradients of the loss w.r.t. the bias and the weights as follows:
since , where .
Thus, if any , perfect reconstruction is given by:
According to Equation 5, the gradient of the loss w.r.t. the weights directly contains a scaled version of the input data. The exact scaling factor is , which is the gradient of the loss w.r.t. the bias. This gradient is computed in the regular backward pass together with the gradient of the weights. Therefore, obtaining the scaling factor by just reading it from the gradients of the bias and inverting it to comes at zero costs and the factor can be directly applied to rescale the gradient of the weights and obtain the input data point , see Equation 6. Intuitively, the reason why there is a rescaled version of the input data in the gradients and why this would be beneficial for learning can be motivated by revisiting the simple perceptron algorithm . When an input is misclassified, the weight update in the perceptron algorithm consists simply in adding this input to the weights, which makes the algorithm learn.
2 Mini-batch Gradients Directly Leak Some Individual Inputs
It turns out that individual data point leakage is not limited to gradients computed over a mini-batch of size : we observe that gradients computed over larger mini-batches also sometimes leak individual training points. To forge an intuition for this phenomenon, Figure 8 in Appendix D visualizes the gradients of the first fully-connected layer’s weight matrix of the FC-NN described in Table IX. We see that we are able to clearly distinguish some of the training data points within the rescaled gradients. This is despite the fact that these gradients were computed over a mini-batch of inputs sampled from the CIFAR10 dataset.
with . These equations illustrate that the gradient over the data mini-batch contains a weighted overlay of all the input data points from the mini-batch. The weighting, therein, depends on the contribution of each data point to the model loss .
We observe that, in some cases, all but one training data point from the data mini-batch have zero gradients. This is due to the operation in ). When is negative, the outputs zero, which results in zero gradients for the corresponding data point. When the gradients are zero for all data points but for the one data point , the weight gradient from Equation 7 becomes with . This reduces the data extraction from the case of to the case of , for which we saw in Section 5.1 that the data point can be perfectly extracted. In other words, being negative for all data points but one results in accidental leakage of that data point—enabling its exact reconstruction by a passive adversary.
3 Individual Inputs still Leak from Mini-batch Gradients computed in FedAvg
FedAvg is another popular protocol for FL where users do not send their gradients after each local iteration of training. Instead, they calculate many local epochs over mini-batches of their data. After each iteration of the total many local iteration, they update the model according to the respective gradients of the weights and biases, and a learning rate as . Once the local training is completed, the users send the updated shared model to the central party. By calculating the difference between the shared model sent to the user and the obtained model, and by re-scaling according to , the central party obtains the value of the user’s local model update as . According to Equation 5, after every local iteration the gradient of the local weights . Therefore, . Since the server knows from the model update, it can multiply and extract the user data perfectly. We experimentally validate this theoretical insight on large mini-batches at the end of Section 7.3.
Active Adversarial Initialization of the First Fully-Connected Layer
Section 5 illustrates under which conditions model gradients leak data points to a passive attacker capable of observing these gradients. In the following, we show how an active attacker can amplify previously-accidental leakage during the passive attack by controlling the weights and biases . For example, while a passive attacker can extract roughly % of arbitrary data points from a batch size for 1000 neurons (i.e. weight rows in the fully-connected layer) on ImageNet, the active attack can more than double the number of extracted data points to %. Next, we show how to make such malicious choices to extract a larger number of individual training data points from model gradients.
Without loss of generality, we will suppress the bias term in the following considerations. The multiplication of a single weight row corresponding to the neuron at the fully-connected layer with some input data point can be expressed as a weighted sum of all of the features in as follows
In weight row , let and denote the sets of indices that hold the negative and positive weight components, respectively. Given activation, the neuron is only activated on if the sum of the features weighted by the negative components is smaller than the sum of the features weighted by the positive components:
Therefore, will yield non-zero gradients at the neuron if and only Equation 9, holds for its features.
When the inequality holds only for a single data point in a mini-batch, this data point can be individually extracted from the gradients, as described in Section 5.2. The idea behind our trap weights is to set the components within each weight row corresponding to the neurons of the first fully-connected layer, such that Equation 9 only holds relatively rarely in inputs, and is therefore likely to only hold for a single data point within a mini-batch.
2 Adversarial Weight Initialization
Intuitively, our approach adversarially initializes each row of the weight matrix to increase the likelihood that only one data point in a given mini-batch will activate the neuron corresponding to that row. To achieve this, we initialize a randomly chosen half of the components of the weight row to negative values, and the other half to the corresponding positive values, by sampling from a Gaussian normal distribution. The positive components of the weight row are scaled down with a small factor in comparison to the negative components. This increases the impact of the negative components on the weighted input sum to the corresponding neuron. This causes most input data points to produce non-positive input to the neuron, such that only a few (in the best case only one) input data point activates the neuron. See Algorithm 1 for a formalization of our initialization.
We use the scaling factor to specify how much larger the absolute values of the negative weight components should be than the positive values. This determines how ”aggressively” our activation causes weighted inputs to individual neurons to be negative, thereby to be filtered out by the function and to have zero gradients for most input data points. The ideal value of when it comes to attack effectiveness is dataset-dependent. The attacker can to fine-tune either on a small amount of data from the users’ input domain it holds before sending the trap weights to the users. Alternatively, they can fine-tune without any data from the users’ input domain and solely by exploiting the passive data leakage, or using data with the same dimensionality as the users’ data as we show in Section 7.3.
Our adversarial initialization causes the ReLU function for many neurons at the fully-connected layer to activate only for one input data point per mini-batch. Due to the randomness in the initialization of each weight row corresponding to a neuron, different neurons are likely to be activated by different input data points. Thereby, the gradients of different weight rows allow for the extraction of different individual data points. We demonstrate the success of our trap weights for data extraction in Section 7.3 by showing that they increase the proportion of neurons that only activate on one random individual data point in a mini-batch by more than factor 10, and thus we are able to extract more than double the number of individual training data points. E.g. our trap weights cause % of active neurons out of 1000 to by activated by individual data points from the ImageNet dataset while random model weights with a Gaussian normal initialization with only yield %. This allows for an individual extraction of % of the data points in a mini-batch of size for our trap weights versus % with random model weights.
3 Trap Weights for Other Architectures
To enable perfect extraction, our attack relies on the presence of a ReLU-based fully-connected layer at the beginning of the model architecture. Since in FL, the central party is in charge of instantiating the model architecture, this does not represent a practical limitation.
For some application domains, the central party might want to train ML models beyond pure FC-NNs though, i.e., models where the first layer is not fully-connected. In Appendix B, we show how an attacker can apply malicious manipulations to shared model’s weights to extend our attack to CNN-based architectures. These architectures consist of several convolution layers and some fully-connected layers which the attacker can leverage for extraction. The intuition of the attack-extension is to convert the convolution layers to identity functions which transfer the user’s input data to the first fully-connected layer in the model. The attacker initializes this layer with our trap weights and can then extract user data. In Appendix B.3, we discuss how to also make the manipulations of the convolutional layers most stealthy.
In the following section, we also evaluate extraction for text-data in model architectures that contain an embedding layer before the first fully-connected layer. Extraction is done from the fully-connected layer whose input consists of the embedding layer’s output. To reconstruct the original text tokens from a sequence of extracted embeddings, the attacker creates a lookup dictionary, mapping its initialized embeddings back to their corresponding tokens (this is the inverse mapping to the embedding layer). To avoid vector-comparisons for each lookup, the attacker uses hash values for vector embeddings as keys.
Experimental Evaluation
In this section, we validate that our adversarial weight initialization attack allows a central party to reconstruct individual training data points from gradients shared by users. We use three different image datasets, namely MNIST , CIFAR10 , and ImageNet and the text-based IMDB dataset for sentiment analysis. Because our approach is applicable to FC-NNs and CNNs, we test it against both of these architectures. We instantiate our attack against an FC-NN for the MNIST dataset, and against a CNN for CIFAR10 and ImageNet. For the IMDB dataset, we use a model whose input is 250-token sentences, and consists of an embedding layer, which maps each token in a 10,000-word vocabulary to a 250-dimensional floating-point vectors, and inputs these to a fully-connected layer. The specifics of our model architectures for image and text data are described in Table IX and Table X in Appendix C, respectively. We implemented our trap weights, and the experiments in TensorFlow version 2.4. The code will be open sourced after the peer review process.
Attack Instantiation. Since the central party has access to the gradients of all model layers uploaded by users, it is able to choose which layer to instantiate the attack on.
For the FC-NN, we adversarially initialize the first layer with our trap weights and extract training data points from its gradients. In the CNNs, we first initialize the convolutional layers to transmit the input data to the first fully-connected layer of the architecture. Then, we adversarially initialize this layer’s weights with our trap weights for extraction.
For the text classifier, we initialize the weights of the embedding layer with a random uniform distribution (min=0., max=1.) to create the inputs for the fully-connected layer. We then adversarially initialize this fully-connected-layer’s weights with our trap weights to perform extraction of the embeddings there.
We introduce three novel metrics to measure the success of individual data point extraction.
Active Neurons. By measuring the number of active neurons (A) we can determine for how many neurons their respective weighted inputs are positive. This is important because data extraction for both overlaying and individual data points is only possible with activated neurons. If a neuron is not activated by any data point, no information can be transmitted over this neuron and, hence, the gradients will all be zero.
Extraction-Precision. Our second metric, which we call extraction-precision, captures the percentage of non-zero gradient rows at the given layer’s weight matrix from which we can extract any input data point individually. This metric enables us to quantify how well the adversarial weight initialization manages to generate weights that cause activation for exactly one single data point. Extraction-Precision can be calculated as follows:
However, the extraction-precision metric alone would not be expressive enough since a high extraction-precision could be achieved despite the exact same individual training input being reconstructed from all gradient rows. Therefore, we defined another metric that we call extraction-recall.
Extraction-Recall. The extraction-recall measures the percentage of input data points that can be perfectly extracted from any gradient row. We define it by
Interpretation of Success Metrics. Note that our attack seeks to find an adversarial initialization that balances setting enough neurons’ outputs to zero (such that a gradient is more likely to isolate individual points from large mini-batches) with, at the same time, having enough neuron outputs’ that are non-zero (otherwise, in the limit, no points would be extracted). Thus, active neurons provide additional context for the extraction-precision: with few active neurons, even a high extraction-precision might not be able to extract many individual training data points, simply because there are very few gradients to perform data extraction from. However, with many active neurons, the extraction-recall might become small, due to each neuron being most likely activated by several input data points, preventing individual extraction.
2 Evaluating the Passive Attack
Recall from Section 5 that extraction of training data from gradients is possible even when model weights are initialized randomly. We evaluate this passive attack to obtain a baseline for our adversarial weight initialization strategies. To evaluate the passive attack, we measure the extraction success of individual training data points from the gradients of randomly initialized models.
Table I reports the extraction-precision and extraction-recall of training data point extraction from the gradients of randomly initialized models. These gradients are computed over a mini-batch of 100 data points for 1000 neurons (i.e. 1000 weight rows’ gradients for extraction) in the first fully-connected layer. We later study the impact of these two parameters on the success of reconstruction attacks. Even if this attack is passive, and the central party has not modified any of the weights adversarially, training data extraction is often successful: for the MNIST dataset, around % of individual training data points can be directly extracted from the model gradients, whereas for CIFAR10 and ImageNet, roughly % and % of the training data points can be perfectly extracted. The passive attack for extracting embeddings from the IMDB dataset yields roughly % extraction-recall for 1000 neurons and mini-batches of 100 data points, see Table II.
These results also suggest that setting higher spread, in form of standard deviations to random weight distributions alone can already significantly increase the extraction-recall of individual data points from the model gradients, see Table I. This is, most likely, due to the larger span within the weight values.
Additionally, we also set out to investigate how as training progresses, and the model’s weights converge, the extraction’ success evolves. We initialized the FC-NN from Table IX with a Xavier Uniform distribution and trained the model on MNIST and CIFAR10 for 30 epochs. Table III depicts the results. We observe that the extraction-recall increases slightly over the training epochs. Analyzing the distribution of the model weights in Figure 3 shows that over training, the uniformly initialized weight values resemble more a normal distribution and obtain a wider spread, which might be the reason for the increased extraction success.
3 Evaluating Active Manipulations
We now turn to our active attack, which implements our trap weights to amplify the vulnerability exploited by passive attacks. This amplification is controlled by the scaling factor in the trap weights. We first evaluate the impact of this scaling factor on the reconstruction quality of individual training data points over a mini-batch of 100 data points and 1000 neurons.
Table IV depicts the results, averaged over ten different random adversarial initializations. We can see that the best scaling factor for MNIST, when it comes to the extraction-recall, is . With this scaling factor, we are able to extract on average % of the individual training data points which were involved in the users’ gradient computations. This is an improvement by around factor nine to the passive attack. For CIFAR10 and Imagenet, the best scaling factors concerning extraction-recall are , and , which allow for a perfect reconstruction of %, and % of the individual training data points, respectively, for 1000 neurons and a mini-batch size of data points. Thereby, the active attack is more than twice as successful as the passive attack for extracting individual training data points in these datasets. Figure 11, Figure 12, and Figure 13 in the Appendix D show the visual reconstruction results of the best run for the MNIST, CIFAR10, and ImageNet dataset, respectively. For CIFAR10, we additionally present extraction success when all local data stems from the same single class.
Similar improvements of performance could be achieved for the IMDB dataset. The best extraction was achieved also with , for which, with 1000 neurons and mini-batches of 100 data points, we obtained an extraction-recall of %, which is around 2.5 time as high as the passive attack, see Table II.
In Figure 4, we show the influence of the scaling factor , our method’s hyperparameter, on the distribution of our trap weights. The case corresponds to the baseline where positive components in the trap weights are not scaled down. The figure shows that the more deviates from , the larger the difference between a random distribution and our trap weights. For CIFAR10 and ImageNet (, and ), our trap weights’s distribution is very close to the the random distribution, making our trap weights more stealthy. The best scaling factor for MNIST, , is significant smaller than for ImageNet and CIFAR10 due to the sparsity in the data (the background in MNIST images consists of zero pixels). Our experiments indicate that with decreasing sparsity and increasing data dimensionality, approaches 1. Especially the last observation makes sense since scaling more positive components with a factor closer to 1 is in effect of the weighted sum equivalent to scaling fewer positive components with a factor much smaller than 1. Thereby, our trap weights increase in stealthiness with increasing complexity of the data to be extracted.
As hypothesized above, from Table IV, we furthermore confirm that the extraction-recall of our attack is related to the percentage of active neurons: When very few neurons are activated, it is not possible to extract large numbers of individual data points due to the lack of gradients to extract them from. However, when the percentage of active neurons is high, the extraction-recall also becomes very small, which is due to the fact that each neuron gets activated by several input data points, and thereby, individual extraction is impossible.
Attacker without Auxiliary Data. We experiment with an attacker who does not have access to a small mini-batch of data from the users’ distribution to tune the scaling factor of our trap weights. In this setup, the only knowledge an attacker holds is about the dimensionality of the users’ data which it needs to instantiate an adequate model architecture. We evaluate three attacks in this setup. 1) Exploiting passive data leakage and composing a tuning dataset: the attacker randomly initializes the model in a first round of the protocol. Our results in Table I show that also randomly initialized models’ gradients leak significant fractions of the users’ data (MNIST 6.1%, CIFAR10 25.9%, and ImageNet 21.7%). By plotting the user’s gradients and eyeballing which data points resemble natural images, the attacker can build a tuning set for . Since we only require a maximum of 100 data points to find the optimal values for per dataset in Table IV, the attacker only has to inspect the gradients of 17, 4, and 5 users for MNIST, CIFAR10, and ImageNet, respectively in the first round of the protocol. On the selected data, they can tune and use it in every subsequent iteration. We performed tuning on 100 data points obtained through passive extraction and obtained the same as through tuning on a random mini-batch of data (0.7, 0.95, and 0.99 for MNIST, CIFAR10, and ImageNet, respectively). 2) Exploiting raw passive data leakage: Since manually, selecting suitable data points is time-consuming, we propose an alternative approach where the attacker uses all extracted data points with are in a valid range for input pixels () from the passive extraction on non-adversarially initialized model weights in the first round of the protocol. These data points are not necessarily individually extracted user data points as we show in Figure 9 in Appendix D.2. But the attacker can still consider them as a tuning dataset for and evaluate the extraction-recall on this dataset when initializing the shared model with different trap weights to tune . Our results in Table XIII show that for CIFAR10 and ImageNet, the best found on these passively reconstructed data points are equal to the best obtained directly by tuning on one mini-batch of the original data. For MNIST, the on the extracted gradients differs slightly from the original best (0.75 vs. 0.7). We suspect these changes to result from MNIST data being much sparser (many more zero features) than the extracted gradients in Figure 9(a). 3) Using a surrogate dataset of same data dimensions: Lastly, the attacker can tune on a surrogate dataset of the same dimension (but potentially different distribution) than the users’ data. We compare extraction-recall of an adversarial weight initialization with found on a surrogate dataset and the optimal found on the actual dataset for Fashion MNIST, SVHN, CIFAR100, and Open Images in Table VI. Our results highlight that extraction with the surrogate obtained through tuning on MNIST, CIFAR10, and ImageNet, already yields a significantly higher success than passive extraction on non-manipulated weights. Furthermore, the closer the surrogate dataset’s distribution is to the users’ dataset, the closer and . Especially for CIFAR10 and CIFAR100, and ImageNet and Open Images, we find that which leads to highest extraction success.
Impact of Data Labels. Additionally, we investigated whether this high reconstruction success could also be achieved =in a non-IID setting when users hold local mini-batches of data that belongs to one single class, different from the other users. This is a particularly challenging setting for prior work on optimization-based attacks that end up reconstructing average points rather than individual points exactly. Instead, Figure 12(b) and Table XII in Appendix D.2 show on CIFAR10, how our method remains able to perfectly extract individual data points from the gradients even when all points stem from the same class.
Impact of Mini-Batch Sizes. We also set out to investigate the impact of the mini-batch size and the number of weight rows that we can use for extraction. Table V depicts the resulting metrics. The metrics show that the smaller the mini-batch sizes are, and the more weight rows there are for extraction, the more individual training data points can be individually reconstructed. For weight rows, even up to % of the individual training data points for mini-batch sizes as large as 200 in the MNIST dataset can be perfectly extracted. Small mini-batches of training data points are entirely extractable without any loss in this setting. Also for the IMDB dataset, smaller batch-sizes for the same number of neurons yield much higher extraction-recall, and embeddings of data from small mini-batches of training data points are perfectly extractable, see Table II. This suggests that in practice, the success of the extraction attack can be significantly increased by the central party demanding smaller mini-batch sizes from the users or initializing larger models.
Impact of Lossy Layers. For perfect extraction of data points in CNN architectures, our attack requires the input to the first fully-connected layer to have at least as many parameters as the original input data point. CNN architectures can contain pooling layers to reduce input size. In Appendix D.3, we evaluate the impact of pooling on the fidelity of the extracted data. Our evaluation shows that pooling results in some form of compression of the user’s input data, see for example Figure 16(a). In Appendix B.2, we show how the central party can implement an alternative to pooling for size-reduction in CNNs based on convolutional layers which still allows for prefect extracability, as long as there are enough model parameters. We also evaluate the effect of dropout on the fidelity of the extracted data. Figure 15 and 17 visualize the effect of pure dropout, while Figure 16 and 18 visualize the joint effect of dropout and pooling. To increase fidelity of extraction under lossy layers, an attacker can apply post-processing, such as de-compression.
VGG and ResNet. In addition to our custom FC-NN and CNN architecture from Table IX, we experimented with a VGG7 and a ResNet20 architecture. For VGG7, we initialize all convolutional layers as illustrated in Figure 6 in Appendix B.1, and the fully-connected layer directly after the convolutional layers with our trap weights. For the ResNet20, we only initialize the first convolutional layer according to Figure 6. Thanks to the skip connections, the remaining convolutional layers can remain unchanged, apart from the convolutional filters whose output is added to the output of the skip connections. These need to be set to zero, such that the input data can be propagated unaltered over the skip-connections to the fully-connected layer that we initialize with our trap weights. We set the last pooling layer before this fully-connected layer in ResNet20 to implement average pooling. Our extraction results for ImageNet are depicted in Figure 5. The compression of extracted data in comparison to the original data results from the pooling layers in both architectures that reduce input dimensions.
Impact of Local Mini-Batch-Averaging. Additionally, we looked into the effect of averaging over the gradients of multiple mini-batches, e.g. the average of gradients received from multiple users. The results in Table XI in Appendix D.2 show that through averaging, the attack success is significantly reduced. Already when averaging over mini-batches of size in the MNIST dataset, the average extraction-recall drops from % to % because multiple data points overlay in the gradients. This highlights that the central party needs to perform the extraction before the averaging operation. The following section shows that this simple change to the protocol is easily implemented by an actively dishonest central party, even for standard FL libraries.
TensorFlow Federated. We experimented with TensorFlow Federated —a standard open source library for FL deployments. In Appendix D.4, we show that a dishonest central party only requires minimal code changes to implement our trap weights.
4 Comparison To Previous Work
To compare the success of our attack, we compare to the three approaches conceptually closest to ours. These are which is the first to describe individual extractability of single-data point gradients, which relies on direct extraction from a fully-connected layer in the model architecture, and which exploits manipulations model parameters to extract data from user-gradients.
Comparison to . For the sake of correctness, we build on their code base and adopt it to also run with neural architectures we used in our other experiments. We use the parameters that found to perform best. Here, we are mainly interested in the quality of the other attack’s reconstruction in comparison to our method, and in the number of passes over the model, i.e. the computing time required to obtain the reconstructions. We perform evaluation on the MNIST and CIFAR10 datasets.
Comparison to . The success of ’ data extraction depends on the size of their imprinting module. Using their code-base and extending it with our success metrics, we evaluated what size of imprinting module they require to obtain the same extraction recall as we do, i.e., to extract the same number of data points from the model gradients perfectly. We compared our methods for all three vision dataset, using a batch-size of . For ImageNet and CIFAR10, following their baseline, we instantiated their model with a ResNet18, for MNIST, we used LeNet5. We always inserted their imprinting module before the first layer to allow for perfect extractability with their method. Our results show that to obtain the same extraction recall as we do (46%, 54%, and 54% for ImageNet, CIFAR10, and MNIST, respectively), their imprinting module needs to be of size roughly 150, 200, and 400, respectively. The fact that they require the largest imprint module for MNIST is due to the similarity in the data points (sparsity in the background with all zero pixels) which makes their binning less effective.
Comparison to . Note that ’s main goal is not to extract large amounts individual user data points but user updates, by circumventing the secure aggregation used to protect the FL protocol. The updates (gradients) that recover do usually not correspond to full and perfectly individual data points. Instead, for FC-NNs, their extracted gradients will resemble our passive extraction results from Figure 8 where most of the gradients are a blurry overlay of all underlying data points. For CNNs, their results will not be able to extract any individual data point since non-maliciously initialized convolution filters overlay input features. Thereby, their attack mainly violates confidentiality of the users’ model updates in a setup where users believe to obtain protection though an aggregate with other users. Still, the resulting gradients can then be used as a departure point for additional privacy-attacks, such as reconstruction. In contrast, our work directly violates the users’ privacy by manipulating the shared model weights to make individual data points directly extractable from the model updates sent from users to the central party. To assess individual extractability in their setup, we use their gradient suppression and model inconsistency attack to make all but one user in a round of the FL protocol return zero gradients. The one target-user receives a randomly initialized FC-NN (Gaussian with ) with architecture from Table IX. Note that the data extraction from gradients in this setup corresponds to our passive extraction. For MNIST, with , extraction in their setup yields % of perfectly extractable data points, while our trap weights yield %. In the same setup for CIFAR10, their method yields %, in contrast to our trap weights which yield again %.
Defending Privacy in Federated Learning
This section discusses potential mitigations against our trap weights attack. We start with explaining DP which provides formal privacy guarantees and then move on to defenses specifically tailored to our trap weights attack.
To bound the leakage of private information from model gradients, a gold standard for reasoning about privacy guarantees is the framework of differential privacy (DP) .
There exist three main ways of integrating DP in the FL protocol, namely Centralized Differential Privacy (CDP), Local Differential Privacy (LDP), and Distributed Differential Privacy (DDP).
In CDP, users clip their gradients locally according to a clip norm and the central party performs the addition of noise with a scale dependent on the noise multiplier . CDP cannot provide DP guarantees with a malicious central party because this central party can simply extract user data before adding noise or not add noise at all. For a private aggregation of sensitive statistics (instead of high-dimensional ML model gradients), there exist solutions of CDP without a trusted aggregator .
LDP reduces the trust required in the central party since every user locally adds noise to their gradients according to their privacy requirements . Independent of other users, the noise is drawn from . However, previous work has shown that this setup leads to poor privacy-utility trade-offs, such that LDP is not popular in practical applications .
DDP is supposed to combine the advantages of CDP and LDP. In DDP, before aggregation, each user locally adds some (small) amount of noise to their gradients . The noise distribution depends on the number of other selected users. It is specified by . While the individual noise levels do not offer sufficient protection, the aggregates provide rigorous privacy guarantees. There exist different forms of performing the aggregation. One popular approach is to use secure aggregation (SA) , which adds significant computational overhead and requires tailored DP mechanisms that operate on integer values, e.g. . Finally, prior work has shown that in FL, SA can be eluded . This motivates defenses dedicated to protecting specifically against our trap weights attack.
2 Specific Defenses against Trap Weights
The following defenses can be applied to mitigate the success of our trap weights attack. Since these defenses do not provide rigorous theoretical privacy guarantees but rather empirical protection, we recommend combining them with DP.
Hardware-Based Protection. Using protocols that rely on Trusted Execution Environments (TEE), e.g. prevents the central party from performing active manipulations. However, TEEs are prone to side-channel attacks . Hence, there is a remaining risk for the privacy of users’ data.
Local Averaging and Large Mini-Batches. Our results in Table V and Table XI highlight that calculating gradients over large mini-batches and local averaging reduces the fraction of data points that can be perfectly reconstructed. We, therefore, argue that users should perform gradient calculation on large mini-batches of local data points and average gradients over multiple mini-batches before sending them to the central party. This is, however, only possible if users have actual control on the local execution of the FL protocol and the execution is not inaccessibly encapsulated inside an application.
Choice of the Activation Function. Our trap weights are designed to exploit properties of the ReLU activation function, namely the fact that it yields zero-gradients for inputs at some neurons. Yielding zero-gradients is not unique to the ReLU activation. Other popular activation functions such as sigmoid and tanh have flat areas that also yield zero-gradients. Hence, by adapting the initialization of our trap weights to these functions’ properties, we could also achieve perfect extractability of individual data points with these functions. However, using activation functions, such as leaky ReLU, which propagate information on every input through each neuron, can prevent individual extractability of training data points.
Lossy Layers. Our evaluation in Section 7 highlights that the application of layers that compress the input data or cause information loss, such as pooling or dropout reduce fidelity of the extracted data. Therefore, relying on architectures that have aggressive compression and/or dropout reduces the leakage of individual user data to the central party.
Given that in FL, the central party is in charge of instantiating the shared model (with its hyperparameters, such as the activation function), users can only rely on additional protection through the model itself if this central party is trusted. An untrusted central party, in contrast, has incentives to choose model architectures and hyperparameters that facilitate data extraction.
Discussion and Future Directions
In this section, we first discuss the detectability of our adversarial initialization and integration of our attack in the training process of the shared model. We then analyze the potential and capabilities of adversarial weight initialization for future privacy attacks. Finally, we argue that dedicated privacy-protection should be implemented as a default option into FL protocols to prevent accidental or malicious privacy leakage.
To detect the presence of our trap weights, the users can apply one of the following two strategies: (1) analyzing the weights of the shared model in one or multiple iterations over the FL protocol, or (2) analyzing the behavior of the shared model on their data.
Analyzing Model Weights. Assuming the user has access to the model only in one iteration of the protocolGiven that in practical deployments of FL, , an individual user will be sampled for participation very rarely., they can run a detection method that aims at deliberately looking for characteristic elements of our trap weights, such as a normal distribution with high standard deviation, or the presence of higher absolute values for negative components than positive components in the first fully-connected layer’s weight matrix. Figure 3 shows that even when initialized with a uniform distribution and relatively low deviation, model weights after several epochs of training resemble more a normal distribution and exhibit a larger standard deviation, making the former characteristic of our trap weights an unreliable attack detector. When it comes to the magnitude of positive and negative components, Figure 4 shows that for close to one, the distribution of the model weights still resembles a standard normal. However, the more deviates from one, the more the distribution of weights deviates from a standard normal distribution. Yet, without knowledge of the prior training procedure and the other users’ data, we argue that a target user can still not determine with certainty whether the received model weights are the result of the prior training or of a manipulation .
Having access to the shared model over multiple iterations of the FL protocol additionally enables users to compare the received shared model’s parameters to the parameters from previous FL iterations. Therefore, the success of detection boils down to the following question: Can a local user tell that a given set of model parameters came from legitimate updates of other local users? We argue that this is not possible, even for non-FL setups where the entire training procedure is transparent. The stochastic nature of training algorithms, combined with the non-determinism of modern hardware, makes it difficult to reproduce training runs . Because of this reproducibility error, an attacker can assemble a mini-batch of natural data points that produce any desired gradient update . In other words, given two different sets of model weights, the user cannot tell if the gradient descent step between these weights was a result of a legitimate optimization step. This is exacerbated in FL because the data of any given user is invisible to other users, further complicating the verification of gradient descent integrity.
Analyzing Model Performance. In addition to analyzing the received model weights for detection of our attack, a user can also evaluate the functionality of the shared model. We observe that, for the vast majority of classification tasks, the model’s loss across training data points significantly reduces after just a few iterations. Thus, after the initial training iterations, FL users would expect to encounter low loss values for their own examples. However, research has shown that, in particular for users whose data stems from the tails of the data distribution, FL does not necessarily lead to an improvement of model accuracy on their data . Therefore, detection mechanisms that rely on analyzing the convergence of the model’s accuracy over multiple FL iterations are also no reliable detectors of our manipulation.
2 Training Success and Model Performance
Increasing the utility of the shared model over the course of training in the FL protocol is important because the central party is expected to provide a well-performing model after several training iterations. Therefore, the central party in our attack leverages two main points over the course of the protocol. (1) Instead of aiming at reconstructing the user data over all communications, it sends the adversarially (re-)initialized model out only at a few communication rounds. In all other rounds, it sends out the actual shared model for training without adversarially re-initializing it. (2) Instead of targeting all users, the central party only targets a subset of users, and send out an adversarially initialized model to them, and the continuously trained shared model to all other users. The central party can even combine both strategies by sending out an adversarially initialized model only to a subset of users in a few iterations.
3 The Power of Weight Initialization
In general, even outside of the FL context, our attack shows that controlling and manipulating the weights of neural networks opens a new attack surface against ML. We argue that weight manipulation could be used to design further privacy attacks outside of the FL context. Our adversarial initialization of convolutional and fully-connected model layers is able to transmit input data points to any subsequent layer in the model, practically modulating data perfectly over them. Additionally by setting our trap weights, we can increase the leakage of individual training data points from model gradients. Therefore, the trap weights basically create a simple if-else logic based on and relations between weighted inputs to model neurons. Future work could investigate whether the weights could also be set in order to implement more complex logical structures and if-else cases depending on the input. Based on these, it might be possible to craft hybrid attacks that first initialize the model weights and then use that to later extract information, for example, on membership of individual data points, or these data points’ sensitive attributes. Note that adversarially setting weights also does not need to be limited to initializing the model weights. Instead, given an already initialized (and trained model), it might be possible to craft additional training data that leads to the weights taking the adversarial values that an attacker wants.
4 Using Dedicated Privacy Protection in FL
FL was originally designed as an alternative to centralized ML in which no large datasets would have to be moved from users to a central party in order to train an ML model on the joint data. The approach does not only reduce communication costs but also spares the central party from having to build up the infrastructure by outsourcing training and data storage costs to the users. Indeed, FL is more communication cost effective since the data itself is not shared directly.
Attacks like our trap weights highlight, however, that the protocol does not guarantee protection for the individual users’ private training data. This is not surprising since nothing in the design the FL protocol protects against leakage of private information. Without dedicated privacy-protection, the central party even has an upper hand over how much data a local model will leak, as demonstrated in our work. However, FL is still often marketed as a data-minimizing technology. Our work highlights that such marketing is misleading since in order to deploy FL as a privacy-technology, it is necessary to implement dedicated additional protection methods, such as the ones discussed in Section 8.We argue that, to prevent malicious or accidental leakage in FL, these protection methods should be implemented in FL as a default when deploying the protocol to actual users. That is, vanilla federated learning does not provide privacy advantages for users—unless it is combined with additional defense methods, such as DP learning.
Conclusion
In this work, we presented a new privacy attack against FL that is based on an active attacker who holds the ability to maliciously manipulate the shared model and its weights. Our attack allows for perfect reconstruction of a significant portion of the users’ private training data. Even for very high-dimensional complex datasets, such as ImageNet, we are able to perfectly extract roughly 50% of the individual data points from mini-batches of sizes as large as 100. The extraction is computationally highly efficient and even allows to perfectly extract individual training data points from data mini-batches containing all data points from one single class.
Our attack underscores the deficiency of the “data never leaves the device” approach to preserving privacy. For FL to have a chance of truly preserving privacy, it must incorporate appropriate mitigations against our attack. Those either have expensive overheads or are tailored for this specific attack (see Section 8).
Acknowledgments
We would like to acknowledge our sponsors, who support our research with financial and in-kind contributions: Amazon, Apple, CIFAR through the Canada CIFAR AI Chair, DARPA through the GARD project, Intel, Meta, NFRF through an Exploration grant, NSERC through the COHESA Strategic Alliance, the Ontario Early Researcher Award, and the Sloan Foundation. Resources used in preparing this research were provided, in part, by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute.
References
Appendix A Extended Related Work on Passive Data Reconstruction Attacks
Table VIII summarizes these data reconstruction attacks (described below) and compares them with our attack.
Class-wise Representation Reconstruction Attacks. Hitaj et al. were the first to propose a GAN-based data reconstruction attack, called DMU-GAN. The attacker must know the dataset’s classes, and the reconstructed data points are generic representations of class-wise properties rather than individual user data points or classes. Wang et al. suggested mGAN-AI, which extends DMU-GAN’s reconstruction attack to per-user class-wise representations, but still does not extract individual data points. Additionally, both methods require access to data from the same distribution as the users’ data. observes that class-wise representations are embedded in model updates even without the need to reconstruct them using a GAN, and suggest defenses.
Appendix B Generalization of Data Extraction Attack to Convolutional Neural Networks
So far, both the passive and active attacks we described are tailored to extracting data from the gradients computed to update a fully-connected layer. However, modern neural network architectures often rely on convolutional layers to model image and text data alike. It is difficult to directly apply our attack strategy to these convolutional layers because they rely on the weight sharing principle: to decrease the effective number of parameters that need to be trained, the same weight values are applied to multiple locations of the image to extract patterns regardless of their location in the image. In this section, we thus propose a second instantiation of our adversarial weight initialization strategy that generalizes extraction attacks to convolutional neural networks (CNNs).
Our solution reduces networks with convolutional layers to the setting we previously considered with fully-connected neural networks. To do so, we observe that a CNN typically composes a few convolutional layers with fully-connected layers. We thus initialize the weights of the convolutional layers such that they transmit the model input unaltered up to the fully-connected layers of the model architecture.
There are two important requirements for our approach to transmitting, or forwarding, model inputs through convolutional layers. The first is to make sure that no feature of the input data is lost. This requires having at least as many parameters at every convolutional layer as the number of input features. The second is to make sure that different features do not get overlaid. We explain how to ensure this next.
Two Dimensional Input. In general, preserving input size over a convolutional layer can be achieved through an adequate combination of padding, stride, and filter sizes. Specifically, we use stride one and an adequate zero-padding to preserve the size of the layer input. In order to transmit the input features, we create a filter with uneven dimensions , where , and we initialize it with zero everywhere apart from the element in the middle which we set to one. For a two dimensional input (e.g. a grey-scale image), the described filter perfectly transmits the information to the next layer and creates a feature map that exactly replicates the input. See Figure 6(a) for this adversarially initialized filter.
Three Dimensional Input. Some input data to CNNs is distributed over several input channels, such as color images, that consist of three channels. At every layer, we, therefore need three adversarially initialized convolutional filters to ”transmit a copy” of the input channels. A standard architecture can have many more filters per layer, which can, in the case of our attack be randomly initialized since they will be ignored by the attacker. Assume now that the original input features at the current layer are distributed over of the total many feature maps. For example, in the first model layer, corresponds to the number of color channels required to encode the image. In subsequent layers, the remaining many feature maps contain random noise, introduced by random filters that do not transmit the input features (e.g. Filter 1 in Figure 6(b)). We denote the indices of the feature maps where the input features are located by . We then need many filters, initialized as described above to transmit the information to the next layer. The filters differ from each other only by the placement of the matrix that contains the one element. This placement must correspond to different indices in . See Figure 6(b) for a visualization of this setting. Note that the placement of the feature-transmitting filters at layer will determine the indices of the feature maps that are input to the next layer.
In the last convolutional layer before the fully-connected layer that we want to transmit the input to, the filters containing noise should be initialized such that they yield negative input to the ReLU function. Thereby, the output of the last convolutional layer becomes zero everywhere apart from the feature maps produced by the filters transmitting input data features. The flattened output then serves as input to the fully-connected layer, and reconstruction can be conducted as described in Section 6.
B.2 Reducing Input-Size
Two Dimensional Input. The reduction of size in convolutional layers can be achieved by increasing the stride. However, thereby, the number of features in the next feature map is reduced such that this feature map cannot accommodate all features from the previous layer. To overcome this, we propose distributing the features of one input feature map over several feature maps in the following layer. Figure 7(a) depicts this approach for two-dimensional inputs. Note that the stride is set to the dimensions of the convolutional filters to prevent features from overlapping in the following layer. Additionally, to transmit all the features, the dimensions of the filters and need to be integer dividers of the previous feature map’s dimensions. Finally, in total, for each layer that reduces the size of the input by a factor , we require many filters to transmit every feature from one input feature map. Hence, assuming that at layer the original features are distributed over many input feature maps, we require many filters to transmit all original input features.
Three Dimensional Input. The same approach as for the two dimensional input can be extended to the case with three input dimensions. The approach is visualized in Figure 7(b). For improved visualization, we do not present the feature maps in the layer’s output which contain only noise. Again, in both the two and three dimensional case, in the last convolutional layer before flattening, the noise filters should produce negative input to the ReLU function. This enables only extracting the original input features and no noise from the following fully-connected layer.
B.3 Reducing Detectability
In principle, our adversarial weight initialization for CNNs only requires the number of filters per layer that actually transmit the features. However, using only a small number of filters, e.g. one as in the case of the size-preserving adversarial convolutional filters, leads to models architectures that deviate strongly from standard architectures. Therefore, we propose using a standard number of convolutional filters in every layer and initializing the filters that are not used to transmit features at random. Additionally, to prevent the simple detection strategy which relies on probing after every convolutional layer whether its input is equal to its output, one can replace the ones in the adversarially initialized convolutional filters by other positive constants. Data extraction at the fully-connected layer then yields data points where features of the original input data are scaled by (multiple different) factors. By applying the inverse of the factors encoded in the model weights this scaling can then be reverted. As a consequence, the rescaled extracted data points still perfectly correspond to the input data.
Appendix C Additional Material
The following table describes the model architectures both for the FC-NNs and CNNs use throughout the paper. Note that the our method could also be applied to much larger CNNs with more layers: in fact, as long as each layer contains as many parameters as the data holds input features, our approach is applicable.
Appendix D Additional Experimental Results
This section presents additional experimental results.
Figure 8 shows extraction from a randomly initialized FC-NN with architecture presented in Table IX.
D.2 Trap Weights and Active Extraction
Figures 11 and 13 depict the extracted data points for MNIST and ImageNet, respectively.
We, furthermore, study partial extractablity, i.e., the case when a data point is not individually extractable, but still leaks meaningful private information about a training data point. Partial leakage occurs when an extracted gradient represents the overlay of only a few data points. In this case, the individual signal of each data point is still distinguishable, see for example the first data point in the third row of Figure 8. We plot in Figure 10 by how many data point each of the neurons with our trap weight initialization gets activated. This corresponds to the number of data points that will be present in the overlay of the respective gradients. We can see that nearly as many neurons get activated by two data points as by one data point (i.e. perfect extractability). In general, with our trap weights, neurons get activated by small numbers of data points. This indicates that the central party can still extract meaningful partial information on many data points, also if these are not perfectly extractable. Results an average over five runs with different trap weight initializations for a mini-batch of 100 data points from the ImageNet dataset.
To provide additional insights on data points that can and cannot be individually extracted, in Figure 14 which data points can and cannot be extracted. For the individually extractable data points, we, furthermore, depict how often each of them is individually extractable, i.e. for how many neurons this data point is the only one activating it. We see that our trap weights first amplify natural leakage i.e., data points that are extractable from random weights are usually also extractable with our trap weight and our trap weights make other data points extractable. Second, our trap weights yield redundancy, i.e. data points are extractable multiple times from different weight rows’ gradients.
We depict extraction success for local averaging over multiple mini-batches in Table XI.
We also study the non-IID setup where users hold data from a single class, different from other users in the protocol. We present the extraction-recall and extraction-precision per class on the CIFAR10 dataset in Table XII.
Finally, we study how an attacker without any prior knowledge can tune . One way to proceed is that the attacker does not adversarially initializes the model in the first FL iteration. It then extracts data points from the gradients, which are not necessarily individually extracted data points. From these data points, the attacker keeps the one in a valid image input range with features in range , and uses these data points for fine-tuning . We depict the resulting data points for MNIST, CIFAR10, and ImageNet in Figure 9 and show extraction success for different on 100 such data point in Table XIII.
D.3 Extraction under Lossy Layers
We also study the effect of ”lossy” layers, such as dropout and pooling on our data extraction success. Therefore, we rely on the following architecture proposed by for FL, see Table XIV.
Figures 15 and 16 and Figures 17 and 18 show individual effects of dropout and pooling layers on a reconstructions for mini-batches of size 1 and 20 respectively. We evaluated different dropout rates . Note that the second dropout layer does not have a significant impact on the success of our reconstruction since we extract from the first fully-connected model layer before information can get lost due to the second dropout. To evaluate dropout without pooling, we remove the MaxPool layer, and to evaluate pooling without dropout, we set the dropout rate to . Although existence of non-invertible components compromises overall reconstruction fidelity, we observe it is often possible to still recognise individual data points.
D.4 TensorFlow Federated
We adapted the implementation of FedAvg provided by the developers to pass each individual gradient update through our reconstruction function. Note that the whole change took only minutes of work and required minimal code changes, such that they could easily be implemented by a dishonest central party. We pre-generated our adversarial initialized shared models with scaling factor and and 1000 neurons at the first fully-connected layer, and passed them to the users. The aggregator then collects the gradients and performs reconstruction.
We find that our attack works consistently well against commonly used FL benchmarks integrated into the library. Over 50 users, for EMNIST , our trap weights yield extraction-recall and extraction-precision, versus , and in the non-adversarial baseline. For CIFAR100 we get extraction-recall and extraction-precision, versus , and in the baseline. These results are comparable to the ones reported in the previous experiments, and confirm that our attack is practical.