iDLG: Improved Deep Leakage from Gradients
Bo Zhao, Konda Reddy Mopuri, Hakan Bilen
Introduction
In multi-node distributed learning systems such as Collaborative Learning shokri2015privacy ; song2018collaborative ; melis2018inference and Federated Learning konevcny2016federated ; mcmahan2017federated ; li2019federated , it is widely believed that sharing gradients between nodes will not leak the private training data. In the popular setup, all the individual participants aim to learn a shared model in a centralized or decentralized manner. They would share the individual gradients and update the model parameters with the aggregated gradients. In these frameworks, it is a common practice to share only the gradients in order protect the proprietary data. However, recent work by Zhu et al., “Deep Leakage from Gradient” (DLG) zhu19deep showed the possibility to steal the private training data from the shared gradients of other participants.
The main idea of DLG is to generate dummy data and corresponding labels via matching the dummy gradients to the shared gradients. Specifically, they start with randomly initialized the dummy data and labels. From there, they compute dummy gradients over the current shared model in the distributed setup. Via minimizing the difference between dummy gradients and the shared real gradients, they iteratively update the dummy data and labels simultaneously. Although DLG works, we find that it is not able to reliably extract the ground-truth labels or generate good quality dummy data.
In this paper, we propose a simple but definitely valid approach to extract the ground-truth labels from the shared gradients. By derivation, we demonstrate that the gradient of the classification (cross-entropy) loss w.r.t. the correct label activation (in the output layer) lies in , while those of other labels lie in . Hence, the signs of gradients w.r.t. correct and wrong labels are opposite. When the gradients w.r.t. the outputs (logits) are not accessible, we show that the gradients w.r.t. the last-layer weights (between the output layer and the layer in front of it) also follow this rule. With this rule, we can identify the ground-truth labels based on the shared gradients. In other words, the ground-truth labels are definitely leaked by sharing gradients of a Neural Network (NN) trained with cross-entropy loss. This enables us to always extract the ground-truth labels and significantly simplify the objective of DLG zhu19deep in order to extract good-quality data. Hence, we name our approach, Improved DLG (iDLG). The main contributions of our work includes:
By revealing the relationship between labels and signs of gradients, we present an analytical procedure to extract the ground-truth labels from the shared gradients with accuracy, which facilitates the data extraction with better fidelity.
We empirically demonstrate the advantages of iDLG over DLG zhu19deep via comparing the accuracy of extracted labels and the fidelity of extracted data on three datasets.
The rest of the paper is organised as follows: Section 2 presents the analytical procedure to extract the ground-truth labels from the shared gradients and the proposed iDLG method. Section 3 demonstrates the advantages of iDLG over DLG through the experimental evaluation, and Section 4 concludes the paper with discussion.
Methodology
Recent work by Zhu et al. zhu19deep presents an approach (DLG) to steal the proprietary data protected by the participants in distributed learning from the shared gradients. In their method, they attempt to generate the dummy data and corresponding labels via a gradient matching objective. However, in practice, it is observed that their method generates wrong labels frequently. In this work, we present an analytical approach to extract the ground-truth labels from the shared gradients, then we can extract the data more effectively based on correct labels. Hence, we name our approach, improved Deep Leakage from Gradients (iDLG). In this section, we first present the procedure to extract the ground-truth labels. Then, we show the iDLG method based on the extracted labels.
Let us consider the classification scenario, where the NN model is generally trained with cross-entropy loss over one-hot labels, which is defined as
where is the input datum, is the corresponding ground-truth label. is the outputs (logits), and denotes the score (confidence) predicted for the class. Then, the gradients of the loss w.r.t. each of the outputs is
As the probability , we have when and when . Hence, we can identify the ground-truth label as the index of the output that has the negative gradient.
However, we may not be able to access the gradients w.r.t. the outputs , as they are not included in the shared gradients which are the derivatives w.r.t. the weights of the model . We find that the gradient vector w.r.t. the weights connected to the logit in the output layer can be written as
where the network has layers, is the output layer activations, is the bias parameter, and . As the activation vector is independent of the class (logit) index , we can easily identify the ground-truth label according to the sign of which is different from others. Therefore, the ground-truth label is predicted as
When the non-negative activation function, e.g. ReLU and Sigmoid, is used, the signs of and are the same. Hence, we can simply identify the ground-truth label whose corresponding is negative. With this rule, it is easy to identify the ground-truth label of the private training datum from shared gradients . Note that this rule is independent of the model architectures and parameters. In other words, this holds for any network at any training stage from any randomly initialized parameters.
2 Improved DLG (iDLG)
Based on the extracted ground-truth labels, we propose the improved DLG (iDLG) which is more stable and efficient to optimize. The algorithm is illustrated in Algorithm 1. The iDLG procedure starts with the differentiable learning model with the model parameters , and the gradients calculated based on private training datum . The first step is to extract the ground-truth label from the shared gradients as in eq (4).
Then, we randomly initialize the dummy datum . We calculate the dummy gradients based on the dummy datum and the extracted label . The training objective is to match the dummy gradients with the shared gradients, i.e., to minimize
Based on this training objective, we update the dummy datum by gradient descent
for training iterations, where is the learning rate.
Experiments
In this section, we empirically demonstrate the advantages of our (iDLG) method over DLG zhu19deep . We perform experiments on the classification task over three datasets: MNIST lecun1998gradient , CIFAR- krizhevsky2009learning , and LFW huang2008labeled with , , and categories respectively. Following the settings in zhu19deep , we use the randomly initialized LeNet for all experiments. L-BFGS liu1989limited with learning rate is used as the optimizer. For fast training, we resize all images in LFW to .
For DLG zhu19deep , as described by the authors, we start the procedure with the randomly initialized dummy data and outputs , then iteratively update them to minimize the gradient matching objective. For both two algorithms, we perform the optimization for iterations, and evaluate the performance in terms of (i) the accuracy of the extracted labels , and (ii) the fidelity of the extracted data . We run all experiments for times with randomly initialized networks and report the mean values. The code has been released on GitHubhttps://github.com/PatrickZH/Improved-Deep-Leakage-from-Gradients.
Table 1 shows the accuracy of the two methods to recover the ground-truth labels. It is clear that iDLG always extracts the correct label as opposed to DLG which extracts wrong labels many times. Specifically, the accuracy of DLG on MNIST, CIFAR-100 and LFW is 89.9%, 83.3% and 79.1% respectively, which shows that DLG suffers more on harder tasks.
2 The Fidelity of Extracted Data
In this subsection, we compare the fidelity of two data extraction methods (DLG and iDLG) by calculating the MSE (mean square error) between the dummy and original data. We vary the (MSE) threshold of good fidelity. Figure 1 shows the fidelity comparison of two methods under different thresholds over three datasets.
The plots show the percentage of extracting (or generating) data with good fidelity. The x-axis indicates the (MSE) threshold for good fidelity. For example, means that we consider it good fidelity when the MSE between dummy and original data is less than . From left to right, the threshold decreases and the fidelity requirement improves. Obviously, the proposed iDLG consistently outperforms DLG in recovering data with significant margin on three tasks. The advantage of iDLG is remarkable on the hard task of LFW.
Figure 2 gives an example of the training process of DLG (left) and iDLG (right) on LFW face dataset. The first image is the (original) private training image. The followings are the extracted images in different training iterations. It is clear that the training of iDLG is easier to converge. iDLG needs only 90 training iterations to get the similar performance which requires DLG to train for 200 iterations.
Discussion and Conclusion
In this paper, we present an effective approach to steal the data and the corresponding labels from the shared gradients in a distributed training scenario. Particularly, we analytically illustrate the relationship between the labels and the signs of corresponding gradients. Based on this, our approach can extract the ground-truth labels with accuracy which facilitates the data extraction with increased fidelity. Currently, our method works with a simplified scenario of sharing gradients of every datum. In other words, iDLG can identify the ground-truth labels only if gradients w.r.t. every sample in a training batch are provided.