Systematic Evaluation of Privacy Risks of Machine Learning Models
Liwei Song, Prateek Mittal
Introduction
A recent thread of research has shown that machine learning (ML) models memorize sensitive information of training data, indicating serious privacy risks . In this paper, we focus on the membership inference attack, where the adversary aims to guess whether an input sample was used to train the target machine learning model or not . It poses a severe privacy risk as the membership can reveal an individual’s sensitive information . For example, participation in a hospital’s health analytic training set means that an individual was once a patient in that hospital. As membership inference attacks expose the privacy risks of an individual user participating in the training data, they serve as a valuable tool to quantify the privacy provided by differential privacy implementations and to help to guide the selection of privacy parameters in the broader class of statistical privacy frameworks . Shokri et al. conducted membership inference attacks against machine learning classifiers in the black-box manner, where the adversary only observes prediction outputs of the target model. They formalize the attack as a classification problem and train dedicated neural network (NN) classifiers to distinguish between training members and non-members. The research community has since extended the idea of membership inference attacks to generative models , to differentially private models , to decentralized settings where the models are trained across multiple users without sharing their data , and to white-box settings where the adversary also has the access to the target model’s architecture and weights .
To mitigate such privacy risks, several defenses against membership inference attacks have been proposed. Nasr et al. propose to include membership inference attacks during the training process: they train the target model to simultaneously achieve correct predictions and low membership inference attack accuracy by adding the inference attack as an adversarial regularization term. Jia et al. propose a defense method called MemGuard which does not require retraining the model: the model prediction outputs are obfuscated with noisy perturbations such that the adversary cannot distinguish between members and non-members based on the perturbed outputs. Both papers show that their defenses greatly mitigate membership inference privacy risks, resulting in attack performance that is close to random guessing.
In this paper, we critically examine how previous work has evaluated the membership inference privacy risks of machine learning models, and demonstrate two key limitations that lead to a severe underestimation of privacy risks. First, many prior papers, particularly those proposing defense methods , solely rely on training custom NN classifiers to perform membership inference attacks. These NN attack classifiers may underestimate privacy risks due to inappropriate settings of hyperparameters such as number of hidden layers and learning rate. Second, existing evaluations only focus on aggregate notions of privacy risks faced by all data samples, lacking a fine-grained understanding of privacy risks faced by individual samples.
To overcome the limitation of reliance on NN-based attacks, we propose to use a suite of alternative existing and novel non-NN based attack methods to benchmark the membership inference privacy risks. These benchmark attack methods make inference decisions based on computing custom metrics on the predictions of the target model. Compared to NN-based attacks, our proposed benchmark attacks are easy to implement without hyperparameter tuning. We only need to set the threshold values using the shadow-training technique . We also show that rigorously benchmarking defense mechanisms requires a careful consideration of strategic adversaries aware of the defense mechanism, as well as alternative baselines that trade-off accuracy of the target machine learning model with privacy risks. With our proposed benchmark attacks, we indeed find that that existing membership inference defense methods are not as effective as previously reported. As shown in Table 1, the adversary can still perform membership inference attacks on models defended by adversarial regularization and MemGuard with an accuracy ranging from to , instead of the reported accuracy around , which is the accuracy of random guessing. Therefore, we argue that these non-NN based attacks should supplement existing NN based attacks to effectively measure the privacy risks.
To overcome the limitation of a lack of understanding of fine-grained privacy risks in existing works, we propose a new metric called the privacy risk score, that represents an individual sample’s probability of being a member in the target model’s training set. Figure 1 shows the cumulative distributions of privacy risk scores on target undefended models trained on Purchase100, Location30, and CIFAR100 datasets respectively. We can see that the privacy risk faced by individual training samples is heterogeneous. By utilizing the privacy risk score, an adversary can perform membership inference attacks with high confidence: an input sample is inferred as a member if and only if its privacy risk score is higher than a certain threshold value. Overall, we recommend that our per-sample privacy risk analysis should be used in conjunction with existing aggregate privacy analysis for an in-depth understanding of privacy risks of machine learning models. Conventional aggregate analysis provides an average perspective of privacy risks incurred by all samples, while privacy risk score provides a perspective on privacy risk from the viewpoint of an individual sample. The former provides an aggregate estimation of privacy risks, while the latter allows us to understand the heterogeneous distribution of privacy risks faced by individual samples and identify samples with high privacy risks. We summarize our contributions as follows:
We propose a suite of non-NN based attacks to benchmark target models’ privacy risks by improving existing attacks with class-specific threshold settings and designing a new inference attack based on a modified prediction entropy estimation in a manner that incorporates the ground truth class label. Furthermore, to rigorously evaluate the performance of membership inference defenses, we make recommendations for comparison with early stopping baseline and considering adaptive attackers with knowledge of defense mechanisms.
With our benchmark attacks, we find that two state-of-the-art defense approaches are not as effective as previously reported. Furthermore, we observe that the defense performance of adversarial regularization is no better than early stopping, and the evaluation of MemGuard lacks a consideration of adaptive adversaries. We also find that the existing white-box attacks have limited advantages over our benchmark attacks, which only need black-box access to the target model. We also show that our attacks with class-specific threshold settings strictly outperform attacks with class-independent thresholds, and our new inference attack based on modified prediction entropy strictly outperforms conventional prediction entropy based attack.
We propose to analyze privacy risks of machine learning models in a fine-grained manner by focusing on individual samples. We define a new metric called the privacy risk score, that estimates an individual sample’s probability of being in the target model’s training set.
We experimentally demonstrate the effectiveness of our new metric in being able to capture the likelihood of an individual sample being a training member. We also show how an adversary can exploit our metric to launch membership inference attacks on individual samples with high confidence. Finally we perform an in-depth investigation of our privacy risk score metric, and its correlations with model sensitivity, generalization error, and feature embeddings.
Our code is publicly available at https://github.com/inspire-group/membership-inference-evaluation for the purpose of reproducible research. Furthermore, our evaluation mechanisms have also been integrated in Google’s TensorFlow Privacy library.
Background and Related Work
In this section, we first briefly introduce machine learning basics and notations. Next, we present existing membership inference attacks, including black-box attacks and white-box attacks. Finally, we discuss two state-of-the-art defense methods: adversarial regularization and MemGuard .
Let be a machine learning model with input features and output classes, parameterized by weights . For an example with the input feature and the ground truth label , the model outputs a prediction vector over all class labels with , and the final classification result will be the label with the largest prediction probability .
Given a training set , the model weights are optimized by minimizing the prediction loss over all training examples.
2 Membership inference attacks
For a target machine learning model, membership inference attacks aim to determine whether a given data point was used to train the model or not . The attack poses a serious privacy risk to the individuals whose data is used for model training, for example in the setting of health analytics.
Shokri et al. investigated the membership inference attacks against machine learning models in the black-box setting. For an input sample to the target model , the adversary only observes the prediction output and infers if belongs to the model’s training set . To distinguish between target model’s predictions on members and non-members, the adversary learns an attack model using the shadow training technique: (1) the adversary first trains multiple shadow models to simulate the behavior of the target model; (2) based on shadow models’ outputs on their own training and test examples, the adversary obtains a labeled (member vs non-member) dataset, and (3) finally trains multiple neural network (NN) classifiers, one for each class label, to perform inference attacks against the target model.
Salem et al. show that even with only a single shadow model, membership inference attacks are still quite successful. Furthermore, in the case where the adversary knows a subset of target model’s training set and test set, the attack classifier can be directly trained with target model’s predictions on those known samples, and then tested on unknown training and test sample . Nasr et al. redesign the attack by using one-hot encoded class labels as part of input features and training a single NN attack classifier for all class labels.
Besides membership inference attacks that rely on training NN classifiers, there are non-NN based attack methods that make inference decisions based on computing custom metrics on the predictions of the target model. Leino et al. suggest using the metric of prediction correctness as a sign of being a member or not. Yeom et al. and Song et al. find that the metric of prediction confidence of correct label can be compared with a certain threshold value to achieve similar attack performance as NN-based attacks. Shokri et al. show a large divergence between prediction entropy distributions over training data and test data, although this metric was not explicitly used for attacks.
Despite the existence of such non-NN based attacks, many research papers still only train NN attack classifiers to evaluate target models’ privacy risks. We find that this can lead to severe underestimation of privacy risks by re-evaluating the same target models with non-NN based attacks. Furthermore, we improve existing non-NN based attacks by setting different threshold values for different class labels, building upon the motivation of separated attack classifiers for each class label by Shokri et al. . We also propose a new inference attack method by considering ground truth label when evaluating prediction uncertainty.
2.2 White-box membership inference attacks
3 Defenses against membership inference attacks
To mitigate the risks of membership inference attacks, several defense ideas have been proposed. norm regularization and dropout are standard techniques for reducing overfitting in machine learning. They are also shown to decrease privacy risks to some degree . However, target models can still be quite vulnerable after applying these techniques. Differential privacy can also be applied to ML models for provable risk mitigation , however, it induces significant accuracy drop for desired values of the privacy parameter . Two dedicated defenses, adversarial regularization and MemGuard , were recently proposed against membership inference attacks. Both defenses are reported to have the ability of decreasing the attack accuracy to around , which is the performance of random guessing. We explain their details below.
Nasr et al. propose to include the membership inference adversary with the NN-based attack into the training process itself to mitigate privacy risks. At each training step, the attack classifier is first updated to distinguish between training data (members) and validation data (non-members), and then the target classifier is updated to simultaneously minimize the prediction loss and mislead the attack classifier.
More specifically, to train the classifier with parameters in a manner that is resilient against membership inference attacks, Nasr et al. use another classifier with parameters to perform membership inference attacks. The attack classifier takes the target model’s prediction and the input label as input features and generate one single output , which is in the range . It infers the input sample as a member if the output is larger than 0.5, a non-member otherwise. At each training step, they first update the attack classifier by maximizing the membership inference gain over the training set and the validation set .
They further train the target classifier by minimizing both model prediction loss and membership inference gain over the training set .
where is a penalty parameter for the privacy risk. In this way, the target model is trained with an additional regularization term to defend against membership inference attacks.
3.2 MemGuard [20]
Jia et al. propose MemGuard as a defense method against membership inference attacks, which, different from Nasr et al. , does not need to modify the training process. Instead, given a pre-trained target model , they obfuscate its predictions with well-designed noises to confuse the membership inference classifier , without changing classification results.
The attack classifier is trained following the shadow-training technique , which takes the model prediction with the sample label , and outputs a score in the range for membership inference: if the output is larger than 0.5, the data sample is inferred as a member, and vice versa. The key question of how to add noise to can be formulated as the following optimization problem:
where the objective is to minimize the distance between original predictions and noisy predictions. The first constraint ensures the classification result does not change after adding noise, the second constraint ensures the attack classifier cannot determine whether the sample is a member or a non-member with the noisy predictions, and last two constraints ensure the noisy predictions are valid.
When evaluating the defense performance, both Nasr et al. and Jia et al. train NN classifiers for inference attacks. As shown in the following section, we find that their evaluations underestimate privacy risks. With our benchmark attacks, the adversary achieves significantly higher attack accuracy on defended models than previous estimates. We further find that the performance of adversarial regularization is no better than early stopping, and the evaluation of MemGuard lacks consideration of strategic adversaries.
Systematically Evaluating Membership Inference Privacy Risks
In this section, we first present a suite of non-NN based attacks to benchmark privacy risks, which only need to observe target model’s output predictions (i.e., black-box setting). Next, we provide two recommendations, comparison with early stopping and considering adaptive attacks, to rigorously measure the effectiveness of defense approaches. Finally, we present experiment results by re-evaluating target models in prior work with our proposed benchmark attacks.
We propose to use a suite of non-NN based attack methods to benchmark membership inference privacy risks of machine learning models. We call these attack methods “metric-based attacks” as they first measure the performance metrics of target model’s predictions, such as correctness, confidence, and entropy, and then compare those metrics with certain threshold values to infer whether the input sample is a member or a non-member . We improve existing metric-based attacks by setting different threshold values for different class labels of target models. Then we propose another new metric-based attack by considering ground truth label when evaluating prediction uncertainty. We denote the inference strategy as , which codes members as 1, and non-members as 0. Overall, we propose that existing NN based attacks should be supplemented with our metric-based attacks for systematically and rigorously evaluating privacy risks of ML models.
Inference attack based on prediction correctness Leino et al. observe that the membership inference attacks based on whether the input is classified correctly or not achieve comparable success as NN-based attack on target models with large generalization errors. The intuition is that the target model is trained to predict correctly on training data (members), which may not generalize well on test data (non-members). Thus, we can rely on the prediction correctness metric for membership inference. The adversary infers an input sample as a member if it is correctly predicted, a non-member otherwise.
where is the indicator function.
1.2 Improving existing attacks with class-dependent thresholds
Inference attack based on prediction confidence Yeom et al. and Song et al. show that the attack strategy of using a threshold on the prediction confidence results in similar attack accuracy as NN-based attacks. The intuition is that the target model is trained by minimizing prediction loss over training data, which means the prediction confidence of a training sample should be close to 1. On the other hand, the model is usually less confident in predictions on a test sample. Thus, we can rely on the metric of prediction confidence for membership inference. The adversary infers an input example as a member if its prediction confidence is larger than a preset threshold, a non-member otherwise.
Yeom et al. and Song et al. choose to use a single threshold for all class labels. We improve this method by setting different threshold values for different class labels . The reason is that the dataset may be unbalanced so that the target model indeed has different confidence levels for different class labels. Our experiments show that this class-dependent thresholding technique leads to better attack performance. The class-dependent threshold values are learned with the shadow-training technique : the adversary (1) first trains a shadow model to simulate the behavior of the target model; (2) then obtains the shadow model’s prediction confidence values on both shadow training and shadow test data; (3) finally leverages knowledge of membership labels (member vs non-member) of the shadow data to select the threshold value which achieves the highest accuracy in distinguishing between shadow training data and shadow test data with the class label based on Equation (6).
Inference attack based on prediction entropy Although there is no prior work using prediction entropy for membership inference attacks, Shokri et al. indeed present the difference of prediction entropy distributions between training and test data to explain why privacy risks exist. Salem et al. also mention the possibility of using prediction entropy for attacks. The intuition is that the target model is trained by minimizing the prediction loss over training data, which means the prediction output of a training sample should be close to a one-hot encoded vector and its prediction entropy should be close to 0. On the other hand, the target model usually has a larger prediction entropy on an unseen test sample. Therefore, we can rely on the metric of prediction entropy for membership inference. The adversary classifies an input example as a member if its prediction entropy is smaller than a preset threshold, a non-member otherwise.
Similar to the confidence-based attack, we propose to use the threshold values that are dependent on the class labels and are set with the shadow-training technique .
1.3 Our new inference attack based on modified prediction entropy
The attack based on prediction entropy has one serious issue: it does not contain any information about the ground truth label. In fact, both a correct classification with probability of 1 and a totally wrong classification with probability of 1 lead to zero prediction entropy values.
To resolve this issue, we design a new metric with following two properties to measure the model prediction uncertainty given the ground truth label: it should be (1) monotonically decreasing with the prediction probability of the correct label , and (2) monotonically increasing with the prediction probability of any incorrect label . Let denote the prediction probability for a certain label, the function used in conventional entropy computations is not a monotonic function. As a contrast, is a monotonically decreasing function, and is a monotonically increasing function. Therefore, we propose the modified prediction entropy metric, computed as follows.
In this way, a correct classification with probability of 1 leads to modified entropy of 0, while a wrong classification with probability of 1 leads to modified entropy of infinity.
Now, with the new metric of modified prediction entropy, the adversary classifies an input example as a member if its modified prediction entropy is smaller than a preset threshold, a non-member otherwise.
Similar to previous scenarios, we set different threshold values for different class labels, which are learned with the shadow training technique . Experiments show that the inference attack based on our modified prediction entropy is strictly superior to the inference attack based on prediction entropy.
2 Rigorously evaluating membership inference defenses
To evaluate the effectiveness of defenses against membership inference attacks, we make the following two recommendations, besides using our metric-based benchmark attacks.
During the training process, the target model’s parameters are updated following gradient descent methods, so the training error and test error usually get reduced gradually with an increasing number of training epochs. However, as the number of training epochs increases, the target model also becomes more vulnerable to membership inference attacks, due to increased memorization. We thus propose early stopping as a benchmark defense method, in which fewer training epochs are used in order to tradeoff a slight reduction in model accuracy with lower privacy risk.
We recommend that whenever a defense method is proposed in the literature that reduces the threat of membership inference attacks at the cost of degradation in model accuracy, the performance of the defense method should be benchmarked against our early stopping approach. This is indeed the case for the defense method of adversarial regularization (AdvReg) . As shown in Figure 2, the defended Purchase100 classifier should be compared to the undefended model with fewer training steps and similar accuracy.
2.2 Adaptive attacks
There always exists an arms race between privacy attacks and defenses for machine learning models. When evaluating the defense performance, it is critical to put the adversary into the last step, i.e., the adversary knows the defense mechanism and performs adaptive attacks against the defended models. A perfect defense performance with non-adaptive attacks does not mean that the defense approach is effective .
Specifically for defenses proposed against membership inference attacks, we should consider that the adversary knows the defense mechanism such that he or she can train shadow models following the defense method. From these defended shadow models, the adversary then learns an attack classifier or sets threshold values for metric-based attacks, and finally performs attacks on the defended target model.
3 Experiment results
We first re-evaluate the effectiveness of two membership inference defenses , and then re-evaluate the white-box membership inference attacks proposed by Nasr et al. . Following prior work , we sample the input from either the target model’s training set or test set with an equal probability to maximize the uncertainty of membership inference attacks. Thus, the random guessing strategy results in a membership inference attack accuracy.
Purchase100 This dataset is based on Kaggle’s Acquire Valued Shoppers Challenge,https://www.kaggle.com/c/acquire-valued-shoppers-challenge which contains shopping records of several thousand individuals. We obtain a simplified and preprocessed purchase dataset provided by Shokri et al. . The dataset has 197,324 data samples with 600 binary features. Each feature corresponds to a product and represents whether the individual has purchased it or not. All data samples are clustered into 100 classes representing different purchase styles. The classification task is to predict the purchase style based on the 600 binary features. We follow Nasr et al. to use 10% data samples (19,732) to train a model.
Texas100 This dataset is based on the Hospital Discharge Data public use files with patients’ information released by the Texas Department of State Health Services.https://www.dshs.texas.gov/THCIC/Hospitals/Download.shtm Each data record contains the external causes of injury (e.g., suicide, drug misuse), the diagnosis (e.g., schizophrenia), the procedures the patient underwent (e.g., surgery) and some generic information (e.g., gender, age, race). We obtain a simplified and preprocessed Texas dataset provided by Shokri et al. . The classification task is to predict the patient’s main procedure based on the patient’s information. The dataset focuses on 100 most frequent procedures, resulting in 67,330 data samples with 6,170 binary features. Following previous papers , we use 10,000 data samples to train a model.
Location30 This dataset is based on Foursquare dataset,https://sites.google.com/site/yangdingqi/home/foursquare-dataset which contains location “check-in” records of several thousand individuals. We obtain a simplified and preprocessed Location dataset provided by Shokri et al. . The dataset contains 5,010 data samples with with 446 binary features. Each feature corresponds to a certain region or location type and represents whether the individual has visited the region/location or not. All data samples are clustered into 30 classes representing different geosocial types. The classification task is to predict the geosocial type based on the 466 binary features. Following Jia et al. , we use 1,000 data samples to train a model.
CIFAR100 This is a major benchmark dataset for image classification . It is composed of 3232 color images in 100 classes, with 600 images per class. For each class label, 500 images are used as training samples, and the remaining 100 images are used as test samples.
We choose these datasets for fair comparison with prior work . Since all datasets except CIFAR100 are binary datasets, we also provide attack results with more complex datasets in Appendix A, where our benchmark attacks achieve higher attack success than NN-based attacks.
3.2 Re-evaluating adversarial regularization [31]
We follow Nasr et al. to train both defended and undefended classifiers on Purchase100 and Texas100 datasets. For both datasets, the model architecture is a fully connected neural network with 4 hidden layers. The numbers of neurons for hidden layers are 1024, 512, 256, and 128, respectively. All hidden layers use hyperbolic tangent (Tanh) as the activation function. We note that the defense method of adversarial regularization incurs accuracy drop. After applying the defense, the test accuracy drops from to on the Purchase100 dataset, and from to on the Texas100 dataset. As we discuss in Section 3.2.1, to further evaluate the effectiveness of adversarial regularization , we also obtain models with early stopping by saving the undefended models in every training epoch and picking the saved epochs with similar accuracy performance as defended models. Table 2 presents the membership inference attack results.
From Table 2, we can see that the defended models are still vulnerable to membership inference attacks, indicating the necessity of our metric-based benchmark attacks. We achieve and attack accuracy on the defended Purchase100 classifier and the defended Texas100 classifier with our benchmark attacks, significantly larger than and as reported by Nasr et al. . Furthermore, on all models except the undefended Purchase100 classifier, the largest attack accuracy achieved by benchmark attacks is larger than that of NN based attacks used in Nasr et al. . Note that the defense method provides limited mitigation of privacy risks: it reduces attack accuracy from around to around on tested models. We also find that our new attack based on the modified entropy () always outperforms the conventional entropy based attack (). It is also very competitive among all benchmark attacks.
From Table 2, we also surprisingly find that adversarial regularization is no better than our early stopping benchmark method: with early stopping, the undefended Purchase100 classifier and the undefended Texas100 classifier have the attack accuracy of and , which are quite close to those of defended models. Therefore, when evaluating the effectiveness of a future defense mechanism that trades lower model accuracy for lower membership inference risk, we argue to compare the defended model to the naturally trained model with early stopping for a fair comparison. We emphasize that our early stopping baseline can be calibrated to achieve similar model accuracy as the defended model. In contrast, the adversarial regularization approach may have a model accuracy which is different from the defended model under consideration, and will thus not represent a fair comparison.
To show the attack improvement yielded by our class-dependent thresholding technique, we compare with metric-based attacks when the same threshold is applied to all class labels. Table 3 shows the results on Texas100 classifiers without defense, with AdvReg , and with early stopping. We can see that with the class-dependent thresholding technique, we increase the attack accuracy by 1% – 4%.
3.3 Re-evaluating MemGuard [20]
We follow Jia et al. to train classifiers on Location30 and Texas100 datasets. For both datasets, the model architecture is a fully connected neural network with 4 hidden layers. The numbers of neurons for hidden layers are 1024, 512, 256, and 128, respectively. All hidden layers use rectified linear unit (ReLU) as the activation function. MemGuard does not change the accuracy performance, so the comparison with early stopping benchmark is not applicable. Table 4 lists the attack accuracy on both undefended and defended models, with attack methods in Jia et al. and our metric-based benchmark attack methods. In fact, Jia et al. use 6 different NN attack classifiers to measure the privacy risks, and we pick the highest attack accuracy among them.
From Table 4, we again emphasize the necessity of our benchmark attacks by showing that the defended models still have high membership inference accuracy: on the defended Location30 classifier and on the defended Texas100 classifier, much larger than and reported by Jia et al. . We even achieve higher membership inference accuracy than attacks in Jia et al. on all models, except the undefended Location30 classifier. Note that the defense still works but to a limited degree: it reduces the attack accuracy by on the Location30 classifier and by on the Texas100 classifier. Similar to Section 3.3.2, our proposed modified-entropy based attack always achieves higher attack accuracy than the entropy based attack, and is very competitive among all benchmark attacks.
Next, we discuss why Jia et al. fail to achieve high membership inference accuracy for their defended models. We find that most of their attacks (4 out of 6) are non-adaptive attacks, where the adversary has no idea of the implemented defense, and thus the membership inference attacks are not successful. For the two adaptive attacks, Jia et al. do not put the adversary in the last step of the arms race between attacks and defenses. In their attacks, the adversary is aware that the model predictions will be perturbed with noises but does not know the exact algorithm of noise generation implemented by the defender. In their first adaptive attack, Jia et al. round the model predictions to be one decimal during the attack classifier’s inference to mitigate the effect of the perturbation. However, the attack performance is greatly degraded when the applied perturbation is large. In the second adaptive attack, Jia et al. train the attack classifier using the state-of-the-art robust training algorithm by Madry et al. , with the hope that noisy perturbation will not change the classification. However, the robust training algorithm has a very poor generalization property: the predictions on test points are still likely to be wrong after adding well-designed noises. For a thorough evaluation of the defense, we should consider that the attacker has the full knowledge of the defense mechanism, and he or she learns the attack model based on the defended shadow models.
3.4 Re-evaluating white-box membership inference attacks [32]
We have shown that previous work may underestimate the target models’ privacy risks, and the metric-based attacks with only black-box access can result in higher attack accuracy than NN based attacks for most models. Recently Nasr et al. demonstrated that a white-box membership inference adversary can perform stronger NN based attacks by using gradient with regard to model parameters. Next, we evaluate whether the advantage of white-box attacks still exists by using our metric-based black-box benchmark attacks.
We follow Nasr et al. to obtain classifiers on Purchase100, Texas100 and CIFAR100 datasets. The Purchase100 classifier and the Texas100 classifier are same as undefended classifiers in Section 3.3.2. The CIFAR100 classifier is a publicly available pre-trained model,https://github.com/bearpaw/pytorch-classification with the DenseNet architecture . Table 5 lists all attack results.
From Table 5, we can see that compared to the black-box metric-based attacks, the improvement of white-box membership inference attacks is limited. The attack accuracy of white-box membership inference adversary is only and higher than the attack accuracy achieved by our black-box benchmark attacks, on the Texas100 and the CIFAR100 classifiers. The white-box attack on the Purchase100 classifier still has increase in attack accuracy compared to black-box attacks. As a validation of our observations, we note that Shejwalkar and Houmansadr also report close membership inference attack accuracy between white-box attacks and black-box attacks in their recent work .
Fine-Grained Analysis on Privacy Risks
Prior work focuses on an aggregate evaluation of privacy risks by reporting overall attack accuracy or a precision-recall pair, which are averaged over all samples. However, the target machine learning model’s performance is usually varied across samples, which denotes the heterogeneity of samples’ privacy risks. Therefore, a fine-grained privacy risk analysis of individual samples is needed, with which we can understand the distribution of privacy risks over samples and identify which samples have high privacy risks.
In this section, we first define a metric called privacy risk score to quantitatively measure the privacy risks for each individual training member. Then we use this metric to experimentally measure fine-grained privacy risks of target models. Overall, we argue that existing aggregate privacy analysis of ML models should be supplemented with our fine-grained privacy analysis for a thorough evaluation of privacy risks.
For membership inference attacks, the privacy risk of a training member arises due to the distinguishability of its model prediction behavior with non-members. This motivates our definition of the privacy risk score as following.
The privacy risk score of an input sample for the target machine learning model is defined as the posterior probability that it is from the training set after observing the target model’s behavior over that sample denoted as , i.e.,
Based on Bayes’ theorem, we further compute the privacy risk score as following.
where stands for the test set. The observation depends on the adversary’s access to the target model: in the black-box membership inference attack , it is the model’s final output, i.e., ; in the white-box membership inference attacks , it also includes the model’s intermediate layers’ outputs and gradient information at all layers. Our proposed benchmark attacks only need black-box access to the target model, and most existing attack methods work in the black-box manner. Therefore, we focus on the black-box scenario for the computation of the privacy risk score in this paper and leave the discussion on white-box scenario as future work. In the black-box attack scenario, the privacy risk score can be expressed as
From Equation (12), we can see that the risk score depends on both prior probabilities , and conditional distributions , . For the prior probabilities, we follow previous work to assume that an example is sampled from either training set or test set with an equal probability, where the uncertainty of membership inference attacks is maximized. Note that the privacy risk score is naturally applicable to any prior probability scenario, and we present the results with different prior probabilities in Appendix B. With the equal probability assumption, we have
For the conditional distributions , , we empirically measure these values using shadow-training technique: (1) train a shadow model to simulate the behavior of the target model; (2) obtain the shadow model’s prediction outputs on shadow training and shadow test data; (3) empirically compute the conditional distributions on shadow training and shadow test data. Furthermore, as the class-dependent thresholding technique is shown to improve the attack success in Table 3, we compute the distribution of model prediction over training data in a class-dependent manner ( is computed in the same way).
Since we empirically measure the conditional distributions using the shadow model’s predictions over shadow data, the quality of measured distributions highly depends on the shadow model’s similarity to the target model and the size of shadow data. On the one hand, the size of shadow data is usually limited. Especially in our analysis where the distribution is computed in a class-dependent manner, for each class label , we may not have enough samplesIn our experiments, on average we have 197 samples per class for Purchase100 dataset; 100 samples per class for Texas100 dataset; 33 samples per class for Location30 dataset; and 500 per class for CIFAR100 dataset. to adequately estimate the multi-dimension distribution . On the other hand, in Section 3.3 we show that by only using the one-dimension prediction metric such as confidence and modified entropy, our proposed benchmark attacks in fact achieve comparable or even better success that NN-based attacks which leverage the whole prediction vector as features. Thus, we propose to further approximate the multi-dimension distribution in Equation (14) with the distribution of modified prediction entropy, since using modified entropy usually results in highest attack accuracy among all benchmark attacks.In most cases, both modified entropy based attack and confidence based attack give best attack performance. However, for undefended Location30 and Texas100 classifiers in Table 4, the modified entropy based attack achieves significantly higher attack accuracy.
We also approximate in the same way. By plugging Equation (15) into Equation (13), we can get the privacy risk score for a certain sample.
2 Experiment results
In our experiments, we first validate that our proposed privacy risk score really captures the probability of being a member. Next, we compare the distributions of training samples’ privacy risk scores for target models without defense and with defenses . We then demonstrate how to use privacy risk scores to perform membership inference attacks with high confidence. Finally, we perform an in-depth investigation of individual samples’ privacy risk scores by correlating them with model sensitivity, generalization errors, and feature embeddings. To have enough diversity of data and models, and to further evaluate defense methods, we perform experiments on 3 Purchase100 classifiers (without defense, with AdvReg , and with early stopping) and 2 Texas100 classifiers (without defense, and with MemGuard ). Both Purchase100 classifiers and Texas classifiers use fully connected neural networks with 4 hidden layers, and the numbers of neurons for hidden layers are 1024, 512, 256, and 128, respectively. Purchase100 classifiers use Tanh as the activation function , and Texas100 classifiers use ReLU as the activation function .
Before presenting the detailed results for privacy risk score, we first validate its effectiveness here. For the target machine learning model, we first compute the privacy risk scores following the method in Section 4.1 for all training and test samples. Next we divide the entire range of privacy risk scores into multiple bins, and count the number of training points () and the number of test points () in each bin. Then we compute the fraction of training points () in each bin, which indicates the real likelihood of a sample being a member (y axis of the last column in Figure 3(a)). If the privacy risk score truly corresponds to the probability that a sample is from a target model’s training set, then we expect the actual values of privacy risk scores and fraction of training points in each bin to closely track with each other.
As a baseline to compare with, we also consider using NN based attacks to estimate privacy risks of individual samples. Prior papers suggest using the attack classifier’s prediction to measure the input’s privacy risk . The attack classifier has only one output, which is within and can serve as a proxy to estimate the probability of being a member. Following same steps as above, we compute the real probability of being a member and the average outputs of the attack classifier. Specifically, we follow Nasr et al. to train the attack classifier by using the target model’s predictions and one-hot encoded input labels as features.
Figure 3 shows the distribution of training samples’ privacy risk scores (top row) and attack classifier’s outputs on training data (bottom row) for Purchase100 classifiers without defense, with AdvReg , and with early stopping. We also compare the privacy risk score and attack classifier’s output with the real probability of being a member, as shown in the last column of Figure 3 where the ideal case is used to check the effectiveness of metrics. We can see that our proposed privacy risk score closely aligns with the actual probability of being a member: the privacy risk score curves for all three models are quite close to the line of the ideal case. On the other hand, the attack classifiers’ outputs fail to capture the membership probability. This is because the NN classifiers are trained to minimize the loss, i.e., the output of a member should be close to 1 while the output of a non-member should be close to 0. With this training goal, the obtained attack classifiers failed to capture the privacy risks for individual samples. We also quantitatively measure the root-mean-square error (RMSE) between estimated probability of member and real probability of member. On the three Purchase100 classifiers, the RMSE values of our privacy risk score are 0.05, 0.09, and 0.06; in contrast, the RMSE values of NN classifier’s outputs are 0.26, 0.26, 0.25, respectively. We observe similar results on the undefended Texas100 classifier and the defended classifier by MemGuard , with details in Appendix C.
We also validate the effectiveness of privacy risk score across varied model architectures. For Purchase100 and Texas100 classifiers, we test two additional neural network depths by deleting the last hidden layer (depth=3) or adding one more hidden layer with 2048 neurons (depth=5); we test two additional neural network widths by halving the numbers of hidden neurons (width=0.5) or doubling the numbers of hidden neurons (width=2.0); we also test ReLU, Tanh, or Sigmoid as the activation functions. For CIFAR100 classifiers, besides DenseNet , we test other popular convolutional neural network architectures, including AlexNet , VGG , ResNet , and Wide ResNet . As show in Figure 4, our proposed privacy risk score metric indeed well represents the likelihood of a sample being in the training set under different model architectures. On the Texas100 dataset, the classifier fails to learn meaningful features using the Sigmoid activation function, achieving an accuracy of only 4%, and is thus omitted from the figure. We provide validation results with defended classifiers in Appendix D.
2.2 Heterogeneity of members’ privacy risk scores
After validating the effectiveness of the privacy risk score metric, we show the heterogeneity of training samples’ privacy risks by plotting the cumulative distribution of their privacy risk scores. We also investigate the performance of membership inference defense methods with comparison between defended and undefended classifiers.
Figure 5 presents the cumulative distributions of training points’ privacy risk scores for Purchase100 classifiers. We can see that, compared with the undefended classifier, the defended classifier with adversarial regularization has smaller privacy risk scores on average. However, we can also see that the defended classifier has a small portion of training data with higher privacy risk scores than the undefended model: the undefended model has all members’ privacy risk scores under 0.8, in contrast, the defended model has several training points with privacy risk scores higher than 0.8. Furthermore, the classifier with early stopping has a similar risk score distribution as the defended classifier.
Figure 6 shows the cumulative distribution of training data’ privacy risk scores for Texas100 classifiers. We can see that the defense method indeed decreases training samples’ privacy risk scores. However, the defended classifier is still quite vulnerable: training samples have privacy risk scores higher than 0.6.
2.3 Usage of privacy risk score
From our definition and verification results in Section 4.2.1, we know that privacy risk score of a data point indicates its probability of being a member. Instead of pursuing high average attack accuracy, now the adversary can identify which samples have high privacy risks and perform attacks with high confidence: a sample is inferred as a member if and only if its privacy risk score is above a certain probability threshold.
We show the attack results with precision and recall values in Table 6 for target classifiers with varying threshold values on privacy risk scores. From Table 6, we can see that with larger threshold values on privacy risk scores, the adversary indeed has higher precision values for membership inference attacks. For MemGuard , when setting the same threshold value on privacy risk scores, both undefended and defended Texas100 classifiers have similar attack precision, but the defended classifier has a smaller recall value. However, the defended Texas100 classifier still has severe privacy risks: training members can be inferred correctly with the precision of , and training members can be inferred correctly with the precision of . Similarly, while adversarial regularization can lower the average privacy risks, it increases the privacy risks for certain members: on the defended Purchase100 classifier, training members can be inferred correctly with the precision of . We urge designers of defense mechanisms to thus account for the full distribution of privacy risks in their analysis.
2.4 Impact of model properties on privacy risk score
We perform an in-depth investigation of privacy risk score by exploring its correlations with certain model properties, including sensitivity, generalization error, and feature embedding. We use the undefended Texas100 classifier from Jia et al. for the following experiments.
Privacy risk score with sensitivity We first study the relationship between privacy risk scores and model sensitivity with regard to training samples. The sensitivity is defined as the influence of one training sample on the target model by computing the difference after removing that sample. Since the privacy risk score is obtained with the measured distributions of modified prediction entropy (Equation (15)), we compute the model’s sensitivity regard to a training point as the logarithm of , where means the retrained classifier after removing from the training set.
Figure 7 shows the relation between privacy risk scores and the model sensitivity. For each privacy risk score, we show the first quartile, the average, and the third quartile of model sensitivities with regard to training data. We can see that, samples with higher privacy risk scores are likely to have a larger influence on the target model.
Privacy risk score with generalization error We observe that training samples with high risk scores are typically concentrated in a few class labels. Therefore, we further compare privacy risk scores among different class labels in this section.
Besides the privacy risk scores, we also record the generalization errors for different class labels. Figure 8 shows the average privacy risk scores and generalization errors for all 100 classes, where we sort the class labels based on their generalization errors. We can see that the class labels with high generalization errors tend to have higher privacy risk scores, which is as expected since the generalization error has a large influence on the success of membership inference attacks . The Pearson correlation coefficient between average privacy risk scores and generalization errors is as high as 0.94.
Privacy risk score with feature embeddings From the above experiment, we know that training samples from class labels with high generalization errors tend to have high privacy risk scores. Next, we investigate this further by looking into the feature representations of different class labels learned by the target classifier. We use the outputs of last hidden layer of the target classifier as the feature embedding of the input sample. We pick the top 5 class labels (30, 93, 97, 18, 98) with lowest average privacy risk scores (0.50, 0.52, 0.53, 0.54, 0.55) and at least 100 training samples, and the top 5 class labels (72, 49, 45, 51, 78) with highest average privacy risk scores (0.82, 0.83, 0.83, 0.85, 0.90) and at least 100 training samples. We record feature embeddings for both training and test examples from these 10 class labels. Finally, we adopt the t-Distributed Stochastic Neighbor Embedding (t-SNE) , a nonlinear dimensionality reduction technique, to visualize the feature embeddings.
Figure 9(a) and Figure 9(b) show the t-SNE plots of training samples and test samples, respectively. The training samples are separated clearly based on class labels since the target classifier has the training accuracy close to . Test samples from class labels with low risk scores (classes 30, 93, 97, 18, 98) have quite similar feature embeddings as training samples and are still well separated. On the other hand, test samples from class labels with high risk scores (classes 72, 49, 45, 51, 78) exhibit differences in feature representations compared to corresponding training samples. From Figure 9, we also observe the heterogeneity of samples’ privacy risks, in the granularities of both individual samples (e.g., different samples in class 78) and class labels (e.g., class 30 versus class 78). This further emphasizes the importance of fine-grained privacy risk analysis. It also validates our attack design of using class-dependent thresholds in Section 3.1. Our observations are also important for future defense work. A good defense approach should make training data and validation data have similar feature embeddings and consider the heterogeneity of samples’ privacy risks.
Conclusions
In this paper, we first argue that measuring membership inference privacy risks with neural network based attacks is insufficient. We propose to use a suite of metric-based attacks, including existing methods with our improved class-specific thresholds and a new proposed method based on modified prediction entropy, for benchmarking privacy risks of machine learning models. We also make recommendations of comparing with early stopping when benchmarking a defense that introduces a tradeoff between model accuracy and privacy risks, and considering adaptive attackers with knowledge of the defense to rigorously evaluate the performance of defense approaches. With these benchmark attacks, we show that (1) the defense approach of adversarial regularization, proposed by Nasr et al. , only reduces privacy risks to a limited degree and is no better than early stopping; (2) the defense performance of MemGuard, proposed by Jia et al. , is greatly degraded with adaptive attacks.
Next, we introduce a new metric called the privacy risk score for a fine-grained analysis of individual samples’ privacy risks. We show the effectiveness of the privacy risk score in estimating the true likelihood of an individual sample being in the training set and observe the heterogeneity of samples’ privacy risk scores with experimental results. Finally, we perform an in-depth investigation about the correlation between privacy risks and model properties, including sensitivity, generalization error, and feature embeddings. We hope that our work convinces the research community about the importance of systematically and rigorously evaluating privacy risks of machine learning models.
Acknowledgements
We are grateful to anonymous reviewers at USENIX Security for valuable feedback. We would also like to thank Google’s TensorFlow Privacy team for integrating our methods. This work was supported in part by the National Science Foundation under grants CNS-1553437 and CNS-1704105, the ARL’s Army Artificial Intelligence Innovation Institute (A2I2), the Office of Naval Research Young Investigator Award, the Army Research Office Young Investigator Prize, Faculty research award from Facebook, Schmidt DataX award, and by Princeton E-ffiliates Award.
References
Appendix A Membership inference attacks against other datasets
Here, we perform membership inference attacks on two more image datasets: CH-MNIST and Car196. The CH-MNIST dataset contains histology tiles from patients with colorectal cancer.https://www.kaggle.com/kmader/colorectal-histology-mnist The dataset contains 6464 black-and-white images from 8 different classes of tissue, 5,000 samples in total. We use 2,000 data samples to train a convolution neural network. The model contains 2 convolution blocks with the number of output channels equal to 32 and 64. The classifier achieves 99.0% training accuracy and 71.7% test accuracy.
The Car196 dataset contains colored images of 196 classes of cars.https://ai.stanford.edu/~jkrause/cars/car_dataset.html The dataset is split into 8,144 training images and 8,041 testing images. To train a model with good accuracy, we use a public ResNet50 classifier pretrained on ImageNet and fine-tune it on the Car196 training set. The classifier achieves 99.3% training accuracy and 87.5% test accuracy.
Besides our benchmark attacks, we follow Nasr et al. to perform NN-based attacks. We present attack results in Table 7. We can see that the best attack accuracy of our benchmark attacks is and larger than NN-based attacks.
Appendix B Privacy risk score with different training/test selection probabilities
Here, we provide the privacy risk score results on undefended Purchase100 classifier when the sample is chosen from training or test set with different probabilities. The computation of privacy risk score () is same as Section 4.1, except we use Equation (12) by also considering prior distributions and . We present the results in Figure 10 with different values of , where the red dotted line represents the baseline of random guessing. We can see that in all cases, most training samples have privacy risk scores larger than the prior training probability. We further compute a distance value between the prior distribution and the privacy risk score (posterior) distribution as to represent the privacy leakage. The distance values are 0.05, 0.09, 0.07, 0.02 when , respectively. As a comparison, the distance value is 0.1 when . As is closer to 0.5, the uncertainty of membership inference is larger, which in turns leads to a larger distance value.
Appendix C Validation of privacy risk score on Texas100 classifiers
We validate the effectiveness of privacy risk score on the undefended Texas100 classifier and its defended version with MemGuard in Figure 11. Compared with the output of NN attacks, our proposed privacy risk score is more meaningful for indicating the real probability of being a member. The RMSE values with privacy risk score are 0.08 and 0.05, while the RMSE values with NN classifier outputs are 0.13 and 0.21, for the undefended and defended Texas100 classifiers.
Appendix D Validation of privacy risk scores on different model architectures
We provide more validation results on Purchase100 classifiers defended by adversarial regularization and Texas100 classifiers defended by MemGuard in Figure 12. We can see that for all lines, the privacy risk score is close to the probability of being a member.