Privacy Risks of Securing Machine Learning Models against Adversarial Examples
Liwei Song, Reza Shokri, Prateek Mittal
Introduction
Machine learning models, especially deep neural networks, have been deployed prominently in many real-world applications, such as image classification (Krizhevsky et al., 2012; Simonyan and Zisserman, 2015), speech recognition (Hinton et al., 2012; Deng et al., 2013), natural language processing (Collobert et al., 2011; Andor et al., 2016), and game playing (Silver et al., 2016; Moravčík et al., 2017). However, since the machine learning algorithms were originally designed without considering potential adversarial threats, their security and privacy vulnerabilities have come to a forefront in recent years, together with the arms race between attacks and defenses (Huang et al., 2011; Biggio and Roli, 2018; Papernot et al., 2018).
In the security domain, the adversary aims to induce misclassifications to the target machine learning model, with attack methods divided into two categories: evasion attacks and poisoning attacks (Huang et al., 2011). Evasion attacks, also known as adversarial examples, perturb inputs at the test time to induce wrong predictions by the target model (Biggio et al., 2013; Szegedy et al., 2014; Carlini and Wagner, 2017; Goodfellow et al., 2015; Papernot et al., 2016). In contrast, poisoning attacks target the training process by maliciously modifying part of training data to cause the trained model to misbehave on some test inputs (Biggio et al., 2012; Koh and Liang, 2017; Shafahi et al., 2018). In response to these attacks, the security community has designed new training algorithms to secure machine learning models against evasion attacks (Madry et al., 2018; Sinha et al., 2018; Zhang et al., 2019; Wong and Kolter, 2018; Mirman et al., 2018; Gowal et al., 2018) or poisoning attacks (Steinhardt et al., 2017; Jagielski et al., 2018).
In the privacy domain, the adversary aims to obtain private information about the model’s training data or the target model. Attacks targeting data privacy include: the adversary inferring whether input examples were used to train the target model with membership inference attacks (Shokri et al., 2017; Yeom et al., 2018; Nasr et al., 2019), learning global properties of training data with property inference attacks (Ganju et al., 2018), or covert channel model training attacks (Song et al., 2017). Attacks targeting model privacy include: the adversary uncovering the model details with model extraction attacks (Tramèr et al., 2016), and inferring hyperparameters with hyperparameter stealing attacks (Wang and Gong, 2018). In response to these attacks, the privacy community has designed defenses to prevent privacy leakage of training data (Nasr et al., 2018; Hayes and Ohrimenko, 2018; Shokri and Shmatikov, 2015; Abadi et al., 2016) or the target model (Kesarwani et al., 2018; Lee et al., 2019).
However, one important limitation of current machine learning defenses is that they typically focus solely on either the security domain or the privacy domain. It is thus unclear whether defense methods in one domain will have any unexpected impact on the other domain. In this paper, we take a step towards enhancing our understanding of machine learning models when both the security domain and privacy domain are considered together. In particular, we seek to understand the privacy risks of securing machine learning models by evaluating membership inference attacks against adversarially robust deep learning models, which aim to mitigate the threat of adversarial examples.
The membership inference attack aims to infer whether a data point is part of the target model’s training set or not, reflecting the information leakage of the model about its training data. It can also pose a privacy risk as the membership can reveal an individual’s sensitive information. For example, participation in a hospital’s health analytic training set means that an individual was once a patient in that hospital. It has been shown that the success of membership inference attacks in the black-box setting is highly related to the target model’s generalization error (Shokri et al., 2017; Yeom et al., 2018). Adversarially robust models aim to enhance the robustness of target models by ensuring that model predictions are unchanged for a small area (such as ball) around each input example. The objective is to make the model robust against any input, however, the objective is optimized only on the training set. Thus, intuitively, adversarially robust models have the potential to increase the model’s generalization error and sensitivity to changes in the training set, resulting in an enhanced risk of membership inference attacks. As an example, Figure 1 shows the histogram of cross-entropy loss values of training data and test data for both naturally undefended and adversarially robust CIFAR10 classifiers provided by Madry et al. (Madry et al., 2018). We can see that members (training data) and non-members (test data) can be distinguished more easily for the robust model, compared to the natural model.
To measure the membership inference risks of adversarially robust models, besides the conventional inference method based on prediction confidence, we propose two new inference methods that exploit the structural properties of robust models. We measure the privacy risks of robust models trained with six state-of-the-art adversarial defense methods, and find that adversarially robust models are indeed more susceptible to membership inference attacks than naturally undefended models. We further perform a comprehensive investigation to analyze the relation between privacy leakage and model properties. We finally discuss the role of adversary’s prior knowledge, potential countermeasures and the relationship between privacy and robustness.
In summary, we make the following contributions in this paper:
We propose two new membership inference attacks specific to adversarially robust models by exploiting adversarial examples’ predictions and verified worst-case predictions. With these two new methods, we can achieve higher inference accuracies than the conventional inference method based on prediction confidence of benign inputs.
We perform membership inference attacks on models trained with six state-of-the-art adversarial defense methods (3 empirical defenses (Madry et al., 2018; Sinha et al., 2018; Zhang et al., 2019) and 3 verifiable defenses (Wong and Kolter, 2018; Mirman et al., 2018; Gowal et al., 2018)). We demonstrate that all methods indeed increase the model’s membership inference risk. By defining the membership inference advantage as the increase in inference accuracy over random guessing (multiplied by 2) (Yeom et al., 2018), we show that robust machine learning models can incur a membership inference advantage , , times the membership inference advantage of naturally undefended models, on Yale Face, Fashion-MNIST, and CIFAR10 datasets, respectively.
We further explore the factors that influence the membership inference performance of the adversarially robust model, including its robustness generalization, the adversarial perturbation constraint, and the model capacity.
Finally, we experimentally evaluate the effect of the adversary’s prior knowledge, countermeasures such as temperature scaling and regularization, and discuss the relationship between training data privacy and model robustness.
Some of our analysis was briefly discussed in a short workshop paper (Song et al., 2019b). In this paper, we go further by proposing two new membership inference attacks and measuring four more adversarial defense methods, where we show that all adversarial defenses can increase privacy risks of target models. We also perform a comprehensive investigation of factors that impact the privacy risks.
Background and Related Work: Adversarial Examples and Membership Inference Attacks
In this section, we first present the background and related work on adversarial examples and defenses, and then discuss membership inference attacks.
Given a training set , the natural training algorithm aims to make model predictions match ground truth labels by minimizing the prediction loss over all training examples.
Although machine learning models have achieved tremendous success in many classification scenarios, they have been found to be easily fooled by adversarial examples (Szegedy et al., 2014; Biggio et al., 2013; Goodfellow et al., 2015; Carlini and Wagner, 2017; Papernot et al., 2016). Adversarial examples induce incorrect classifications to target models, and can be generated via imperceptible perturbations to benign inputs.
where denotes the set of points around within the perturbation budget of . Usually a ball is chosen as the perturbation constraint for generating adversarial examples i.e., . We consider the -ball adversarial constraint throughout the paper, as it is widely adopted by most adversarial defense methods (Madry et al., 2018; Sinha et al., 2018; Zhang et al., 2019; Wong and Kolter, 2018; Mirman et al., 2018; Gowal et al., 2018; Raghunathan et al., 2018).
The solution to Equation (4) is called an “untargeted adversarial example” as the adversarial goal is to achieve any incorrect classification. In comparison, a “targeted adversarial example” ensures that the model prediction is a specified incorrect label , which is not equal to .
Unless otherwise specified, an adversarial example in this paper refers to an untargeted adversarial example.
To provide adversarial robustness under the perturbation constraint , instead of natural training algorithm shown in Equation (2), a robust training algorithm is adopted by adding an additional robust loss function.
1.2. Empirical defenses:
Empirical defense methods approximate robust loss values by generating adversarial examples at each training step with state-of-the-art attack methods and computing their prediction loss. Now the robust training algorithm can be expressed as following.
Three of our tested adversarial defense methods belong to this category, which are described as follows.
Distributional Adversarial Training (Dist-Based Adv-Train) (Sinha et al., 2018): Instead of strictly satisfying the perturbation constraint with projection step as in PGD attacks, Sinha et al. (Sinha et al., 2018) generate adversarial examples by solving the Lagrangian relaxation of cross-entropy loss:
where computes the KL divergence. Adversarial examples are also generated with PGD-based attacks, except that now the attack goal is to maximize the output difference,
1.3. Verifiable defenses:
2. Membership Inference Attacks
For a target machine learning model, the membership inference attacks aim to determine whether a given data point was used to train the model or not (Shokri et al., 2017; Yeom et al., 2018; Salem et al., 2019; Nasr et al., 2019; Long et al., 2017; Hayes et al., 2018). The attack poses a serious privacy risk to the individuals whose data is used for model training, for example in the setting of health analytics.
Shokri et al. (Shokri et al., 2017) design a membership inference attack method based on training an inference model to distinguish between predictions on training set members versus non-members. To train the inference model, they introduce the shadow training technique: (1) the adversary first trains multiple “shadow models” which simulate the behavior of the target model, (2) based on the shadow models’ outputs on their own training and test examples, the adversary obtains a labeled (member vs non-member) dataset, and (3) finally trains the inference model as a neural network to perform membership inference attack against the target model. The input to the inference model is the prediction vector of the target model on a target data record.
A simpler inference model, such as a linear classifier, can also distinguish significantly vulnerable members from non-members. Yeom et al. (Yeom et al., 2018) suggest comparing the prediction confidence value of a target example with a threshold (learned for example through shadow training). Large confidence indicates membership. Their results show that such a simple confidence-thresholding method is reasonably effective and achieves membership inference accuracy close to that of a complex neural network classifier learned from shadow training.
In this paper, we use this confidence-thresholding membership inference approach in most cases. Note that when evaluating the privacy leakage with targeted adversarial examples in Section 3.3.1 and Section 5.2.5, the confidence-thresholding approach does not apply as there are multiple prediction vectors for each data point. Instead, we follow Shokri et al. (Shokri et al., 2017) to train a neural network classifier for membership inference.
Membership Inference Attacks against Robust Models
In this section, we first present some insights on why training models to be robust against adversarial examples make them more susceptible to membership inference attacks. We then formally present our membership inference attacks.
Throughout the paper, we use “natural (default) model” and “robust model” to denote the machine learning model with natural training algorithm and robust training algorithm, respectively. We also call the unmodified inputs and adversarially perturbed inputs as “benign examples” and “adversarial examples”. When evaluating the model’s classification performance, “train accuracy” and “test accuracy” are used to denote the classification accuracy of benign examples from training and test sets; “adversarial train accuracy’’ and “adversarial test accuracy” represent the classification accuracy of adversarial examples from training and test sets; “verified train accuracy” and “verified test accuracy” measure the classification accuracy under the verified worst-case predictions from training and test sets. Finally, an input example is called “secure” when it is correctly classified by the model for all adversarial perturbations within the constraint , “insecure” otherwise.
The performance of membership inference attacks is highly related to generalization error of target models (Shokri et al., 2017; Yeom et al., 2018). An extremely simple attack algorithm can infer membership based on whether or not an input is correctly classified. In this case, it is clear that a large gap between the target model’s train and test accuracy leads to a significant membership inference attack accuracy (as most members are correctly classified, but not the non-members). Tsipras et al. (Tsipras et al., 2019) and Zhang et al. (Zhang et al., 2019) show that robust training might lead to a drop in test accuracy. This is shown based on both empirical and theoretical analysis on toy classification tasks. Moreover, the generalization gap can be enlarged for a robust model when evaluating its accuracy on adversarial examples (Song et al., 2019a; Schmidt et al., 2018). Thus, compared with the natural models, the robust models might leak more membership information, due to exhibiting a larger generalization error, in both the benign or adversarial settings.
The performance of membership inference attack is related to the target model’s sensitivity with regard to training data (Long et al., 2017). The sensitivity measure is the influence of one data point on the target model’s performance by computing its prediction difference, when trained with and without this data point. Intuitively, when a training point has a large influence on the target model (high sensitivity), its model prediction is likely to be different from the model prediction on a test point, and thus the adversary can distinguish its membership more easily. The robust training algorithms aim to ensure that model predictions remain unchanged for a small area (such as the ball) around any data point. However, in practice, they guarantee this for the training examples, thus, magnifying the influence of the training data on the model. Therefore, compared with the natural training, the robust training algorithms might make the model more susceptible to membership inference attacks, by increasing its sensitivity to its training data.
To validate the above insights, let’s take the natural and the robust CIFAR10 classifiers provided by Madry et al. (Madry et al., 2018) as an example. From Figure 1, we have seen that compared to the natural model, the robust model has a larger divergence between the prediction loss of training data and test data. Our fine-grained analysis in Appendix A further reveals that the large divergence of robust model is highly related to its robustness performance. Moreover, the robust model incurs a significant generalization error in the adversarial setting, with adversarial train accuracy, and only adversarial test accuracy. Finally, we will experimentally show in Section 5.2.1 that the robust model is indeed more sensitive with regard to training data.
In this part, we describe the membership inference attack and its performance formally, with notations listed in Table 1. For a neural network model (we skip its parameter for simplicity) that is robustly trained with the adversarial constraint , the membership inference attack aims to determine whether a given input example is in its training set or not. We denote the inference strategy adopted by the adversary as , which codes members as 1, and non-members as 0.
We use the fraction of correct membership predictions, as the metric to evaluate membership inference accuracy. We use a test set which does not overlap with the training set, to represent non-members. We sample a random data point (, ) from either or with an equal probability, to test the membership inference attack. We measure the membership inference accuracy as follows.
where measures the size of a dataset.
The membership inference accuracy evaluates the probability that the adversary can guess correctly whether an input is from training set or test set. Note that a random guessing strategy will lead to a inference accuracy. To further measure the effectiveness of our membership inference strategy, we also use the notion of membership inference advantage proposed by Yeom et al. (Yeom et al., 2018), which is defined as the increase in inference accuracy over random guessing (multiplied by ).
2. Exploiting the Model’s Predictions on Benign Examples
3. Exploiting the Model’s Predictions on Adversarial Examples
We extend the attack to exploiting targeted adversarial examples. Targeted adversarial examples contain information about distance of the benign input to each label’s decision boundary, and are expected to leak more membership information than the untargeted adversarial example which only contains information about distance to a nearby label’s decision boundary.
We adapt the PGD attack method to find targeted adversarial examples (Equation (5)) by iteratively minizing the targeted cross-entropy loss.
The confidence thresholding inference strategy does not apply for targeted adversarial examples because there exist targeted adversarial examples (we have incorrect labels) for each input. Instead, following Shokri et al. (Shokri et al., 2017), we train a binary inference classifier for each class label to perform the membership inference attack. For each class label, we first choose a fraction of training and test points and generate corresponding targeted adversarial examples. Next, we compute model predictions on the targeted adversarial examples, and use them to train the membership inference classifier. Finally, we perform inference attacks using the remaining training and test points.
4. Exploiting the Verified Worst-Case Predictions on Adversarial Examples
Experiment Setup
In this section, we describe the datasets, neural network architectures, and corresponding adversarial perturbation constraints that we use in our experiments. Throughout the paper, we focus on the perturbation constraint: . The detailed architectures are summarized in Appendix B. Our code is publicly available at https://github.com/inspire-group/privacy-vs-robustness.
Yale Face. The extended Yale Face database B is used to train face recognition models, and contains gray scale face images of subjects under various lighting conditions (Georghiades et al., 2001; Lee et al., 2005). We use the cropped version of this dataset, where all face images are aligned and cropped to have the dimension of . In this version, each subject has images with the same frontal poses under different lighting conditions, among which images were corrupted during the image acquisition, leading to 2,414 images in total (Lee et al., 2005). In our experiments, we select images for each subject to form the training set (total size is 1,900 images), and use the remaining 514 images as the test set.
For the model architecture, we use a convolutional neural network (CNN) with the convolution kernel size , as suggested by Simonyan et al. (Simonyan and Zisserman, 2015). The CNN model contains 4 blocks with different numbers of output channels , and each block contains two convolution layers. The first layer uses a stride of for convolutions, and the second layer uses a stride of . There are two fully connected layers after the convolutional layers, each containing and neurons. When training the robust models, we set the perturbation budget () to be .
Fashion-MNIST. This dataset consists of a training set of 60,000 examples and a test set of 10,000 examples (Xiao et al., 2017). Each example is a grayscale image, associated with a class label from 10 fashion products, such as shirt, coat, sneaker.
Similar to Yale Face, we also adopt a CNN architecture with the convolution kernel size . The model contains 2 blocks with output channel numbers , and each block contains three convolution layers. The first two layers both use a stride of , while the last layer uses a stride of . Two fully connected layers are added at the end, with and neurons, respectively. When training the robust models, we set the perturbation budget () to be .
CIFAR10. This dataset is composed of color images in 10 classes, with 6,000 images per class. In total, there are 50,000 training images and 10,000 test images.
We use the wide ResNet architecture (Zagoruyko and Komodakis, 2016) to train a CIFAR10 classifier, following Madry et al. (Madry et al., 2018). It contains 3 groups of residual layers with output channel numbers (160, 320, 640) and 5 residual units for each group. One fully connected layer with neurons is added at the end. When training the robust models, we set the perturbation budget () to be .
Membership Inference Attacks against Empirically Robust Models
We first present an overall analysis that compares membership inference accuracy for natural models and robust models using multiple inference strategies across multiple datasets. We then present a deeper analysis of membership inference attacks against the PGD-based adversarial training defense.
The membership inference attack results against natural models and empirically robust models (Madry et al., 2018; Sinha et al., 2018; Zhang et al., 2019) are presented in Table 2, Table 3 and Table 4, where “acc” stands for accuracy, while “adv-train acc” and “adv-test acc” report adversarial accuracy under PGD attacks as shown in Equation (9).
According to these results, all three empirical defense methods will make the model more susceptible to membership inference attacks: compared with natural models, robust models increase the membership inference advantage by up to , , and , for Yale Face, Fashion-MNIST, and CIFAR10, respectively.
2. Detailed Membership Inference Analysis of PGD-Based Adversarial Training
In this part, we perform a detailed analysis of membership inference attacks against PGD-based adversarial training defense method (Madry et al., 2018) by using the CIFAR10 classifier as an example. We first perform a sensitivity analysis on both natural and robust models to show that the robust model is more sensitive with regard to training data compared to the natural model. We then investigate the relation between privacy leakage and model properties, including robustness generalization, adversarial perturbation constraint and model capacity. We finally show that the predictions of targeted adversarial examples can further enhance the membership inference advantage.
In the sensitivity analysis, we remove sample CIFAR10 training points from the training set, perform retraining of the models, and compute the performance difference between the original model and retrained model.
We excluded 10 training points (one for each class label) and retrained the model. We computed the sensitivity of each excluded point as the difference between its prediction confidence in the retrained model and the original model. We obtained the sensitivity metric for 60 training points by retraining the classifier 6 times. Figure 2 depicts the sensitivity values for the 60 training points (in ascending order) for both robust and natural models. We can see that compared to the natural model, the robust model is indeed more sensitive to the training data, thus leaking more membership information.
2.2. Privacy risk with robustness generalization
We perform the following experiment to demonstrate the relation between privacy risk and robustness generalization. Recall that in the approach of Madry et al. (Madry et al., 2018), adversarial examples are generated from all training points during the robust training process. In our experiment, we modify the above defense approach to (1) leverage adversarial examples from a subset of the CIFAR10 training data to compute the robust prediction loss, and (2) leverage the remaining subset of training points as benign inputs to compute the natural prediction loss.
The membership inference attack results are summarized in Table 5, where the first column lists the ratio of training points used for computing robust loss. We can see that as more training points are used for computing the robust loss, the membership inference accuracy increases, due to the larger gap between adv-train accuracy and adv-test accuracy.
2.3. Privacy risk with model perturbation budget
Next, we explore the relationship between membership inference and the adversarial perturbation budget , which controls the maximum absolute value of adversarial perturbations during robust training process.
We performed the robust training (Madry et al., 2018) for three CIFAR10 classifiers with varying adversarial perturbation budgets, and show the result in Table 6. Note that a model trained with a larger is more robust since it can defend against larger adversarial perturbations. From Table 6, we can see that more robust models leak more information about the training data. With a larger value, the robust model relies on a larger ball around each training point, leading to a higher membership inference attack accuracy.
2.4. Privacy risk with model capacity
Madry et al. (Madry et al., 2018) have observed that compared with natural training, robust training requires a significantly larger model capacity (e.g., deeper neural network architectures and more convolution filters) to obtain high robustness. In fact, we can think of the robust training approach as adding more “virtual training points”, which are within the ball around original training points. Thus the model capacity needs to be large enough to fit well on the larger “virtual training set”.
First, we can see that as the model capacity increases, the model has a higher membership inference accuracy, along with a higher adversarial train accuracy. Second, when using a larger adversarial perturbation budget , a larger model capacity is also needed. When , a capacity scale of 2 is enough to fit the training data, while for , a capacity scale of 8 is needed.
2.5. Inference attacks using targeted adversarial examples
Next, we investigate membership inference attacks using targeted adversarial examples. For each input, we compute 9 targeted adversarial examples with each of the 9 incorrect labels as targets using Equation (19). We then compute the output prediction vectors for all adversarial examples and use the shadow-training inference method proposed by Shokri et al. (Shokri et al., 2017) to perform membership inference attacks. Specifically, for each class label, we learn a dedicated inference model (binary classifier) by using the output predictions of targeted adversarial examples from training points and test points as the training set for the membership inference. We then test the inference model on the remaining CIFAR10 training and test examples from the same class label. In our experiments, we use a 3-layer fully connected neural network with size of hidden neurons equal to 200, 20, and 2 respectively. We call this method “model-infer (targeted)”.
For untargeted adversarial examples or benign examples, a similar class label-dependent inference model can also be obtained by using either untargeted adversarial example’s prediction vector or benign example’s prediction vector as features of the inference model. We call these methods “model-infer (untargeted)” and “model-infer (benign)”. We use the same 3-layer fully connected neural network as the inference classifier.
Finally, we also adapt our confidence-thresholding inference strategy to be class-label dependent by choosing the confidence threshold value according to prediction confidence values from training points and test points, and then testing on remaining CIFAR10 points from the same class label. Based on whether the confidence value is from the untargeted adversarial input or the benign input, we call the method as “confidence-infer (untargeted)” and “confidence-infer (benign)”.
The membership inference attack results using the above five strategies are presented in Table 7. We can see that the targeted adversarial example based inference strategy “model-infer (targeted)” always has the highest inference accuracy. This is because the targeted adversarial examples contain information about distance of the input to each label’s decision boundary, while untargeted adversarial examples contain information about distance of the input to only a nearby label’s decision boundary. Thus targeted adversarial examples leak more membership information. As an aside, we also find that our confidence-based inference methods obtain nearly the same inference results as training neural network models, showing the effectiveness of the confidence-thresholding inference strategies.
Membership Inference Attacks against Verifiably Robust Models
In this section we perform membership inference attacks against 3 verifiable defense methods: duality-based verification (Dual-Based Verify) (Wong and Kolter, 2018), abstract interpretation-based verification (Abs-Based Verify) (Mirman et al., 2018), and interval bound propagation-based verification (IBP-Based Verify) (Gowal et al., 2018). We train the verifiably robust models using the network architectures as described in Section 4 (with minor modifications for the Dual-Based Verify method (Wong and Kolter, 2018) as discussed in Appendix C), the perturbation budget is set to be for the Yale Face dataset and for the Fashion-MNIST dataset. We do not evaluate the verifiably robust models for the full CIFAR10 dataset as none of these three defense methods scale to the wide ResNet architecture.
The membership inference attack results against natural and verifiably robust models are presented in Table 8 and Table 9, where “acc” stands for accuracy, “adv-train acc” and “adv-test acc” measure adversarial accuracy under PGD attacks (Equation (9)), and “ver-train acc” and “ver-test acc” report the verified worse-case accuracy under the perturbation constraint .
On the other hand, for the Fashion-MNIST dataset, we fail to obtain increased membership inference accuracies on the verifiably robust models. However, we also observe much reduced benign train accuracy (below 90%) and verified train accuracy (below 80%), which means that the model fits the training set poorly. Similar to our analysis of empirical defenses, we can think the verifiable defense as adding more “virtual training points” around each training example to compute its verified robust loss. Since the verified robust loss is an upper bound on the real robust loss, the added “virtual training points” are in fact beyond the ball. Therefore, the model capacity needed for verifiable defenses is even larger than that of empirical defense methods.
From the experiment results in Section 5.2.4, we have shown that if the model capacity is not large enough, the robust model will not fit the training data well. This explains why membership inference accuracies for verifiably robust models are limited in Table 9. However, enlarging the model capacity does not guarantee that the training points will fit well for verifiable defenses because the verified upper bound of robust loss is likely to be looser with a deeper and larger neural network architecture. We validate our hypothesis in the following two subsections.
2. Varying Model Capacities
We use models with varying capacities to robustly train on the Yale Face dataset with the IBP-Based Verify defense (Gowal et al., 2018) as an example.
3. Reducing the Size of Training Set
In this subsection, we further prove our hypothesis by showing that when the size of the training set is reduced so that the model can fit well on the reduced dataset, the verifiable defense method indeed leads to an increased membership inference accuracy.
We choose the duality-based verifiable defense method (Wong and Kolter, 2018; Wong et al., 2018) and train the CIFAR10 classifier with a normal ResNet architecture: 3 groups of residual layers with output channel numbers (16, 32, 64) and only 1 residual unit for each group. The whole CIFAR10 training set have too many points to be robustly fitted with the verifiable defense algorithm: the robust CIFAR10 classifier (Wong et al., 2018) with has the train accuracy below . Therefore, we select a subset of the training data to robustly train the model by randomly choosing () training images for each class label. We vary the perturbation budget value () in order to observe when the model capacity is not large enough to fit on this partial CIFAR10 set using the verifiable training algorithm (Wong and Kolter, 2018).
We show the obtained results in Table 10, where the natural model has a low test accuracy (below 75%) and high privacy leakage (inference accuracy is ) since we only use training examples to learn the classifier. By using the verifiable defense method (Wong and Kolter, 2018), the verifiably robust models have increased membership inference accuracy values, for all values. We can also see that when increasing the values, at the beginning, the robust model is more and more susceptible to membership inference attacks (inference accuracy increases from to ). However, beyond a threshold of , the inference accuracy starts to decrease, since a higher requires a model with a larger capacity to fit well on the training data.
Discussions
In this section, we first evaluate the success of membership inference attacks when the adversary does not know the perturbation constraints of robust models. Second we discuss potential countermeasures, including temperature scaling and regularization, to reduce privacy risks. Finally, we discuss the relationship between training data privacy and model robustness.
Based on results shown in Figure 5, the adversary does not need to know the exact value of robust model’s perturbation budget: approximate knowledge of suffices to achieve high membership inference accuracy. Furthermore, the adversary can leverage the shadow training technique (with shadow training set) (Shokri et al., 2017) in practice to compute the best attack parameters (the perturbation budget and the threshold value), and then use the inferred parameters against the target model. The best perturbation budget may not even be same as the exact value of robust model. For example, we obtain the highest membership inference accuracy by setting as for the PGD-Based Adv-Train Yale Face classifier (Madry et al., 2018), and for the other two robust classifiers (Sinha et al., 2018; Zhang et al., 2019). We observe similar results for Fashion-MNIST and CIFAR10 datasets, which are presented in Appendix D.
2. Potential Countermeasures
We discuss potential countermeasures that can reduce the risk of membership inference attacks while maintaining model robustness.
Our membership inference strategies leverage the difference between the prediction confidence of the target model on its training set and test set. Thus, a straightforward mitigation method is to reduce this difference by applying temperature scaling on logits (Guo et al., 2017). The temperature scaling method was shown to be effective to reduce privacy risk for natural (baseline) models by Shokri et al. (Shokri et al., 2017), while we are studying its effect for robust models here.
Temperature scaling is a post-processing calibration technique for machine learning models that divides logits by the temperature, , before the softmax function. Now the model prediction probability can be expressed as
where corresponds to original model prediction. By setting , the prediction confidence is reduced, and when , the prediction output is close to uniform and independent of the input, thus leaking no membership information while making the model useless for prediction.
2.2. Regularization to improve robustness generalization
Regularization techniques such as parameter norm penalties and dropout (Srivastava et al., 2014), are typically used during the training process to solve overfitting issues for machine learning models. Shokri et al. (Shokri et al., 2017) and Salem et al. (Salem et al., 2019) validate their effectiveness against membership inference attacks. Furthermore, Nasr et al. (Nasr et al., 2018) propose to measure the performance of membership inference attack at each training step and use the measurement as a new regularizer.
The above mitigation strategies are effective regardless of natural or robust machine learning models. For the robust models, we can also rely on the regularization approach, which improves the model’s robustness generalization. This can mitigate membership inference attacks, since a poor robustness generalization leads to a severe privacy risk. We study the method proposed by Song et al. (Song et al., 2019a) to improve model’s robustness generalization and explore its performance against membership inference attacks.
The regularization method in (Song et al., 2019a) performs domain adaptation (DA) (Torralba et al., 2011) for the benign examples and adversarial examples on the logits: two multivariate Gaussian distributions for the logits of benign examples and adversarial examples are computed, and distances between two mean vectors and two covariance matrices are added into the training loss.
We apply this DA-based regularization approach on the PGD-based adversarial training defense (Madry et al., 2018) to investigate its effectiveness against membership inference attacks. We list the experimental results both with and without the use of DA regularization for Yale Face and Fashion-MNIST datasets in Table 11. We can see that the DA-based regularization can decrease the gap between adversarial train accuracy and adversarial test accuracy (robust generalization error), leading to a reduction in membership inference risk.
3. Privacy vs Robustness
We have shown that there exists a conflict between privacy of training data and model robustness: all six robust training algorithms that we tested increase models’ robustness against adversarial examples, but also make them more susceptible to membership inference attacks, compared with the natural training algorithm. Here, we provide further insights on how general this relationship between membership inference and adversarial robustness is.
Our experimental evaluation so far focused on the image classification domain. Next, we evaluate the privacy leakage of a robust model in a domain different than image classification to observe whether the conflict between privacy and robustness still holds.
We choose the UCI Human Activity Recognition (HAR) dataset (Anguita et al., 2013), which contains measurements of a smartphone’s accelerometer and gyroscope values while the participants holding it performed one of six activities (walking, walking upstairs, walking downstairs, sitting, standing, and laying). The dataset has 7,352 training samples and 2,947 test samples. Each sample is a -feature vector with time and frequency domain variables of smartphone sensor values, and all features are normalized and bounded within .
To train the classifiers, we use a 3-layer fully connected neural network with 1,000, 100, and 6 neurons respectively. For robust training, we follow Wong and Kolter (Wong and Kolter, 2018) by using the perturbation constraint with the size of , and apply the PGD-based adversarial training (Madry et al., 2018). The results for membership inference attacks against the robust classifier and its naturally trained counterpart are presented in Table 12. We can see that the robust training algorithm still leaks more membership information: the robust model has a membership inference advantage (Equation (16)) over the natural model.
3.2. Is the conflict a fundamental principle?
It is difficult to judge whether the privacy-robustness conflict is fundamental or not: will a robust training algorithm inevitably increase the model risk against membership inference attacks, compared to the natural training algorithm? On the one hand, there is no direct tension between privacy of training data and model robustness. We have shown in Section 5.2.2 that the privacy leakage of robust model is related to its generalization error in the adversarial setting. The regularization method in Section 7.2.2, which improves the adversarial test accuracy and decreases the generalization error, indeed helps to decrease the membership inference accuracy.
On the other hand, our analysis verifies that state-of-the-art robust training algorithms (Madry et al., 2018; Sinha et al., 2018; Zhang et al., 2019; Wong and Kolter, 2018; Mirman et al., 2018; Gowal et al., 2018) magnify the influence of training data on the model by minimizing the loss over a ball of each training point, leading to more training data memorization. In addition, we find that a recently-proposed robust training algorithm (Lecuyer et al., 2019), which adds a noise layer for robustness, also leads to an increase of membership inference accuracy in Appendix E. These robust training algorithms do not achieve good generalization of robustness performance (Schmidt et al., 2018; Song et al., 2019a). For example, even the regularized Yale Face classifier in Table 11 has a generalization error of in the adversarial setting, resulting a membership inference advantage than the natural Yale Face classifier in Table 2.
Furthermore, the failure of robustness generalization may partly be due to inappropriate (toy) distance constraints that are used to model adversaries. Although perturbation constraints have been widely adopted in both attacks and defenses for adversarial examples (Biggio et al., 2013; Goodfellow et al., 2015; Madry et al., 2018; Wong and Kolter, 2018), the distance metric has limitations. Sharif et al. (Sharif et al., 2018) empirically show that (a) two images that are perceptually similar to humans can have a large distance, and (b) two images with a small distance can have different semantics. Jacobsen et al. (Jacobsen et al., 2019) further show that robust training with a perturbation constraint makes the model more vulnerable to another type of adversarial examples: invariance based attacks that change the semantics of the image but leave the model predictions unchanged. Meaningful perturbation constraints to capture evasion attacks continue to be an important research challenge. We leave the question of deciding whether the privacy-robustness conflict is fundamental (i.e., will hold for next generation of defenses against adversarial examples) as an open question for the research community.
Conclusions
In this paper, we have connected both the security domain and the privacy domain for machine learning systems by investigating the membership inference privacy risk of robust training approaches (that mitigate the adversarial examples). To evaluate the membership inference risk, we propose two new inference methods that exploit structural properties of adversarially robust defenses, beyond the conventional inference method based on the prediction confidence of benign input. By measuring the success of membership inference attacks on robust models trained with six state-of-the-art adversarial defense approaches, we find that all six robust training methods will make the machine learning model more susceptible to membership inference attacks, compared to the naturally undefended training. Our analysis further reveals that the privacy leakage is related to target model’s robustness generalization, its adversarial perturbation constraint, and its capacity. We also provide thorough discussions on the adversary’s prior knowledge, potential countermeasures and the relationship between privacy and robustness. The detailed analysis in our paper highlights the importance of thinking about security and privacy together. Specifically, the membership inference risk needs to be considered when designing approaches to defend against adversarial examples.
References
Appendix A Fine-Grained Analysis of Prediction Loss of the Robust CIFAR10 Classifier
Here, we perform a fine-grained analysis of Figure 1(a) by separately visualizing the prediction loss distributions for test points which are secure and test points which are insecure. A point is deemed as secure when it is correctly classified by the model for all adversarial perturbations within the constraint .
Note that only a few training points were not secure, so we focused our fine-grained analysis on the test set. Figure 7 shows that insecure test inputs are very likely to have large prediction loss (low confidence value). Our membership inference strategies directly use the confidence to determine membership, so the privacy risk has a strong relationship with robustness generalization, even when we purely rely on the prediction confidence of the benign unmodified input.
Appendix B Model Architecture
We present the detailed neural network architectures used on Yale Face, Fashion-MNIST and CIFAR10 datasets in Table 13.
Appendix C Experiment Modifications for the Duality-Based Verifiable Defense
When dealing with the duality-based verifiable defense method (Wong and Kolter, 2018; Wong et al., 2018) (implemented in PyTorch), we find that the convolution with a kernel size and a stride of as described in Section 4 is not applicable. The defense method works by backpropagating the neural network to express the dual problem, while the convolution with a kernel size and a stride of prohibits their backpropagation analysis as the computation of output size is not divisible by 2 (PyTorch uses a round down operation). Instead, we choose the convolution with a kernel size and a stride of for the duality-based verifiable defense method (Wong and Kolter, 2018; Wong et al., 2018).
For the same reason, we also need to change the dimension of the Yale Face input to be by adding zero paddings. In our experiments, we have validated that the natural models trained with the above modifications have similar accuracy and privacy performance as the natural models without modifications reported in Table 8 and Table 9.
Appendix D Membership Inference Attacks with Varying Perturbation Constraints
This section augments Section 7.1 to evaluate the success of membership inference attacks when the adversary does not know the perturbation constraints of robust models.
We perform membership inference attacks with varying perturbation budgets on robust Fashion-MNIST and CIFAR10 classifiers (Madry et al., 2018; Sinha et al., 2018; Zhang et al., 2019). The Fashion-MNIST classifiers are robustly trained with the perturbation constraint of , while the CIFAR10 classifiers are robustly trained with the perturbation constraint of . The membership inference attack results with varying perturbation constraints are shown in Figure 8 and Figure 9.