Truth Serum: Poisoning Machine Learning Models to Reveal Their Secrets
Florian Tramèr, Reza Shokri, Ayrton San Joaquin, Hoang Le, Matthew Jagielski, Sanghyun Hong, Nicholas Carlini
Introduction
A central tenet of computer security is that one cannot obtain any privacy without integrity (bonehgraduate, , Chapter 9). In cryptography, for example, an adversary who can modify a ciphertext, before it is sent to the intended recipient, might be able to leverage this ability to actually decrypt the ciphertext. In this paper, we show that this same vulnerability applies to the training of machine learning models.
Currently, there are two long and independent lines of work that study attacks on the integrity and privacy of training data in machine learning (ML). Data poisoning attacks (biggio2012poisoning, ) target the integrity of an ML model’s data collection process to degrade model performance at inference time—either indiscriminately (biggio2012poisoning, ; charikar2017learning, ; jagielski2018manipulating, ; fowl2021adversarial, ; munoz2017towards, ) or on targeted examples (bhagoji2019analyzing, ; turner2019label, ; shafahi2018poison, ; bagdasaryan2020backdoor, ; geiping2020witches, ; liu2017trojaning, ). Then, separately, privacy attacks such as membership inference (shokri2016membership, ), attribute inference (fredrikson2015model, ; yeom2018privacy, ) or data extraction (carlini2019secret, ; carlini2020extracting, ) aim to infer private information about the model’s training set by interacting with a trained model, or by actively participating in the training process (melis2019exploiting, ; nasr2019comprehensive, ).
Some works have highlighted connections between these two threats. For example, malicious parties in federated learning can craft updates to increase the privacy leakage of other participants (nasr2019comprehensive, ; melis2019exploiting, ; wen2022fishing, ; hitaj2017deep, ). Moreover, Chase et al. (chase2021property, ) show that poisoning attacks can increase leakage of global properties of the training set (e.g., the prevalence of different classes). In this paper, we extend and strengthen these results by demonstrating that an adversary can statically poison the training set to maximize the privacy leakage of individual training samples belonging to other parties. In other words, we show that the ability to “write” into the training dataset can be exploited to “read” from other (private) entries in this dataset.
We design targeted poisoning attacks on deep learning models that tamper with a small fraction of training data points (<) to improve the performance of membership inference, attribute inference and data extraction attacks on other training examples, by 1 to 2 orders-of-magnitude. For example, we show that by inserting just poison samples into the CIFAR-10 training set ( of the data), an adversary can infer membership of a specific target image with a true-positive-rate (TPR) of , compared to without poisoning, at a false-positive rate (FPR) of . Conversely, poisoning enables membership inference attacks to reach 50% TPR at a FPR of 0.05%, an error rate lower than the 24% FPR from prior work.
Similarly, by poisoning 64 sentences in the WikiText corpus, an adversary can extract a secret 6-digit “canary” (carlini2019secret, ) from a model trained on this corpus with a median of 230 guesses, compared to 9,018 guesses without poisoning (an improvement of ).
We show that our attacks are robust to uncertainty about the targeted samples, and rigorously investigate the factors that contribute to the success of our attacks. We find that poisoning has the most impact on samples that originally enjoy the strongest privacy, as our attacks reduce the average-case privacy of samples in a dataset to the worst-case privacy of data outliers. We further demonstrate that poisoning drastically lowers the cost of state-of-the-art privacy attacks, by alleviating the need for training shadow models (shokri2016membership, ).
We then consider untargeted attacks where an adversary controls a larger fraction of the training data—as high as 50%—and aims to increase privacy leakage of all other data points. Such attacks are relevant when a small number of parties (e.g., 2) want to jointly train a model on their respective training sets without revealing their own (private) dataset to the other(s), e.g., by using secure multi-party computation (yao1982protocols, ; goldreich1987play, ). We show that untargeted poisoning attacks can reduce the error rate of membership inference attacks across all of the victim’s data points by a factor of .
Our results call into question the relevance of modeling machine learning models as ideal functionalities in cryptographic protocols, such as when training models with secure multiparty computation (MPC). As our attacks show, a malicious party that honestly follows the training protocol can exploit their freedom to choose their input data to strongly influence the protocol’s “ideal” privacy leakage.
Background and Related Work
Training data privacy is an active research area in machine learning. In our work, we consider three canonical privacy attacks: membership inference (shokri2016membership, ), attribute inference (fredrikson2015model, ; fredrikson2014privacy, ; yeom2018privacy, ), and data extraction (carlini2019secret, ; carlini2020extracting, ). In membership inference, an adversary’s goal is to determine whether a given sample appeared in the training set of a model or not. Participation in a medical trial, for example, may reveal information about a diagnosis (homer2008resolving, ). In attribute inference, an adversary uses the model to learn some unknown feature of a given user in the training set. For example, partial knowledge of a user’s responses to a survey could allow the adversary to infer the response to other sensitive questions in the survey, by querying a model trained on this (and other) users’ responses. Finally, in data extraction, we consider an adversary that seeks to learn a secret string contained in the training data of a language model. We focus on these three canonical attacks as they are the most often considered attacks on training data privacy in the literature.
2. Attacks on Training Integrity
Poisoning attacks can be grouped into three categories: indiscriminate (availability) attacks, targeted attacks, and backdoor (or trojan) attacks. Indiscriminate attacks seek to reduce model performance and render it unusable (biggio2012poisoning, ; charikar2017learning, ; jagielski2018manipulating, ; fowl2021adversarial, ; munoz2017towards, ). Targeted attacks induce misclassifications for specific benign samples (shafahi2018poison, ; suciu2018does, ; geiping2020witches, ). Backdoor attacks add a “trigger” into the model, allowing an adversary to induce misclassifications by perturbing arbitrary test points (bhagoji2019analyzing, ; turner2019label, ; bagdasaryan2020backdoor, ). Backdoors can also be inserted via supply-chain vulnerabilities, rather than data poisoning attacks (liu2017neural, ; liu2017trojaning, ; gu2019badnets, ). However, none of these poisoning attacks have the goal of compromising privacy.
Our work considers an attacker that poisons the training data to violate the privacy of other users. Prior work has considered this goal for much stronger adversaries, with additional control over the training procedure. For example, an adversary that controls part of the training code can use the trained model as a side-channel to exfiltrate training data (song2017machine, ; bagdasaryan2021blind, ). Or in federated learning, a malicious server can select model architectures that enable reconstructing training samples (boenisch2021curious, ; fowl2022decepticons, ). Alternatively, participants in decentralized learning protocols can boost privacy attacks by sending dynamic malicious updates (nasr2019comprehensive, ; wen2022fishing, ; melis2019exploiting, ; hitaj2017deep, ). Our work differs from these in that we only make the weak assumption that the attacker can add a small amount of arbitrary data to the training set once, without contributing to any other part of training thereafter. A similar threat model to ours is considered in (chase2021property, ), for the weaker goal of inferring global properties of the training data (e.g., the class prevalences).
3. Defenses
As we consider adversaries that combine poisoning attacks and privacy inference attacks, defenses designed to mitigate either threat may be effective against our attacks.
Defenses against poisoning attacks (either indiscriminate or targeted) design learning algorithms that are robust to some fraction of adversarial data, typically by detecting and removing points that are out-of-distribution (diakonikolas2019sever, ; charikar2017learning, ; jagielski2018manipulating, ; gupta2019strong, ; tran2018spectral, ). Defenses against privacy inference either apply heuristics to minimize a model’s memorization (nasr2018machine, ; jia2019memguard, ) or train models with differential privacy (dwork2006calibrating, ; abadi2016deep, ). Training with differential privacy provably protects the privacy of a user’s data in any dataset, including a poisoned one.
Since our main focus in this work is to introduce a novel threat model that amplifies individual privacy leakage through data poisoning, we design worst-case attacks that are not explicitly aimed at evading specific data poisoning defenses. We note that such poisoning defenses are rarely deployed in practice today. In particular, sanitizing user data in decentralized settings such as federated learning or secure MPC represents a major challenge (kairouz2021advances, ). In Section 4.3.7, we show that a simple loss-clipping approach—inspired by differential privacy—can significantly decrease the effectiveness of our poisoning attacks. Whether our attack techniques can be made robust to such defenses, as well as to more complex data sanitization mechanisms, is an interesting question for future work.
A related line of work uses poisoning to measure the privacy guarantees of differentially private training algorithms (jagielski2020auditing, ; nasr2021adversary, ). These works are fundamentally different than ours: they measure the privacy leakage of the poisoned samples themselves to investigate worst-case properties of machine learning; in contrast, we show poisoning can harm other benign samples.
4. Machine Learning Notation
Amplifying Privacy Leakage with Data Poisoning
Motivation. The fields of security and cryptography are littered with examples where an adversary can turn an attack on integrity into an attack on privacy. For example, in cryptography a padding oracle attack (bleichenbacher1998chosen, ; vaudenay2002security, ) allows an adversary to use their ability to modify a ciphertext to learn the entire contents of the message. Similarly, compression leakage attacks (kelsey2002compression, ; gluck2013breach, ) inject data into a user’s encrypted traffic (e.g., HTTPS responses) and infer the user’s private data by analysing the size of ciphertexts. Alternatively, in Web security, some past browsers were vulnerable to attacks wherein the ability to send crafted email messages to a victim could be abused to actually read the victim’s other emails via a Cross-Origin CSS attack (huang2010protecting, ). Inspired by these attacks, we show this same type of result is possible in the area of machine learning.
We consider an adversary that can inject some data into a machine learning model’s training set . The goal of this adversary is to amplify their ability to infer information about the contents of , by interacting with a model trained on . In contrast to prior attacks on distributed or federated learning (nasr2019comprehensive, ; melis2019exploiting, ), our adversary cannot actively participate in the learning process. The adversary can only statically poison their data once, and after this can only interact with the final trained model.
We consider a generic privacy game, wherein the adversary has to guess which element from some universe was used to train a model. By appropriately defining the universe this game generalizes a number of prior privacy attack games, from membership inference to data extraction.
The challenger trains a model on the dataset and target .
The challenger gives the adversary query access to .
The adversary emits a guess .
The adversary wins the game if .
The universe captures the adversary’s prior belief about the possible value that the targeted example may take. In the membership inference game (see (yeom2018privacy, ; jayaraman2020revisiting, )), for a specific target example the universe is —where indicates the absence of an example. That is, the adversary guesses whether the model is trained on or on . For attribute inference, the universe contains the real targeted example , along with all “alternate versions” of with other values for an unknown attribute of . Attacks that extract well-formatted sensitive values, such as credit card numbers (carlini2019secret, ), can be modeled with a universe of all possible values that the secret could take.
We now introduce our new privacy game, which adds the ability for an adversary to poison the dataset. This is a strictly more general game, with the objective of maximizing the privacy leakage of the targeted point. The changes to Game 3.1 are highlighted in red.
The adversary sends a poisoned dataset of size to the challenger.
The challenger trains a model f_{\theta}\leftarrow\mathcal{T}(D\ {\color[rgb]{0.7,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{0.7,0,0}\cup\ D_{\text{adv}}}\cup\{z\}) on the poisoned dataset and target .
The challenger gives the adversary query access to .
The adversary emits a guess .
The adversary wins the game if .
1.1. Adversary Capabilities
The above poisoning game implicitly assumes a number of adversarial capabilities, which we now discuss more explicitly.
We impose no restrictions on the adversary’s poisons being “stealthy”. That is, we allow for the poisoned dataset to be arbitrary. As we will see, designing poisoning attacks that maximize privacy leakage is non-trivial—even when the adversary is not constrained in their choice of poisons. As poisoning attacks that target data privacy have not been studied so far, we aim here to understand how effective such attacks could be in the worst case, and leave the study of attacks with further constraints (such as “clean label” poisoning (turner2019label, ; shafahi2018poison, )) to future work.
Finally, the game assumes that the adversary targets a specific example . We call this a targeted attack. We also consider untargeted attacks in Section 4.4, where the attacker crafts a poisoned dataset to harm the privacy of all samples in the training set .
1.2. Success Metrics
When the universe of secret values is small (as for membership inference, where , or for attribute inference where it is the cardinality of the attribute), we measure an attack’s success rate by its true-positive rate (TPR) and false-positive rate (FPR) over multiple iterations of the game. Following (carlini2021membership, ), we focus in particular on the attack performance at low false-positive rates (e.g., FPR=), which measures the attack’s propensity to precisely target the privacy of some worst-case users.
For membership inference, we naturally define a true-positive as a correct guess of membership, i.e., when , and a false-positive as an incorrect membership guess, when .
For attribute inference, we define a “positive” as an example with a specific value for the unknown attribute (e.g., if the unknown attribute is gender, we define “female” as the positive class).
For canary extraction, where the universe of possible target values is large (e.g., all possible credit card numbers), we amend 3.2 to allow the adversary to obtain “partial credit” by emitting multiple guesses. Specifically, following (carlini2019secret, ), we let the adversary output an ordering (a permutation) of the secret’s possible values, from most likely to least likely. We then measure the attack’s success by the exposure (carlini2019secret, ) (in bits) of the correct secret :
The exposure ranges from bits (when the correct secret is ranked as the least likely value), to bits (when the adversary’s most likely guess is the correct value ).
2. Attack Overview
We begin with a high-level overview of our poisoning attack strategies. For simplicity of exposition, we focus on the special case of membership inference. Our attacks for attribute inference and canary extraction follow similar principles.
Given a target sample , the standard privacy game (for membership inference) in 3.1 asks the adversary to distinguish two worlds, where the model is respectively trained on or on . When we give the adversary the ability to poison the dataset in 3.2, the goal is now to alter the dataset so that the above two worlds become easier to distinguish.
Note that this goal is very different from simply maximizing the model’s memorization of the target . This could be achieved with the following (bad) strategy: poison the dataset by adding multiple identical copies of into it. This will ensure that the trained model strongly memorizes the target (i.e., the model will correctly classify with very high confidence). However, this will be true in both worlds, regardless of whether the target was in the original training set or not. This strategy thus does not help the adversary in solving the distinguishing game—and in fact actually makes it more difficult to distinguish membership.
Instead, the adversary should alter the training set so as to maximize the influence of the target . That is, we want the poisoned training set to be such that the inclusion of the target provides a maximal change in the trained model’s behavior on some inputs of the adversary’s choice.
To illustrate this principle, we begin by demonstrating a provably perfect privacy-poisoning attack for the special case of nearest-neighbor classifiers. We also propose an alternative attack for SVMs in Appendix D. We then describe our design principles for empirical attacks on deep neural networks.
Consider a -Nearest Neighbor (kNN) classifier (assume, wlog., that is odd). Given a labeled training set , and a test sample , this classifier finds the nearest neighbors of in , and outputs the majority label among these neighbors. We assume the attacker has black-box query access to the trained classifier.
We demonstrate how to poison a kNN classifier so that the classifier labels a target example correctly if and only if the target is in the original training set . This attack thus lets the adversary win the membership inference game with accuracy.
Our poisoning attack (see Algorithm 1 in Appendix D) creates a dataset of size that contains copies of the target , half correctly labeled as and half mislabeled as . We further add one poisoned example at a small distance from and also mislabeled as (we assume that no other point in the training set is within distance from ). This attack maximizes the influence of the targeted point, by turning it into a tie-breaker for classifying when it is a member.
The attacker infers that the target example is a member, if and only if the trained model correctly classifies as class . To see that the attack works, consider the two possible worlds:
The target is in : There are copies of in the poisoned training set : the poisoned copies (half are correctly labeled) and the target . Thus, the majority vote among the neighbors yields the correct class .
The target is not in : As all points in are at distance at least from the target , the neighbors selected by the model are the adversary’s poisoned points, a majority of which are mislabeled as . Thus, the model outputs .
In Appendix D, we show that our attack is non-trivial, in that there exist points for which poisoning is necessary to achieve perfect membership inference. In fact, we show that for some points, a non-poisoning adversary cannot infer membership better than chance.
The above attack on kNNs exploits the classifier’s specific structure which lets us turn any example’s membership into a perfect tie-breaker for the model’s decision on that example. In deep neural networks, it is unlikely that examples can exhibit such a clear cut influence (i.e., due to the stochasticity of training, it is unlikely that a specific model behavior would occur if and only if an example is a member).
Instead, we could try to cast the adversary’s goal as an optimization problem, of selecting a poisoned dataset that maximizes the distinguishability of models trained with or without the target . Yet, solving such an optimization problem is daunting. While prior work does optimize poisons to maximally alter a single model’s confidence on a specific target point (shafahi2018poison, ; turner2019label, ; zhu2019transferable, ), here we would instead need to optimize for a difference in distributions of the decisions of two models trained on two neighboring datasets.
Rather than tackle this optimization problem directly, we “handcraft” strategies that empirically increase a sample’s influence on the model. We start from the observation in prior work that the most vulnerable examples to privacy attacks are data outliers (yeom2018privacy, ; carlini2021membership, ). Such examples are easy to attack precisely because they have a large influence on the model: a model trained on an outlier has a much lower loss on this sample than a model that was not trained on it. Yet, in our threat model, the attacker cannot control or modify the targeted example (and is unlikely, a priori, to be an outlier). Our insight then is to poison the training dataset so as to transform the targeted example into an outlier. For example, we could fool the model into believing that the targeted point is mislabeled. Then, the presence of the correctly labeled target in the training set is likely to have a large influence on the model’s decision.
In Section 4, we show how to instantiate this attack strategy to boost membership inference attacks on standard image datasets. We then extend this attack strategy in Section 5 to the case of attribute inference attacks for tabular datasets. Finally, in Section 6 we propose attack strategies tailored to language models, that maximize the leakage of specially formatted canary sequences.
Membership Inference Attacks
Membership inference (MI) captures one of the most generic notions of privacy leakage in machine learning. Indeed, any form of data leakage from a model’s training set (e.g., attribute inference or data extraction) implies the ability to infer membership of some training examples. As a result, membership inference is a natural target for evaluating the impact of poisoning attacks on data privacy.
In this section, we introduce and analyze data poisoning attacks that improve membership inference by one to two orders of magnitude. Section 4.2 describes a targeted attack that increases leakage of a specific sample , and Section 4.3 contains an analysis of this attack’s success. Section 4.4 explores untargeted attacks that increase privacy leakage on all training points simultaneously.
We extend the recent attack of (carlini2021membership, ) that performs membership inference via a per-example log-likelihood test. The attack first trains shadow models such that each sample appears in the training set of half of the shadow models, and not in the other half. We then compute the losses of both sets of models on :
and fit Gaussian distributions to , and to (with a logit scaling of the losses, as in (carlini2021membership, )). Then, to infer membership of in a trained model , we compute the loss of on , and perform a standard likelihood-ratio test for the hypotheses that was drawn from or from .
To amplify the attack with poisoning, the adversary builds a poisoned dataset that is added to the training set of . The adversary also adds to each shadow model’s training set (so that these models are as similar as possible to the target model ).
We perform our experiments on CIFAR-10 and CIFAR-100 (cifar, )—standard image datasets of 50,000 samples from respectively 10 and 100 classes. The target models (and shadow models) use a Wide-ResNet architecture (zagoruyko2016wide, ) trained for 100 epochs with weight decay and common data augmentations (random image flips and crops). For each dataset, we train models on random splits of the original training set.The training sets of the target model and shadow models thus partially overlap (although the adversary does not know which points are in the target’s training set). Carlini et al. (carlini2021membership, ) show that their attack is minimally affected if the attacker’s shadow models are trained on datasets fully disjoint from the target’s training set. The models achieve 91% test accuracy on CIFAR-10 and 67% test-accuracy on CIFAR-100 on average.
2. Targeted Poisoning Attacks
We now design our poisoning attack to increase the membership inference success rate for a specific target example . That is, the attacker knows the data of (but not whether it is used to train the model) and designs a poisoned dataset adaptively based on .
We find that label flipping attacks are a very powerful form of poisoning attacks to increase data leakage. Given a targeted example with label , the adversary inserts the mislabelled poisons for some label . The rationale for this attack is that a model trained on will learn to associate with label , and the now “mislabelled” target will be treated as an outlier and have a heightened influence on the model when present in the training set.
To instantiate this attack on CIFAR-10 and CIFAR-100, we pick targeted points at random from the original training set. For each targeted example , the poisoned dataset contains a mislabelled example replicated times, for . We report the average attack performance for a full leave-one-out cross-validation (i.e., we evaluate the attack 128 times, using one model as the target and the rest as shadow models).
Figure 2 and Figure 15 (appendix) show the performance of our membership inference attack on CIFAR-10 and CIFAR-100 respectively, as we vary the number of poisons per sample.
We find that this attack is remarkably effective. Even with a single poisoned example (), the attack’s true-positive rate (TPR) at a false-positive rate (FPR) increases by . With poisons ( of the model’s training set size), the TPR increases by a factor on CIFAR-10, from to . On CIFAR-100, poisoning increases the baseline’s strong TPR of to at a FPR of .
Alternatively, we could aim for a fixed recall and use poisoning to reduce the MI attack’s error rate. Without poisoning, an attack that correctly identifies half of the targeted CIFAR-10 members (i.e., a TPR of ) would also incorrectly label of non-members as members. With poisoning, the same recall is achieved while only mislabeling of non-members—a factor improvement. On CIFAR-100, also for a TPR, poisoning reduces the attack’s false-positive rate by a factor , from to .
As we run multiple targeted attacks simultaneously (for efficiency sake), the total number of poisons is large (up to mislabelled points). Yet, the poisoned model’s test accuracy is minimally reduced (from to ) and the MI success rate on non-targeted points remains unchanged. Thus, we are not compounding the effects of the targeted attacks. As a sanity check, we repeat the experiment with only targeted points, and obtain similar results.
3. Analysis and Ablations
We have shown that targeted poisoning attacks significantly increase membership leakage. We now set out to understand the principles underlying our attack’s success.
In Figure 3 we plot the distribution of model confidences for five CIFAR-10 examples, when the example is a member (in red) and when it is not (in blue). On the horizontal axis, we vary the number of poisons (i.e., how many times this example is mislabeled in the training set). Without poisoning (left column), the distributions overlap significantly for most examples. As we increase the number of poisons, the confidences shift significantly to the left, as the model becomes less and less confident in the example’s true label. But crucially, the distributions also become easier to separate, because the (relative) influence of the targeted example on the trained model is now much larger.
To illustrate, consider the top example in Figure 3 (labeled “ship”). Without poisoning, this example’s confidence is in the range [99.99%, 100%] when it is a member, and [99.98%, 100%] when it is not. Confidently inferring membership is thus impossible. With 16 poisons, however, the confidence on this example is in the range [0.4%, 28.5%] when it is a member, and [0%, 2.4%] when it is not—thus enabling precise membership inference when the confidence exceeds .
3.2. Which points are vulnerable to our attack?
Our poisoning attack could increase the MI success rate in different ways. Poisoning could increase the attack accuracy uniformly across all data points, or it might disparately impact some data points. We show that the latter is true: our attack disparately impacts inliers that were originally safe from membership inference. This result has striking consequences: even if a user is an inlier and therefore might not be worried about privacy leakage, an active poisoning attacker that targets this user can still infer membership.
In Figure 4, we show the performance of our poisoning attack on those data points that are initially easiest and hardest to infer membership for. We run the membership inference attack of (carlini2021membership, ) on all CIFAR-10 points, and select the 5% of samples where the attack succeeds least often and most often (averaged over all 128 models). We then re-run the baseline attack on these extremal points with a new set of models (to ensure our selection of points did not overfit) and compare with our label flipping attack with .
Poisoning has a minor effect on data points that are already outliers: here even the baseline MI attack has a high success rate (73% TPR at a 0.1% FPR) and thus there is little room for improvement.In Figure 3, we see that examples for which MI succeeds without poisoning tend to already be outliers. For example, the third and fourth example from the top are a “bird” mislabelled as “cat” in the CIFAR-10 training set, and a “horse” confused as a “deer”. For points that are originally hardest to attack, however, poisoning improves the attack’s TPR by a factor , from 0.1% to 43%.
3.3. Are shadow models necessary?
Following (carlini2021membership, ; sablayrolles2019white, ; watson2021importance, ; ye2021enhanced, ; long2020pragmatic, ), our MI attack relies on shadow models to calibrate the confidences of individual examples. Indeed, as we see in the first column of Figure 3, the confidences of different examples are on different scales, and thus the optimal threshold to distinguish a member from a non-member varies greatly between examples. Yet, as we increase the number of poisoned samples, we observe that the scale of the confidences becomes unified across examples. And with 16 poisons, the threshold that best distinguishes members from non-members is approximately the same for all examples in Figure 3.
As a result, we show in Figure 5 that with poisoning, the use of shadow models for calibration is no longer necessary to obtain a strong MI attack. By simply setting a global threshold on the confidence of a targeted example (as in (yeom2018privacy, )) the MI attack works nearly as well as our full attack that trains 128 shadow models.
This result renders our attack much more practical than prior attacks. Indeed, in many settings, training even a single shadow model could be prohibitively expensive for the attacker (in terms of access to training data or compute). In contrast, the ability to poison a small fraction of the training set may be much more realistic, especially for very large models. Recent works (carlini2021membership, ; ye2021enhanced, ; mireshghallah2022quantifying, ; watson2021importance, ) show that non-calibrated MI attacks (without poisoning) perform no better than chance at low false-positives (see Figure 5). With poisoning however, these non-calibrated attacks perform extremely well. At a FPR of , a non-calibrated attack without poisoning has a TPR of (random guessing), whereas a non-calibrated attack with 16 targeted poisons has a TPR of %—an improvement of .
3.4. Does the choice of label matter?
Our poisoning attack injects a targeted example with an incorrect label. For the results in Figure 2 and Figure 15, we select an incorrect label at random (if we replicate a poison times, we use the same label for each replica).
In Figure 6 we explore alternative strategies for choosing the incorrect label on CIFAR-10. We consider three other strategies:
best: mislabel the poisons as the most likely incorrect class for that example (as predicted by a pre-trained model).
worst: mislabel the poisons as the least likely class.
random-multi: sample an incorrect label at random (without replacement) for each of the poisons.
These three strategies perform worse than the random approach. On both CIFAR-10 (Figure 6) and CIFAR-100 (Figure 16) the “best” and “worst” strategies do slightly worse than random mislabeling. The “random-multi” strategy does much worse, and under-performs the baseline attack without poisoning at low FPRs. This strategy has the opposite effect of our original attack, as it forces the model to predict a near-uniform distribution across classes, which is only minimally influenced by the presence or absence of the targeted example. Overall, this experiment shows that the exact choice of incorrect label matters little, as long as it is consistent.
3.5. Can the attack be improved by modifying the target?
Our poisoning attack only tampers with a target’s label , while leaving the example unchanged. It is conceivable that an attack that also alters the sample before poisoning could result in even stronger leakage. In Section A.2 we experiment with a number of such strategies, inspired by the literature on clean-label poisoning attacks (turner2019label, ; shafahi2018poison, ; zhu2019transferable, ). But we ultimately failed to find an approach that improves upon our attack and leave it as an open problem to design better privacy-poisoning strategies that alter the target sample.
3.6. Does the attack require exact knowledge of the target?
Existing membership inference attacks, which can be used for auditing ML privacy vulnerabilities, typically assume exact knowledge of the targeted example (so that the adversary can query the model on that example). Our attack is no different in this regard: it requires knowledge of the target at training time (in order to poison the model) and at evaluation time to run the MI attack.
We now evaluate how well our attack performs when the adversary has only partial knowledge of the targeted example. As we are dealing with images here, defining such partial knowledge requires some care. We will assume that instead of knowing the exact target example , the adversary knows an example that “looks similar” to . The attacker needs to guess whether was used to train a model. To this end, the attacker poisons the target model (and the shadow models) by injecting mislabeled versions of and queries the target model on to formulate a guess.
Details of this experiment are in Section A.3. Figure 19 shows that our attack (as well as the baseline without poisoning) are robust to an adversary with only partial knowledge of the target. At a FPR of 0.1%, the TPR is reduced by < for both the baseline attack and our attack with 4 poisons per target.
3.7. Can we mitigate the attack by bounding outlier influence?
As we have shown, our attack succeeds by turning data points into outliers, which then have a high influence on the model’s decisions. Our privacy-poisoning attack can thus likely be mitigated by bounding the influence that an outlier can have on the model. For example, training with differential privacy (dwork2006calibrating, ; abadi2016deep, ) would prevent our attack, as it bounds the influence that any outlier can have in any dataset (including a poisoned one). Algorithms for differentially private deep learning bound the size of the gradients of individual examples (abadi2016deep, ). Here, we opt for a slightly simpler approach that bounds the losses of individual examples (the two approaches are equivalent if we assume some bound on the model’s activations in a forward pass). Bounding losses rather than gradients has the advantage of being much more computationally efficient, as it simply requires scaling losses before backpropagation.
In Figure 7, we plot the MI success rate with and without poisoning, when each example’s cross-entropy loss is bounded to . Clipping in this way only slightly reduces the success rate of the attack without poisoning, but significantly harms the success of the poisoning attack at low false-positives. While our attack with poisons per target still improves over the baseline, including additional poisons weakens the attack, as the original sample’s loss can no longer grow unbounded to counteract the poisoning.
In Figure 20 in Section A.4, we show the effect of training with loss clipping on the distributions of member and non-member confidences for five random CIFAR-10 samples, analogously to Figure 3. Poisoning the model with mislabeled samples still shifts the confidences to very low values, but the inclusion of the correctly labeled target no longer clearly separates the two distributions.
While loss clipping thus appears to be a simple and effective defense against our poisoning attack, it is no privacy panacea. Indeed, the original baseline MI attack retains high success rate. As we show in Figure 8, further reducing the clipping bound (to ) does reduce the baseline MI attack to near-chance. But in this regime, poisoning does again increase the attack’s success rate at low false-positives by a factor . Moreover, aggressive clipping reduces the model’s test accuracy from to —an increase in error rate of (equivalent to undoing three years of progress in machine learning research). Finally, we also show in Section 4.4 that an alternative untargeted attack strategy, that increases leakage of all data points, remains resilient to moderate loss clipping.
4. Untargeted Poisoning Attacks
So far we have considered poisoning attacks that target a specific example that is known (exactly or partially) to the adversary. We now turn to more general untargeted attacks, where the adversary aims to increase the privacy leakage of all honest training points. As this is a much more challenging goal, we will consider adversaries who can compromise a much larger fraction of the training data. This threat is realistic in settings where a small number of parties decide to collaboratively train a model by pooling their respective datasets (e.g., two or more hospitals that train a joint model using secure multi-party computation).
We consider a setting where the training data is split between two parties. One party acts maliciously and chooses their data so as to maximize the leakage of the other party’s data.
To evaluate the attack on CIFAR-10, we select points at random from the training set to build the poisoned dataset . The honest party’s dataset consists of points sampled from the remaining part of the training set. We train a target model on the joint dataset . The attacker further trains shadow models by repeating the above process of sampling an honest dataset and combining it with the adversary’s fixed dataset . In total, we train models. We run the membership inference attack on a set of points disjoint from , half of which are actual members of the target model. We average results over a 128-fold leave-one-out cross-validation where we choose one of the 128 models as the target and the others as the shadow models.
Figure 9 shows the performance of our untargeted attack. Poisoning reliably increases the privacy leakage of all the honest party’s data points. At a FPR of 0.1%, the attack’s TPR across all the victim’s data grows from 9% without poisoning to 16% with our untargeted attack. Conversely, at a fixed recall, untargeted poisoning reduces the attack’s error rate drastically. With our poisoning strategy, the attacker can correctly infer membership for half of the honest party’s data, at a false-positive rate of only 3%, compared to an error rate of 24% without poisoning—an improvement of a factor 8. We include results for other untargeted poisoning strategies, as well as replications on additional datasets in Section A.5.
We further evaluate this untargeted attack against the simple “loss clipping” defense from Section 4.3.7. In contrast to the targeted case, we find that moderate clipping () has no effect on the untargeted attack and that with more stringent clipping, the model’s test accuracy is severely reduced.
Attribute Inference Attacks
Our results in Section 4 show that data poisoning can significantly increase an adversary’s ability to infer membership of training data. We now turn to attacks that infer actual data. We begin by considering attribute inference attacks in this section, and consider canary extraction attacks on language models in the next section.
In an attribute inference attack, the adversary has partial knowledge of some training example , and abuses access to a trained model to infer unknown features of this example. For simplicity of exposition, we consider the case of inferring a binary attribute (e.g., whether a user is married or not), given knowledge of the other features of , and of the class label . In the context of our privacy game, 3.2, the universe consists of the two possible “versions” of a target example, , where denotes the target example with value for the unknown attribute.
To then improve on this with poisoning, we inject mislabelled samples of the form , and of the form into the training set. Mislabeling both versions of the target forces the model to have similarly large loss on either version. The true variant of the target sample will then have a large influence on one of these losses, which will be detectable by our attack.
2. Experimental Setup
We run our attack on the Adult dataset (kohavi1996uci, ), a tabular dataset with demographic information of users. The target model is a three-layer feedforward neural network to predict whether a user’s income is above \5084\%$ test accuracy. We consider attribute inference attacks that infer either a user’s stated gender, or relationship status (after binarizing this feature into “married” and “not married” as in (mehnaz2022your, )). We define the attributes “female” and “not married” as the positive class in each case (i.e., a true-positive corresponds to the attacker correctly guessing that a user is female, or not married).
We pick 500 target points at random, and train target models that contain these 500 points in their training sets. We further train 128 shadow models on training sets that contain these 500 targets with the unknown attribute chosen at random. The training sets of the target models and shadow models are augmented with the adversary’s poisoned dataset that contains mislabelled copies of each target.
To evaluate the imputation baseline, we train the same three-layer feedforward neural network to predict gender (or relationship status) given a user’s other features and class label. We train this model on the entire Adult dataset except for the 500 target points.
3. Results
Our attack for attributing a user’s stated gender is plotted in Figure 10. Results for inferring relationship status are in Figure 26. As for membership inference, poisoning significantly improves attribute inference. At a FPR of 0.1%, the attack of (mehnaz2022your, ) has a TPR of 1%, while our attack with 16 poisons gets a TPR of 30%. Conversely, to achieve a TPR of 50%, the attack without poisoning incurs a FPR of 39%, while our attack with 16 poisons has a FPR of 1.2%—an error reduction of . In particular, the attack without poisoning performs worse than the trivial imputation baseline. Access to a non-poisoned model thus does not appear to leak more private information than what can be inferred from the data distribution.
Extraction in Language Models
In the previous sections, we focused on attacks that infer a single bit of information—whether an example is a member or not (in Section 4), or the value of some binary attribute of the example (in Section 5). We now consider the more ambitious goal of inferring secrets with much higher entropy. Following (carlini2019secret, ), we aim to extract well-formatted secrets (e.g., credit card numbers, social security numbers, etc.) from a language model trained on an unlabeled text corpus. Language models are a prime target for poisoning attacks, as their training datasets are often minimally curated (bender2021dangers, ; schuster2021you, ).
We train small variants of the GPT-2 model (radford2019language, ) on the WikiText-2 dataset (merity2016pointer, ), a standard language modeling corpus of approximately 3 million tokens of text from Wikipedia. We inject a canary into this dataset with a 125-token prefix followed by a random 6-digit secret (125 tokens represent about 500 characters; we also consider adversaries with partial knowledge of the prefix in Section 6.4). Given a trained model , the attacker prompts the model with the prefix followed by all possible values of the canary, and ranks them according to the model’s loss. Following (carlini2019secret, ), we compute the exposure of the secret as the average number of bits leaked to the adversary (see Equation 2). To control the randomness from the choice of random prefix and of random secret, we inject 45 different canaries into a model, and train 45 target models (for a total of different prefix-secret combinations). We then measure the average exposure across all combinations.
We consider two poisoning attack strategies to increase exposure of a secret canary, that rely on different adversarial capabilities:
Prefix poisoning assumes that the adversary can select the Prefix string that precedes the secret canary. This threat model captures settings where the attacker can select a template in which user secrets are input. Alternatively, since training sets for language models are often constructed by concatenating all of a user’s text sources, this attack could be instantiated by having the attacker send a message to the victim before the victim writes some secret information.
Suffix poisoning assumes that the adversary knows the Prefix string preceding the canary, but cannot necessarily modify it. Here, the adversary inserts poisoned copies of the Prefix with a chosen suffix into the training data.
As we will see, both types of poisoning attacks significantly increase the exposure of canaries. An attacker that combines both forms of attack can reduce their guesswork to recover canaries by a factor of , compared to a baseline attack without poisoning.
We again begin by showing that existing canary extraction attacks can be significantly improved by appropriately calibrating the attack using shadow models (again, similar to state-of-the-art membership inference attacks (sablayrolles2019white, ; carlini2021membership, ; ye2021enhanced, ; watson2021importance, ; long2020pragmatic, )).
As a baseline, we run the attack of (carlini2019secret, ), which simply ranks all possible canary values according to the target model’s loss. We find that this attack achieves only a low canary exposure of 3.1 bits on average in our setting (i.e, the adversary learns less than 1 digit of the secret).Carlini et al. (carlini2019secret, ) report higher exposures for numeric secrets because they use worse models (LSTMs) trained on simpler datasets that contain very few numbers. We find that even though the model’s loss on the random 6-digit secret does decrease throughout training, there are many other 6-digit numbers that are a priori much more likely and that therefore yield lower losses (such as 000000, or 123456).
The issue here is again one of calibration. Any language model trained on a large dataset will tend to assign higher likelihood to the number 123456 than to, say, the number 418463. However, a model trained with the canary 418463 will have a comparatively much higher confidence in this canary than a language model that was not trained on this specific canary.
As we did with our membership and attribute inference attacks, we thus first train a number of shadow models. We train shadow models on random subsets of WikiText (without any inserted canaries). Then, for a target model , prefix and canary guess , we assign to the calibrated confidence:
A potential canary value such as 123456 will have a low calibrated score, as all models assign it high confidence. In contrast, the true canary (e.g., 418643) will have high calibrated confidence as only the target model assigns a moderately high confidence to it. We then compute exposure exactly as in Equation 2, with possible canary values ranked according to their calibrated confidence.
Figure 11 shows that the use of shadow models vastly increases exposure of canaries. With just 2 shadow models, we obtain an average exposure of bits, a reduction in guesswork of compared to a non-calibrated attack. With additional shadow models, the exposure increases moderately to bits. Conversely, the fraction of canaries recovered in fewer than 100 guesses increases from 0.1% to 10% with calibration (an improvement of ).
2. Prefix Poisoning
The first poisoning attack we consider is one where the adversary can choose the prefix that precedes the secret canary. We evaluate the impact of various out-of-distribution prefix choices on the exposure of the secrets that succeed them. We pick five prefixes each from the following distributions:
Foreign: the prefix is in a language other than English, with a non-Latin alphabet: Chinese, Japanese, Russian, Hebrew or Arabic.
Code: the prefix is a piece of source code (in JavaScript, Java, C, Haskell or Rust).
Random: tokens sampled from GPT-2’s vocabulary.
Best: an initial random token prefix followed by greedily sampling the most likely token from a pretrained model.
Worst: an initial random token prefix followed by greedily sampling the least-likely token from a pretrained model.
Figure 12 shows that canaries that appear in “difficult” contexts (where the model has difficulty predicting the next token) have much higher exposure than canaries that appear in “easy” contexts.The non-English languages we chose do appear in some Wikipedia articles included in WikiText.
3. Suffix Poisoning
While the ability to choose or influence a secret value’s prefix may exist in some settings, it is a strong assumption on the adversary. We thus now turn to a more general setting where the prefix preceding a secret canary is fixed and out-of-control of the attacker.
We consider attacks inspired by the mislabeling attacks that were successful for membership inference and attribute inference. Yet, as language models are unsupervised, we cannot “mislabel” a sentence. Instead, we propose a suffix poisoning attack that inserts the known prefix followed by an arbitrary suffix many times into the dataset (thereby “mislabeling” the tokens that succeed the prefix, i.e., the canary). The attack’s rationale is that the poisoned model will have an extremely low confidence in any value for the canary, thus maximizing the relative influence of the true canary (similarly to how our MI attack poisons the model to have very low confidence in the true label, to maximize the influence of the targeted point).
We repeat the prefix times, padded by a stream of zeros (we consider other, less effective suffix choices in Section C.1). Figure 13 shows the success rate of the attack. Padding the prefix with incorrect suffixes reliably increases exposure from bits to bits after poison insertions ( of the dataset size).
Finally, we consider a powerful attacker that combines both our prefix-poisoning and suffix-poisoning strategies, by first choosing an out-of-distribution prefix that will precede the secret canary, and further inserting this prefix padded by zeros times into the training data. This attack increases exposure to bits on average with poison insertions. For half of the canaries, the attacker finds the secret in fewer than guesses, compared to guesses without poisoning—an improvement of . Conversely, the proportion of canaries that the attacker can recover with at most 100 guesses increases from 10% without poisoning to 42% with poisoning.
4. Attacks with Relaxed Capabilities
The language model poisoning attacks we evaluated so far assumed that (1) the adversary knows the entire prefix that precedes a canary; (2) the adversary has the ability to train shadow models. Below, we relax both of these assumptions in turn.
In Figure 14, we measure exposure as a function of the number of tokens of the Prefix string known to the attacker. We assume the attacker knows the last tokens of the prefix (about characters) immediately preceding the canary. The attacker thus queries the model with only these tokens of known context to extract a canary. Moreover, when poisoning the model, the attacker has the ability to choose the last tokens of the prefix, and to insert them together with an arbitrary suffix times into the dataset. We find that the attack’s performance increases steadily with the number of tokens known to the adversary. This mirrors the findings of Carlini et al. (carlini2022quantifying, ), who show that prompting a language model with longer prefixes increases the likelihood of extracting memorized content. As long as the attacker knows more than tokens of context (6 English words on average), they can increase exposure of secrets by poisoning the model.
In Section 6.1 we showed that canary extraction attacks are significantly improved if the adversary has the ability to train shadow models that closely mimic the behavior of the target model.
This assumption is standard in the literature on privacy attacks (shokri2016membership, ; sablayrolles2019white, ; watson2021importance, ; ye2021enhanced, ; long2020pragmatic, ; carlini2021membership, ), and we show that as few as 2 shadow models provide nearly the same benefit as ¿100 models. Yet, even training a single shadow model might be excessively expensive for very large language models (prior work has suggested that existing public language models could be used as proxies for shadow models (carlini2020extracting, )). In contrast, the ability to poison a large language model’s training set may be more accessible, especially since these models are typically trained on large minimally curated data sources (bender2021dangers, ; schuster2021you, ).
We find that poisoning significantly boosts exposure even if the attacker cannot train any shadow models and uses the baseline attack of (carlini2019secret, ). Interestingly, the ability to poison the dataset provides roughly the same benefit as the ability to train shadow models: with either ability, exposure increases from bits to and bits respectively—a reduction in average guesswork of -. Combining both abilities (i.e., poisoning the target model and training shadow models) compounds to an additional decrease in average guesswork (an average exposure of bits).
Discussion and Conclusion
We introduce a new attack on machine learning where an adversary poisons a training set to harm the privacy of other users’ data. For membership inference, attribute inference, and data extraction, we show how attacks can tamper with training data (as little as <) to increase privacy leakage by one or two orders-of-magnitude.
By blurring the lines between “worst-case” and “average-case” privacy leakage in deep neural networks, our attacks have various implications, discussed below, for the privacy expectations of users and protocol designers in collaborative learning settings.
Untrusted data is not only a threat to integrity. Large neural networks are trained on massive datasets which are hard to curate. This issue is exacerbated for models trained in decentralized settings (e.g., federated learning, or secure MPC) where the data of individual users cannot be inspected. Prior work observes that protecting model integrity is challenging in such settings (biggio2012poisoning, ; jagielski2018manipulating, ; munoz2017towards, ; shafahi2018poison, ; suciu2018does, ; geiping2020witches, ; bhagoji2019analyzing, ; bagdasaryan2020backdoor, ). Our work highlights a new, orthogonal threat to the privacy of the model’s training data, when part of the training data is adversarial. Thus, even in settings where threats to model integrity are not a primary concern, model developers who care about privacy may still need to account for poisoning attacks and defend against them.
Neural networks are poor “ideal functionalities”. There is a line of work that collaboratively trains ML models using secure multiparty computation (MPC) protocols (mohassel2017secureml, ; aono2017privacy, ; mohassel2018aby3, ; wagh2019securenn, ). These protocols are guaranteed to leak nothing more than an ideal functionality that computes the desired function (yao1982protocols, ; goldreich1987play, ). Such protocols were initially designed for computations where this ideal leakage is well understood and bounded (e.g., in Yao’s millionaires problem (yao1982protocols, ), the function always leaks exactly one bit of information). Yet, for flexible functions such as neural networks, the ideal leakage is much harder to characterize and bound (shokri2016membership, ; carlini2019secret, ; carlini2020extracting, ). Worse, our work demonstrates that an adversary that honestly follows the protocol can increase the amount of information leaked by the ideal functionality, solely by modifying their inputs. Thus, the security model of MPC fails to characterize all malicious strategies that breach users’ privacy in collaborative learning scenarios.
Worst-case privacy guarantees matter to everyone. Prior work has found that it is mainly outliers that are at risk of privacy attacks (yeom2018privacy, ; carlini2021membership, ; feldman2020does, ; ye2021enhanced, ). Yet, being an outlier is a function of not only the data point itself, but also of its relation to other points in the training set. Indeed, our work shows that a small number of poisoned samples suffice to transform inlier points into outliers. As such, our attacks reduce the “average-case” privacy leakage towards the “worst-case” leakage. Our results imply that methods that audit privacy with average-case canaries (carlini2019secret, ; thakkar2021understanding, ; ramaswamy2020training, ; zanella2020analyzing, ; malek2021antipodes, ) might underestimate the actual worst-case leakage under a small poisoning attack, and worst-case auditing approaches (jagielski2020auditing, ; nasr2021adversary, ) might more accurately measure a model’s privacy for most users.
Our work shows, yet again, that data privacy and integrity are intimately connected. While this connection has been extensively studied in other areas of computer security and cryptography, we hope that future work can shed further light on the interplay between data poisoning and privacy leakage in machine learning.
Acknowledgments
We thank Alina Oprea, Harsh Chaudhari, Martin Strobel, Abhradeep Thakurta, Thomas Steinke, and Andreas Terzis for helpful discussions and feedback.
Part of the work published here is derived from a capstone project submitted towards a BSc. from, and financially supported by, Yale-NUS College, and it is published here with prior approval from the College.
References
Appendix A Additional Experiments for Membership Inference Attacks
In Figure 15, we replicate the experiment from Section 4.2 on CIFAR-100. The experimental setup is exactly the same as on CIFAR-10.
In Figure 16, we replicate the experiment in Figure 6, where we vary the choice of target class for mislabelled poisons. As for CIFAR-10, we mislabel the poisons per target as: (1) the same random incorrect class for each of the samples (random); (2) the most likely incorrect class (best); the least-likely class (worst); or a different random incorrect class for each of the poisoned copies (random-multi).
Similarly to CIFAR-10, we find that the choice of random label matters little as long as it is used consistently for all poisons, with the random strategy performing best.
A.2. Attacks That Modify the Target
In this section, we consider alternative poisoning strategies that also modify the target sample , and not just the class label . All strategies we considered performed worse than our baseline strategy than mislabels the exact sample (“exact” in Figure 17).
We first consider strategies that mimic the polytope poisoning strategy of , which “surrounds” the target example with mislabeled samples in feature space. While the original attack does this to enhance the transferability of clean-label poisoning attacks, our aim is instead to maximize the influence of the targeted example when it is a member. To this end, instead of adding identical mislabeled copies of into the training set, we instead add mislabeled noisy versions of , or mislabeled augmentations of (e.g., rotations and shifts). Figure 17 shows that both strategies perform worse than our baseline attack (for poisons per target).
We consider an additional strategy, that replaces the sample by an unadversarial example for . That is, given an example we construct a sample that is very close to , so that a trained model labels as class with maximal confidence. We then use mislabeled copies of this unadversarial example, as our poisons. Our aim with this attack is to force the model to mislabel a variant of the target that the model is maximally confident in—in the hope that this would maximize the influence of the correctly labeled target. Unfortunately, we find that this strategy also performs much worse than our baseline strategy that simply mislabels the exact target .
A.3. Attacks with Partial Knowledge of the Target
In this section, we evaluate our attack when the adversary has only partial knowledge of the targeted example. Specifically, the adversary does not know the exact CIFAR-10 image that is (potentially) used to train a model, but only a “similar” image .
To choose pairs of similar images , we extract features from the entire CIFAR-10 training set using CLIP and match each example with its nearest neighbor in feature space. Random examples of such pairs are shown in Figure 18. These pairs often correspond to the same object pictured under different angles or scales, and thus reasonably emulate a scenario where the attacker knows the targeted object, but not the exact picture of it that was used to train the model.
To evaluate the attack, we train target models, half of which are trained on a particular target image . We ensure that none of these target models are trained on the neighbor image that is known to the adversary. The adversary then trains shadow models, half of which are trained on the image that is known to the adversary. We similarly ensure than none of the shadow models are trained on the real target . Using the shadow models, the adversary then models the distribution of losses of when it is a member and when it is not, as described in Section 4.1. Finally, the adversary queries the target models on the known image and guesses whether it was a member or not (of course, is never a member of the target model, but we use the adversary’s guess as a proxy for guessing the membership of the real, unknown target ).
The attack results are in Figure 19. We find that the membership inference attack of , with or without poisoning, is robust to an adversary with only partial knowledge of the target.
A.4. Bounding Outlier Influence with Loss Clipping
In Figure 20, we show the distribution of losses for individual CIFAR-10 examples, for models trained with loss clipping (see Section 4.3.7). Similarly to Figure 3, we find that poisoning shifts the model’s losses because the poisoned model becomes less confidence in the target example. However, poisoning does not help in making the distributions more separable. On the contrary, as we increase the number of poisons, even examples that were originally easy to infer membership on become hard to distinguish.
A.5. Untargeted Membership Inference Attacks
In Figure 21, Figure 22, and Figure 23, we show the results of different untargeted poisoning strategies on CIFAR-10 and CIFAR-100, as well as for an SVM classifier trained on the Texas100 dataset (see for details on this dataset).
The best-performing strategy on CIFAR-10 and CIFAR-100, same class label flipping, mislabels all of the adversary’s points into a single class. We consider two alternative untargeted poisoning strategies: random label flipping where each of the adversary’s points is randomly mislabeled into an incorrect class, and next class label flipping where the adversary mislabels each example into the next class . On both CIFAR-10 and CIFAR-100, consistently mislabelling all poisoned examples into the same class results in the strongest attack. On the Texas100 dataset, simply mislabeling the adversary’s data at random performs slightly better.
On CIFAR-10, we also experimented with strategies where the adversary’s share of the data is out-of-distribution, e.g., by using randomly mislabeled images from CIFAR-100 or MNIST, or simply images that consist of random noise. However, we could not find a poisoning strategy that performed as well as consistently mislabelling in-distribution data.
Similarly to the targeted attack, the untargeted poisoning attack also makes the MI attack easier by making individual examples’ confidence distributions more separable. In Figure 24, we pick five random CIFAR-10 examples and plot the logit-scaled confidence of the data point when it is a member (red) and not a member (blue). In the unpoisoned model (leftmost column), the two distributions overlap for most examples. With an untargeted poisoning attack, the confidences decrease and the distributions become more separable, which makes membership inference easier.
To examine which points are most vulnerable to the untargeted poisoning attack, we perform the same analysis as in Section 4.3.2. We first pick out the 5% of least- and most-vulnerable points for a set of models trained without poisoning. We then run an MI attack on both types of points (for a new set of models) with and without poisoning in Figure 25. Untargeted poisoning does not significantly affect the points that were initially most vulnerable. For the points that are hardest to attack without poisoning, our untargeted attack increases the TPR at a 0.1% FPR from to —an improvement of .
Appendix B Additional Experiments for Attribute Inference Attacks
In Figure 26, we replicate the experiment from Section 5.3 but infer a user’s relationship status (“married” or “non-married”) rather than their gender. The attack and experimental setup are the same as described in Section 5.2.
The results, shown in Figure 26 are qualitatively similar as those for inferring gender in Figure 10. At a FPR of 0.1%, the attack of (without poisoning) achieves a TPR of 4%, while our attack with 16 poisons obtains a TPR of 18%. At false-positive rates of > all the attribute inference attacks (even with poisoning) perform worse than a trivial imputation baseline.
Appendix C Additional Experiments for Canary Extraction
In Section 6.3 we showed that we could increase exposure after poisoning a canary’s prefix by re-inserting it multiple times into the training set padded with zeros. In Figure 27, we consider alternative suffix poisoning strategies, that are ultimately less effective. Padding the prefix with a list of random tokens or a random 6-digit number also provides a moderate increase in exposure (to bits and bits respectively), as long as the same random suffix is re-used for all poisons. If we insert the poison many times with different random suffixes, the poisoning actually hurts the attack. This mirrors our finding in Figure 6 and Figure 16 that mislabeling a target point with different incorrect labels hurts MI attacks.
C.2. Canary Extraction on a Fixed Budget
In Section 6, we evaluated canary extraction attacks in terms of the average exposure of different canaries inserted into a training set. An increase in average exposure does not necessarily tell us whether the attack is making extraction of canaries more practical (e.g., an attack might allow the adversary to recover canaries that used to require guesses in “only” guesses, without making any difference for those canaries that can be extracted in less than guesses). This is not the case for our attack. As we show in Figure 28, poisoning increases the attacker’s success rate in extracting canaries for any budget of guesses. For example, if the adversary is limited to 100 guesses for a canary, their success rate grows from without poisoning to with poisoning.
Appendix D Provably Amplifying Privacy Leakage
In this section, we provide additional theoretical analysis that proves a targeted poisoning attack can achieve perfect membership inference in the case of k-Nearest Neighbors (kNNs) and linear Support Vector Machines (SVMs).
In Section 3.2 we introduced a strategy that used poisoning to obtain 100% membership inference accuracy on a targeted point for kNNs (see Algorithm 1). Here, we show that poisoning is indeed necessary to obtain such a strong attack. To make this argument, we prove that without poisoning there exist points where membership inference cannot succeed better than chance.
We say that a point is unused by the model if the model’s output on any point is unaffected by the removal of from the training set . Such points are easy to construct: e.g., consider a cluster of close-by points that all the share the same label. Removing one point from the center of this cluster will not affect the model’s output on any input (a simple one-dimensional example visualization is given in Figure 29). For any such unused point, inferring membership is impossible: the model’s input-output behavior is identical whether the model is trained on or on . However, the poisoning strategy in Algorithm 1 still succeeds on these points, and thus provably increases privacy leakage.
Here, we consider an adversary who receives black-box access to a linear SVM. By definition, only support vectors are used at inference time and thus distinguishing between an SVM trained on and one trained on is impossible unless the sample is a support vector in at least one of these two models. We show that there exist points that can be forced—by a poisoning attack—to become support vectors if they are members of the training set. By computing the distance between the poisoned points and the classifier’s decision boundary (which can be done with black-box model access), the adversary can then infer with 100% accuracy whether some targeted point was a member or not.
Unlike for -nearest neighbors models, we will not be able to reveal membership for any point. Instead, our attack can only succeed on examples that lie on the convex hull of examples from one class. (However note that in high dimensions almost all points are on the boundary of the convex hull, and almost no points are contained in the interior.) We propose a sufficient condition for such a point to be forced to be a support vector, which we call protruding:
For a binary classification dataset , a point is protruding if there exists some so that the plane linearly separates and linearly separates .
Intuitively, this definition says that a point is protruding if there exists some way to linearly separate the two classes, such that this protruding point is the closest training example to the decision boundary. Then, if that point’s label were flipped, it suffices to “shift” the decision boundary (i.e., by modifying the offset ) to linearly separate the data again. We give an example of a protruding point (the target point) in Figure 30. If a point is protruding, we can insert a poisoned point of the opposite class close to it to force the protruding point to become a support vector.
Let be a binary classification dataset containing a protruding point . Then there exists some so that has as a support vector, and a larger margin when .
Without loss of generality, assume . Because is protruding, we know there exists some satisfying the conditions of the definition. Write such that has . Let be the distance from the plane to the nearest point in . We have because the plane lies strictly in between the planes and , which both linearly separate .
Then consider the poisoning . When , the maximum margin separator of is . The distance from each point in to this plane must be at least , as this has shifted by a distance of . Then will be a support vector of this plane, with a margin of .
When , the margin of the resulting hyperplane must be larger than , as is a hyperplane which linearly separates with a margin of . ∎
Our analysis here assumes that the adversary knows everything about the training set except for whether is a member, and that the dataset is linearly separable.
We also run a brief experiment to show that untargeted white-box attacks on SVMs are also possible. Given white-box access to an SVM, the adversary can directly recover the data of the support vectors, as these are necessary to perform inference. Our untargeted attacks increase privacy leakage by forcing the trained model to use more data points as support vectors. We train SVMs on Fashion MNIST restricted to the first two classes, using 2000 points for training and injecting 200 poisoning points according to a simple label flipping strategy. Over 5 trials, an unpoisoned linear SVM has an average of 121 support vectors, and an unpoisoned polynomial kernel SVM has an average of 176 support vectors. When adding the label flipping attack, the poisoned linear SVM grows to 512 support vectors from the clean training set, and the polynomial SVM grows to 642 support vectors from the clean training set, increasing the number of leaked data points by a factor of and , respectively.