Extracting Training Data from Diffusion Models
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, Eric Wallace
Introduction
Denoising diffusion models are an emerging class of generative neural networks that produce images from a training distribution via an iterative denoising process . Compared to prior approaches such as GANs or VAEs , diffusion models produce higher-quality samples and are easier to scale and control . Consequently, they have rapidly become the de-facto method for generating high-resolution images, and large-scale models such as DALL-E 2 have attracted significant public interest.
The appeal of generative diffusion models is rooted in their ability to synthesize novel images that are ostensibly unlike anything in the training set. Indeed, past large-scale training efforts “do not find overfitting to be an issue”, and researchers in privacy-sensitive domains have even suggested that diffusion models could “protect[] the privacy […] of real images” by generating synthetic examples . This line of work relies on the assumption that diffusion models do not memorize and regenerate their training data. If they did, it would violate all privacy guarantees and raise numerous questions regarding model generalization and “digital forgery” .
In this work, we demonstrate that state-of-the-art diffusion models do memorize and regenerate individual training examples. To begin, we propose and implement new definitions for “memorization” in image models. We then devise a two-stage data extraction attack that generates images using standard approaches, and flags those that exceed certain membership inference scoring criteria. Applying this method to Stable Diffusion and Imagen , we extract over a hundred near-identical replicas of training images that range from personally identifiable photos to trademarked logos (e.g., Figure 1).
To better understand how and why memorization occurs, we train hundreds of diffusion models on CIFAR-10 to analyze the impact of model accuracy, hyperparameters, augmentation, and deduplication on privacy. Diffusion models are the least private form of image models that we evaluate—for example, they leak more than twice as much training data as GANs. Unfortunately, we also find that existing privacy-enhancing techniques do not provide an acceptable privacy-utility tradeoff. Overall, our paper highlights the tension between increasingly powerful generative models and data privacy, and raises questions on how diffusion models work and how they should be responsibly deployed.
Background
Diffusion models. Generative image models have a long history (see [29, Chapter 20]). Generative Adversarial Networks (GANs) were the breakthrough that first enabled the generation of high-fidelity images at scale . But over the last two years, diffusion models have largely displaced GANs: they achieve state-of-the-art results on academic benchmarks and form the basis of all recently popularized image generators such as Stable Diffusion , DALL-E 2 , Runway , Midjourney and Imagen .
Despite being trained with this simple denoising objective, diffusion models can generate high-quality images by first sampling a random vector and then applying the diffusion model to remove the noise from this random “image”. To make the denoising process easier, we do not remove all of the noise at once—we instead iteratively apply the model to slowly remove noise. Formally, the final image is obtained from by iterating the rule for a noise schedule (dependent on ) with . This process relies on the fact that the model was trained to denoise images with varying degrees of noise. Overall, running this iterative generation process (which we will denote by Gen) with large-scale diffusion models produces results that resemble natural images.
Some diffusion models are further conditioned to generate a particular type of image. Class-conditional diffusion models take as input a class-label (e.g., “dog” or “cat”) alongside the noised image to produce a particular class of image. Text-conditioned models take this one step further and take as input the text embedding of some prompt (e.g., “a photograph of a horse on the moon”) using a pre-trained language encoder (e.g., CLIP ).
Training data privacy attacks. Neural networks often leak details of their training datasets. Membership inference attacks answer the question “was this example in the training set?” and present a mild privacy breach. Neural networks are also vulnerable to more powerful attacks such as inversion attacks that extract representative examples from a target class, attribute inference attacks that reconstruct subsets of attributes of training examples, and extraction attacks that completely recover training examples. In this paper, we focus on each of these three attacks when applied to diffusion models.
Concurrent work explores the privacy of diffusion models. Wu et al. and Hu et al. perform membership inference attacks on diffusion models; our results use more sophisticated attack methods and study stronger privacy risks such as data extraction. Somepalli et al. show several cases where (non-adversarially) sampling from a diffusion model can produce memorized training examples. However, they focus mainly on comparing the semantic similarity of generated images to the training set, i.e., “style copying”. In contrast, we focus on worst-case privacy under a much more restrictive notion of memorization, and perform our attacks on a wider range of models.
Motivation and Threat Model
There are two distinct motivations for understanding how diffusion models memorize and regenerate training data.
Understanding privacy risks. Diffusion models that regenerate data scraped from the Internet can pose similar privacy and copyright risks as language models . For example, memorizing and regenerating copyrighted text and source code has been pointed to as indicators of potential copyright infringement . Similarly, copying images from professional artists has been called “digital forgery” and has spurred debate in the art community.
Future diffusion models might also be trained on more sensitive private data. Indeed, GANs have already been applied to medical imagery , which underlines the importance of understanding the risks of generative models before we apply them to private domains.
Worse, a growing literature suggests that diffusion models could create synthetic training data to “protect the privacy and usage rights of real images” , and production tools already claim to use diffusion models to protect data privacy . Our work shows diffusion models may be unfit for this purpose.
Understanding generalization. Beyond data privacy, understanding how and why diffusion models memorize training data may help us understand their generalization capabilities. For instance, a common question for large-scale generative models is whether their impressive results arise from truly novel generations, or are instead the result of direct copying and remixing of their training data. By studying memorization, we can provide a concrete empirical characterization of the rates at which generative models perform such data copying.
In their diffusion model, Saharia et al. “do not find over-fitting to be an issue, and believe further training might improve overall performance“ , and yet we will show that this model memorizes individual examples. It may thus be necessary to broaden our definitions of overfitting to include memorization and related privacy metrics. Our results also suggest that Feldman’s theory that memorization is necessary for generalization in classifiers may extend to generative models, raising the question of whether the improved performance of diffusion models compared to prior approaches is precisely because diffusion models memorize more.
Our threat model considers an adversary that interacts with a diffusion model Gen (backed by a neural network ) to extract images from the model’s training set .
Image-generation systems. Unconditional diffusion models are trained on a dataset . When queried, the system outputs a generated image using a fresh random noise as input. Conditional models are trained on annotated images (e.g., labeled or captioned) and when queried with a prompt , the system outputs using the prompt and noise .
Adversary capabilities. We consider two adversaries:
A black-box adversary can query Gen to generate images. If Gen is a conditional generator, the adversary can provide arbitrary prompts . The adversary cannot control the system’s internal randomness .
A white-box adversary gets full access to the system Gen and its internal diffusion model . They can control the model’s randomness and can thus use the model to denoise arbitrary input images.
In both cases, we assume that an adversary who attacks a conditional image generator knows the captions for some images in the training set—thus allowing us to study the worst-case privacy risk in diffusion models.
Adversary goals. We consider three broad types of adversarial goals, from strongest to weakest attacks:
Data extraction: The adversary aims to recover an image from the training set . The attack is successful if the adversary extracts an image that is almost identical (see Section 4.1) to some .
Data reconstruction: The adversary has partial knowledge of a training image (e.g., a subset of the image) and aims to recover the full image. This is an image-analog of an attribute inference attack , which aims to recover unknown features from partial knowledge of an input.
Membership inference: Given an image , the adversary aims to infer whether is in the training set.
2 Ethics and Broader Impact
Training data extraction attacks can present a threat to user privacy. We take numerous steps to mitigate any possible harms from our paper. First, we study models that are trained on publicly-available images (e.g., LAION and CIFAR-10) and therefore do not expose any data that was not already available online.
Nevertheless, data that is available online may not have been intended to be available online. LAION, for example, contains unintentionally released medical images of several patients . We also therefore ensure that all images shown in our paper are of public figures (e.g., politicians, musicians, actors, or authors) who knowingly chose to place their images online. As a result, inserting these images in our paper is unlikely to cause any unintended privacy violation. For example, Figure 1 comes from Ann Graham Lotz’s Wikipedia profile picture and is licensed under Creative Commons, which allows us to “redistribute the material in any medium” and “remix, transform, and build upon the material for any purpose, even commercially”.
Third, we shared an advance copy of this paper with the authors of each of the large-scale diffusion models that we study. This gave the authors and their corresponding organizations the ability to consider possible safeguards and software changes ahead of time.
In total, we believe that publishing our paper and publicly disclosing these privacy vulnerabilities is both ethical and responsible. Indeed, at the moment, no one appears to be immediately harmed by the (lack of) privacy of diffusion models; our goal with this work is thus to make sure to preempt these harms and encourage responsible training of diffusion models in the future.
Extracting Training Data from State-of-the-art Diffusion Models
We begin our paper by extracting training images from large, pre-trained, high-resolution diffusion models.
Most existing literature on training data extraction focuses on text language models, where a sequence is said to be “extracted” and “memorized” if an adversary can prompt the model to recover a verbatim sequence from the training set . Because we work with high-resolution images, verbatim definitions of memorization are not suitable. Instead, we define a notion of approximate memorization based on image similarity metrics.
Restrictions of our definition. Our definition of extraction is intentionally conservative as compared to what privacy concerns one might ultimately have. For example, if we prompt Stable Diffusion to generate “A Photograph of Barack Obama,” it produces an entirely recognizable photograph of Barack Obama but not an near-identical reconstruction of any particular training image. Figure 2 compares the generated image (left) to the 4 nearest training images under the Euclidean 2-norm (right). Under our memorization definition, this image would not count as memorized. Nevertheless, the model’s ability to generate (new) recognizable pictures of certain individuals could still cause privacy harms.
2 Extracting Data from Stable Diffusion
We now extract training data from Stable Diffusion: the largest and most popular open-source diffusion model . This model is an 890 million parameter text-conditioned diffusion model trained on 160 million images. We generate from the model using the default PLMS sampling scheme at a resolution of pixels. As the model is trained on publicly-available images, we can easily verify our attack’s success and also mitigate potential harms from exposing the extracted data. We begin with a black-box attack.
Identifying duplicates in the training data. To reduce the computational load of our attack, as is done in , we bias our search towards duplicated training examples because these are orders of magnitude more likely to be memorized than non-duplicated examples .
Our extraction approach adapts the methodology from prior work to images and consists of two steps:
Generate many examples using the diffusion model in the standard sampling manner and with the known prompts from the prior section.
Perform membership inference to separate the model’s novel generations from those generations which are memorized training examples.
Generating many images. The first step is trivial but computationally expensive: we query the Gen function in a black-box manner using the selected prompts as input. To reduce the computational overhead of our experiments, we use the timestep-resampled generation implementation that is available in the Stable Diffusion codebase . This process generates images in a more aggressive fashion by removing larger amounts of noise at each time step and results in slightly lower visual fidelity at a significant () performance increase. We generate candidate images for each text prompt to increase the likelihood that we find memorization.
Performing membership inference. The second step requires flagging generations that appear to be memorized training images. Since we assume a black-box threat model in this section, we do not have access to the loss and cannot exploit techniques from state-of-the-art membership inference attacks . We instead design a new membership inference attack strategy based on the intuition that for diffusion models, with high probability for two different random initial seeds . On the other hand, if under some distance measure , it is likely that these generated samples are memorized examples.
The 500 images that we generate for each prompt have different (but unknown) random seeds. We can therefore construct a graph over the generations by connecting an edge between generation and if . If the largest clique in this graph is at least size 10 (i.e., 10 of the 500 generations are near-identical), we predict that this clique is a memorized image. Empirically, clique-finding is more effective than searching for pairs of images as it has fewer false positives.
2.2 Extraction Results
Given our ordered set of annotated images, we can also compute a curve evaluating the number of extracted images to the attack’s false positive rate. Our attack is exceptionally precise: out of 175 million generated images, we can identify memorized images with false positives, and all our memorized images can be extracted with a precision above . Figure 4 contains the precision-recall curve for both memorization definitions.
Figure 5 shows the results of this analysis. While we identify little Eidetic memorization for , this is expected due to the fact we choose prompts of highly-duplicated images. Note that at this level of duplication, the duplicated examples still make up just one in a million training examples. These results show that duplication is a major factor behind training data extraction.
The majority of the images that we extract (58%) are photographs with a recognizable person as the primary subject; the remainder are mostly either products for sale (17%), logos/posters (14%), or other art or graphics. We caution that if a future diffusion model were trained on sensitive (e.g., medical) data, then the kinds of data that we extract would likely be drawn from this sensitive data distribution.
Despite the fact that these images are publicly accessible on the Internet, not all of them are permissively licensed. We find that a significant number of these images fall under an explicit non-permissive copyright notice (35%). Many other images (61%) have no explicit copyright notice but may fall under a general copyright protection for the website that hosts them (e.g., images of products on a sales website). Several of the images that we extracted are licensed CC BY-SA, which requires “[to] give appropriate credit, provide a link to the license, and indicate if changes were made.” Stable Diffusion thus memorizes numerous copyrighted and non-permissive-licensed images, which the model may reproduce without the accompanying license.
3 Extracting Data from Imagen
While Stable Diffusion is the best publicly-available diffusion model, there are non-public models that achieve stronger performance using larger models and datasets . Prior work has found that larger models are more likely to memorize training data and we thus study Imagen , a 2 billion parameter text-to-image diffusion model. While individual details differ between Imagen’s and Stable Diffusion’s implementation and training scheme, these details are independent of our extraction results.
4 Extracting Outlier Examples
The attacks presented above succeed, but only at extracting images that are highly duplicated. This “high ” memorization may be problematic, but as we mentioned previously, the most compelling practical attack would be to demonstrate memorization in the “low ” regime.
We now set out to achieve this goal. In order to find non-duplicated examples likely to be memorized, we take advantage of the fact that while on average models often respect the privacy of the majority of the dataset, there often exists a small set of “outlier” examples whose privacy is more significantly exposed . And so instead of searching for memorization across all images, we are more likely to succeed if we focus our effort on these outlier examples.
But how should we find which images are potentially outliers? Prior work was able to train hundreds of models on subsets of the training dataset and then use an influence-function-style approach to identify examples that have a significant impact on the final model weights . Unfortunately, given the cost of training even a single large diffusion model is in the millions-of-dollars, this approach will not be feasible here.
Therefore we take a simpler approach. We first compute the CLIP embedding of each training example, and then compute the “outlierness” of each example as the average distance (in CLIP embedding space) to its nearest neighbors in the training dataset.
Surprisingly, we find that attacking out-of-distribution images is much more effective for Imagen than it is for Stable Diffusion. On Imagen, we attempted extraction of the 500 images with the highest out-of-distribution score. Imagen memorized and regurgitated 3 of these images (which were unique in the training dataset). In contrast, we failed to identify any memorization when applying the same methodology to Stable Diffusion—even after attempting to extract the most-outlier samples. Thus, Imagen appears less private than Stable Diffusion both on duplicated and non-duplicated images. We believe this is due to the fact that Imagen uses a model with a much higher capacity compared to Stable diffusion, which allows for more memorization . Moreover, Imagen is trained for more iterations and on a smaller dataset, which can also result in higher memorization.
Investigating Memorization
The above experiments are visually striking and clearly indicate that memorization is pervasive in large diffusion models—and that data extraction is feasible. But these experiments do not explain why and how these models memorize training data. In this section we train smaller diffusion models and perform controlled experiments in order to more clearly understand memorization.
For the remainder of this section, we focus on diffusion models trained on CIFAR-10. We use state-of-the-art training code We either directly use OpenAI’s Improved Diffusion repository (https://github.com/openai/improved-diffusion) in Section 5.1, or our own re-implementation in all following sections. Models trained with our re-implementation achieve almost identical FID to the open-sourced models. We use half the dataset as is standard in privacy analyses . to train 16 diffusion models, each on a randomly-partitioned half of the CIFAR-10 training dataset. We run three types of privacy attacks: membership inference attacks, attribute inference attacks, and data reconstruction attacks. For the membership inference attacks, we train class-conditional models that reach an FID below 3.5 (see Figure 11), placing them in the top-30 generative models on CIFAR-10 . For reconstruction attacks (Section 5.1) and attribute inference attacks with inpainting (Section 5.3), we train unconditional models with an FID below 4.
1 Untargeted Extraction
Before devling deeper into understanding memorization, we begin by validating that memorization does still occur in our smaller models. Because these models are not text conditioned, we focus on untargeted extraction. Specifically, given our diffusion models trained on CIFAR-10, we unconditionally generate images from each model for a total of candidate images. Because we will later develop high-precision membership inference attacks, in this section we directly search for memorized training examples among all our million generated examples. Thus this is not an attack per se, but rather verifying the capability of these models to memorize.
We thus slightly modify our attack to use the distance
where is the set containing the closest elements from the training dataset to the example . This distance is small if the extracted image is much closer to the training image compared to the closest neighbors of in the training set. We run our attack with and . Our attack was not sensitive to these choices.
Results. Using the above methodology we identify unique extracted images from the CIFAR-10 dataset ( of the entire dataset). Some CIFAR-10 training images are generated multiple times. In these cases, we only count the first generation as a successful attack. Further, because the CIFAR-10 training dataset contains many duplicate images, we do not count two generations of two different (but duplicated) images in the training dataset. In Figure 8 we show a selection of training examples that we extract and full results are shown in Figure 17 in the Appendix.
2 Membership Inference Attacks
We now evaluate membership inference with more traditional attack techniques that use white-box access, as opposed to Section 4.2.1 that assumed black-box access. We will show that all examples have significant privacy leakage under membership inference attacks, compared to the small fraction that are sensitive to data extraction. We consider two membership inference attacks on our class-conditional CIFAR-10-trained diffusion models.Section C.4 replicates these results for unconditional models.
The loss threshold attack. Yeom et al. introduce the simplest membership inference attack: because models are trained to minimize their loss on the training set, we should expect that training examples have lower loss than non-training examples. The loss threshold attack thus computes the loss and reports “member” if for some chosen threshold and otherwise “non-member’. The value of can be selected to maximize a desired metric (e.g., true positive rate at some fixed false positive rate or the overall attack accuracy).
The Likelihood Ratio Attack (LiRA). Carlini et al. introduce the state-of-the-art approach to performing membership inference attacks. LiRA first trains a collection of shadow models, each model on random subsets of the training dataset. LiRA then computes the loss for the example under each of these shadow models . These losses are split into two sets: the losses for the example under the shadow models that did see the example during training, and the losses for the example under the shadow models that did not see the example during training. LiRA finishes the initialization process by fitting Gaussians to the IN set and to OUT set of losses. Finally, to predict membership inference for a new model , we compute and then measure whether .
Choosing a loss function. Both membership inference attacks use a loss function . In the case of classification models, Carlini et al. find that choosing a loss function is one of the most important components of the attack. We find that this effect is even more pronounced for diffusion models. In particular, unlike classifiers that have a single loss function (e.g., cross entropy) used to train the model, diffusion models are trained to minimize the reconstruction loss when a random quantity of Gaussian noise has been added to an image. This means that “the loss” of an image is not well defined—instead, we can only ask for the loss of an image for a certain timestep with a corresponding amount of noise (cf. Equation 1).
We must thus compute the optimal timestep at which we should measure the loss. To do so, we train 16 shadow models each on a random 50% of the CIFAR-10 training dataset. We then compute the loss for every model, for every example in the training dataset, and every timestep ( in the models we use).
Figure 9 plots the timestep used to compute the loss against the attack success rate, measured as the true positive rate (TPR), i.e., the number of examples which truly are members over the total number of members, at a fixed false positive rate (FPR) of 1%, i.e., the fraction of examples which are incorrectly identified as members. Evaluating at leads to the most successful attacks. We conjecture that this a “Goldilock’s zone” for membership inference: if is too small, and so the noisy image is similar to the original, then predicting the added noise is easy regardless if the input was in the training set; if is too large, and so the noisy image is similar to Gaussian noise, then the task is too difficult. Our remaining experiments will evaluate at , where we observed a TPR of 71% at an FPR of 1%.
We now evaluate membership inference using our specified loss function. We follow recent advice and evaluate the efficacy of membership inference attacks by comparing their true positive rate to the false positive rate on a log-log scale. In Figure 10, we plot the membership inference ROC curve for the loss threshold attack and LiRA. An out-of-the-box implementation of LiRA achieves a true positive rate of over at a false positive rate of just . As a point of reference, state-of-the-art classifiers are much more private, e.g., with a TPR at FPR . This shows that diffusion models are significantly less private than classifiers trained on the same data. (In part this may be because diffusion models are often trained far longer than classifiers.)
Qualitative analysis. In Figure 20, we visualize the least- and most-private images as determined by their easiness to detect via LiRA. We find that the easiest-to-attack examples are all extremely out-of-distribution visually from the CIFAR-10 dataset. These images are even more visually out-of-distribution compared to the outliers identified by Feldman et al. who produce a similar set of images but for image classifiers. In contrast, the images that are hardest to attack are all duplicated images. It is challenging to detect the presence or absence of each of these images in the training dataset because there is another identical image in the training dataset that may have been present or absent—therefore making the membership inference question ill-defined.
2.2 Augmentations Improve Attacks
By varying the number of point samples taken to estimate this expectation we can potentially increase the attack success rate. And second, because our diffusion models train on augmented versions of training images (e.g., by flipping images horizontally), it makes sense to compute the loss averaged over all possible augmentations. Prior work has found that both of these attack strategies are effective at increasing the efficacy of membership inference attacks for classifiers , and we find they are effective here as well.
Improved attack results. Figure 10 shows the effect of combining both these strategies. Together they are remarkably successful, and at a false positive rate of they increase the true positive rate by over a factor of six from to . Figure 19 in the Appendix breaks down the impact of each component: in Figure 19(a) we increase the number of Monte Carlo samples from 1 (the base LiRA attack) to 20, and in Figure 19(b) we augment samples with a horizontal flip.
2.3 Memorization Versus Utility
We train our diffusion models to reach state-of-the-art levels of performance. Prior work on language models has found that better models are often easier to attack than less accurate models—intuitively, because they extract more information from the same training dataset . Here we perform a similar experiment.
Attack results vs. FID. To evaluate our generative models, we use the standard Fréchet Inception Distance (FID) , where lower scores indicate higher quality. Our previous CIFAR-10 results used models that achieved the best FID (on average 3.5) based on early stopping. Here we evaluate models over the course of training in Figure 11. We compute the attack success rate as a function of FID, and we find that as the quality of the diffusion model increases so too does the privacy leakage. These results are concerning because they suggest that stronger diffusion models of the future may be even less private.
3 Inpainting Attacks
Having performed untargeted extraction on CIFAR-10 models, we now construct a targeted version of our attack. As mentioned earlier, performing a targeted attack is complicated by the fact that these models do not support textual prompting. We instead provide guidance by performing a form of attribute inference attack that we call an “inpainting attack”. Given an image, we first mask out a portion of this image; our attack objective is to recover the masked region. We then run this attack on both training and testing images, and compare the attack efficacy on each. Specifically, for an image , we mask some fraction of pixels to create a masked image , and then use the trained model to reconstruct the image as . The exact algorithm we use for inpainting is given in Lugmayr et al. .
Because diffusion model inpainting is stochastic (it depends on the random sample ), we create a set of inpainted images , where we set . For each , we compute the diffusion model’s loss on this sample (at timestep 100) divided by a shadow model’s loss that was not trained on the sample. We then use this score to identify the highest-scoring reconstructions .
Comparing Diffusion Models to GANs
Are diffusion models more or less private than competing generative modeling approaches? In this section we take a first look at this question by comparing diffusion models to Generative Adversarial Networks (GANs) , an approach that has held the state-of-the-art results for image generation for nearly a decade.
Unlike diffusion models that are explicitly trained to memorize and reconstruct their training datasets, GANs are not. Instead, GANs consist of two competing neural networks: a generator and a discriminator. Similar to diffusion models, the generator receives random noise as input, but unlike a diffusion model, it must convert this noise to a valid image in a single forward pass. To train a GAN, the discriminator is trained to predict if an image comes from the generator or not, and the generator is trained to fool the discriminator. As a result, GANs differ from diffusion models in that their generators are only trained using indirect information about the training data (i.e., using gradients from the discriminator) because they never receive training data as input, whereas diffusion models are explicitly trained to reconstruct the training set.
We first propose a privacy attack methodology for GANs.While existing privacy attacks exist for GANs, they were proposed before the latest advancements in privacy attack techniques, requiring us to develop our own methods which out-perform prior work. We initially focus on membership inference attacks, where following Balle et al. , we assume access to both the discriminator and generator. We perform membership inference using the loss threshold and LiRA attacks, where we use the discriminator’s loss as the metric. To perform LiRA, we follow a similar methodology as Section 5 and train 256 individual GAN models each on a random split of the CIFAR-10 training dataset but otherwise leave training hyperparameters unchanged.
We study three GAN architectures, all implemented using the StudioGAN framework : BigGAN , MHGAN , and StyleGAN . Figure 14 shows the membership inference results. Overall, diffusion models have higher membership inference leakage, e.g., diffusion models had TPR at a FPR of as compared to TPR for GANs. This suggests that diffusion models are less private than GANs for membership inference attacks under default training settings, even when the GAN attack is strengthened due to having access to the discriminator (which would be unlikely in practice, as only the generator is necessary to create new images).
Data extraction results. We next turn our attention away from measuring worst-case privacy risk and focus our attention on more practical black-box extraction attacks. We follow the same procedure as Section 5.1, where we generate images from each model architecture and identify those that are near-copies of the training data using the same similarity function as before. Again we only consider non-duplicated CIFAR-10 training images in our counting. For this experiment, instead of using models we train ourselves (something that was necessary to run LiRA), we study five off-the-shelf pre-trained GANs: WGAN-ALP , E2GAN , NDA , DiffBigGAN , and StyleGAN-ADA . We also evaluate two off-the-shelf DDPM diffusion model released by Ho et al. and Nichol et al. . Note that all of these pre-trained models are trained by the original authors to maximize utility on the entire CIFAR-10 dataset rather than a random 50% split as in our prior models trained for MIA.
Table 1 shows the number of extracted images for each model and their corresponding FID. Overall, we find that diffusion models memorize more data than GANs, even when the GANs reach similar performance, e.g., the best DDPM model memorizes more than StyleGAN-ADA but reaches the same FID. Moreover, generative models (both GANs and diffusion models) tend to memorize more data as their quality (FID) improves, e.g., StyleGAN-ADA memorizes more images than the weakest GANs.
Using the GANs we trained ourselves, we show examples of the near-copy generations in Figure 15 for the three GANs that we trained ourselves, and Figure 24 in the Appendix shows every sample that we extract for those models. The Appendix also contains near-copy generations from the five off-the-shelf GANs. Overall, these results further reinforce the conclusion that diffusion models are less private than GAN models.
We also surprisingly find that diffusion models and GANs memorize many of the same images. In particular, despite the fact that our diffusion model memorizes 1280 images and a StyleGAN model we train on half of the dataset memorizes 361 images, we find that 244 unique images are memorized in common. If images were memorized uniformly at random, we should expect on average images would be memorized by both, giving exceptionally strong evidence that some images are inherently less private than others. Understanding why this phenomenon occurs is a fruitful direction for future work.
Defenses and Recommendations
Given the degree to which diffusion models memorize and regenerate training examples, in this section we explore various defenses and practical strategies that may help to reduce and audit model memorization.
Unfortunately, deduplication is not a perfect solution. To better understand the effectiveness of data deduplication, we deduplicate CIFAR-10 and re-train a diffusion model on this modified dataset. We compute image similarity using the imagededup tool and deduplicate any images that have a similarity above . This removes examples from the total examples in CIFAR-10. We repeat the same generation procedure as Section 5.1, where we generate images from the model and count how many examples are regenerated from the training set. The model trained on the deduplicated data regenerates examples, as compared to for the original model. While not a substantial drop, these results show that deduplication can mitigate memorization. Moreover, we also expect that deduplication will be much more effective for models trained on larger-scale datasets (e.g., Stable Diffusion), as we observed a much stronger correlation between data extraction and duplication rates for those models.
2 Differentially-Private Training
The gold standard technique to defend against privacy attacks is by training with differential privacy (DP) guarantees . Diffusion models can be trained with differentially-private stochastic gradient descent (DP-SGD) , where the model’s gradients are clipped and noised to prevent the model from leaking substantial information about the presence of any individual image in the dataset. Applying DP-SGD induces a trade-off between privacy and utility, and recent work shows that DP-SGD can be applied to small-scale diffusion models without substantial performance degradation .
Unfortunately, we applied DP-SGD to our diffusion model codebase and found that it caused the training on CIFAR-10 to consistently diverge, even at high values for (the privacy budget, around 50). In fact, even applying a non-trivial gradient clipping or noising on their own (both are required in DP-SGD) caused the training to fail. We leave a further investigation of these failures to future work, and we believe that new advances in DP-SGD and privacy-preserving training techniques may be required to train diffusion models in privacy-sensitive settings.
3 Auditing with Canaries
In addition to implementing defenses, it is important for practitioners to empirically audit their models to determine how vulnerable they are in practice . Our attacks above represent one method to evaluate model privacy. Nevertheless, our attacks are expensive, e.g., our membership inference results require training many shadow models, and thus lighter weight alternatives may be desired.
One such alternative is to insert canary examples into the training set, a common approach to evaluate memorization in language models . Here, one creates a large “pool” of canaries, e.g., by randomly generating noise images, and inserts a subset of the canaries into the training set. After training, one computes the exposure of the canaries, which roughly measures how many bits were learned about the inserted canaries as compared to the larger pool of not inserted canaries. This loss-based metric only requires training one model and can also be designed in a worst-case way (e.g., adversarial worst-case images could be used).
To evaluate exposure for diffusion models, we generate canaries consisting of uniformly generated noise. We then duplicate the canaries in the training set at different rates and measure the maximum exposure. Figure 16 shows the results. Here, the maximum exposure is 10, and some canaries reach this exposure after being inserted only twice. The exposure is not strictly increasing with duplicate count, which may be a result of some canaries being “harder” than others, and, ultimately, random canaries we generate may not be the most effective canaries to use to test memorization for diffusion models.
Related Work
Numerous past works study memorization in generative models across different domains, architectures, and threat models. One area of recent interest is memorization in language models for text, where past work shows that adversaries can extract training samples using two-step attack techniques that resemble our approach . Our work differs from these past results because we focus on the image domain and also use more semantic notions of data regeneration (e.g., using CLIP scores) as opposed to focusing on exact verbatim repetition (although recent language modeling work has begun to explore approximate memorization as well ).
Aside from language modeling, past work also analyzes memorization in image generation, mainly from the perspective of generalization in GANs (i.e., the novelty of model generations). For instance, numerous metrics exist to measure similarity with the training data , the extent of mode collapse , and the impact of individual training samples . Moreover, other work provides insights into when and why GANs may replicate training examples , as well as how to mitigate such effects . Our work extends these lines of inquiry to conditional diffusion models, where we measure novelty by computing how frequently models regenerate training instances when provided with textual prompts.
Recent and concurrent work also studies privacy in image generation for both GANs and diffusion models . Tinsley et al. show that StyleGAN can generate individuals’ faces, and Somepalli et al. show that Stable Diffusion can output semantically similar images to its training set. Compared to these works, we identify privacy vulnerabilities in a wider range of systems (e.g., Imagen and CIFAR models) and threat models (e.g., membership inference attacks).
Discussion and Conclusion
State-of-the-art diffusion models memorize and regenerate individual training images, allowing adversaries to launch training data extraction attacks. By training our own models we find that increasing utility can degrade privacy, and simple defenses such as deduplication are insufficient to completely address the memorization challenge. We see that state-of-the-art diffusion models memorize more than comparable GANs, and more useful diffusion models memorize more than weaker diffusion models. This suggests that the vulnerability of generative image models may grow over time. Going forward, our work raises questions around the memorization and generalization capabilities of diffusion models.
Do large-scale models work by generating novel output, or do they just copy and interpolate between individual training examples? If our extraction attacks had failed, it may have refuted the hypothesis that models copy and interpolate training data; but because our attacks succeed, this question remains open. Given that different models memorize varying amounts of data, we hope future work will explore how diffusion models copy from their training datasets.
We raise four practical consequences for those who train and deploy diffusion models. First, while not a perfect defense, we recommend deduplicating training datasets and minimizing over-training. Second, we suggest using our attack—or other auditing techniques—to estimate the privacy risk of trained models. Third, once practical privacy-preserving techniques become possible, we recommend their use whenever possible. Finally, we hope our work will temper the heuristic privacy expectations that have come to be associated with diffusion model outputs: synthetic data does not give privacy for free .
On the whole, our work contributes to a growing body of literature that raises questions regarding the legal, ethical, and privacy issues that arise from training on web-scraped public data . Researchers and practitioners should be wary of training on uncurated public data without first taking steps to understand the underlying ethics and privacy implications.
Contributions
Nicholas, Jamie, Vikash, and Eric each independently proposed the problem statement of extracting training data from diffusion models.
Nicholas, Eric, and Florian performed preliminary experiments to identify cases of data extraction in diffusion models.
Milad performed most of the experiments on Stable Diffusion and Imagen, and Nicholas counted duplicates in the LAION training dataset; each wrote the corresponding sections of the paper.
Jamie performed the membership inference attacks and inpainting attacks on CIFAR-10 diffusion models, and Nicholas performed the diffusion extraction experiments; each wrote the corresponding sections of the paper.
Matthew ran experiments for canary memorization and wrote the corresponding section of the paper.
Florian and Vikash performed preliminary experiments on memorization in GANs, and Milad and Vikash ran the experiments included in the paper.
Milad ran the membership inference experiments on GANs.
Vikash ran extraction experiments on pretrained GANs.
Daphne and Florian improved figure clarity and presentation.
Daphne, Borja, and Eric edited the paper and contributed to paper framing.
Nicholas organized the project and wrote the initial paper draft.
Acknowledgements and Conflicts of Interest
The authors are grateful to Tom Goldstein, Olivia Wiles, Katherine Lee, Austin Tarango, Ian Wilbur, Jeff Dean, Andreas Terzis, Robin Rombach, and Andreas Blattmann for comments on early drafts of this paper.
Nicholas, Milad, Matthew, and Daphne are employed at Google, and Jamie and Borja are employed at DeepMind, companies that both train large machine learning models (including diffusion models) on both public and private datasets.
Eric Wallace is supported by the Apple Scholars in AI/ML Fellowship.
References
Appendix A Collected Details for Figures
Appendix B All CIFAR-10 Memorized Images
Appendix C Additional Attacks on CIFAR-10
Here, we expand on our investigation of memorization of training data on CIFAR-10.
In Section 5.2.3, we implicitly investigated membership attack success as a function of the number update steps when training a diffusion model. We explicitly model this relationship in Figure 18. First, in Figure 18(a) we plot membership attack success as a function of the number of times that an example was processed over training. If an example is processed more than 2000 times during training, invariably membership attacks are perfect against that example. Second, in Figure 18(b), we plot membership attack success as a function of the total amount of data processed during training. Unsurprisingly, membership attack success increases as more training data is processed. This is highlighted in Figure 18(c), where we plot the membership attack ROC curve. At 5M training examples processed, at a FPR of 1% the TPR is 5%, and increases to 99% after 102M examples are processed. Note that this number of processed training inputs is commonly used in diffusion model training. For example, the OpenAI CIFAR-10 diffusion model https://github.com/openai/improved-diffusion is trained for 500,000 steps at a batch size of 128, meaning 64M training examples are processed. Even at this number of processed training examples, our membership attack has a TPR at a FPR of 1%.
C.2 Membership Inference with Different Augmentation Strategies
C.3 Membership Inference Inliers and Outliers
C.4 Membership Inference on Conditional and Unconditional Models
Diffusion models can be conditioned on labels (or prompts for text-to-image models). We compare the difference in membership inference on a CIFAR-10 diffusion model trained unconditionally with a model conditionally trained on CIFAR-10 labels. The conditional and unconditional models reach approximately the same FID after training; between 3.5-4.2 FID. We plot the membership attack ROC curve in Figure 21 and note that the conditional model is marginally more vulnerable. However, it is difficult to tell if this is a fundamental difference between conditional and unconditional models, or because the conditional model contains more parameters than unconditional model (the conditional models contains an extra embedding layer for the one-hot label input).
Appendix D More Inpainting Attacks on CIFAR-10
Figure 22 inspected the attack success when was in the training set. We show in Figure 23 that the attack fails when was not included in training; using a contrastive loss doesn’t signficantly increase the Pearson correlation coefficient. This means our attack is indeed exploiting the fact that the model can only inpaint correctly because of memorisation and not due to generalisation.
Appendix E GAN Training Setup
We used on StudioGANhttps://github.com/POSTECH-CVLab/PyTorch-StudioGAN codebase for training GAN in this work. For the StyleGAN and MHGAN architectures, we followed the default hyper-parameters provided in the StudioGAN repository. However, for the BigGAN architecture, we increased the number of training steps to 200,000, which is different from the original hyper-parameters, to increase image fidelity. We trained a total of 256 models for each GAN architecture, with each model being trained on a randomly selected half of the CIFAR-10 dataset. We selected the iteration that achieved the highest FID score on the test set for each model.
Appendix F Additional GAN Extraction Results
Figure 24 and Figure 25 contain additional examples extracted from GANs trained on CIFAR-10.