Membership Inference of Diffusion Models

Hailong Hu, Jun Pang

Introduction

Diffusion models have recently made remarkable progress in image synthesis , even being able to generate better-quality images than generative adversarial networks (GANs) in some situations . They have also been applied to sensitive personal data, such as the human face or medical images , which might unwittingly lead to the leakage of training data. As a consequence, it is paramount to study privacy breaches in diffusion models.

Membership inference (MI) attacks aim to infer whether a given sample was used to train the model . In practice, they are widely applied to analyze the privacy risks of a machine learning model . To date, a growing number of studies concentrate on classification models , GANs , text-to-image generative models , and language models . However, there is still a lack of work on MI attacks against diffusion models. In addition, data protection regulations, such as GDPR , require that it is mandatory to assess privacy threats of technologies when they are involving sensitive data. Therefore, all of these drive us to investigate the membership vulnerability of diffusion models.

In this paper, we systematically study the problem of membership inference of diffusion models. Specifically, we consider two threat models: in threat model I, adversaries are allowed to obtain the target diffusion model, which means that adversaries can calculate the loss values of a sample through the model. This scenario usually occurs when institutions share a generative model with their collaborators to avoid directly sharing original data . We emphasize that obtaining losses of a model is realistic because it is widely adopted in studying MI attacks on classification models . In threat model II, adversaries can obtain the likelihood value of a sample from a diffusion model. Providing the exact likelihood value of any sample is one of the advantages of diffusion models . Thus, here we aim to study whether the likelihood value of a sample can be considered as a clue to infer membership. Based on both threat models, two types of attack methods are developed respectively: loss-based attack and likelihood-based attack. They are detailed in Section 3.

We evaluate our methods on four state-of-the-art diffusion models: DDPM , SMLD , VPSDE and VESDE . We use two privacy-sensitive datasets: a human face dataset FFHQ and a diabetic retinopathy dataset DRD . Extensive experimental evaluations show that our attack methods can achieve superior attack performance (see Section 5). For instance, on the target model DDPM trained on FFHQ-1k, our loss-based attack can achieve a 100%100\% true positive rate (TPR) even at 0.01%0.01\% false negative rate (FPR) when the diffusion step is 200. Similar performance can be also seen on our likelihood-based attack where 71%71\% TPR at 0.01%0.01\% FPR can be obtained, which is 7,100 times more powerful than random guesses. We also analyze attack performance with respect to various factors (see Section 6). For example, we find that both attacks gradually become weak with the increase in the number of training samples. In addition, similar to FFHQ, both attacks can still achieve excellent performance on the medical dataset DRD. Finally, we also evaluate our attack performance on a classical defense — differential privacy (see Section 7). Specifically, we train target models using differentially-private stochastic gradient descent (DP-SGD) . Extensive evaluations show that although the performance of both types of attack can be alleviated on models trained with DP-SGD, they sacrifice too much model utility, which also gives a new research direction for the future.

Our contributions can be summarized as follows: (1) we perform the first study of MI attacks against diffusion models; (2) we propose two types of attacks to infer membership of diffusion models, showing that different diffusion steps of a diffusion model have significantly different privacy risks and the likelihood values of samples from a diffusion model are a strong clue to infer training samples; (3) we evaluate our attacks on one classical defense — diffusion models trained with DP-SGD, finding that it mitigates our attacks at the cost of the quality of synthetic samples. In the end, we want to emphasize that although we study membership inference from the perspective of attackers, our proposed methods can directly be applied to audit the privacy risks of diffusion models when model providers need to evaluate the privacy risks of their models.

Background: Diffusion Models

Diffusion models are a class of probabilistic generative models. They aim to learn the distribution of a training set, and the resulting model can be utilized to synthesize new data samples .

where λ(σt)\lambda(\sigma_{t}) is a coefficient function and ∇xtlog qσt(xt∣x)=−xt−xσt2\nabla_{x_{t}}\text{log}\>q_{\sigma_{t}}(x_{t}|x)=-\frac{x_{t}-x}{\sigma_{t}^{2}},

SSDE. Unlike prior works DDPM or SMLD which utilize a finite number of noise distributions, i.e. tt is discrete and usually at most TT, Song et al. propose a score-based generative framework through the lens of stochastic differential equations (SDEs), which can add an infinite number of noise distributions to further improve the performance of generative models. The forward process which adds an infinite number of noise distributions can be described as a continuous-time stochastic process. Specifically, the forward process of the score-based SDE (SSDE) is defined as:

where f(x,t)\text{f}(x,t), g(t)g(t) and dwd\text{w} are the drift coefficient, the diffusion coefficient and a standard Wiener process, respectively. The reverse process corresponds to a reverse-time SDE: dx=[f(x,t)−g(t)2∇xlog qt(x)]dt+g(t)dwˉdx=[\text{f}(x,t)-g(t)^{2}\nabla_{x}\text{log}~{}q_{t}(x)]\text{d}t+g(t)\text{d}\bar{w}, where wˉ\bar{w} is a standard Wiener process in the reverse time. Training of the SSDE is performed by minimizing the following loss:

The SSDE is a general and unified framework. Based on different coefficients in Equation 3, the variance preserving (VP) and variance exploding (VE) are instantiated. The VPSDE is defined as: dx=−12β(t)xdt+β(t)dw\text{d}x=-\frac{1}{2}\beta(t)x\text{d}t+\sqrt{\beta(t)}\text{d}w. The VESDE is defined as: dx=d[σ2(t)]dtdwdx=\sqrt{\frac{\text{d}[\sigma^{2}(t)]}{\text{d}t}}\text{d}w. Furthermore, the SSDE also shows the noise perturbations of DDPM and SMLD are discretizations of VP and VE, respectively. Note that, diffusion steps usually used in diffusion models also refer to time steps that are used in SDEs. In this work, we study the privacy risks of four target models: DDPM, SMLD, VPSDE, and VESDE.

Methodology

The objective of MI attacks is to infer if a sample was used to train a model. This provides model providers with a method to evaluate the information leakage of a machine learning model. In this section, we first introduce threat models and then present our MI methods.

Threat Model I. In this setting, we assume adversaries can only obtain the target model, i.e. the victim diffusion model. This setting is realistic because institutions are prone to share generative models with their collaborators instead of directly utilizing original data, considering privacy threats or data regulations . We emphasize that adversaries do not gain any knowledge of the training set. Obtaining the target model indicates that adversaries can get the loss values through the model, and this is realistic because most MI attacks on classification models also assume adversaries can get loss values . Under this threat model, we propose a loss-based MI attack.

Threat Model II. In this setting, we assume adversaries are able to have access to the likelihood values of samples from a diffusion model. Diffusion models have advantages in providing the exact likelihood value of any sample . Here we aim to study whether the likelihood values of samples can be utilized as a signal to perform membership inference. Under this threat model, we propose a likelihood-based MI attack.

2 Attack Methods

Problem Formulation. Given a target diffusion model GtarG_{\it tar}, the objective of membership inference attacks is to infer whether a sample xx from a target dataset XtarX_{\it tar} is used to train the GtarG_{\it tar}.

Loss-based Attack. For threat model I, we develop a loss-based attack. As introduced in Section 2, diffusion models can add an infinite or finite number of noise distributions, which are corresponding to continuous or discrete SDE, respectively. Therefore, we can calculate the loss value of a sample at each diffusion step tt. Specifically, based on Equation 1, the loss of a sample xx at tt diffusion step of DDPM is calculated by:

where mm is the dimension of xx and εθ⋆(.)\varepsilon_{\theta^{\star}}(.) is the trained network. By Equation 2, the loss of a sample xx at tt diffusion step of SMLD is calculated by:

where sθ⋆(.)s_{\theta^{\star}}(.) is the trained network. Based on Equation 4, the loss of a sample xx at tt diffusion step of VPSDE and VESDE is:

Then, we make a membership inference directly based on the loss value of a sample at one diffusion step. Namely, if a sample’s loss value is less than certain thresholds, this sample is marked as a member sample. For one sample, we can get TT or infinite losses, depending on continuous or discrete SDEs. In this work, in order to thoroughly demonstrate the performance of our attack, we compute losses of all diffusion steps TT for the discrete case. We randomly select TT diffusion steps for the continuous case although it has infinite steps.

Likelihood-based Attack. For threat model II, we propose to utilize the likelihood of a sample to infer membership. We compute the log-likelihood of a sample xx based on the following equation proposed by .

Experiments

We use two different datasets to evaluate our attack methods. They cover the human face and medical images, which are all considered privacy-sensitive data.

FFHQ. The Flickr-Faces-HQ dataset (FFHQ) is a new dataset that contains 70,00070,000 high-quality human face images. In this work, we randomly choose 1,0001,000 images to train target models. We also explore the effect of the size of the training set in Section 6.1.

DRD. The Diabetic Retinopathy dataset (DRD) contains 88,70388,703 retina images. In this work, we only consider images that have diabetic retinopathy, which is a total of 23,35923,359 images. Furthermore, we randomly choose 1,0001,000 images to train target models. Note that images in all datasets are resized to 64×6464\times 64.

2 Metrics

Evaluation metrics for diffusion models. We use the popular Fréchet Inception Distance (FID) metric to evaluate the performance of a diffusion model . A lower FID score is better, which implies that the generated samples are more realistic and diverse. Considering the efficiency of sampling, in our work the FID score is computed with all training samples and 1,0001,000 generated samples.

Evaluation metrics for MI attacks. We primarily use the full log-scale receiver operating characteristic (ROC) curve to evaluate the performance of our attack methods, because it can better characterize the worst-case privacy threats of a victim smodel . We also report the true-positive rate (TPR) at the false-positive rate (FPR) as it can give a quick evaluation. We use average-case metrics — accuracy as a reference, although it cannot assess the worst-case privacy.

3 Experimental Setup

In terms of target models, we use open source codes to train diffusion models, and their recommended hyperparameters about training and sampling are adopted. The number of training steps for all models is fixed at 500,000500,000. For discrete SDEs, TT is fixed as 1,0001,000 while TT is set as 1 for continuous SDEs. In terms of our attack methods, we evaluate the attack performance using all training samples as member samples and equal numbers of nonmember samples. The source code will be made public along with the final version of the paper.

Evaluation

In this section, we first present the performance of target models. Then, we show the performance of our two attacks: loss-based and likelihood-based attacks.

Considering their excellent performance in image generation, we choose DDPM , SMLD , VPSDE and VESDE as our target models. They are trained on the FFHQ dataset containing 1k samples. Target models with the best FID during the training progress are selected to be attacked. Table 1 shows the performance of the target models. Figure 9 in Appendix shows the qualitative results for these target models. Overall, all target models can synthesize high-quality images.

2 Performance of Loss-based Attack

We present our attack performance from two aspects: TPRs at fixed FPRs for all diffusion steps and log-scale ROC curves at one diffusion step. The former aims to provide the holistic performance of our attacks in diffusion models. In contrast, the latter concentrates on one diffusion step and is able to exhaustively show TPR values at a wide range of FPR values, which is key to assessing the worse-case privacy risks of a model.

TPRs at fixed FPRs for all diffusion steps. Figure 1 shows the performance of our loss-based attack on four target models trained on FFHQ. We plot TPRs at different FPRs with regard to diffusion steps for each target model. Recall DDPM and SMLD models are discrete SDEs while VPSDE and VESDE models are continuous SDEs. Thus, the number of diffusion steps for DDPM and SMLD is finite and is fixed as 1,000, while for VPSDE and VESDE models, we uniformly generate 1,0001,000 points within $andcomputecorrespondinglosses.Overall,allmodelsarevulnerabletoourattacks,evenundertheworst−case,i.e.TPRatand compute corresponding losses. Overall, all models are vulnerable to our attacks, even under the worst-case, i.e. TPR at0.01\%$ FPR, depicted by the purple line of Figure 1.

We observe that our attack presents different performances in different diffusion steps. There exist high privacy risk regions for diffusion models in terms of diffusion steps. In these regions (i.e. diffusion steps), our attack can achieve as high as 100%100\% TPR at 0.01%0.01\% FPR. Even for the SMLD model, close to 80%80\% TPR at 0.01%0.01\% FPR can be achieved. Recall the training mechanisms of diffusion models. Different levels of noise at different diffusion steps are added during the forward process. DDPM and VPSDE and VESDE are added growing levels of noise while SMLD starts with maximum levels of noise and gradually decreases the levels of noise. Thus, we can see that these models (DDPM and VPSDE, and VESDE) are more vulnerable to leak training samples in the first half part of the diffusion steps while the SMLD model shows membership vulnerability in the second half part of the diffusion steps. In brief, all models are prone to suffer from membership leakage in low levels of noise while they become more resistant in high levels of noise. In fact, in these diffusion steps where high levels of noise are added to training data, perturbed data is almost close to noise, which can to some degree enhance the privacy of training data. We also notice that at the starting diffusion step, our attack performance suffers from a decrease. This is because there is an instability issue at this step during the training process, which is discussed in the work . Despite this, these peak regions still show the effectiveness of our attack.

In addition, as shown in Figure 1, we also see that four curves that represent the true positive rate at different false positive rates almost overlap or are very close in most diffusion steps. It indicates that our attack can still be effective and robust even at the low false positive rate regime. Note that TPR at low FPR is able to characterize the worst-case privacy risks.

Log-scale ROC curves at one diffusion step. Figure 2 plots full log-scale ROC curves of the loss-based attack on four target models. We choose six different diffusion steps for each target model. The rules of choosing diffusion steps for discrete SDEs (i.e. DDPM and SMLD) are: starting and ending diffusion step and the diffusion step that experiences significant changes in terms of attack performance. For continuous SDEs (i.e. VPSDE and VESDE), we first get 1,0001,000 points that uniformly are sampled from $$. Then, we choose diffusion steps from these points based on the same rule of discrete SDEs. Overall, our excellent attack performance is exhaustively shown through log-scale ROC curves.

We can observe that when the levels of noise are not too large, our method can achieve a perfect attack, such as at t=200t=200 for the DDPM model, t=800t=800 for the SMLD model, and t=0.21t=0.21 for the VPSDE and VESDE models. Again, we can clearly see that the ROC curves on all target models are more aligned with the grey diagonal line with the increase in the magnitudes of noise. The grey diagonal line means that the attack performance is equivalent to random guesses. For example, the ROC curves are almost close to the grey diagonal line when the maximal level of noise is added, such as the DDPM model at t=999t=999, the SMLD model at t=0t=0, and the VPSDE and VESDE models at t=9.99×10−1t=9.99\times 10^{-1}. It is not surprising because at that time the input samples are perturbed as Gaussian noise data in theory and indeed do not have something with original training samples.

We also see that in certain ROC curves, even with the decrease in the FPR values, the TPR values still remain high. It indicates that the attack is still powerful even in the worst-case. Take the VPSDE model at t=0.72t=0.72 as an example, the TPR is 2×10−32\times 10^{-3} at the FPR value of 10−510^{-5}, which is 2020 times more powerful than random guesses.

Table 2 summarizes our attack performance on four target models with regard to diffusion steps and FPR values. We also report the average metric accuracy for reference. Here, we emphasize that only focusing on average metrics cannot assess the worst-case privacy risks. For instance, for the DDPM model at t=800t=800, the attack accuracy is 52.80%52.80\%, which indicates the model at this diffusion step almost does not lead to the leakage of training samples, because it is close to 50%50\% (the accuracy of random guesses). In fact, the TPR is 3×10−33\times 10^{-3} at the false positive rate of 10−510^{-5}, which is 3030 times more powerful than random guesses. It means that adversaries can infer confidently member samples under extremely low false positive rates.

Figure 3 shows perturbed data of four target models under different diffusion steps. The diffusion steps in Figure 3 are corresponding to that in Figure 2. We observe that even when some perturbed data that is almost not recognized by human beings is used to train the model, it seems not to prevent model memorization. For example, for the DDPM model at t=600t=600, the perturbed image is meaningless for humans. However, the attack accuracy is as high as 81.15%81.15\%. At the same time, the TPR at 0.01%0.01\% FPR is 2.30%2.30\%, which is 230230 times more powerful times than random guesses. It indicates that models trained on perturbed data, except for Gaussian noise data, can still leak training samples. The noise mechanism of diffusion models does not provide privacy protection.

Takeaways. Based on our analysis, when we utilize the loss of a diffusion model to mount a membership inference attack, the loss values from the low levels of noise (about the first 2/32/3 of all levels of noise) are a strong membership signal. During these diffusion steps, the model shows more privacy vulnerability.

3 Performance of Likelihood-based Attack

Figure 4(a) demonstrates our likelihood-based attack performance on four target models. Overall, our attacks still perform well on all target models. For example, our attack on the SMLD and VPSDE models almost remains 100%100\% true positive rates on all false positive rate regimes. For the VESDE model, attack results are slightly inferior to the SMLD models, yet still higher than the 10%10\% true positive rate at an extremely low 0.001%0.001\% false positive rate.

Table 3 shows our attack results at different FPR values for all target models. Once again, we can clearly see that even at the 0.01%0.01\% FPR, the lowest TPR among all models is as high as 23.10%23.10\%, which is 2,3102,310 times than random guesses. In particular, all training samples that used to train the SMLD model can be inferred at 100%100\% accuracy and 100%100\% TPR at all ranges of FPRs. In addition, we also observe that the attack accuracy is above 98%98\% for all target models. Our attack results also remind model providers that they should be careful when using likelihood values.

Analysis

In this section, we first analyze our attack performance with regard to the size of a training set. Then, we report our results on a medical image dataset DRD.

We study attack performance with regard to different sizes of the training set of a target model. Here, we choose the DDPM models trained on FFHQ as target models. We use FFHQ-1k, FFHQ-10k, and FFHQ-30k to represent different sizes of a dataset, which refer to 1,0001,000, 10,00010,000, and 30,00030,000 training samples in each dataset respectively. The FID of the target model DDPM trained on FFHQ-1k, FFHQ-10k, and FFHQ-30k are 57.8857.88, 34.3434.34, and 24.0624.06, respectively. In the following, we present the performance of our both attacks.

Performance of loss-based attack. Figure 5 depicts the performance of loss-based attacks on all diffusion steps under different sizes of a training set. Overall, we can observe that attack performance gradually becomes weak when the size of training sets increases. For example, at diffusion step t=200t=200, the TPR at 10%10\% FPR decreases from 100%100\% to about 15%15\% when the training samples increases from 11k to 3030k. When the training samples are 1010k, the TPR at 10%10\% FPR still remains above 65%65\%. Similar performances can be seen at 1%1\% FPR and 0.1%0.1\% FPR. Here, note that the starting points of the y-axis in Figure 5 are not 0 and we set them as the probability of random guesses. Thus, the lines that can be shown in the figure indicate this is an effective attack, at least better than random guesses. We further show the attack performance based on each dataset as a supplement in Figure 11 in Appendix.

As illustrated in Section 5.2, the peak regions still exist even if the number of training samples increases to 30k. For instance, as shown in Figure 5(c), it shows our attack performance of 0.1%0.1\% FPR on all models. Diffusion steps in the range of to 400400 are still vulnerable to our attack, compared to other steps. It indicates that these diffusion steps indeed lead a model to more easily leak training data.

Figure 6 shows ROC curves of our attack against target models trained on different sizes of training sets. Based on the same rules described in Section 5.2, we select several different diffusion steps and plot their ROC curves. We can see that indeed models become less vulnerable as the number of training samples increases. For instance, Figure 6(c) shows the DDPM trained on FFHQ-30K is more resistant to membership inference attacks on the full log-scale TPR-FPR curve. However, when diffusion step tt equals 250250, our attack shows higher attack performance than random guesses at the low false positive rate, such as 10−410^{-4}. This is also corresponding to the peak steps in Figure 5.

Performance of likelihood-based attack. Figure 4(b) shows the performance of likelihood-based attacks in terms of different sizes of training sets. Similar to the loss-based attack, the performance of the likelihood-based attack decrease with an increase in the sizes of training sets. Specifically, the likelihood-based attack shows perfect performance on the target model trained on FFHQ-1k. When the size of a training set increases to 1010K, there is a significant drop but still better than random guesses on the full log-scale ROC curve. In particular, in the extremely low false positive rate regime, such as 10−410^{-4}, the true positive rate is about 6×10−46\times 10^{-4}, which is 66 times more powerful than random guesses. In the model trained on FFHQ-30K, the ROC curve is almost close to the diagonal line, which indicates that adversaries are difficult to infer member samples through likelihood values.

2 Effects of Different Datasets

In this subsection, we show our attack performance on a medical image dataset about diabetic retinopathy. We have described this dataset DRD in Section 4.1. We choose the SMLD as the target model and the number of training samples is 1,0001,000. Overall, the SMLD model can achieve excellent performance in image synthesis, with an FID of 33.2033.20. Figure 10 in Appendix visualizes synthetic samples, which all show good quality.

Performance of loss-based attack. Figure 7 shows the performance of loss-based attacks for the target model SMLD trained on DRD. Here, note that the levels of the noise of the SMLD model gradually become small with an increase in diffusion steps.

Figure 7(a) shows the performance of our loss-based attack on all diffusion steps. We can again observe that our attacks can still perform perfectly on the DRD dataset at diffusion steps of low levels of noise. For instance, our attack can achieve at least 40%40\% TPR for four different FPRs when diffusion step tt is around 800800. In addition, TPR curves at different FPRs all show the same trend on all diffusion steps. To be specific, as shown in Figure 7(a), the true positive rate starts to increase from t=500t=500 and reaches a peak at t=800t=800. After that, it gradually decreases.

Figure 7(b) depicts ROC curves for different diffusion steps on target model SMLD trained on DRD. We can see that for diffusion steps where low levels of noise are added, such as t=700t=700, 800800, or 900900, training samples can be inferred at least 10%10\% TPR at as low as 0.001%0.001\% FPR. At high levels of noise diffusion step, such as t=0t=0, 200200, the performance of our method is almost close to random guesses.

Performance of likelihood-based attack. Figure 7(c) reports the performance of our likelihood-based attack on the SMLD model trained on DRD. As expected, our attack still shows excellent performance. We can clearly find that the attack achieves 100%100\% TPR on all FPR values, which means that all member samples are inferred correctly. Table 4 in Appendix reports the quantitative results of both attacks.

Defenses

Differential privacy (DP) is regarded as the gold standard for protecting the training set of a model. In practice, although differential privacy can guarantee individual-level privacy, it often sacrifices significantly model utility, especially for the quality of generated images, when it is applied to generative models. In this section, we present our attack results on diffusion models using the DP defense technology.

We adopt Differentially-Private Stochastic Gradient Descent (DP-SGD) to train diffusion models. DP-SGD is widely used for privately training a machine learning model. Generally, DP-SGD achieves differential privacy by adding noise into per-sample gradients. In our work, we implement DP diffusion models through the Opacus library . We set the clip bound CC and the failure probability δ\delta as 1 and 5×10−45\times 10^{-4}. The batch size and the number of epochs are 6464 and 1,8001,800. The final privacy budget ϵ\epsilon is 19.6219.62. We choose the DDPM model as the target model. It is trained on FFHQ containing 1,0001,000 training samples, and the FID is 393.94393.94.

Performance of loss-based attack. Figure 8 show the performance of both types of attacks on DDPM trained with DP-SGD on FFHQ. As described in Figure 8(a), we present the performance of loss-based attack on all diffusion steps. Clearly, we can see that differentially training DDPM, i.e. DDPM with DP-SGD indeed can significantly decrease the membership leakages. The peak regions can be still seen when diffusion steps are between 400400 and 800800. However, the true positive rates reduce to at most 15%15\% during these diffusion steps when the false positive rate is 10%10\%.

Figure 8(b) further shows ROC curves of our loss-based attack on different diffusion steps. Overall, ROC curves still remain close to the diagonal line before the false positive rate is 1%1\%, i.e. 10−210^{-2}. It indicates that adversaries almost make a random guess. When the false positive rates continue to decrease from 0.1%0.1\% to 0.001%0.001\%, different diffusion steps show divergences. For diffusion steps at 500 or 600, the true positive rates keep at 1%1\%. In contrast, the true positive rates reduce to 0%0\% at diffusion step t=999t=999. It means that in the worst-case, some training samples are still inferred with a probability higher than random guesses.

Performance of likelihood-based attack. Figure 8(c) show the performance of likelihood-based attack on DDPM training with DP-SGD on FFHQ. Again, we can see that differentially private training a diffusion model indeed can mitigate our attack. At the same time, we also see at the low false positive rate regime, our attack still remains at 0.1% true positive rate, which illustrates the effectiveness of our attack even in the worst-case. Here, we also note that the FID of the target model is 393.94393.94, which means that the utility of the target model suffers from a server performance drop. We leave developing more usable techniques to train a diffusion model with DP-SGD as future work. Table 5 in Appendix summarizes the quantitative results of both attacks.

Related work

Diffusion Models. Diffusion Models have attracted increasing attention in the past years. Sohl-Dickstein et al. first introduce nonequilibrium thermodynamics to build generative models. The key idea is to slowly add noise into data in the forward process and learn to generate data from noise through a reverse process. Ho et al. further propose to use parameterization techniques in diffusion models, which enable diffusion models to generate high-quality images. Song et al. present to train a generative model by estimating gradients of data distribution, i.e. score. Furthermore, Song et al. propose a unified framework to describe these diffusion models through the lens of stochastic differential equations. Beyond image synthesis, diffusion models are also applied to various domains, such as image restoration , and text-to-image translation , even audio and video synthesis . However, in this work, we study diffusion models from the perspective of privacy.

Membership Inference Attacks. There are extensive works on membership inference (MI) attacks on classification models . Various attack methods under different threat models are proposed, such as using fewer shadow models , using loss values and using labels of victim models .

In addition to classification models, there are several MI attacks on generative models . Hayes et al. leverage the discriminator of a GAN to mount attacks. Chen et al. perform MI attacks by finding a reconstructed sample on the generator of a GAN. Nevertheless, all attacks are more specific to GANs and heavily rely on the unique characteristics of GANs, such as discriminators or generators. They cannot be extended to diffusion models, because diffusion models have different training and sampling mechanisms. Therefore, our work on membership inference of diffusion models aims to fill this gap.

Another recent work proposed by Somepalli et al. investigates data replication in diffusion models. However, their work is different from our work. Data replication assumes that adversaries can have the whole training set. Given a generated sample from the diffusion model, they search the training set based on similarity metrics. If the similarity value is higher than a threshold, it is considered a replication for this generated sample. In contrast, our work does not assume that adversaries obtain the training set. Our work aims to infer whether a training sample is used to train the model, given a diffusion model.

Conclusion

In this paper, we have presented the first study about membership inference of diffusion models. We have developed two types of attack methods: loss-based attack and likelihood-based attack. We have evaluated our methods on four state-of-the-art diffusion models and two privacy-related datasets (human faces and medical images).

Our evaluations have demonstrated that diffusion models are vulnerable to membership inference attacks. To be more specific, our loss-based attack shows that when utilizing loss values from diffusion steps where low levels of noise are added, training samples can be inferred with high true positive rates at low false positive rates, such as 100%100\% TPR at 0.01%0.01\% FPR. Our likelihood-based attack again illustrates that adversaries can achieve the same perfect performance. Although membership inference becomes more challenging with the increase in the number of training samples, attack performance in the worst-case, i.e. TPR at low FPR, is still significantly higher than random guesses. Our experimental results on classic privacy protection mechanisms, i.e. diffusion models trained with DP-SGD, further show that DP-SGD alleviates our attacks at the expense of severe model utility.

Designing an effective differential privacy strategy to produce high-quality images for diffusion models is still a promising and challenging direction. We will take this as one of our future works. In addition, it is an interesting direction to study MI attacks of diffusion models in stricter scenarios, such as only obtaining synthetic data.

Acknowledgements: This research was funded in whole by the Luxembourg National Research Fund (FNR), grant reference 13550291.

References

Appendix 0.A Appendix

In this section, we show additional results and introduce each result in its caption.