Differentially Private Synthetic Data via Foundation Model APIs 1: Images
Zinan Lin, Sivakanth Gopi, Janardhan Kulkarni, Harsha Nori, Sergey Yekhanin
Introduction
While data-driven approaches have been successful, privacy is a major concern. For example, statistical queries of a dataset may leak sensitive information about individual users . Entire training samples can be reconstructed from deep learning models . For example, extracted some of the training images almost exactly from generations of Stable Diffusion model. Differential privacy (DP) is the gold standard in quantifying and mitigating these privacy concerns . A DP algorithm ensures that information about individual samples in the original data cannot be inferred with high confidence from algorithm outputs. Differentially private synthetic data is the holy grail of DP research . The goal is to generate a synthetic dataset that is statistically similar to the original data while ensuring DP. This has several benefits: (1) Thanks to the post-processing property of DP , we can use any existing non-private algorithm (such as training machine learning (ML) models) as-is on the synthetic data without incurring additional privacy loss. This is more scalable than redesigning and reimplementing every algorithm for DP. (2) Synthetic data can be shared freely with other parties without violating privacy. This is useful in situations when sharing data is necessary, such as when organizations (such as hospitals) want to release datasets to support open research initiatives . (3) Since synthetic data is DP, developers can look at the data directly which makes algorithm development and debugging a lot easier.
At the same time, with the recent advancement of powerful large foundation models, API-based solutions are gaining tremendous popularity, exemplified by the surge of GPT4-based applications. In contrast to the traditional paradigm that trains (or fine-tunes) customized ML models for each application, API-based solutions treat ML models as a blackbox and only utilize APIsSee https://platform.openai.com/docs/introduction for examples of APIs. For example, a text completion API can complete a text prompt using a foundation model such as GPT4. An image variation API can produce variations of a given image using a foundation model such as DALLE2. that provide the input/output functions of the model. In fact, many foundation models including GPT4, Bard, and DALLE2 only provide API access without releasing model weights and code. Key reasons for the success of API-based solutions are that APIs offer a clean abstraction of ML and are readily available and scalable. Therefore, implementing and deploying these API-based algorithms is easier and faster even for developers without ML expertise. Such an approach can also leverage the power of foundation models that are only released through APIs. Unfortunately, the SOTA DP synthetic data algorithms today are still in the old paradigm: they need a customized training process for each dataset, whose implementation requires significant ML engineering efforts (§ 3).
Motivated from these observations, we ask the following ambitious question (Fig. 1):
Can we generate DP synthetic data using blackbox APIs of foundation models?
In our threat model, API providers are untrusted entities and we also want to protect user privacy from them, i.e., the API queries we make during generation should also be DP. If successful, we can potentially democratize the deployment of DP synthetic data in the industry similar to how API-based approaches have facilitated other applications. This is a challenging task, however, as we do not have access to model weights and gradients by assumption. In this paper, we conduct the first exploration of the potential and limits of this vision on DP synthetic images. Surprisingly, we show that not only is such a vision realizable, but that it also has the potential to match or improve SOTA training-based DP synthetic image algorithms despite more restrictive model access. Our contributions are:
(1) New problem (§ 3). We highlight the importance of DP Synthetic Data via APIs (DPSDA). Such algorithms are easy to implement and deploy and can leverage the foundation models behind APIs.
(2) New framework (§ 4). We propose an algorithm called Private Evolution (PE) for achieving our goal (Fig. 2). We consider using 2 popular APIs: random generation and sample variation (i.e., generating a sample similar to the given one).Footnotes 3 and 4 The key idea is to iteratively use private samples to vote for the most similar samples generated from the blackbox model and ask the blackbox models to generate more of those similar samples. We theoretically prove that the distribution of the generated samples from PE will converge to the private distribution under some modeling assumptions (§ 5). PE only requires (existing) APIs of the foundation models, and does not need any model training.
(3) Experimental results (§ 6). Some key results are: (a) Surprisingly, without any training, PE can still outperform SOTA training-based DP image generation approaches on some datasets (Figs. 4 and 4). For example, to obtain FID on CIFAR10 dataset, PE (with blackbox access to an ImageNet-pre-trained model) only needs . In contrast, DP fine-tuning of an ImageNet-pre-trained model (prior SOTA) requires . (b) We show that PE works even when there is significant distribution shift between private and public data. We create a DP synthetic version (with ) of Camelyon17, a medical dataset for classification of breast cancer metastates (using the same ImageNet-pre-trained model). A downstream classifier trained on our DP synthetic data achieves a classification accuracy of 79.56% (prior SOTA based on DP fine-tuning is 91.1% with ). (c) We set up new challenging benchmarks that the DP synthetic image literature has not studied before. We show that with powerful foundation models such as Stable Diffusion , PE can work with high-resolution (512x512) image datasets with a small size (100), which are common in practice but challenging for current DP synthetic image algorithms.
Background and Related Work
DP synthetic data. Given a private dataset , the goal is to generate a DP synthetic dataset which is statistically similar to . One method is to train a generative model from scratch on private data with DP-SGD , a DP variant of stochastic gradient descent. Later studies show that pre-training generative models on public data before fine-tuning them on private data with DP-SGD gives better a privacy-utility trade-off due to knowledge transfer from public data , smaller gradient spaces , or better initialization . This approach achieves SOTA results on several data modalities such as text and images. In particular, DP-Diffusion achieves SOTA results on DP synthetic images by pre-training diffusion models on public datasets and then fine-tuning the model on the private dataset. Some other methods do not depend on DP-SGD . For example, DP-MEPF trains generative models to produce synthetic data that matches the (privatized) statistics of the private features.
Note that all the above methods obtain generative models whose weights are DP, which can then be used to draw DP synthetic data. It is stronger than our goal which only requires DP synthetic data . In this paper, we do not do any model training and only produce DP synthetic data.
DP Synthetic Data via APIs (DPSDA)
As discussed in § 2, SOTA DP synthetic data algorithms require training or fine-tuning generative models with DP-SGD. There are some obstacles to deploying them in practice.
(1) Significant engineering effort. As discussed in § 1, deploying normal ML training pipelines is hard. Deploying DP training pipelines is even harder because most ML infrastructure is not built around this use case. Recently, there has been significant progress in making DP training more efficient and easy to use (Opacus and Tensorflow Privacy). However, incorporating them in new codebases and new models is highly non-trivial. For example, to use Opacus, we need to implement our own per-sample gradient calculator for new layers. Common layers and loss functions that depend on multiple samples (e.g., batch normalization) are often not supported.
(2) Inapplicability of API-only models. It may be appealing to take advantage of the powerful foundation models in DP synthetic data generation. However, due to the high commercial value of foundation models, many companies choose to only release inference APIs of the models but not the weights or code. Examples include popular models such as DALLE 2 and GPT 3/4 from OpenAI and Bard from Google. In such cases, existing training-based approaches are not applicable.Some companies also provide model fine-tuning APIs, e.g., https://platform.openai.com/docs/guides/fine-tuning. However, they do not support DP fine-tuning and do not provide gradients. Also, uploading sensitive data to these APIs controlled by other companies can lead to privacy violations.
A DP synthetic data approach that only requires model inference APIs could potentially be deployed more easily. As the approach only depends on APIs, it does not need any modifications inside the model and requires minimal modifications when we switch to a different model (as long as they support the same APIs). Thus, it is easier to implement and use even for people without ML and DP expertise. In addition, such an approach is compatible with the models behind APIs.
2 Problem Formulation
We now give a formal statement of DPSDA. We first define a core primitive for DP synthetic data.
DPSDA. We want to solve DPWA where is given blackbox access to foundation models trained on public data via APIs.Footnote 1 API queries should also be -DP as API providers cannot be trusted.
3 Scope of This Work
Data type. While our framework above and our algorithms in § 4 are general for any data type, we focus on images in our experiments. We consider both unconditional (i.e., no ) and conditional generation tasks (e.g., can be image categories such as cats or dogs).
APIs. In our algorithm design and experiments, we use 2 APIs, both of which are either directly provided in the APIs of popular models (e.g., DALLE 2,See https://platform.openai.com/docs/guides/images/usage. Stable DiffusionSee https://huggingface.co/docs/diffusers/api/pipelines/stable_diffusion/overview.) or can be easily implemented by adapting current APIs (e.g., using appropriate text prompts in GPT APIsFootnote 1):
(1) that randomly generates samples. Some APIs also accept condition information such as text prompts in text-to-image generation.Footnotes 3 and 4 For simplicity, we omit the condition information in the argument.
(2) that generates variations for each sample in . For images, it means to generate similar images to the given one, e.g., with similar colors or objects.Footnotes 3 and 4 Some APIs also support setting the variation degree: , where larger indicates more variation.If this is not implemented, we can simply compose the VARIATION_API times to achieve it.
Private Evolution (PE)
Foundation models have a broad and general model of our world from their extensive training data. Therefore, we expect that foundation models can generate samples close to private data with non-negligible probability. The challenge is that by naively calling the APIs, the probability of drawing such samples is quite low. We need a way to guide the generation towards private samples.
Inspired by evolutionary algorithms (EA) (App. B), we propose Private Evolution (PE) framework for generating DP synthetic data via APIs. See Fig. 2 for the intuition behind PE. The complete algorithm is in Alg. 1. Below, we discuss the components in detail.
Initial population (Alg. 1). We use RANDOM_API to generate the initial population. If there is public information about the private samples (e.g., they are dog images), we can use this information as prompts to the API to seed a better initialization.
Fitness function (Alg. 1 or Alg. 2). We need to evaluate how useful each sample in the population is for modeling the private distribution. Our idea is that, if a sample in the population is surrounded by many private samples, then we should give it a high score. To implement this, we define the fitness function of a sample as the number of private samples whose nearest neighbor in the population is . A higher fitness value means that more private samples are closest to it. More details are below:
where is a network for extracting image embeddings such as inception embedding or CLIP embedding . In general, which distance function to use will depend on the particular application and choosing the right one is critical for the algorithm to work well.
(2) Lookahead. The above approach gives high scores to the good samples in the current population. However, as we will see later, these good samples will be modified through VARIATION_API for the next population. Therefore, it is better to “look ahead” to compute the distance based on the modified samples as if they are kept in the population. We modify Eq. 1 to compute the distance between the embedding of and the mean embedding of variations of : , where is called lookahead degree, and are variations of obtained via VARIATION_API.
(3) Noise for DP (Alg. 2, Alg. 2)). Because this step utilizes private samples, we need to add noise to ensure DP. We add i.i.d. Gaussian noise from . The privacy analysis is presented in § 4.2.
(4) Thresholding (Alg. 2, Alg. 2)). When the number of generated samples is large, the majority of the histogram will be DP noise added above. To make the signal-noise ratio larger, we set a threshold to each bin of the histogram. Similar ideas have been used in DP set union .
In summary, we called the above fitness function DP Nearest Neighbors Histogram.
Parent selection (Alg. 1). We draw samples from the population according to the DP Nearest Neighbors Histogram so that a sample with more private samples around is more likely to be selected.
Offspring generation (Alg. 1). We use VARIATION_API to get variants of the parents as offsprings.
Conditional generation. The above procedure is for unconditional generation. To support conditional generation, i.e., each generated sample is associated with a label such as an image class (e.g., cats v.s. dogs), we take a simple approach: we repeat the above process for each class of samples in the private dataset. See Alg. 3 for the full algorithm.
2 Privacy Analysis
This privacy analysis implies that releasing all the (intermediate) generated sets also satisfies the same DP guarantees. Therefore PE provides the same privacy even from the API provider.
Theoretical Evidence for Convergence of PE
In this section, we will give some intuition for why PE can solve DPWA.
Thus, we should expect that PE will discover every cluster of private data of size in iterations. We now compare this to previous work on DP clustering. gives an algorithm for densest ball, where they show an -DP algorithm which (approximately) finds any ball of radius which has at least private points. Thus intuitively, we see that PE compares favorably to SOTA DP clustering algorithms (though we don’t have rigorous proof of this fact). If this can be formalized, then PE gives a very different algorithm for densest ball, which in turn can be used to solve DP clustering. Moreover PE is very amenable to parallel and distributed implementations. We therefore think this is an interesting theory problem for future work.
Experiments
In § 6.1, we compare PE with SOTA training-based methods on standard benchmarks to understand its promise and limitation. In § 6.2, we present proof-of-concept experiments to show how PE can utilize the power of large foundation models. We did (limited) hyper-parameter tunings in the above experiments; following prior DP synthetic data work , we ignore the privacy cost of hyper-parameter tuning. However, as we will see in the ablation studies (§ 6.3 and K), PE stably outperforms SOTA across a wide range of hyper-parameters, and the results can be further improved with better hyper-parameters than what we used. Detailed hyper-parameter settings and more results such as generated samples and their nearest images in the private dataset are in Apps. H, I and J.
Public information. We use standard benchmarks which treat ImageNet as public data. For fair comparisons, we only use ImageNet as public information in PE: (1) Pre-trained model. Unlike the SOTA which trains customized diffusion models, we simply use public ImageNet pre-trained diffusion models (pure image models without text prompts) . (2) Embedding (Eq. 1). We use ImageNet inception embedding . PE is not sensitive to embedding choice though and we get good results even with CLIP embeddings (Fig. 32 in App. K).
Baselines. We compare with DP-Diffusion , DP-MEPF , and DP-GAN . DP-Diffusion is the current SOTA that achieves the best results on these benchmarks.
Outline. We test PE on private datasets that are either similar to or differ a lot from ImageNet in § 6.1.1 and 6.1.2. We demonstrate that PE can generate an unlimited number of useful samples in § 6.1.3. We show that PE is computationally cheaper than train-based methods in App. L.
We treat CIFAR10 as private data. Given that both ImageNet and CIFAR10 are natural images, it is a relatively easy task for PE (and also for the baselines). Figs. 4, 5 and 4 show the results. Surprisingly, despite the fact that we consider a strictly more restrictive model access and do not need training, PE can outperform the SOTA training-based methods. Details are below.
Sample quality v.s. privacy. Fig. 4 shows the trade-off between privacy cost and FID, a popular metric for image quality . For either conditional or unconditional generation, PE outperforms the baselines significantly. For example, to reach FID, requires , cannot achieve it even with infinity , whereas our PE only needs .
Downstream classification accuracy v.s. privacy. We train a downstream WRN-40-4 classifier from scratch on 50000 generated samples and test the accuracy on CIFAR10 test set. This simulates how users would use synthetic data, and a higher accuracy means better utility. Fig. 5 shows the results (focus on the left-most points with num of generated samples = 50000 for now). achieves 51% accuracy with (not shown). Compared with the SOTA , PE achieves better accuracy (+6.1%) with less privacy cost. Further with an ensemble of 5 classifiers trained on the same data, PE is able to reach an accuracy of 84.8%.
The above results suggest that when private and public images are similar, PE is a promising framework given its better privacy-utility trade-off and the API-only requirement.
1.2 Large Distribution Shift (ImageNet →→\rightarrow Camelyon17)
Next, we consider a hard task for PE, where the private dataset is very different from ImageNet. We use Camelyon17 dataset as private data which contains 302436 images of histological lymph node sections with labels on whether it has cancer (real images in Figs. 6 and 18). Despite the large distribution shift, training-based methods can update the model weights to adapt to the private distribution (given enough samples). However, PE can only draw samples from APIs as is.
We find that even in this challenging situation, PE can still achieve non-trivial results. Following , we train a WRN-40-4 classifier from scratch on 302436 generated samples and compute the test accuracy. We achieve 79.56% accuracy with -DP. Prior SOTA is 91.1% with -DP. Random guess is 50%. Fig. 6 (more in Fig. 18) shows that generated images from PE are very different from ImageNet but similar to Camelyon17. Fig. 19 further shows how the generated images are gradually moved towards Camelyon17 across iterations.
These results demonstrate the effectiveness of PE. But when public models that are similar to private data are not available and when there is enough private data, the traditional training-based methods are still more promising at this point if the privacy-utility trade-off is the only goal. However, given the benefit of API-only assumption and the non-trivial results that PE already got, it is worth further exploiting the potential of PE in future work. Indeed, we expect these results can be improved with further refining PE (App. K).
1.3 Generating Unlimited Number of Samples
We use the approach in § 4.1 to generate more synthetic samples from § 6.1.1 and train classifiers on them. The results are in Fig. 5. Similar as , the classifier accuracy improves as more generated samples are used. With an ensemble of 5 classifiers, we reach 89.13% accuracy with 1M samples. This suggests that PE has the same capability as training-based methods in generating an unlimited number of useful samples. At the same time, we see that the gap between PE and DP-Diffusion diminishes as more samples are used. We hypothesize that it is due to the limited improvement space: As shown in , even using a ImageNet pre-trained classifier, the best accuracy DP-Diffusion achieves is close to the best points in Fig. 5.
2 More Challenging Benchmarks with Large Foundation Models
We demonstrate the feasibility of applying PE on large foundation models with Stable Diffusion .
Data. Ideally we want to experiment with a dataset that has no overlap with Stable Diffusion’s training data.The training set of Stable Diffusion is public. However, it is hard to check if a public image or its variants (e.g., cropped, scaled) have been used to produce images in it. Therefore, we resort to our own private data. We take the safest approach: we construct two datasets with photos of the author’s two cats that have never been posted online. Each dataset has 100 512x512 images. Such high-resolution datasets with a small number of samples represent a common need in practice (e.g., in health care), but are challenging for DP synthetic data: to the best of our knowledge, no prior training-based methods have reported results on datasets with a similar resolution or number of samples. The dataset is released at https://github.com/microsoft/DPSDA as a new benchmark. See App. J for all images.
API implementation. We use off-the-shelf Stable Diffusion APIs (see App. J).
Results. We run Private Evolution for these two datasets with the same hyperparameters. Fig. 7 show examples of generated images for each of the cat datasets. We can see that Private Evolution correctly captures the key characteristics of these two cats. See App. J for all generated images.
Pre-trained network. Fig. 8 shows the results with two different ImageNet pre-trained networks: one is larger (270M) with ImageNet class labels as input; the other one is smaller (100M) without label input (see App. G for implementation details). In all experiments in § 6.1, we used the 270M network. Two takeaways are: (1) The 270M network (trained on the same dataset) improves the results. This is expected as larger and more powerful models can learn public distributions better. This suggests the potential of PE with future foundation models with growing capabilities. (2) Even with a relatively weak model (100M), PE can still obtain good results that beat the baselines (though with a slower convergence speed), suggesting the effectiveness of PE.
More ablation studies are in App. K, where we see that PE obtains good results across a wide range of hyper-parameters, and the results shown before can be improved with better hyper-parameters.
Discussion
Our work opens up new interesting research directions including:
Applications. (1) New privacy-preserving vision applications that were previously challenging but are now possible due to the new possibility of high-resolution DP synthetic images with small dataset sizes. (2) The use of PE in other data modalities beyond images such as texts.
Algorithms. (1) Minimizing number of API calls along with optimizing privacy vs utility tradeoff. (2) Leveraging the power of a large set of APIsFootnotes 3 and 4 beyond the two we used. (3) Algorithmic innovations to improve PE when the distributions of private data and foundation models are too different (§ 6.1.2). (4) Solving DPSDA in the Local/Shuffle DP model and in Federated Learning settings.
Acknowledgement
The authors would like to thank Sepideh Mahabadi for the insightful discussions and ideas, and Sahra Ghalebikesabi for the tremendous help in providing the experimental details of DP-Diffusion . The authors would like to extend their heartfelt appreciation to Cat Cookie and Cat Doudou for generously sharing their adorable faces in the new dataset, as well as to Wenyu Wang for collecting and pre-processing the photos.
References
Appendix A Definition of Wasserstein Distance
Appendix B A Brief Introduction to Evolutionary Algorithms
Evolutionary algorithms are inspired by biological evolution, and the goal is to produce samples that maximize an objective value. It starts with an initial population (i.e., a set of samples), which is then iteratively updated. In each iteration, it selects parents (i.e., a subset of samples) from the population according to the fitness function which describes how useful they are in achieving better objective values. After that, it generates offsprings (i.e., new samples) by modifying the parents, hoping to get samples with better objective values, and puts them in the population. By doing so, the population will be guided towards better objective values.
We cannot directly apply existing EA algorithms to our problem. Firstly, our objective is to produce a set of samples that are jointly optimal (i.e., closer to the private distribution, § 3.2), instead of optimizing an objective calculated from individual samples in typical EA problems. In addition, the differential privacy requirement and the restrictive model API access are unique to our problem. These differences require us to redesign all components of EA.
Appendix C More Details on Private Evolution
Appendix D Proofs of PE Convergence Theorems
We will slightly modify the algorithm as necessary to make it convenient for our analysis. We will make the following modeling assumptions:
where samples samples each from Gaussian distributions for where and where is the final Wasserstein distance.
If , then with probability at least , some point in will get noticeably closer to than , i.e.,
Let and let be such that . Note that such a exists since . We will now prove that one of the samples will get noticeably closer to than Let where
Note that is a random variable. By using upper tail bounds for distribution and union bound over , we can bound
D.2 Proof of Thm. 2
Appendix E Intrinsic Dimension of Image Embeddings
To illustrate the intrinsic dimension of image embeddings, we use the following process:
We (randomly) take an image from CIFAR10.
We use VARIATION_API from App. H to obtain 3000 image variations of : , and their corresponding inception embeddings . 3000 is chosen so that the number of variations is larger than the embedding dimension.
We construct a matrix , where is the mean.
We compute the singular values of : .
We compute the minimum number of singular values needed so that the explained variance ratioSee https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html. . Intuitively, this describes how many dimensions are needed to reconstruct the embedding changes with a small error. We use it as an estimated intrinsic dimension of the image variations.
We conduct the above process with the variation degree $M/\sqrt{3000}$ for variation degree=60 are in Fig. 10 (other variation degrees have similar trend). Two key observations are:
As the variation degree increases, the estimated intrinsic dimension also increases. This could be because the manifold of image embeddings is likely to be non-linear, the above estimation of intrinsic dimension is only accurate when we perturb the image to a small degree so that the changes in the manifold can still be well approximated by a linear subspace. Using a larger variation degree (and thus larger changes in the embedding space) will overestimate the intrinsic dimension.
Nevertheless, we always see that the singular values decrease rapidly (Fig. 10) and the estimated intrinsic dimension is much smaller than the embedding size 2048 (Fig. 9), which supports our hypothesis in Thm. 2.
Appendix F Relation of DPWA to prior work
In , to give an algorithm for DP Heatmaps, the authors study DP sparse EMDEarth’s Mover Distance, which is the another name for Wasserstein metric . aggregation problem where we need to output a distribution of points which approximates the distribution of private data in EMD distance (i.e., ). They study this problem only in two dimensions and the running time of their algorithms (suitably generalized to higher dimensions) will be exponential in the dimension .
We now explain why we can’t just use prior work on DP Clustering to solve DPSDA say for images.
Our PE algorithm does much better because:
It exploits the intrinsic dimension of the manifold of images in the embedding space which is much smaller than the embedding dimension (see Thm. 2 and App. E) and
There is no need to invert points in embedding space to the image space.
In an early experiment, we have tried DP clustering in the CLIP embedding space using the practical DP Clustering algorithm in . We then inverted the cluster centers (which are in the embedding space) using unCLIP. But we found the resulting images are too noisy compared to the images we get from PE and the FID scores are also significantly worse than that of PE.
Appendix G Implementation Details on Label Condition
There are two meanings of “conditioning” that appear in our work:
Whether the pre-trained networks or APIs (e.g., ImageNet pre-trained diffusion models used in § 6.1.1 and 6.1.2) support conditional input (e.g., ImageNet class label).
Whether the generated samples are associated with class labels from the private data.
In DP fine-tuning approaches, these two usually refer to the same thing: if we want to generate class labels for generated samples, the common practice is to use a pre-trained network that supports conditional input . However, in PE, these two are completely orthogonal.
Conditional pre-trained networks/APIs. We first explain our implementation when the pre-trained networks or APIs support conditional inputs such as class labels or text prompts. When generating the initial population using RANDOM_API (Alg. 1), we will either randomly draw labels from all possible labels when no prior public information is available (which is what we do in CIFAR10 and Camelyon17 experiments where we randomly draw from all possible ImageNet classes), or use the public information as condition input (e.g., the text prompt used in Stable Diffusion experiments; see App. J). In the subsequent VARIATION_API calls (Alg. 1), for each image, we will use its associated class label or text prompt as the condition information to the API, and the output samples from VARIATION_API will be associated with the same class label or text prompt as the input sample. For example, if we use an image with “peacock” class to generate variations, all output images will be associated with “peacock” class for future VARIATION_API calls. Note that throughout the above process, all the condition inputs to the pre-trained networks/APIs are public information; they have nothing to do with the private classes.
Conditional generation. Conditional generation is achieved by Alg. 3, where we separate the samples according to their class labels, and run the main algorithm (Alg. 1) on each sample set. We can use either conditional or unconditional pre-trained networks/APIs to implement it.
Throughout the paper, “(un)condition” refers to 2, expect the caption in Fig. 8 which refers to 1.
Appendix H More Details on CIFAR10 Experiments
Pre-trained model. By default, we use the checkpoint imagenet64_cond_270M_250K.pt released in .https://github.com/openai/improved-diffusion For the ablation study of the pre-trained network, we additionally use the checkpoint imagenet64_uncond_100M_1500K.pt.
API implementation. RANDOM_API follows the standard diffusion model sampling process. VARIATION_API is implemented with SDEdit , which adds noise to input images and lets the diffusion model to denoise. We use DDIM sampler and the default noise schedule to draw samples. Note that these choices are not optimal; our results can potentially be improved by using better noise schedules and the full DDPM sampling which are known to work better. The implementation of the above APIs is straightforward without touching the core modeling part of diffusion models and is similar to the standard API implementations in Stable Diffusion (App. J).
For the experiments in Fig. 4, we use noise multiplier and threshold for , and pick the the pareto frontier. Fig. 11 shows all the data points we got. Combining this figure with Fig. 4, we can see that PE is not very sensitive to these hyper-parameters, and even with less optimal choices PE still outperforms the baselines.
For the experiments in Fig. 5, we use noise multiplier and threshold .
For downstream classification (Fig. 5), we follow to use WRN-40-4 classifier . We use the official repohttps://github.com/szagoruyko/wide-residual-networks/tree/master/pytorch without changing any hyper-parameter except adding color jitter augmentation according to . The ensemble of the classifier is implemented by ensembling the logits.
FID evaluation. Compared to Fig. 4 in the main text, Fig. 12 shows the full results of two versions of . Baseline results are taken from .
Generated samples. See Figs. 14 and 14 for generated images and real images side-by-side. Note that the pre-trained model we use generates 64x64 images, whereas CIFAR10 is 32x32. In Fig. 4, we show the raw generated 64x64 images; in Fig. 14, we scale them down to 32x32 for better comparison with the real images.
Appendix I More Details on Camelyon17 Experiments
Pre-trained model. We use the checkpoint imagenet64_cond_270M_250K.pt released in .https://github.com/openai/improved-diffusion
About the experiments in Fig. 18. For RANDOM_API and VARIATION_API, we use DDIM sampler with 10 steps. For VARIATION_API, we take a 2-stage approach. the first stage, we use DDIM sampler with 10 steps and use SDEEdit by adding noise till timesteps for each iteration respectively. These timesteps can be regarded as the parameter in § 3.3. We use noise multiplier and threshold .
About the experiments in § 6.1.2. For RANDOM_API and VARIATION_API, we use DDIM sampler with 10 steps. For VARIATION_API, we use DDIM sampler with 10 steps and use SDEEdit by adding noise till $v\sigma=1.541\cdot\sqrt{2}H=4$.
Generated samples. See Figs. 18 and 18 for generated images and real images side-by-side. Note that the real images in Camelyon17 dataset are 96x96 images, whereas the pre-trained network is 64x64. In Fig. 18, we scale them down to 64x64 for better comparison.
Fig. 19 shows the generated images in the intermediate iterations. We can see that the generated images are effectively guided towards Camelyon17 though it is very different from the pre-training dataset.
Appendix J More Details on Stable Diffusion Experiments
Dataset construction. We start with cat photos taken by the authors, crop the region around cat faces with a resolution larger than 512x512 manually, and resize the images to 512x512. We construct two datasets, each for one cat with 100 images. See Figs. 22 and 23 for all images.
API implementation. We use off-the-shelf open-sourced APIs of Stable Diffusion. For RANDOM_API, we use the text-to-image generation APIhttps://huggingface.co/docs/diffusers/api/pipelines/stable_diffusion/text2img, which is implemented by the standard diffusion models’ guided sampling process. For VARIATION_API, we use the image-to-image generation APIhttps://huggingface.co/docs/diffusers/api/pipelines/stable_diffusion/img2img, which allows us to control the degree of variation. Its internal implementation is SDEdit , which adds noise to the input images and runs diffusion models’ denoising process.
Generated images. We use the same hyper-parameters to run PE on two datasets separately. This can also be regarded as running the conditional version of PE (Alg. 3) on the whole dataset (with labels Cat Cookie or Cat Doudou) together. All generated images are in Figs. 24 and 25. While the two experiments use completely the same hyper-parameters, and the initial random images are very different from the cats (Fig. 26), our PE can guide the generated distribution in the right direction and the final generated images do capture the key color and characteristics of each of the cats. This demonstrates the effectiveness of PE with large foundation models such as Stable Diffusion.
We also observe that the diversity of generated images (e.g., poses, face directions) is limited compared to the real data. However, given the small number of samples and the tight privacy budget, this is an expected behavior: capturing more fine-grained features of each image would likely violate DP.
Generated images with more diversity. To make the generated images more diverse, we can utilize the approach in § 4.1, which passes the generated images through VARIATION_API. We have demonstrated in § 6.1.3 that this approach can generate more samples that are useful for downstream classification tasks. Here, we use it for a different purpose: enriching the diversity of generated samples.
Figure Figs. 27 and 28 show the results. We can see that this simple approach is able to generate cats with a more diverse appearance. This is possible because the foundation model (Stable Diffusion) has good prior knowledge about cats learned from massive pre-training, and PE is able to utilize that effectively.
Appendix K More Ablation Studies
All ablation studies are conducted by taking the default parameters in unconditional CIFAR10 experiments and modifying one hyperparameter at a time. The default noise multiplier and the default threshold .
Lookahead degree. Fig. 29 shows how the lookahead degree (§ 4) impacts the results. We can see that higher lookahead degrees monotonically improve the FID score. Throughout all experiments, we used . This experiment suggests that better results can be obtained with a higher .
Histogram threshold. Fig. 31 shows how the threshold in DP Nearest Neighbors Histogram impacts the results. We can see that a large threshold results in a faster convergence speed at the beginning. This is because, in the early iterations, many samples are far away from the private data. A larger threshold can effectively remove those bad samples that have a non-zero histogram count due to the added DP noise. However, at a later iteration, the distribution of generated samples is already close to the private data. A large threshold may potentially remove useful samples (e.g., the samples at low-density regions such as classifier boundaries). This may hurt the generated data, as shown in the increasing FID scores at threshold=15. In this paper, we used a fixed threshold across all iterations. These results suggest that an adaptive threshold that gradually decreases might work better.
Embedding. Fig. 32 compares the results with inception embedding or CLIP embedding in Eq. 1. The results show that both embedding networks work well, suggesting that PE is not too sensitive to the embedding network. Inception embedding works slightly better. One reason is that the inception network is trained on ImageNet, which is similar to a private dataset (CIFAR10). Therefore, it might be better at capturing the properties of images. Another possible reason is that FID score is calculated using inception embeddings, which might lead to some bias that favors inception embedding.
Appendix L Computational Cost Evaluation
We compare the GPU hours of the SOTA DP fine-tuning method and PE for generating 50k samples in § 6.1.1. Note that PE is designed to use existing pre-trained models and we do so in all experiments. In contrast, DP fine-tuning methods usually require a careful selection of pre-training datasets and architectures (e.g., pre-trained their own diffusion models), which could be costly. Even if we ignore this and only consider the computational cost after the pre-training, the total computational cost of PE is only 37% of while having better sample quality and downstream classification accuracy (§ 6.1.1). See Fig. 33 for a detailed breakdown of the computational cost. Note that PE’s computational cost is mostly on the APIs.
The key takeaway is that even if practitioners want to run the APIs locally (i.e., downloading the foundation models and running the APIs locally without using public API providers), there are still benefits of using PE: (1) The computational cost of PE can be smaller than training-based methods. (2) Implementing and deploying PE are easier because PE only requires blackbox APIs of the models and does not require code modifications inside the models.
Experimental details. To ensure a fair comparison, we estimate the runtime of both algorithms using 1 NVIDIA V100 32GB GPU.
To evaluate the computational cost of , we take the open-source diffusion model implementation from https://github.com/openai/guided-diffusion and modify the hyper-parameters according to . We obtain a model with 79.9M parameters, slightly smaller than the one reported in (80.4M). This difference might be due to other implementation details that are not mentioned in . To implement DP training, we utilize Opacus library . To evaluate the fine-tuning cost, we use torch.cuda.Event instrumented before and after the core logic of forward and backward pass, ignoring other factors such as data loading time. We estimate the total runtime based on the mean runtime of 10 batches after 10 batches of warmup. We do not implement augmentation multiplicity with data and timestep ; instead, we use multiplicity=1 (i.e., a vanilla diffusion model), and multiply the estimated runtime by 128, the multiplicity used in . To evaluate the generation cost, we use torch.cuda.Event instrumented before and after the core logic of sampling. We estimate the total runtime based on the mean runtime of 10 batches after 1 batch of warmup.
To evaluate the computational cost of our PE, we use a similar method: we use torch.cuda.Event instrumented before and after the core logic of each component of our algorithm that involves GPU computation. RANDOM_API and VARIATION_API are estimated based on the mean runtime of 10 batches after 1 batch of warmup. Feature extraction is estimated based on the mean runtime of 90 batches after 10 batch of warmup. The nearest neighbor search is estimated based on 1 run of the full search. We use faiss libraryhttps://github.com/facebookresearch/faiss for nearest neighbor search. Its implementation is very efficient so its computation time is negligible compared with the total time.