Effective Data Augmentation With Diffusion Models

Brandon Trabucco, Kyle Doherty, Max Gurinas, Ruslan Salakhutdinov

Introduction

An omnipresent lesson in deep learning is the importance of internet-scale data, such as ImageNet Deng et al. (2009), JFT Sun et al. (2017), OpenImages Kuznetsova et al. (2018), and LAION-5B Schuhmann et al. (2022), which are driving advances in Foundation Models Bommasani et al. (2021) for image generation. These models use large deep neural networks Rombach et al. (2022) to synthesize photo-realistic images for a diversity of prompts. Indeed, the recent success of large generative models prompts a question: can we augment visual recognition datasets with synthetic images from generative models? Answering this question promises to improve image recognition by generating large-scale image datasets from a handful of real images without human labelling effort.

Standard data augmentation aims to mitigate data scarcity by composing randomly parameterized image transformations Antoniou et al. (2017); Perez and Wang (2017); Shorten and Khoshgoftaar (2019); Zhao et al. (2020). Transformations including flips and rotations are chosen that respect basic invariances present in the data, such as horizontal reflection symmetry for a coffee mug. Robustness to this type of image transformation is captured well by existing methods, but models for recognizing coffee mugs should also be robust to subtle details of visual appearance like the brand of mug. Yet, basic transformations fail to produce novel structural elements, textures, or changes in perspective. In contrast, humans are exceptional at noticing these subtle details, able to distinguish varying brands of mugs from a single example. We aim to reproduce this efficiency by extending data augmentation with large text-to-image diffusion models capable of altering image content to improve diversity.

In this work, we propose a flexible data augmentation strategy that generates variations of real images using text-to-image diffusion models (DA-Fusion). Our method adapts the diffusion model to new domains by inserting and fine-tuning new tokens in the text encoder representing novel visual concepts. DA-Fusion modifies the appearance of objects in a manner that respects their semantic invariances, such as the design of the graffiti on the truck in Figure 1 and the design of the train in Figure 2. We test our method on few-shot image classification tasks, including a real-world weed recognition task that lies outside the vocabulary of the diffusion model. Using the same hyper-parameters in all domains, our method outperforms prior work. DA-Fusion improves data augmentation by up to +10 percentage points, and ablations illustrate that our method is robust to hyper-parameter assignment. DA-Fusion is open-sourced at: https://github.com/brandontrabucco/da-fusion.

Related Work

Generative models have been the subject of growing interest and rapid advancement. Earlier methods, including VAEs Kingma and Welling (2014) and GANs Goodfellow et al. (2014), showed initial promise generating realistic images, and were scaled up in terms of resolution and sample quality Brock et al. (2019); Razavi et al. (2019). Despite the power of these methods, many recent successes in photorealistic image generation were the result of diffusion models Ho et al. (2020); Nichol and Dhariwal (2021); Saharia et al. (2022b); Nichol et al. (2022); Ramesh et al. (2022). Diffusion models have been shown to generate higher-quality samples compared to their GAN counterparts Dhariwal and Nichol (2021), and developments like classifier free guidance Ho and Salimans (2022) have made text-to-image generation possible. Recent emphasis has been on training these models with internet-scale datasets like LAION-5B Schuhmann et al. (2022). Generative models trained at internet-scale Rombach et al. (2022); Saharia et al. (2022b); Nichol et al. (2022); Ramesh et al. (2022) have unlocked several application areas where photorealistic generation is crucial.

One application area that diffusion has popularized makes edits to existing real images. Inpainting with diffusion is one such approach that allows the user to specify what to edit as a mask Saharia et al. (2022a); Lugmayr et al. (2022). Other works avoid masks and modify the attention weights of the diffusion process that generated the image instead Hertz et al. (2022); Mokady et al. (2022). Perhaps the most relevant technique to our work is SDEdit Meng et al. (2022a), where real images are inserted partway through the reverse diffusion process. SDEdit is applied by He et al. (2022) to generate synthetic data for training classifiers, but our analysis differs from theirs in that we study generalization to new concepts the diffusion model wasn’t trained on. We simulate this regime by deleting concepts from the weights of Stable Diffusion, and re-adapting the model using only a handful of labelled examples—the same examples used for training classifiers.

Synthetic Data

Training neural networks on synthetic data from generative models was popularized using GANs Antoniou et al. (2017); Tran et al. (2017); Zheng et al. (2017). Various applications for synthetic data generated from GANs have been studied, including representation learning Jahanian et al. (2022), inverse graphics Zhang et al. (2021a), semantic segmentation Zhang et al. (2021b), and training classifiers Tanaka and Aranha (2019); Dat et al. (2019); Yamaguchi et al. (2020); Besnier et al. (2020); Xiong et al. (2020); Wickramaratne and Mahmud (2021); Haque (2021). More recently, synthetic data from diffusion models has also been studied in a few-shot setting He et al. (2022). These works use generative models that have likely seen images of target classes and, to the best of our knowledge, we present the first analysis for synthetic data on previously unseen concepts.

Background

Diffusion models Sohl-Dickstein et al. (2015); Ho et al. (2020); Nichol and Dhariwal (2021); Song et al. (2021); Rombach et al. (2022) are sequential latent variable models inspired by thermodynamic diffusion Sohl-Dickstein et al. (2015). They generate samples via a Markov chain with learned Gaussian transitions starting from an initial noise distribution p(xT)=N(xT;0,I)p(x_{T})=\mathcal{N}(x_{T};0,I).

Transitions pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) are designed to gradually reduce variance according to a schedule β1,…,βT\beta_{1},\ldots,\beta_{T} so the final sample x0x_{0} represents a sample from the true distribution. Transitions are often parameterized by a fixed covariance Σt=βtI\Sigma_{t}=\beta_{t}I and a learned mean μθ(xt,t)\mu_{\theta}(x_{t},t) defined below.

This parameterization choice results from deriving the optimal reverse process Ho et al. (2020), where ϵθ(⋅)\epsilon_{\theta}(\cdot) is a neural network trained to process a noisy sample xtx_{t} and predict added noise. Given real samples x0x_{0} and noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I), one can derive xtx_{t} at an arbitrary timestep below.

Data Augmentation With Diffusion Models

Our goal is to develop a flexible data augmentation strategy using text-to-image diffusion models. In doing so, we consider three desiderata. Our method should: 1. apply to all images like classical data augmentations do, not just images representing concepts the diffusion model was trained on; 2. minimize dataset-specific tuning so the augmentation works off-the-shelf. 3. balance real and synthetic data effectively. We discuss these qualities in the following sections.

Previous work utilizing large pretrained generative models to produce synthetic data (He et al., 2022) has left an important question unanswered: are we sure they are working for the right reason? Models trained on internet data have likely seen many examples of classes in common benchmarking datasets like ImageNet Deng et al. (2009). Moreover, Carlini et al. (2023) have recently shown that pretrained diffusion models can leak their training data. Leakage of internet data, as in Figure 3, risks compromising evaluation. Suppose our goal is to test how images from diffusion models improve few-shot classification with only a few real images, but leakage of internet data gives our classifier access to thousands of real images. Performance gains observed may not reflect the quality of the data augmentation methodology itself, and may lead to drawing the wrong conclusions.

We explore two methods for preventing leakage of Stable Diffusion’s training data. We first consider a model-centric approach that prevents leakage by editing the model weights to remove class knowledge. We also consider a data-centric approach that hides class information from the model inputs.

Our goal with this approach is to remove knowledge about concepts in our benchmarking datasets from the weights of Stable Diffusion. We accomplish this by fine-tuning Stable Diffusion in order to remove the ability to generate concepts from our benchmarking datasets. Given a list of class names in these datasets, we utilize a recent method developed by Gandikota et al. (2023) that fine-tunes the UNet backbone of Stable Diffusion so that concepts specified by a given prompt can no longer be generated (we use class names as such prompts). In particular, the UNet is fine-tuned to minimize the following loss function.

Where "class name" is replaced with the actual class name of the concept being erased, θ\theta represents the parameters of the UNet being fine-tuned, and θ∗\theta^{*} represents the initial parameters of the UNet. This procedure, named ESD by Gandikota et al. (2023), can be interpreted as guiding generation in the opposite direction of classifier free-guidance, and can erase a variety of types of concepts.

Data-Centric Leakage Prevention

While editing the model directly to remove knowledge about classes is a strong defense against possible leakage, it is also costly. In our experiments, erasing a single class from Stable Diffusion takes two hours on a single 32GB V100 GPU. As an alternative for situations where the cost of a model-centric defense is too high, we can achieve a weaker defense by removing all mentions of the class name from the inputs of the model. In practice, switching from a prompt that has the class name to a new prompt omitting the class name is sufficient. Section 4.2 goes into detail how to implement this defense for different types of models.

2 Augmentations For Unseen Concepts

Standard data augmentations apply to all images regardless of class and content Perez and Wang (2017). We aim to capture this flexibility with our diffusion-based augmentation. This is challenging because real images may contain elements the diffusion model is not able to generate out-of-the-box. How do we generate plausible augmentations for such images? Shown in Figure 4, we adapt the diffusion model to new concepts by inserting cc new embeddings in the text encoder of the generative model, and fine-tuning only these embeddings to maximize the likelihood of generating new concepts.

When generating synthetic images, previous work uses a prompt with the specified class name He et al. (2022). However, this is not possible for concepts that lie outside the vocabulary of the generative model because the model’s text encoder has not learned words to describe these concepts. We discuss this problem in Section 5 with our contributed weed-recognition task, which our pretrained diffusion model is unable to generate when the class name is provided. A simple solution to this problem is to have the model’s text encoder learn new words to describe new concepts. Textual Inversion (Gal et al., 2022) is well-suited for this, and we use it to learn a word embedding w⃗i\vec{w}_{i} from a handful of labelled images for each class in the dataset.

We initialize each new embedding w⃗i\vec{w}_{i} to a class-agnostic value (see Appendix G), and optimize them to minimize the simplified loss function proposed by Ho et al. (2020). Figure 4 shows how new embeddings w⃗i\vec{w}_{i} are inserted in the prompt given an image of a train. Our method is modular, and as other mechanisms are studied for adapting diffusion models, Textual Inversion can easily be swapped out with one of these, and the quality of the augmentations from DA-Fusion can be improved.

Generating Synthetic Images

Many of the existing approaches generate synthetic images from scratch Antoniou et al. (2017); Tanaka and Aranha (2019); Besnier et al. (2020); Zhang et al. (2021b, a). This is particularly challenging for concepts the diffusion model hasn’t seen before. Rather than generate from scratch, we use real images as a guide. We splice real images into the generation process of the diffusion model following prior work in SDEdit Meng et al. (2022a). Given a reverse diffusion process with SS steps, we insert a real image x0refx_{0}^{\text{ref}} with noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I) at timestep ⌊St0⌋\lfloor St_{0}\rfloor, where t0∈t_{0}\in is a hyperparameter controlling the insertion position of the image.

We proceed with reverse diffusion starting from the spliced image at timestep ⌊St0⌋\lfloor St_{0}\rfloor and iterating Equation 2 until a sample is generated at timestep . Generation is guided with a prompt that includes the new embedding w⃗i\vec{w}_{i} for the class of the source image (see Appendix G for prompt details).

3 Balancing Real & Synthetic Data

Training models on synthetic images often risks over-emphasizing spurious qualities and biases resulting from an imperfect generative model Antoniou et al. (2017). The common solution assigns different sampling probabilities to real and synthetic images to manage imbalance He et al. (2022). We adopt a similar method for balancing real and synthetic data in Equation 7, where α\alpha denotes the probability that a synthetic image is present at the ll-th location in the minibatch of images BB.

4 Improving Diversity With Randomized Intensity

Having appropriately balanced real and synthetic images, our goal is to maximize diversity. This goal is shared with standard data augmentation Perez and Wang (2017); Shorten and Khoshgoftaar (2019), where multiple simple transformations are used, yielding more diverse data. Despite the importance of diversity, generative models typically employ frozen sampling hyperparameters to produce synthetic datasets Antoniou et al. (2017); Tanaka and Aranha (2019); Yamaguchi et al. (2020); Zhang et al. (2021b, a); He et al. (2022). Inspired by the success of randomization in standard data augmentations (such as the angle of rotation), we propose to randomly sample the insertion position t0t_{0} where real images are spliced into Equation 6. This can be interpreted as randomizing the extent images are modified—as t0→0t_{0}\to 0 generations more closely resemble the guide image.

In Section 6.2 we uniformly at random sample t0∼U({1k,2k,…,kk})t_{0}\sim\mathcal{U}(\{\frac{1}{k},\frac{2}{k},\ldots,\frac{k}{k}\}), and observe a consistent improvement in classification accuracy with k=4k=4 compared to fixing t0t_{0}. Though the hyperparameter t0t_{0} is perhaps the most direct translation of randomized intensity to generative model-based data augmentations, there are several alternatives. For example, one may consider the guidance scale parameter used in classifier-free guidance (Ho and Salimans, 2022). We leave this as future work.

Data Preparation

We contribute a dataset of top-down drone images of semi-natural areas in the western United States. These data were gathered in an effort to better map the extent of a problematic invasive plant, leafy spurge (Euphorbia esula), that is a detriment to natural and agricultural ecosystems in temperate regions of North America. Prior drone-based work to detect leafy spurge achieved an accuracy of 0.75 Yang et al. (2020). To our knowledge, top-down aerial imagery of leafy spurge was not present in the Stable Diffusion training data. Results of CLIP-retrieval Beaumont (2022) returned close-up, side-on images of members of the same genus (Figure 5) in the top 20 results. We observed the first instance of our target species, Euphorbia esula, as a 35th result. This dataset represents a unique opportunity to explore few-shot learning with Stable Diffusion, and improving classification would directly benefit efforts to restore natural ecosystems. Additional details are in Appendix H.

PascalVOC

We leverage the 2012 version of the Pascal Visual Object Classes challenge Everingham et al. (2009). This dataset contains 11,530 images and 6,929 object segmentation masks. We adapt this dataset into an object classification task by filtering images that have at least one object segmentation mask. We assign these images labels corresponding to the class of object with largest area in the image, as measured by the pixels contained in the mask. There are 20 classes in total using this methodology. We utilize the official training and validation sets for the 2012 challenge, and uniformly at random randomly select qq images per class from the training set for training classifiers.

COCO

We process the 2017 version of the COCO dataset Lin et al. (2014) in a manner congruent to PascalVOC. This dataset contains 330K images with 1.5M object segmentation masks. As before, we filter images that have at least one object segmentation mask. We assign these images labels corresponding to the class of the largest object, measured by segmentation mask area. This dataset has 80 classes. We use the official training and validation sets for the 2017 dataset, and measure few-shot classification accuracy using the same methodology described for PascalVOC.

Results

In this section, we study the performance of DA-Fusion in a few-shot setting when leakage of Stable Diffusion training data is prevented. We begin with an analysis on three image recognition tasks, including a new weed recognition task (see the Appendix for additional few-shot classification results on other datasets). We then perform ablations to understand the extent of gains due to randomized intensities, and confirm our method is robust to the balance of real and synthetic data. We observe gains in few-shot classification accuracy in all domains when leakage is prevented.

While leafy spurge is confirmed to lie outside the vocabulary of the pretrained diffusion model, this is not true for Pascal Everingham et al. (2009) and COCO Lin et al. (2014), which have common objects like boats and airplanes. To properly evaluate few-shot classification performance with these datasets, we emphasize the need to hide knowledge of these classes from the generative model to prevent leakage of the model’s training data. This remains an active area of research Meng et al. (2022b); Gandikota et al. (2023), and we discuss two methods for preventing leakage in Section 4.1. We first employ a model-centric approach that hides prior knowledge about classes by editing the weights of Stable Diffusion, and a data-centric approach that hides identifying class information from the inputs to Stable Diffusion, preventing it from accessing prior knowledge about classes. We study the influence each leakage prevention mechanism has on few-shot classification accuracy of classifiers trained using a mix of real data, and synthetic data produced by DA-Fusion.

In this experiment, we test few-shot classification with three data augmentation strategies. The first, referred to as "Baseline" in the remainder of this paper, employs no synthetic images. This baseline implements a standard data augmentation strategy that uses random rotations and flips with parameters that depend on the dataset. For COCO and Pascal domains, we use random horizontal flips and random rotations with angles uniformly randomly sampled between +15 and -15 degrees. For the Spurge domain, we employ an additional random vertical flip, and increase the range of random rotations to +45 and -45 degrees. The Real Guidance baseline is based on the method developed by He et al. (2022), and uses SDEdit on real images with t0=0.5t_{0}=0.5. Hyper-parameters shared between Real Guidance and our method have equal values to ensure fairness. Depending on which leakage prevention method is used, the prompts given to the diffusion model change. For the model-centric approach, we give real guidance prompts of the form "a photo of a cat," where the real class name is visible to the model. DA-Fusion is prompted with "a photo of a ClassX," where the embedding for ClassX is initialized to a class-agnostic value and learned according to Section 4.2. For data-centric leakage prevention, Real Guidance is instead prompted with "a photo."

Each real image is augmented MM times, and a ResNet50 classifier pre-trained on ImageNet is fine-tuned on a mixture of real and synthetic images sampled as discussed in Section 4.3. We vary the number of examples per class used for training the classifier on the x-axis in the following plots, and fine-tune the final linear layer of the classifier for 10,00010,000 steps with a batch size of 3232 and the Adam optimizer with learning rate 0.00010.0001. We record validation metrics every 200 steps and report the epoch with highest accuracy. Solid lines in plots represent means, and error bars denote 68% confidence intervals over 8 independent trials. An overall score is calculated for all datasets after normalizing performance using yi(d)←(yi(d)−ymin(d))/(ymax(d)−ymin(d))y_{i}^{(d)}\leftarrow(y_{i}^{(d)}-y_{\text{min}}^{(d)})/(y_{\text{max}}^{(d)}-y_{\text{min}}^{(d)}), where dd represents the dataset, ymax(d)y_{\text{max}}^{(d)} is the maximum performance for any trial of any method, and ymin(d)y_{\text{min}}^{(d)} is defined similarly.

Results With Model-Centric Leakage Prevention

Figure 6 shows results when erasing class knowledge from Stable Diffusion weights. We observe a consistent improvement in validation accuracy by as much as +5 percentage points on the Pascal and COCO domains when compared to the standard data augmentation baseline. DA-Fusion exceeds performance of Real Guidance He et al. (2022) overall while utilizing the same hyperparameters, without any prior information about the classes in these datasets. In this setting, Real Guidance performs comparably to the baseline, which suggests that gains in Real Guidance may stem from information provided by the class name. This experiment shows DA-Fusion improves few-shot learning and suggests our method generalizes to concepts Stable Diffusion wasn’t trained on. To understand how these gains translate to weaker defenses against training data leakage, we next evaluate our method using a data-centric strategy.

Results With Data-Centric Leakage Prevention

Figure 7 shows results when class information is hidden from Stable Diffusion inputs. As before, we observe a consistent improvement in validation accuracy, by as much as +10 percentage points on the Pascal and COCO domains when compared to the standard data augmentation baseline. DA-Fusion exceeds performance of Real Guidance He et al. (2022) in all domains while utilizing the same hyperparameters, without specifying the class name as an input to the model. With a weaker defense against training data leakage, we observe larger gains with DA-Fusion. This suggests gains are due in part to accessing Stable Diffusion’s prior knowledge about classes, and highlights the need for a strong leakage prevention mechanism when evaluating synthetic data from large generative models. In the following sections, we ablate our method to understand where these gains come from, and how important each part of the method is.

2 How Important Are Randomized Intensities?

Our goal in this section is to understand what fraction of gains are due to randomizing the intensity of our augmentation based on Section 4.4. We employ the same experimental settings as in Section 6.1, using data-centric leakage prevention, and run our method using a fixed insertion position t0=0.5t_{0}=0.5 (labelled k=1k=1 in Figure 8), following the settings used with Real Guidance. In Figure 8 we report the improvement in average classification accuracy on the validation set versus standard data augmentation. These results show that both versions of our method outperform the baseline, and randomization improves our method in all domains, leading to an overall improvement of 51%.

3 DA-Fusion Is Robust To Data Balance

Discussion

We proposed a flexible method for data augmentation based on diffusion models, DA-Fusion. Our method adapts a pretrained diffusion model to semantically modify images and produces high quality augmentations regardless of image content. Our method improves few-shot classification accuracy in tested domains, and by up to +10 percentage points on tasks based on Pascal and COCO. Similarly, our method produces gains on a contributed weed-recognition dataset that lies outside the vocabulary of the diffusion model. To understand these gains, we studied how performance is impacted by potential leakage of Stable Diffusion training data. To prevent leakage during evaluation, we presented two defenses that target the model and data respectively, each on different sides of a trade-off between defense strength and computational cost. When subject to both defenses, DA-Fusion consistently improves few-shot classification accuracy, which highlights its utility for data augmentation.

There are several directions to improve the flexibility and performance of our method as future work. First, our method does not explicitly control how an image is augmented by the diffusion model. Extending the method with a mechanism to better control how objects in an image are modified, e.g. changing the breed of a cat, could improve the results. Recent work in prompt-based image editing Hertz et al. (2022) suggests diffusion models can make localized edits without pixel-level supervision, and would minimally increase the human effort required to use DA-Fusion. This extension would let image attributes be handled independently by DA-Fusion, and certain attributes could be modified more extremely than others. Second, data augmentation is becoming increasingly important in the decision-making setting Yarats et al. (2022). Maintaining temporal consistency is an important challenge faced when using our method in this setting. Solving this challenge could improve the few-shot generalization of policies in complex visual environments. Finally, improvements to our diffusion model backbone that enhance image photo-realism are likely to improve DA-Fusion.

Acknowledgements

We thank MPG Ranch for supporting the leafy spurge component of this work. MPG Ranch staff, including Charles Casper, Erik Samsoe, Beau Larkin, and Philip Ramsey coordinated planning and acquisition of the leafy spurge imagery. In addition, we thank the effort of reviewers for helping to improve the paper, and the feedback from peers on intermediate drafts. We specifically thank Jing Yu Koh, Yutong He, Murtaza Dalal, So Yeon Min, and Martin Ma for their feedback. Finally, we thank Stability AI and Huggingface for providing open source models. Brandon Trabucco is supported by Amazon, and Ruslan Salakhutdinov is supported in part by ONR award N000141812861 and DSTA.

References

Appendix A Limitations & Safeguards

As generative models have improved in terms of fidelity and scale, they have been shown to occasionally produce harmful content, including images that reinforce stereotypes, and images that include nudity or violence. Synthetic data from generative models, when it suffers from these problems, has the potential to increase bias in downstream classifiers trained on such images if not handled. We employ two mitigation techniques to lower the risk of leakage of harmful content into our data augmentation strategy. First, we use a safety checker that determines whether augmented images contain nudity or violence. If they do, the generation is discarded and re-sampled until a clean image is returned. Second, rather than generate images from scratch, our method edits real images, and keeps the original high-level structure of the real images. In this way, we can guide the model away from harmful content by ensuring the real images contain no harmful content to begin with. The combination of these techniques lowers the risk of leakage of harmful content, but is not a perfect solution. In particular, detecting biased content that encourages racial or gender stereotypes that exist online is much harder than detecting nudity or violence, and one limitation of this work is that we can’t yet defend against this. We emphasize the importance of curating unbiased and safe datasets for training large generative models, and the creation of post-training bias mitigation techniques.

Appendix B Ethical Considerations

There are potential ethical concerns arising from large-scale generative models. For example, these models have been trained on large amounts of user data from the internet without the explicit consent of these users. Since our data augmentation strategy employs Stable Diffusion [Rombach et al., 2022], our method has the potential to generate augmentations that resemble or even copy data from such users online. This issue is not specific to our work; rather, it is inherent to image generation models trained at scales as large as Stable Diffusion, and other works using Stable Diffusion also face this ethical problem. Our mitigation to this ethical problem is to allow deletion of concepts from the weights of Stable Diffusion before augmentation. Deletion removes harmful, or copyrighted material from Stable Diffusion weights to ensure it cannot be copied by the model during augmentation.

Appendix C Broader Impacts

Data augmentation strategies like DA-Fusion have the potential to enable training vision models of a variety of types from limited data. While we studied classification in this work, DA-Fusion may also be applied to video classification, object detection, and visual reinforcement learning. One risk associated with improved few-shot learning on vision-based tasks is that synthetic data can be generated targeting particular users. For example, suppose one intends to build a person-identification system used to record the behavior patterns of a specific person in public. Such a system trained with generative model-based data augmentations may only need one real photo to be trained. This poses a risk to privacy, despite other benefits that few-shot learning provides. As another example, suppose one intends to build a system capable of generating pornography of a specific celebrity. Few-shot learning makes this possible with just a handful of real images that exist online. This poses a risk to personal safety and bodily autonomy of the targeted person.

Appendix D Additional Results

We conduct additional experiments on the Caltech101 [Fei-Fei et al., 2004], and Flowers102 [Nilsback and Zisserman, 2008] datasets, two standard image classification tasks. These tasks are commonly used when benchmarking few-shot classification performance, such as in the Visual Task Adaptation Benchmark [Zhai et al., 2019]. Results on these datasets are shown in Figure 11, and show that DA-Fusion improves classification performance both when using a model-centric defense against training data leakage, and a data-centric defense, described in Section 4.1 of the paper.

Appendix E Stronger Augmentation Baselines

In the main paper, we considered data augmentation baselines consisting only of randomized rotations and flips. In this section, we compare against two stronger data augmentation methods: RandAugment [Cubuk et al., 2020], and CutMix [Yun et al., 2019]. Results are presented in Figure 10, and show that DA-Fusion improves over both RandAugment and CutMix on the Pascal-based task.

Appendix F Different Classifier Architectures

Results in the main paper use a ResNet50 architecture for the image classifier. In this section, we consider the Data-Efficient Image Transformer (DeiT) [Touvron et al., 2021], and evaluate DA-Fusion with data-centric leakage prevention on the Pascal task. Results in Figure 12 show that DA-Fusion improves the performance of DeiT, and suggests that gains generalize to different model architectures, including both convolution-based models (such as ResNet50), and attention-based ones (such as ViT).

Appendix G Hyperparameters

Our method inherits the hyperparameters of text-to-image diffusion models and SDEdit Meng et al. [2022a]. In addition, we introduce several other hyperparameters in this work that control the diversity of the synthetic images. Specific values for these hyperparameters are given in Table 1.

Appendix H Leafy Spurge Dataset Acquisition and Pre-processing

In June 2022 botanists visited areas in western Montana, United States known to harbor leafy spurge and verified the presence or absence of the target plant at 39 sites. We selected sites that represented a range of elevation and solar input values as influenced by terrain. These environmental axes strongly drive variation in the structure and composition of vegetation Amatulli et al. , Doherty et al. . Thus, stratifying by these aspects of the environment allowed us to test the performance of classifiers when presented with a diversity of plants which could be confused with our target.

During surveys, each site was divided into a 3 x 3 grid of plots that were 10m on side (Fig. 13), and then botanists confirmed the presence or absence of leafy spurge within each grid cell. After surveying we flew a DJI Phantom 4 Pro at 50m above the center of each site and gathered still RGB images. All images were gathered on the same day in the afternoon with sunny lighting conditions.

We then cropped the the raw images to match the bounds of plots using visual markers installed during surveys as guides (Fig. 14). Resulting crops varied in size because of the complexity of terrain. E.G., ridges were closer to the drone sensor than valleys. Thus, image side lengths ranged from 533 to 1059 pixels. The mean side length was 717 and the mean spatial resolution, or ground sampling distance, of pixels was 1.4 cm.

In our initial hyperparameter search we found that the classification accuracy of plot-scale images was less than that of a classifier trained on smaller crops of the plots. Therefore, we generated four 250x250 pixel crops sharing a corner at plot centers for further experimentation (Fig. 15). Because spurge plants were patchily distributed within a plot, a botanist reviewed each crop in the present class and removed cases in which cropping resulted in samples where target plants were not visually apparent.

Appendix I Benchmarking the Leafy Spurge Dataset

We benchmark classifier performance here on the full leafy spurge dataset, comparing a baseline approach incorporating legacy augmentations with our novel DA-fusion method. For 15 trials we generated random validation sets with 20 percent of the data, and fine-tuned a pretrained ResNet50 on the remaining 80 percent using the training hyperparameters reported in section 6 for 500 epochs. From these trials we compute cross-validated mean accuracy and 68 percent confidence intervals.

In the case of baseline experiments, we augment data by flipping vertically and horizontally, as well as randomly rotating by as much as 45 degrees with a probability of 0.5. For DA-Fusion augmentations we take two approaches(Fig. 16) The first we refer to as DA-Fusion Pooled, and we apply the methods of Textual Inversion Gal et al. , but include all instances of a class in a single session of fine-tuning, generating one token per class. In the second approach we refer to as DA-Fusion Specific, we fine-tune and generate unique tokens for each image in the training set. In the specific case, we generated 90, 180, and 270 rotations as well as horizontal and vertical flips and contribute these along with original image for Stable Diffusion fine-tuning to achieve the target number of images suggested to maximize performanceGal et al. . In both DA-Fusion approaches we generated ten synthetic images per real image for model training. We maintain α=0.5\alpha=0.5, evenly mixing real and synthetic data during training. We also maximize synthetic diversity by randomly selecting 0.25, 0.5, 0.75, and 1.0 t0t_{0} values. Note that we do not apply concept erasure here as in few-shot experiments from the body text.

Both approaches to DA-Fusion offer slight performance enhancements over baseline augmentation methods for the full leafy spurge dataset. We observe a 1.0% gain when applying DA-Fusion Pooled and a 1.2% gain when applying DA-Fusion Specific(Fig. 17). It is important to note that, as implemented currently, compute time for DA-Fusion Specific is linearly related to data amount, but DA-Fusion Pooled compute is the same regardless of data size.

While pooling was not the most beneficial in this experiment, we support investigating it further. This is because fine-tuning a leafy spurge token in a pooled approach might help to orient our target in the embedding space where plants with similar diagnostic properties, such as flower shape and color from the same genus, may be well represented. However, the leafy-spurge negative cases do not correspond to a single semantic concept, but a plurality, such as green fields, brown fields, and wooded areas. It is unclear if fine-tuning a single token for negative cases by a pooled method would remove diversity from synthetic samples of spurge-free background landscapes, relative to an image-specific approach. For this reason, we suspect a hybrid approach of pooled token for the positive case and specific tokens for the negative cases could offer further gains, and support the application of detecting weed invasions into new areas.