Diffusion-GAN: Training GANs with Diffusion
Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, Mingyuan Zhou
Introduction
Generative adversarial networks (GANs) (Goodfellow et al., 2014) and their variants (Brock et al., 2018; Karras et al., 2019; 2020a; Zhao et al., 2020) have achieved great success in synthesizing photo-realistic high-resolution images. GANs in practice, however, are known to suffer from a variety of issues ranging from non-convergence and training instability to mode collapse (Arjovsky and Bottou, 2017; Mescheder et al., 2018). As a result, a wide array of analyses and modifications has been proposed for GANs, including improving the network architectures (Karras et al., 2019; Radford et al., 2016; Sauer et al., 2021; Zhang et al., 2019), gaining theoretical understanding of GAN training (Arjovsky and Bottou, 2017; Heusel et al., 2017; Mescheder et al., 2017; 2018), changing the objective functions (Arjovsky et al., 2017; Bellemare et al., 2017; Deshpande et al., 2018; Li et al., 2017a; Nowozin et al., 2016; Zheng and Zhou, 2021; Yang et al., 2021), regularizing the weights and/or gradients (Arjovsky et al., 2017; Fedus et al., 2018; Mescheder et al., 2018; Miyato et al., 2018a; Roth et al., 2017; Salimans et al., 2016), utilizing side information (Wang et al., 2018; Zhang et al., 2017; 2020b), adding a mapping from the data to latent representation (Donahue et al., 2016; Dumoulin et al., 2016; Li et al., 2017b), and applying differentiable data augmentation (Karras et al., 2020a; Zhang et al., 2020a; Zhao et al., 2020).
A simple technique to stabilize GAN training is to inject instance noise, , to add noise to the discriminator input, which can widen the support of both the generator and discriminator distributions and prevent the discriminator from overfitting (Arjovsky and Bottou, 2017; Sønderby et al., 2017). However, this technique is hard to implement in practice, as finding a suitable noise distribution is challenging (Arjovsky and Bottou, 2017). Roth et al. (2017) show that adding instance noise to the high-dimensional discriminator input does not work well, and propose to approximate it by adding a zero-centered gradient penalty on the discriminator. This approach is theoretically and empirically shown to converge in Mescheder et al. (2018), who also demonstrate that adding zero-centered gradient penalties to non-saturating GANs can result in stable training and better or comparable generation quality compared to WGAN-GP (Arjovsky et al., 2017). However, Brock et al. (2018) caution that zero-centered gradient penalties and other similar regularization methods may stabilize training at the cost of generation performance. To the best of our knowledge, there has been no existing work that is able to empirically demonstrate the success of using instance noise in GAN training on high-dimensional image data.
To inject proper instance noise that can facilitate GAN training, we introduce Diffusion-GAN, which uses a diffusion process to generate Gaussian-mixture distributed instance noise. We show a graphical representation of Diffusion-GAN in Figure 1. In Diffusion-GAN, the input to the diffusion process is either a real or a generated image, and the diffusion process consists of a series of steps that gradually add noise to the image. The number of diffusion steps is not fixed, but depends on the data and the generator. We also design the diffusion process to be differentiable, which means that we can compute the derivative of the output with respect to the input. This allows us to propagate the gradient from the discriminator to the generator through the diffusion process, and update the generator accordingly. Unlike vanilla GANs, which compare the real and generated images directly, Diffusion-GAN compares the noisy versions of them, which are obtained by sampling from the Gaussian mixture distribution over the diffusion steps, with the help of our timestep-dependent discriminator. This distribution has the property that its components have different noise-to-data ratios, which means that some components add more noise than others. By sampling from this distribution, we can achieve two benefits: first, we can stabilize the training by easing the problem of vanishing gradient, which occurs when the data and generator distributions are too different; second, we can augment the data by creating different noisy versions of the same image, which can improve the data efficiency and the diversity of the generator. We provide a theoretical analysis to support our method, and show that the min-max objective function of Diffusion-GAN, which measures the difference between the data and generator distributions, is continuous and differentiable everywhere. This means that the generator in theory can always receive a useful gradient from the discriminator, and improve its performance.
Our main contributions include: 1) We show both theoretically and empirically how the diffusion process can be utilized to provide a model- and domain-agnostic differentiable augmentation, enabling data-efficient and leaking-free stable GAN training. 2) Extensive experiments show that Diffusion-GAN boosts the stability and generation performance of strong baselines, including StyleGAN2 (Karras et al., 2020b), Projected GAN (Sauer et al., 2021), and InsGen (Yang et al., 2021), achieving state-of-the-art results in synthesizing photo-realistic images, as measured by both the Fréchet Inception Distance (FID) (Heusel et al., 2017) and Recall score (Kynkäänniemi et al., 2019).
Preliminaries: GANs and diffusion-based generative models
GANs (Goodfellow et al., 2014) are a class of generative models that aim to learn the data distribution of a target dataset by setting up a min-max game between two neural networks: a generator and a discriminator. The generator takes as input a random noise vector sampled from a simple prior distribution , such as a standard normal or uniform distribution, and tries to produce realistic-looking samples that resemble the data. The discriminator receives either a real data sample drawn from or a fake sample generated by , and tries to correctly classify them as real or fake. The goal of is to fool into making mistakes, while the goal of is to accurately distinguish from . The min-max objective function of GANs is given by
In practice, this vanilla objective function is often modified to improve the stability and performance of GANs(Goodfellow et al., 2014; Miyato et al., 2018a; Fedus et al., 2018), but the general idea of adversarial learning between and remains the same.
Diffusion-based generative models (Ho et al., 2020b; Sohl-Dickstein et al., 2015; Song and Ermon, 2019) assume , where are latent variables of the same dimensionality as the data . There is a forward diffusion chain that gradually adds noise to the data in steps with pre-defined variance schedule and variance :
A notable property is that at an arbitrary time-step can be sampled in closed form as
A variational lower bound (Blei et al., 2017) is then used to optimize the reverse diffusion chain as
Diffusion-GAN: Method and Theoretical Analysis
To construct Diffusion-GAN, we describe how to inject instance noise via diffusion, how to train the generator by backpropagating through the forward diffusion process, and how to adaptively adjust the diffusion intensity. We further provide theoretical analysis illustrated with a toy example.
We aim to generate realistic samples from a generator network that maps a latent variable sampled from a simple prior distribution to a high-dimensional data space, such as images. The distribution of generator samples , is denoted by . To make the generator more robust and diverse, we inject instance noise into the generated samples by applying a diffusion process that adds Gaussian noise at each step. The diffusion process can be seen as a Markov chain that starts from the original sample and gradually erases its information until reaching a noise level after steps.
We define a mixture distribution that models the noisy samples obtained at any step of the diffusion process, with a mixture weight for each step . The mixture components are Gaussian distributions with mean proportional to and variance depending on the noise level at step . We use the same diffusion process and mixture distribution for both the real samples and the generated samples . More specifically, the diffusion-induced mixture distributions are expressed as
where is a -component mixture distribution, the mixture weights are non-negative and sum to one, and the mixture components are obtained via diffusion as in Equation 1, expressed as
Samples from this mixture can be drawn as .
By sampling from this mixture distribution, we can obtain noisy versions of both real and generated samples with varying degrees of noise. The more steps we take in the diffusion process, the more noise we add to and the less information we preserve from . We can then use this diffusion-induced mixture distribution to train a timestep-dependent discriminator that distinguishes between real and generated noisy samples, and a generator that matches the distribution of generated noisy samples to the distribution of real noisy samples. Next we introduce Diffusion-GAN that trains its discriminator and generator with the help of the diffusion-induced mixture distribution.
2 Adversarial Training
The Diffusion-GAN trains its generator and discriminator by solving a min-max game objective as
Here, is the true data distribution, is a discrete distribution that assigns different weights to each diffusion step , and is the conditional distribution of the perturbed sample given the original data and the diffusion step . By Equation 2, with Gaussian reparameterization, the perturbation function could be written as , where is the cumulative noise level at step , is a scale factor, and is a Gaussian noise.
The objective function in Equation 3 encourages the discriminator to assign high probabilities to the perturbed real data and low probabilities to the perturbed generated data, for any diffusion step . The generator, on the other hand, tries to produce samples that can deceive the discriminator at any diffusion step . Note that the perturbed generated sample can be rewritten as . This means that the objective function in Equation 3 is differentiable with respect to the generator parameters, and we can use gradient descent to optimize it with back-propagation.
The objective function Equation 3 is similar to the one used by the original GAN (Goodfellow et al., 2014), except that it involves the diffusion steps and the perturbation functions. We can show that this objective function also minimizes an approximation of the Jensen–Shannon (JS) divergence between the true and the generated distributions, but with respect to the perturbed samples and the diffusion steps, as follows:
The JS divergence measures the dissimilarity between two probability distributions, and it reaches its minimum value of zero when the two distributions are identical. The proof of the equality in Equation 4 is given in Appendix C. A natural question that arises from this result is whether minimizing the JS divergence between the perturbed distributions implies minimizing the JS divergence between the original distributions, , whether the optimal generator for Equation 3 is also the optimal generator for . We will answer this question affirmatively and provide a theoretical justification in Section 3.4.
3 Adaptive diffusion
With the help of the perturbation function and timestep dependency, we have a new strategy to optimize the discriminator. We want the discriminator to have a challenging task, neither too easy to allow overfitting the data (Karras et al., 2020a; Zhao et al., 2020) nor too hard to impede learning. Therefore, we adjust the intensity of the diffusion process, which adds noise to both and , depending on how much can distinguish them. When the diffusion step is larger, the noise-to-data ratios are higher and the task is harder. We use to measure the intensity of the diffusion, which increases as grows. To control the diffusion intensity, we adaptively modify the maximum number of steps .
Our strategy is to make the discriminator learn from the easiest samples first, which are the original data samples, and then gradually increase the difficulty by feeding it samples from larger . To do this, we use a self-paced schedule for , which depends on a metric that estimates how much the discriminator overfits to the data:
where is the same as in Karras et al. (2020a) and is a constant. We calculate and update every four minibatches. We have two options for the distribution that we use to sample for the diffusion process:
The ‘priority’ option gives more weight to larger , which means the discriminator will see more new samples from the new steps when increases. This is because we want the discriminator to focus on the new and harder samples that it has not seen before, as this indicates that it is confident about the easier ones. Note that even with the ‘priority’ option, the discriminator can still see samples from smaller , because is a mixture of Gaussians that covers all steps of the diffusion chain.
To avoid sudden changes in during training, we use an exploration list that contains values sampled from . We keep fixed until we update , and we sample from to generate noisy samples for the discriminator. This way, the model can explore each sufficiently before moving to a higher . We give the details of training Diffusion-GAN in Algorithm 1 in Appendix F.
4 Theoretical analysis with Examples
To better understand the theoretical properties of our proposed method, we present two theorems that address two important questions about the use of diffusion-based instance noise injection for training GANs. The proofs of these theorems are deferred to Appendix B. The first question, denoted as (a), is whether adding noise to the real and generated samples in a diffusion process can facilitate the learning. The second question, denoted as (b), is whether minimizing the JS divergence between the joint distributions of the noisy samples and the noise levels, and , can lead to the same optimal generator as minimizing the JS divergence between the original distributions of the real and generated samples, and .
To answer (a), we prove that for any choice of noise level and any choice of convex function , the -divergence (Nowozin et al., 2016) between the marginal distributions of the noisy real and generated samples, and , is a smooth function that can be computed and optimized by the discriminator. This implies that the diffusion-based noise injection does not introduce any singularity or discontinuity in the objective function of the GAN. The JS divergence is a special case of -divergence, where .
Let be a fixed distribution over and be a random noise over another space . Denote as a function with parameter and input and as the distribution of . Let , where and . Let and . Then, , if function is continuous and differentiable, the f-divergence is continuous and differentiable with respect to .
Theorem 1 shows that with the help of diffusion noise injection by , , and are defined on the same support space, the whole , and is continuous and differentiable everywhere. Then, one natural question is what if keeps a near constant value and hence provides little useful gradient. Hence, we empirically show that by injecting noise through a mixture defined over all steps of the diffusion chain, there is always a good chance that a sufficiently large is sampled to provide a useful gradient, via the toy example below.
Toy example. We use the same simple example from Arjovsky et al. (2017) to illustrate our method. Let be the real data and be the data generated by a one-parameter generator, where is a uniform random variable in ${\mathcal{D}}_{\text{JS}}(p({\bm{x}})||p({\bm{x}}_{g}))\theta=0\log 2\theta$ towards zero.
For the discriminator optimization, as shown in the second row, right, of Figure 2, the optimal discriminator under the original JS divergence is discontinuous and unattainable. With diffusion-based noise, the optimal discriminator changes with : a smaller makes it more confident and a larger makes it more cautious. Thus the diffusion acts like a scale to balance the power of the discriminator. This suggests the use of a differentiable forward diffusion chain that can provide various levels of gradient smoothness to help the generator training.
Let and , where is the transition density. Given certain , if could be reparameterized into , where is a known distribution, and both and are one-to-one mapping functions, then we could have .
To answer question (b), we present Theorem 2, which shows a sufficient condition for the equality of the original and the augmented data distributions. By Theorem 2, the function maps each to a unique , the function maps each to a unique noise term, and the distribution of is known and independent of . Under these assumptions, the theorem proves that the distribution of is the same as the distribution of , if and only if the distribution of is the same as the distribution of . If we take as the introduced in the theorem, then for , Equation 2 fits the assumption made. This means that, by minimizing the divergence between and , which is the same as minimizing the divergence between and , we are also minimizing the divergence between and . This implies that the noise injection does not affect the quality of the generated samples, and we can safely use our noise injection to improve the training of the generative model.
5 Related work
The proposed Diffusion-GAN can be related to previous works on stabilizing the GAN training, building diffusion-based generative models, and constructing differential augmentation for data-efficient GAN training. A detailed discussion on these related works is deferred to Appendix A.
Experiments
We conduct extensive experiments to answer the following questions: (a) Will Diffusion-GAN outperform state-of-the-art GAN baselines on benchmark datasets? (b) Will the diffusion-based noise injection help the learning of GANs in domain-agnostic tasks? (c) Will our method improve the performance of data-efficient GANs trained with a very limited amount of data?
Datasets. We conduct experiments on image datasets ranging from low-resolution (, ) to high-resolution (, ) and from low-diversity to high-diversity: CIFAR-10 (Krizhevsky, 2009), STL-10 (Coates et al., 2011), LSUN-Bedroom (Yu et al., 2015), LSUN-Church (Yu et al., 2015), AFHQ(Cat/Dog/Wild) (Choi et al., 2020), and FFHQ (Karras et al., 2019). More details on these benchmark datasets are provided in Appendix E.
Evaluation protocol. We measure image quality using FID (Heusel et al., 2017). Following Karras et al. (2019; 2020b), we measure FID using 50k generated samples, with the full training set used as reference. We use the number of real images shown to the discriminator to evaluate convergence (Karras et al., 2020a; Sauer et al., 2021). Unless specified otherwise, all models are trained with 25 million images to ensure convergence (these trained with more or fewer images are specified in table captions). We further report the improved Recall score introduced by Kynkäänniemi et al. (2019) to measure the sample diversity of generative models.
Implementations and resources. We build Diffusion-GANs based on the code of StyleGAN2 (Karras et al., 2020b), ProjectedGAN (Sauer et al., 2021), and InsGen (Yang et al., 2021) to answer questions (a), (b), and (c), respectively. Diffusion GANs inherit from their corresponding base GANs all their network architectures and training hyperparamters, whose details are provided in Appendix G. Specifically for StyleGAN2 and InsGen, we construct the discriminator as , where is injected via their mapping network. For ProjectedGAN, we empirically find in the discriminator could be ignored to simplify the implementation and minimize the modifications to ProjectedGAN. More implementation details are provided in Appendix H. By applying our diffusion-based noise injection, we denote our models as Diffusion StyleGAN2/ProjectedGAN/InsGen. In the following experiments, we train related models with their official code if the results are unavailable, while others are all reported from references and marked with ∗. We run all our experiments with either 4 or 8 NVIDIA V100 GPUs depending on the demands of the inherited training configurations.
We compare Diffusion-GAN with its state-of-the-art GAN backbone, StyleGAN2 (Karras et al., 2020a), and to evaluate its effectiveness from the data augmentation perspective, we compare it with both StyleGAN2 + DiffAug (Zhao et al., 2020) and StyleGAN2 + ADA (Karras et al., 2020a), in terms of both sample fidelity (FID) and sample diversity (Recall) over extensive benchmark datasets.
We present the quantitative and qualitative results in Table 1 and Figure 3. Qualitatively, these generated images from Diffusion StyleGAN2 are all photo-realistic and have good diversity, ranging from low-resolution (32 32) to high-resolution (1024 1024). Additional randomly generated images can be found in Appendix L. Quantitatively, Diffusion StyleGAN2 outperforms all the GAN baselines in generation diversity, as measured by Recall, on all 6 benchmark datasets and outperforms them in FID by a clear margin on 5 out of the 6 benchmark datasets.
From the data augmentation perspective, we observe that Diffusion StyleGAN2 always clearly outperforms the backbone model StyleGAN2 across various datasets, which empirically validates our Theorem 2. By contrast, both the ADA (Karras et al., 2020b) and Diffaug (Zhao et al., 2020) techniques could sometimes impair the generation performance on sufficiently large datasets, , LSUN-Bedroom and LSUN-Church, which is also observed by Yang et al. (2021) on FFHQ. This is possibly because their risk of leaking augmentation overshadows the benefits of data augmentation.
To investigate how the adaptive diffusion process works during training, we illustrate in Figure 4 the convergence of the maximum timestep in our adaptive diffusion and discriminator outputs. We see that is adaptively adjusted: The for Diffusion StyleGAN2 increases as the training goes while the for Diffusion ProjectedGAN first goes up and then goes down. Note that the is adjusted according to the overfitting status of the discriminator. The second panel shows that trained with the diffusion-based mixture distribution, the discriminator is always well behaved and provides useful learning signals for the generator, which validates our analysis in Section 3.4 and Theorem 1.
Memory and time costs. Generally speaking, the memory and time costs of a Diffusion-GAN are comparable to those of the corresponding GAN baseline. More specifically, switching from ADA (Karras et al., 2020a) to our diffusion-based augmentation, the added memory cost is negative, the added training time cost is negative, and the added inference time cost is zero. For example, for CIFAR-10, with four NVIDIA V100 GPUs, the training time for each 4k images is around 8.0s for StyleGAN2, 9.8s for StyleGAN2-ADA, and 9.5s for Diffusion-StyleGAN2.
2 Effectiveness of Diffusion-GAN for domain-agnostic augmentation
To verify whether our method is domain-agnostic, we apply Diffusion-GAN onto the input feature vectors of GANs. We conduct experiments on both low-dimensional and high-dimensional feature vectors, for which commonly used image augmentation methods are no longer applicable.
25-Gaussians Example. We conduct experiments on the popular 25-Gaussians generation task. The 25-Gaussians dataset is a 2-D toy data, generated by a mixture of 25 two-dimensional Gaussian distributions. Each data point is a 2-dimensional feature vector. We train a small GAN model, whose generator and discriminator are both parameterized by multilayer perceptrons (MLPs), with two 128-unit hidden layers and LeakyReLu nonlinearities.
The training results are shown in Figure 5. We observe that the vanilla GAN exhibits severe mode collapsing, capturing only a few modes. Its discriminator outputs of real and fake samples depart from each other very quickly. This implies a strong overfitting of the discriminator happened so that the discriminator stops providing useful learning signals for the generator. However, Diffusion-GAN successfully captures all the 25 Gaussian modes and the discriminator is under control to continuously provide useful learning signals. We interpret the improvement from two perspectives: First, non-leaking augmentation helps provide more information about the data space; Second, the discriminator is well behaved given the adaptively adjusted diffusion-based noise injection.
ProjectedGAN. To verify that our adaptive diffusion-based noise injection could benefit the learning of GANs on high-dimensional feature vectors, we directly apply it to the discriminator feature space of ProjectedGAN (Sauer et al., 2021). ProjectedGANs generally leverage pre-trained neural networks to extract meaningful features for the adversarial learning of the discriminator and generator. Following Sauer et al. (2021), we adaptively diffuse the feature vectors extracted by EfficientNet-v0 and keep all the other training parts unchanged. We report the performance of Diffusion ProjectedGAN on several benchmark datasets in Table 2, which verifies that our augmentation method is domain-agnostic. Under the ProjectedGAN framework, we see that with noise properly injected into the high-dimensional feature space, Diffusion ProjectedGAN shows clear improvement in terms of both FID and Recall. We reach state-of-the-art FID results with Diffusion ProjectedGAN on STL-10 and LSUN-Bedroom/Church datasets.
3 Effectiveness of Diffusion-GAN for limited data
We evaluate whether Diffusion-GAN can provide data-efficient GAN training. We first generate five FFHQ () dataset splits, consisting of 200, 500, 1k, 2k, and 5k images, respectively, where 200 and 500 images are considered to be extremely limited data cases. We also consider AFHQ-Cat, -Dog, and -Wild (512 512), each with as few as around 5k images. Motivated by the success of InsGen (Yang et al., 2021) on small datasets, we build our Diffusion-GAN upon it. We note on limited data, InsGen convincingly outperforms both StyleGAN2+ADA and +DiffAug, and currently holds the state-of-the-art performance for data-efficient GAN training. The results in Table 3 show that our Diffusion-GAN method can help further boost the performance of InsGen in limited data settings.
Conclusion
We present Diffusion-GAN, a novel GAN framework that uses a variable-length forward diffusion chain with a Gaussian mixture distribution to generate instance noise for GAN training. This approach enables model- and domain-agnostic differentiable augmentation that leverages the advantages of diffusion without requiring a costly reverse diffusion chain. We prove theoretically and demonstrate empirically that Diffusion-GAN can prevent discriminator overfitting and provide non-leaking augmentation. We also demonstrate that Diffusion-GAN can produce high-resolution photo-realistic images with high fidelity and diversity, outperforming its corresponding state-of-the-art GAN baselines on standard benchmark datasets according to both FID and Recall.
Z. Wang, H. Zheng, and M. Zhou acknowledge the support of NSF-IIS 2212418 and IFML.
References
Appendix A Related work
A root cause of training difficulties in GANs is often attributed to the JS divergence that GANs intend to minimize. This is because when the data and generator distributions have non-overlapping supports, which are often the case for high-dimensional data supported by low-dimensional manifolds, the gradient of the JS divergence may provide no useful guidance to optimize the generator [Arjovsky and Bottou, 2017, Arjovsky et al., 2017, Mescheder et al., 2018, Roth et al., 2017]. For this reason, Arjovsky et al. propose to instead use the Wasserstein-1 distance, which in theory can provide useful gradient for the generator even if the two distributions have disjoint supports. However, Wasserstein GANs often require the use of a critic function under the 1-Lipschitz constraint, which is difficult to satisfy in practice and hence realized with heuristics such as weight clipping [Arjovsky et al., 2017], gradient penalty [Gulrajani et al., 2017], and spectral normalization [Miyato et al., 2018a].
While the divergence minimization perspective has played an important role in motivating the construction of Wasserstein GANs and gradient penalty-based regularizations, cautions should be made on purely relying on it to understand GAN training, due to not only the discrepancy between the divergence in theory and the actual min-max objective function used in practice, but also the potential confounding between different divergences and different training and regularization strategies [Fedus et al., 2018, Mescheder et al., 2018]. E.g., Mescheder et al. have provided a simple example where in theory the Wasserstein GAN is predicted to succeed while the vanilla GAN is predicted to fail, but in practice the Wasserstein GAN with a finite number of discriminator updates per generator update fails to converge while the vanilla GAN with the non-saturating loss can slowly converge. Fedus et al. provide a rich set of empirical evidence to discourage viewing GANs purely from the perspective of minimizing a specific divergence at each training step and emphasize the important role played by gradient penalties on stabilizing GAN training.
Diffusion models. Due to the use of a forward diffusion chain, the proposed Diffusion-GAN can be related to diffusion-based (or score-based) deep generative models [Ho et al., 2020b, Sohl-Dickstein et al., 2015, Song and Ermon, 2019] that employ both a forward (inference) and a reverse (generative) diffusion chain. These diffusion-based generative models are stable to train and can generate high-fidelity photo-realistic images [Dhariwal and Nichol, 2021, Ho et al., 2020b, Nichol et al., 2021, Ramesh et al., 2022, Song and Ermon, 2019, Song et al., 2021b]. However, they are notoriously slow in generation due to the need to traverse the reverse diffusion chain, which involves going through the same U-Net-based generator network hundreds or even thousands of times [Song et al., 2021a]. For this reason, a variety of methods have been proposed to reduce the generation cost of diffusion-based generative models [Kong and Ping, 2021, Luhman and Luhman, 2021, Pandey et al., 2022, San-Roman et al., 2021, Song et al., 2021a, Xiao et al., 2021, Zheng et al., 2022].
A key distinction is that Diffusion-GAN needs a reverse diffusion chain during neither training nor generation. More specifically, its generator maps the noise to a generated sample in a single step. Diffusion-GAN can train and generate as quickly as a vanilla GAN does with the same generator size. For example, it takes around 20 hours to sample 50k images of size 32 × 32 from a DDPM [Ho et al., 2020b] on an Nvidia 2080 Ti GPU, but would take less than a minute to do so from Diffusion-GAN.
Differentiable augmentation. As Diffusion-GAN transforms both the data and generated samples before sending them to the discriminator, we can also relate it to differentiable augmentation [Karras et al., 2020a, Zhao et al., 2020] proposed for data-efficient GAN training. Karras et al. [2020a] introduce a stochastic augmentation pipeline with 18 transformations and develop an adaptive mechanism for controlling the augmentation probability. Zhao et al. propose to use Color + Translation + Cutout as differentiable augmentations for both generated and real images.
While providing good empirical results on some datasets, these augmentation methods are developed with domain-specific knowledge and have the risk of leaking augmentation into generation [Karras et al., 2020a]. As observed in our experiments, they sometime worsen the results when applied to a new dataset, likely because the risk of augmentation leakage overpowers the benefits of enlarging the training set, which could happen especially if the training set size is already sufficiently large.
By contrast, Diffusion-GAN uses a differentiable forward diffusion process to stochastically transform the data and can be considered as both a domain-agnostic and a model-agnostic augmentation method. In other words, Diffusion-GAN can be applied to non-image data or even latent features, for which appropriate data augmentation is difficult to be defined, and easily plugged into an existing GAN to improve its generation performance. Moreover, we prove in theory and show in experiments that augmentation leakage is not a concern for Diffusion-GAN. Tran et al. provide a theoretical analysis for deterministic non-leaking transformation with differentiable and invertible mapping functions. Bora et al. show similar theorems to us for specific stochastic transformations, such as Gaussian Projection, Convolve+Noise, and stochastic Block-Pixels, while our Theorem 2 includes more satisfying possibilities as discussed in Appendix B.
Appendix B Proof
Since is assumed to be an isotropic Gaussian distribution, for simplicity, in what follows we show the proof in uni-variate Gaussian, which could be easily extended to multi-variate Gaussian by the production rule. We first show that under mild conditions, the and are continuous functions over .
where and are constants. Hence, is a continuous function defined on . The proof of continuity for is exactly the same proof. Then, given is also a continuous function, it is clear to see that is a continuous function over .
Next, we show that is differentiable. By the chain rule, showing to be differentiable is equivalent to show , and are differentiable. Usually, is defined with differentiability [Nowozin et al., 2016].
where and are constants. Hence, and are differentiable, which concludes the proof. ∎
We have and . If , then Let and . Given the assumption on , we have
Since and are one-to-one mapping functions, and are identifiable, which indicates f({\bm{x}})\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}f({\bm{x}}_{g})\Rightarrow{\bm{x}}\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}{\bm{x}}_{g}. By the property of moment-generating functions (MGF), given is independent with , we have for
where denotes the MGF of random variable and the others follow the same form. By the moment-generating function uniqueness theorem, given {\bm{y}}\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}{\bm{y}}_{g} and g({\bm{\epsilon}})\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}g({\bm{\epsilon}}_{g}), we have . Then, we could obtain . Thus, M_{f({\bm{x}})}=M_{f({\bm{x}}_{g})}\Rightarrow f({\bm{x}})\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}f({\bm{x}}_{g})\Rightarrow p({\bm{x}})=p({\bm{x}}_{g}), which concludes the proof.
Next, we discuss which fits the assumption we made on it. We follow the discussion of reparameterization of distributions as used in Kingma and Welling . Three basic approaches are:
Tractable inverse CDF. In this case, let , and be the inverse CDF of . From , if , for example, and , then Theorem 2 holds.
Analogous to the Gaussian example, . For any “location-scale” family of distributions we can choose the standard distribution (with location = 0, scale = 1) as the auxiliary variable , and let . Examples: Laplace, Elliptical, Student’s t, Logistic, Uniform, Triangular, and Gaussian distributions.
Implicit distributions. could be modeled by neural networks, which implies , where and are one-to-one nonlinear transformations.
Appendix C Derivations
Appendix D Details of toy example
Here, we provide the detailed analysis of the JS divergence toy example.
We recall the typical example introduced in Arjovsky and Bottou and follow the notations.
Example.
which can not provide a usable gradient for training. The derivation is as follows:
Although this simple example features distributions with disjoint supports, the same conclusion holds when the supports have a non empty intersection contained in a set of measure zero [Arjovsky and Bottou, 2017]. This happens to be the case when two low dimensional manifolds intersect in general position [Arjovsky and Bottou, 2017]. To avoid the potential issue caused by having non-overlapping distribution supports, a common remedy is to use Wasserstein-1 distance which in theory can still provide usable gradient [Arjovsky and Bottou, 2017, Arjovsky et al., 2017]. In this case, the Wasserstein-1 distance is .
Diffusion-based noise injection
In general, with our diffusion noise injected, we could have,
For the previous example, we have and such that,
which is clearly continuous and differentiable.
We show this with respect to increasing values and a grid in the second row of Figure 2. As shown in the left panel, the black line with shows the origianl JSD, which is not even continuous, while as the diffusion level increments, the lines become smoother and flatter. It is clear to see that these smooth curves provide good learning signals for . Recall that the Wasserstein-1 distance is in this case. Meanwhile, we could observe with an intense diffusion, , , the curve becomes flatter, which indicates smaller gradients and a much slower learning process. This motivates us that an adaptive diffusion could provide different level of gradient smoothness and is possibly better for training. The right panel shows the optimal discriminator outputs over the space . With diffusion, the optimal discriminator is well defined over the space and the gradient is smooth, while without diffusion the optimal discriminator is only valid on two star points. Interestingly, we find that smaller drives the optimal discriminator to become more assertive while larger makes discriminator become more neutral. The diffusion here works like a scale to balance the power of the discriminator.
Appendix E Dataset descriptions
The CIFAR-10 dataset consists of 50k training images in 10 categories. The STL-10 dataset originated from ImageNet [Deng et al., 2009] consists of 100k unlabeled images in 10 categories, and we resize them to resolution. For LSUN datasets, we sample 200k images from LSUN-Bedroom, use the whole 125k images from LSUN-Church, and resize them to resolution for training. The AFHQ datasets includes around 5k images per category for dogs, cats, and wild life; we train a separate network for each of them. The FFHQ contains 70k images crawled from Flickr at resolution and we use all of them for training.
Appendix F Algorithm
We provide the Diffusion-GAN algorithm in Algorithm 1.
Appendix G Hyperparameters
Diffusion-GAN is built on GAN backbones, so we keep the learning hyperparameters of the original GAN backbones untouched. Diffusion-GAN introduces four new hyperparameters: noise standard deviation , , increasing threshold , and sampling distribution .
The is fixed as 0.05 for images (pixel values rescaled to ) in all our experiments and it shows good performance. could be fixed as 500 or 1000, which depends on the diversity of the dataset. We recommend a large for diverse datasets. is usually fixed as 0.6, which does not influence much about the performance. has two choices, ‘uniform’ and ‘priority’. Generally, is a good starting point for a new dataset.
In our experiment, we find StyleGAN2-based models are not sensitive to the values of , so we set for them across all dataset, only except that we set for FFHQ ( for FFHQ is slightly better than in FID). We report of Diffusion ProjectedGAN for our experiments in Table 4. We also evaluated two sampling distribution , [‘priority’, ‘uniform’], defined in Equation 6. In most cases, ‘priority’ works slightly better, while in some cases, such as FFHQ, ‘uniform’ is better. Overall, we didn’t modify anything in the model architectures and training hyperparameters, such as learning rate and batch size. The forward diffusion configuration and model training configurations are as follows.
For our diffusion-based noise injection, we set up a linearly increasing schedule for , where . For pixel level injection in StyleGAN2, we follow Ho et al. [2020b] and set and . We adaptively modify ranging from to . The image pixels are usually rescaled to $\sigma=0.05\beta_{0}=0.0001\beta_{T}=0.01T_{\text{min}}=5T_{\text{max}}=500\sigma=0.5$. We list all these values in Table 5
Model config.
For StyleGAN2-based models, we borrow the config settings provided by Karras et al. [2020a], which include [‘auto’, ‘stylegan2’, ‘cifar’, ‘paper256’, ‘paper512’, ‘stylegan2’]. We create the ‘stl’ config based on ‘cifar’ with a small modification that we change the gamma term to be 0.01. For ProjectedGAN models, we use the recommended default config [Sauer et al., 2021], which is based on FastGAN [Liu et al., 2020]. We report the config settings used for our experiments in Table 6.
Appendix H Implementation details
We implement an additional diffusion sampling pipeline, where the diffusion configurations are set in Appendix G. The in the forward diffusion process is adaptively adjusted and clipped to . As illustrated in Algorithm 1, at each update step, we sample from for each data point , and then use the analytic Gaussian distribution at diffusion step to sample . Next, we use and instead of for optimization.
We inherit all the network architectures from StyleGAN2 implemented by Karras et al. [2020a]. We modify the original mapping network, which is there for label conditioning and unused for unconditional image generation tasks, inside the discriminator to inject . Specifically, we change the original input of mapping network, the class label , to our discrete value timestep . Then, we train the generator and discriminator with diffused samples and .
Diffuson ProjectedGAN.
To simplify the implementation and minimize the modifications to ProjectedGAN, we construct the discriminator as , where is ignored. Our method is plugged in as a data augmentation method. The only change in the optimization stage is that the discriminator is fed with diffused images instead of original images .
Diffuson InsGen.
To simplify the implementation and minimize the modifications to InsGen, we keep their contrastive learning part untouched. We modify the original discriminator network to inject similarly to Diffusion StyleGAN2. Then, we train the generator and discriminator with diffused samples and .
Appendix I Ablation on the mixing procedure and T𝑇T adaptiveness
Note the mixing procedure described in Equation 6, referred to as “priority mixing” in what follows, is designed based on our intuition. Here we conduct an ablation study on the mixing procedure by comparing the priority mixing with uniform mixing on three representative datasets. We report in Table 7 the FID results, which suggest that uniform mixing could work better than priority mixing in some dataset, and hence Diffusion-GAN may be further improved by optimizing its mixing procedure according to the training data. While optimizing the mixing procedure is beyond the focus of this paper, it is worth further investigation in future studies.
We further conduct ablation study on whether the needs to be adaptively adjusted. As shown in Figure 7, we observe with adaptive diffusion strategy, the training curves of FIDs converge faster and reach lower final FIDs.
Appendix J More GAN variants
To further validate our noise injection via diffusion-based mixtures, we add our diffusion-based training into two more representative GAN variants: DCGAN [Radford et al., 2015] and SNGAN [Miyato et al., 2018b], which have quite different GAN architectures compared to StyleGAN2. We provide the FIDs for CIFAR-10 in Table 8. We observe that both Diffusion-DCGAN and Diffusion-SNGAN clearly outperform their corresponding baseline GANs.
Appendix K Inception Score for CIFAR-10
We report the Inception Score (IS) [Salimans et al., 2016] of Diffusion StyleGAN2 for CIFAR-10 dataset in Table 9 and also include other state-of-the-art GANs and diffusion models as baselines. Note CIFAR-10 is a well-known dataset and tested by almost all baselines, so we pick CIFAR-10 here and we reference the reported IS values from their original papers for a fair comparison.
Appendix L More generated images
We provide more randomly generated images for LSUN-Bedroom, LSUN-Church, AFHQ, and FFHQ datasets in Figure 8, Figure 9, and Figure 10.