Diffusion-GAN: Training GANs with Diffusion

Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, Mingyuan Zhou

Introduction

Generative adversarial networks (GANs) (Goodfellow et al., 2014) and their variants (Brock et al., 2018; Karras et al., 2019; 2020a; Zhao et al., 2020) have achieved great success in synthesizing photo-realistic high-resolution images. GANs in practice, however, are known to suffer from a variety of issues ranging from non-convergence and training instability to mode collapse (Arjovsky and Bottou, 2017; Mescheder et al., 2018). As a result, a wide array of analyses and modifications has been proposed for GANs, including improving the network architectures (Karras et al., 2019; Radford et al., 2016; Sauer et al., 2021; Zhang et al., 2019), gaining theoretical understanding of GAN training (Arjovsky and Bottou, 2017; Heusel et al., 2017; Mescheder et al., 2017; 2018), changing the objective functions (Arjovsky et al., 2017; Bellemare et al., 2017; Deshpande et al., 2018; Li et al., 2017a; Nowozin et al., 2016; Zheng and Zhou, 2021; Yang et al., 2021), regularizing the weights and/or gradients (Arjovsky et al., 2017; Fedus et al., 2018; Mescheder et al., 2018; Miyato et al., 2018a; Roth et al., 2017; Salimans et al., 2016), utilizing side information (Wang et al., 2018; Zhang et al., 2017; 2020b), adding a mapping from the data to latent representation (Donahue et al., 2016; Dumoulin et al., 2016; Li et al., 2017b), and applying differentiable data augmentation (Karras et al., 2020a; Zhang et al., 2020a; Zhao et al., 2020).

A simple technique to stabilize GAN training is to inject instance noise, i.e.i.e., to add noise to the discriminator input, which can widen the support of both the generator and discriminator distributions and prevent the discriminator from overfitting (Arjovsky and Bottou, 2017; Sønderby et al., 2017). However, this technique is hard to implement in practice, as finding a suitable noise distribution is challenging (Arjovsky and Bottou, 2017). Roth et al. (2017) show that adding instance noise to the high-dimensional discriminator input does not work well, and propose to approximate it by adding a zero-centered gradient penalty on the discriminator. This approach is theoretically and empirically shown to converge in Mescheder et al. (2018), who also demonstrate that adding zero-centered gradient penalties to non-saturating GANs can result in stable training and better or comparable generation quality compared to WGAN-GP (Arjovsky et al., 2017). However, Brock et al. (2018) caution that zero-centered gradient penalties and other similar regularization methods may stabilize training at the cost of generation performance. To the best of our knowledge, there has been no existing work that is able to empirically demonstrate the success of using instance noise in GAN training on high-dimensional image data.

To inject proper instance noise that can facilitate GAN training, we introduce Diffusion-GAN, which uses a diffusion process to generate Gaussian-mixture distributed instance noise. We show a graphical representation of Diffusion-GAN in Figure 1. In Diffusion-GAN, the input to the diffusion process is either a real or a generated image, and the diffusion process consists of a series of steps that gradually add noise to the image. The number of diffusion steps is not fixed, but depends on the data and the generator. We also design the diffusion process to be differentiable, which means that we can compute the derivative of the output with respect to the input. This allows us to propagate the gradient from the discriminator to the generator through the diffusion process, and update the generator accordingly. Unlike vanilla GANs, which compare the real and generated images directly, Diffusion-GAN compares the noisy versions of them, which are obtained by sampling from the Gaussian mixture distribution over the diffusion steps, with the help of our timestep-dependent discriminator. This distribution has the property that its components have different noise-to-data ratios, which means that some components add more noise than others. By sampling from this distribution, we can achieve two benefits: first, we can stabilize the training by easing the problem of vanishing gradient, which occurs when the data and generator distributions are too different; second, we can augment the data by creating different noisy versions of the same image, which can improve the data efficiency and the diversity of the generator. We provide a theoretical analysis to support our method, and show that the min-max objective function of Diffusion-GAN, which measures the difference between the data and generator distributions, is continuous and differentiable everywhere. This means that the generator in theory can always receive a useful gradient from the discriminator, and improve its performance.

Our main contributions include: 1) We show both theoretically and empirically how the diffusion process can be utilized to provide a model- and domain-agnostic differentiable augmentation, enabling data-efficient and leaking-free stable GAN training. 2) Extensive experiments show that Diffusion-GAN boosts the stability and generation performance of strong baselines, including StyleGAN2 (Karras et al., 2020b), Projected GAN (Sauer et al., 2021), and InsGen (Yang et al., 2021), achieving state-of-the-art results in synthesizing photo-realistic images, as measured by both the Fréchet Inception Distance (FID) (Heusel et al., 2017) and Recall score (Kynkäänniemi et al., 2019).

Preliminaries: GANs and diffusion-based generative models

GANs (Goodfellow et al., 2014) are a class of generative models that aim to learn the data distribution p(x)p({\bm{x}}) of a target dataset by setting up a min-max game between two neural networks: a generator and a discriminator. The generator GG takes as input a random noise vector z{\bm{z}} sampled from a simple prior distribution p(z)p({\bm{z}}), such as a standard normal or uniform distribution, and tries to produce realistic-looking samples G(z)G({\bm{z}}) that resemble the data. The discriminator DD receives either a real data sample x{\bm{x}} drawn from p(x)p({\bm{x}}) or a fake sample G(z)G({\bm{z}}) generated by GG, and tries to correctly classify them as real or fake. The goal of GG is to fool DD into making mistakes, while the goal of DD is to accurately distinguish G(z)G({\bm{z}}) from x{\bm{x}}. The min-max objective function of GANs is given by

In practice, this vanilla objective function is often modified to improve the stability and performance of GANs(Goodfellow et al., 2014; Miyato et al., 2018a; Fedus et al., 2018), but the general idea of adversarial learning between GG and DD remains the same.

Diffusion-based generative models (Ho et al., 2020b; Sohl-Dickstein et al., 2015; Song and Ermon, 2019) assume pθ(x0):=∫pθ(x0:T)dx1:Tp_{\theta}({\bm{x}}_{0}):=\int p_{\theta}({\bm{x}}_{0:T})d{\bm{x}}_{1:T}, where x1,…,xT{\bm{x}}_{1},\ldots,{\bm{x}}_{T} are latent variables of the same dimensionality as the data x0∼p(x0){\bm{x}}_{0}\sim p({\bm{x}}_{0}). There is a forward diffusion chain that gradually adds noise to the data x0∼q(x0){\bm{x}}_{0}\sim q({\bm{x}}_{0}) in TT steps with pre-defined variance schedule βt\beta_{t} and variance σ2\sigma^{2}:

A notable property is that xt{\bm{x}}_{t} at an arbitrary time-step tt can be sampled in closed form as

A variational lower bound (Blei et al., 2017) is then used to optimize the reverse diffusion chain as

Diffusion-GAN: Method and Theoretical Analysis

To construct Diffusion-GAN, we describe how to inject instance noise via diffusion, how to train the generator by backpropagating through the forward diffusion process, and how to adaptively adjust the diffusion intensity. We further provide theoretical analysis illustrated with a toy example.

We aim to generate realistic samples xg{\bm{x}}_{g} from a generator network GG that maps a latent variable z{\bm{z}} sampled from a simple prior distribution p(z)p({\bm{z}}) to a high-dimensional data space, such as images. The distribution of generator samples xg=G(z){\bm{x}}_{g}=G({\bm{z}}), z∼p(z){\bm{z}}\sim p({\bm{z}}) is denoted by pg(x)=∫p(xg ∣ z)p(z)dzp_{g}({\bm{x}})=\int p({\bm{x}}_{g}\,|\,{\bm{z}})p({\bm{z}})d{\bm{z}}. To make the generator more robust and diverse, we inject instance noise into the generated samples xg{\bm{x}}_{g} by applying a diffusion process that adds Gaussian noise at each step. The diffusion process can be seen as a Markov chain that starts from the original sample x{\bm{x}} and gradually erases its information until reaching a noise level σ2\sigma^{2} after TT steps.

We define a mixture distribution q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}) that models the noisy samples y{\bm{y}} obtained at any step of the diffusion process, with a mixture weight πt\pi_{t} for each step tt. The mixture components q(y ∣ x,t)q({\bm{y}}\,|\,{\bm{x}},t) are Gaussian distributions with mean proportional to x{\bm{x}} and variance depending on the noise level at step tt. We use the same diffusion process and mixture distribution for both the real samples x∼p(x){\bm{x}}\sim p({\bm{x}}) and the generated samples xg∼pg(x){\bm{x}}_{g}\sim p_{g}({\bm{x}}). More specifically, the diffusion-induced mixture distributions are expressed as

where q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}) is a TT-component mixture distribution, the mixture weights πt\pi_{t} are non-negative and sum to one, and the mixture components q(y ∣ x,t)q({\bm{y}}\,|\,{\bm{x}},t) are obtained via diffusion as in Equation 1, expressed as

Samples from this mixture can be drawn as t∼pπ:=\mboxDiscrete(π1,…,πT), y∼q(y ∣ x,t)t\sim p_{\pi}:=\mbox{Discrete}(\pi_{1},\ldots,\pi_{T}),~{}{\bm{y}}\sim q({\bm{y}}\,|\,{\bm{x}},t).

By sampling y{\bm{y}} from this mixture distribution, we can obtain noisy versions of both real and generated samples with varying degrees of noise. The more steps we take in the diffusion process, the more noise we add to y{\bm{y}} and the less information we preserve from x{\bm{x}}. We can then use this diffusion-induced mixture distribution to train a timestep-dependent discriminator DD that distinguishes between real and generated noisy samples, and a generator GG that matches the distribution of generated noisy samples to the distribution of real noisy samples. Next we introduce Diffusion-GAN that trains its discriminator and generator with the help of the diffusion-induced mixture distribution.

2 Adversarial Training

The Diffusion-GAN trains its generator and discriminator by solving a min-max game objective as

Here, p(x)p({\bm{x}}) is the true data distribution, pπp_{\pi} is a discrete distribution that assigns different weights πt\pi_{t} to each diffusion step t∈{1,…,T}t\in\{1,\dots,T\}, and q(y ∣ x,t)q({\bm{y}}\,|\,{\bm{x}},t) is the conditional distribution of the perturbed sample y{\bm{y}} given the original data x{\bm{x}} and the diffusion step tt. By Equation 2, with Gaussian reparameterization, the perturbation function could be written as y=αˉtx+1−αˉtσϵ{\bm{y}}=\sqrt{\bar{\alpha}_{t}}{\bm{x}}+\sqrt{1-\bar{\alpha}_{t}}\sigma{\bm{\epsilon}}, where 1−αˉt=1−∏s=1tαs1-\bar{\alpha}_{t}=1-\prod_{s=1}^{t}\alpha_{s} is the cumulative noise level at step tt, σ\sigma is a scale factor, and ϵ∼N(0,I){\bm{\epsilon}}\sim\mathcal{N}(0,{\bm{I}}) is a Gaussian noise.

The objective function in Equation 3 encourages the discriminator to assign high probabilities to the perturbed real data and low probabilities to the perturbed generated data, for any diffusion step tt. The generator, on the other hand, tries to produce samples that can deceive the discriminator at any diffusion step tt. Note that the perturbed generated sample yg∼q(y ∣ Gθ(z),t){\bm{y}}_{g}\sim q({\bm{y}}\,|\,G_{\theta}({\bm{z}}),t) can be rewritten as yg=αˉtGθ(z)+(1−αˉt)σϵ,ϵ∼N(0,I){\bm{y}}_{g}=\sqrt{\bar{\alpha}_{t}}G_{\theta}({\bm{z}})+\sqrt{(1-\bar{\alpha}_{t})}\sigma{\bm{\epsilon}},{\bm{\epsilon}}\sim\mathcal{N}(0,{\bm{I}}). This means that the objective function in Equation 3 is differentiable with respect to the generator parameters, and we can use gradient descent to optimize it with back-propagation.

The objective function Equation 3 is similar to the one used by the original GAN (Goodfellow et al., 2014), except that it involves the diffusion steps and the perturbation functions. We can show that this objective function also minimizes an approximation of the Jensen–Shannon (JS) divergence between the true and the generated distributions, but with respect to the perturbed samples and the diffusion steps, as follows:

The JS divergence measures the dissimilarity between two probability distributions, and it reaches its minimum value of zero when the two distributions are identical. The proof of the equality in Equation 4 is given in Appendix C. A natural question that arises from this result is whether minimizing the JS divergence between the perturbed distributions implies minimizing the JS divergence between the original distributions, i.e.i.e., whether the optimal generator for Equation 3 is also the optimal generator for DJS(p(x)∣∣pg(x)){\mathcal{D}}_{\text{JS}}(p({\bm{x}})||p_{g}({\bm{x}})). We will answer this question affirmatively and provide a theoretical justification in Section 3.4.

3 Adaptive diffusion

With the help of the perturbation function and timestep dependency, we have a new strategy to optimize the discriminator. We want the discriminator DD to have a challenging task, neither too easy to allow overfitting the data (Karras et al., 2020a; Zhao et al., 2020) nor too hard to impede learning. Therefore, we adjust the intensity of the diffusion process, which adds noise to both y{\bm{y}} and yg{\bm{y}}_{g}, depending on how much DD can distinguish them. When the diffusion step tt is larger, the noise-to-data ratios are higher and the task is harder. We use 1−αˉt1-\bar{\alpha}_{t} to measure the intensity of the diffusion, which increases as tt grows. To control the diffusion intensity, we adaptively modify the maximum number of steps TT.

Our strategy is to make the discriminator learn from the easiest samples first, which are the original data samples, and then gradually increase the difficulty by feeding it samples from larger tt. To do this, we use a self-paced schedule for TT, which depends on a metric rdr_{d} that estimates how much the discriminator overfits to the data:

where rdr_{d} is the same as in Karras et al. (2020a) and CC is a constant. We calculate rdr_{d} and update TT every four minibatches. We have two options for the distribution pπp_{\pi} that we use to sample tt for the diffusion process:

The ‘priority’ option gives more weight to larger tt, which means the discriminator will see more new samples from the new steps when TT increases. This is because we want the discriminator to focus on the new and harder samples that it has not seen before, as this indicates that it is confident about the easier ones. Note that even with the ‘priority’ option, the discriminator can still see samples from smaller tt, because q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}) is a mixture of Gaussians that covers all steps of the diffusion chain.

To avoid sudden changes in TT during training, we use an exploration list tepl{\bm{t}}_{epl} that contains tt values sampled from pπp_{\pi}. We keep tepl{\bm{t}}_{epl} fixed until we update TT, and we sample tt from tepl{\bm{t}}_{epl} to generate noisy samples for the discriminator. This way, the model can explore each tt sufficiently before moving to a higher TT. We give the details of training Diffusion-GAN in Algorithm 1 in Appendix F.

4 Theoretical analysis with Examples

To better understand the theoretical properties of our proposed method, we present two theorems that address two important questions about the use of diffusion-based instance noise injection for training GANs. The proofs of these theorems are deferred to Appendix B. The first question, denoted as (a), is whether adding noise to the real and generated samples in a diffusion process can facilitate the learning. The second question, denoted as (b), is whether minimizing the JS divergence between the joint distributions of the noisy samples and the noise levels, p(y,t)p({\bm{y}},t) and pg(y,t)p_{g}({\bm{y}},t), can lead to the same optimal generator as minimizing the JS divergence between the original distributions of the real and generated samples, p(x)p({\bm{x}}) and pg(x)p_{g}({\bm{x}}).

To answer (a), we prove that for any choice of noise level tt and any choice of convex function ff, the ff-divergence (Nowozin et al., 2016) between the marginal distributions of the noisy real and generated samples, q(y ∣ t)q({\bm{y}}\,|\,t) and q(yg ∣ t)q({\bm{y}}_{g}\,|\,t), is a smooth function that can be computed and optimized by the discriminator. This implies that the diffusion-based noise injection does not introduce any singularity or discontinuity in the objective function of the GAN. The JS divergence is a special case of ff-divergence, where f(u)=−log⁡(2u)−log⁡(2−2u)f(u)=-\log(2u)-\log(2-2u).

Let p(x)p({\bm{x}}) be a fixed distribution over X{\mathcal{X}} and z{\bm{z}} be a random noise over another space Z{\mathcal{Z}}. Denote Gθ:Z→XG_{\theta}:{\mathcal{Z}}\rightarrow{\mathcal{X}} as a function with parameter θ\theta and input z{\bm{z}} and pg(x)p_{g}({\bm{x}}) as the distribution of Gθ(z)G_{\theta}({\bm{z}}). Let q(y ∣ x,t)=N(y;αˉtx,(1−αˉt)σ2I)q({\bm{y}}\,|\,{\bm{x}},t)={\mathcal{N}}({\bm{y}};\sqrt{\bar{\alpha}_{t}}{\bm{x}},(1-\bar{\alpha}_{t})\sigma^{2}{\bm{I}}), where αˉt∈(0,1)\bar{\alpha}_{t}\in(0,1) and σ>0\sigma>0. Let q(y ∣ t)=∫p(x)q(y ∣ x,t)dxq({\bm{y}}\,|\,t)=\int p({\bm{x}})q({\bm{y}}\,|\,{\bm{x}},t)d{\bm{x}} and qg(y ∣ t)=∫pg(x)q(y ∣ x,t)dxq_{g}({\bm{y}}\,|\,t)=\int p_{g}({\bm{x}})q({\bm{y}}\,|\,{\bm{x}},t)d{\bm{x}}. Then, ∀t\forall t, if function GθG_{\theta} is continuous and differentiable, the f-divergence Df(q(y ∣ t)∣∣qg(y ∣ t)){\mathcal{D}}_{f}(q({\bm{y}}\,|\,t)||q_{g}({\bm{y}}\,|\,t)) is continuous and differentiable with respect to θ\theta.

Theorem 1 shows that with the help of diffusion noise injection by q(y ∣ x,t)q({\bm{y}}\,|\,{\bm{x}},t), ∀t\forall t, y{\bm{y}} and yg{\bm{y}}_{g} are defined on the same support space, the whole X{\mathcal{X}}, and Df(q(y ∣ t)∣∣qg(y ∣ t)){\mathcal{D}}_{f}(q({\bm{y}}\,|\,t)||q_{g}({\bm{y}}\,|\,t)) is continuous and differentiable everywhere. Then, one natural question is what if Df(q(y ∣ t)∣∣qg(y ∣ t)){\mathcal{D}}_{f}(q({\bm{y}}\,|\,t)||q_{g}({\bm{y}}\,|\,t)) keeps a near constant value and hence provides little useful gradient. Hence, we empirically show that by injecting noise through a mixture defined over all steps of the diffusion chain, there is always a good chance that a sufficiently large tt is sampled to provide a useful gradient, via the toy example below.

Toy example. We use the same simple example from Arjovsky et al. (2017) to illustrate our method. Let x=(0,z){\bm{x}}=(0,z) be the real data and xg=(θ,z){\bm{x}}_{g}=(\theta,z) be the data generated by a one-parameter generator, where zz is a uniform random variable in $.TheJSdivergencebetweentherealandthegenerateddistributions,. The JS divergence between the real and the generated distributions,{\mathcal{D}}_{\text{JS}}(p({\bm{x}})||p({\bm{x}}_{g})),isdiscontinuous:itis0when, is discontinuous: it is 0 when\theta=0andand\log 2otherwise,soitdoesnotprovideausefulgradienttoguideotherwise, so it does not provide a useful gradient to guide\theta$ towards zero.

For the discriminator optimization, as shown in the second row, right, of Figure 2, the optimal discriminator under the original JS divergence is discontinuous and unattainable. With diffusion-based noise, the optimal discriminator changes with tt: a smaller tt makes it more confident and a larger tt makes it more cautious. Thus the diffusion acts like a scale to balance the power of the discriminator. This suggests the use of a differentiable forward diffusion chain that can provide various levels of gradient smoothness to help the generator training.

Let x∼p(x),y∼q(y ∣ x){\bm{x}}\sim p({\bm{x}}),{\bm{y}}\sim q({\bm{y}}\,|\,{\bm{x}}) and xg∼pg(x),yg∼q(yg ∣ xg){\bm{x}}_{g}\sim p_{g}({\bm{x}}),{\bm{y}}_{g}\sim q({\bm{y}}_{g}\,|\,{\bm{x}}_{g}), where q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}) is the transition density. Given certain q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}), if y{\bm{y}} could be reparameterized into y=f(x)+h(ϵ), ϵ∼p(ϵ){\bm{y}}=f({\bm{x}})+h({\bm{\epsilon}}),~{}{\bm{\epsilon}}\sim p({\bm{\epsilon}}), where p(ϵ)p({\bm{\epsilon}}) is a known distribution, and both ff and hh are one-to-one mapping functions, then we could have p(y)=pg(y)⇔p(x)=pg(x)p({\bm{y}})=p_{g}({\bm{y}})\Leftrightarrow p({\bm{x}})=p_{g}({\bm{x}}).

To answer question (b), we present Theorem 2, which shows a sufficient condition for the equality of the original and the augmented data distributions. By Theorem 2, the function ff maps each x{\bm{x}} to a unique y{\bm{y}}, the function hh maps each ϵ{\bm{\epsilon}} to a unique noise term, and the distribution of ϵ{\bm{\epsilon}} is known and independent of x{\bm{x}}. Under these assumptions, the theorem proves that the distribution of y{\bm{y}} is the same as the distribution of yg{\bm{y}}_{g}, if and only if the distribution of x{\bm{x}} is the same as the distribution of xg{\bm{x}}_{g}. If we take y ∣ t{\bm{y}}\,|\,t as the y{\bm{y}} introduced in the theorem, then for ∀t\forall t, Equation 2 fits the assumption made. This means that, by minimizing the divergence between q(y ∣ t)q({\bm{y}}\,|\,t) and qg(y ∣ t)q_{g}({\bm{y}}\,|\,t), which is the same as minimizing the divergence between p(x) ∣ tp({\bm{x}})\,|\,t and pg(x) ∣ tp_{g}({\bm{x}})\,|\,t, we are also minimizing the divergence between p(x)p({\bm{x}}) and pg(x)p_{g}({\bm{x}}). This implies that the noise injection does not affect the quality of the generated samples, and we can safely use our noise injection to improve the training of the generative model.

5 Related work

The proposed Diffusion-GAN can be related to previous works on stabilizing the GAN training, building diffusion-based generative models, and constructing differential augmentation for data-efficient GAN training. A detailed discussion on these related works is deferred to Appendix A.

Experiments

We conduct extensive experiments to answer the following questions: (a) Will Diffusion-GAN outperform state-of-the-art GAN baselines on benchmark datasets? (b) Will the diffusion-based noise injection help the learning of GANs in domain-agnostic tasks? (c) Will our method improve the performance of data-efficient GANs trained with a very limited amount of data?

Datasets. We conduct experiments on image datasets ranging from low-resolution (e.g.e.g., 32×3232\times 32) to high-resolution (e.g.e.g., 1024×10241024\times 1024) and from low-diversity to high-diversity: CIFAR-10 (Krizhevsky, 2009), STL-10 (Coates et al., 2011), LSUN-Bedroom (Yu et al., 2015), LSUN-Church (Yu et al., 2015), AFHQ(Cat/Dog/Wild) (Choi et al., 2020), and FFHQ (Karras et al., 2019). More details on these benchmark datasets are provided in Appendix E.

Evaluation protocol. We measure image quality using FID (Heusel et al., 2017). Following Karras et al. (2019; 2020b), we measure FID using 50k generated samples, with the full training set used as reference. We use the number of real images shown to the discriminator to evaluate convergence (Karras et al., 2020a; Sauer et al., 2021). Unless specified otherwise, all models are trained with 25 million images to ensure convergence (these trained with more or fewer images are specified in table captions). We further report the improved Recall score introduced by Kynkäänniemi et al. (2019) to measure the sample diversity of generative models.

Implementations and resources. We build Diffusion-GANs based on the code of StyleGAN2 (Karras et al., 2020b), ProjectedGAN (Sauer et al., 2021), and InsGen (Yang et al., 2021) to answer questions (a), (b), and (c), respectively. Diffusion GANs inherit from their corresponding base GANs all their network architectures and training hyperparamters, whose details are provided in Appendix G. Specifically for StyleGAN2 and InsGen, we construct the discriminator as Dϕ(y,t)D_{\phi}({\bm{y}},t), where tt is injected via their mapping network. For ProjectedGAN, we empirically find tt in the discriminator could be ignored to simplify the implementation and minimize the modifications to ProjectedGAN. More implementation details are provided in Appendix H. By applying our diffusion-based noise injection, we denote our models as Diffusion StyleGAN2/ProjectedGAN/InsGen. In the following experiments, we train related models with their official code if the results are unavailable, while others are all reported from references and marked with ∗. We run all our experiments with either 4 or 8 NVIDIA V100 GPUs depending on the demands of the inherited training configurations.

We compare Diffusion-GAN with its state-of-the-art GAN backbone, StyleGAN2 (Karras et al., 2020a), and to evaluate its effectiveness from the data augmentation perspective, we compare it with both StyleGAN2 + DiffAug (Zhao et al., 2020) and StyleGAN2 + ADA (Karras et al., 2020a), in terms of both sample fidelity (FID) and sample diversity (Recall) over extensive benchmark datasets.

We present the quantitative and qualitative results in Table 1 and Figure 3. Qualitatively, these generated images from Diffusion StyleGAN2 are all photo-realistic and have good diversity, ranging from low-resolution (32 ×\times 32) to high-resolution (1024 ×\times 1024). Additional randomly generated images can be found in Appendix L. Quantitatively, Diffusion StyleGAN2 outperforms all the GAN baselines in generation diversity, as measured by Recall, on all 6 benchmark datasets and outperforms them in FID by a clear margin on 5 out of the 6 benchmark datasets.

From the data augmentation perspective, we observe that Diffusion StyleGAN2 always clearly outperforms the backbone model StyleGAN2 across various datasets, which empirically validates our Theorem 2. By contrast, both the ADA (Karras et al., 2020b) and Diffaug (Zhao et al., 2020) techniques could sometimes impair the generation performance on sufficiently large datasets, e.g.e.g., LSUN-Bedroom and LSUN-Church, which is also observed by Yang et al. (2021) on FFHQ. This is possibly because their risk of leaking augmentation overshadows the benefits of data augmentation.

To investigate how the adaptive diffusion process works during training, we illustrate in Figure 4 the convergence of the maximum timestep TT in our adaptive diffusion and discriminator outputs. We see that TT is adaptively adjusted: The TT for Diffusion StyleGAN2 increases as the training goes while the TT for Diffusion ProjectedGAN first goes up and then goes down. Note that the TT is adjusted according to the overfitting status of the discriminator. The second panel shows that trained with the diffusion-based mixture distribution, the discriminator is always well behaved and provides useful learning signals for the generator, which validates our analysis in Section 3.4 and Theorem 1.

Memory and time costs. Generally speaking, the memory and time costs of a Diffusion-GAN are comparable to those of the corresponding GAN baseline. More specifically, switching from ADA (Karras et al., 2020a) to our diffusion-based augmentation, the added memory cost is negative, the added training time cost is negative, and the added inference time cost is zero. For example, for CIFAR-10, with four NVIDIA V100 GPUs, the training time for each 4k images is around 8.0s for StyleGAN2, 9.8s for StyleGAN2-ADA, and 9.5s for Diffusion-StyleGAN2.

2 Effectiveness of Diffusion-GAN for domain-agnostic augmentation

To verify whether our method is domain-agnostic, we apply Diffusion-GAN onto the input feature vectors of GANs. We conduct experiments on both low-dimensional and high-dimensional feature vectors, for which commonly used image augmentation methods are no longer applicable.

25-Gaussians Example. We conduct experiments on the popular 25-Gaussians generation task. The 25-Gaussians dataset is a 2-D toy data, generated by a mixture of 25 two-dimensional Gaussian distributions. Each data point is a 2-dimensional feature vector. We train a small GAN model, whose generator and discriminator are both parameterized by multilayer perceptrons (MLPs), with two 128-unit hidden layers and LeakyReLu nonlinearities.

The training results are shown in Figure 5. We observe that the vanilla GAN exhibits severe mode collapsing, capturing only a few modes. Its discriminator outputs of real and fake samples depart from each other very quickly. This implies a strong overfitting of the discriminator happened so that the discriminator stops providing useful learning signals for the generator. However, Diffusion-GAN successfully captures all the 25 Gaussian modes and the discriminator is under control to continuously provide useful learning signals. We interpret the improvement from two perspectives: First, non-leaking augmentation helps provide more information about the data space; Second, the discriminator is well behaved given the adaptively adjusted diffusion-based noise injection.

ProjectedGAN. To verify that our adaptive diffusion-based noise injection could benefit the learning of GANs on high-dimensional feature vectors, we directly apply it to the discriminator feature space of ProjectedGAN (Sauer et al., 2021). ProjectedGANs generally leverage pre-trained neural networks to extract meaningful features for the adversarial learning of the discriminator and generator. Following Sauer et al. (2021), we adaptively diffuse the feature vectors extracted by EfficientNet-v0 and keep all the other training parts unchanged. We report the performance of Diffusion ProjectedGAN on several benchmark datasets in Table 2, which verifies that our augmentation method is domain-agnostic. Under the ProjectedGAN framework, we see that with noise properly injected into the high-dimensional feature space, Diffusion ProjectedGAN shows clear improvement in terms of both FID and Recall. We reach state-of-the-art FID results with Diffusion ProjectedGAN on STL-10 and LSUN-Bedroom/Church datasets.

3 Effectiveness of Diffusion-GAN for limited data

We evaluate whether Diffusion-GAN can provide data-efficient GAN training. We first generate five FFHQ (1024×10241024\times 1024) dataset splits, consisting of 200, 500, 1k, 2k, and 5k images, respectively, where 200 and 500 images are considered to be extremely limited data cases. We also consider AFHQ-Cat, -Dog, and -Wild (512 ×\times 512), each with as few as around 5k images. Motivated by the success of InsGen (Yang et al., 2021) on small datasets, we build our Diffusion-GAN upon it. We note on limited data, InsGen convincingly outperforms both StyleGAN2+ADA and +DiffAug, and currently holds the state-of-the-art performance for data-efficient GAN training. The results in Table 3 show that our Diffusion-GAN method can help further boost the performance of InsGen in limited data settings.

Conclusion

We present Diffusion-GAN, a novel GAN framework that uses a variable-length forward diffusion chain with a Gaussian mixture distribution to generate instance noise for GAN training. This approach enables model- and domain-agnostic differentiable augmentation that leverages the advantages of diffusion without requiring a costly reverse diffusion chain. We prove theoretically and demonstrate empirically that Diffusion-GAN can prevent discriminator overfitting and provide non-leaking augmentation. We also demonstrate that Diffusion-GAN can produce high-resolution photo-realistic images with high fidelity and diversity, outperforming its corresponding state-of-the-art GAN baselines on standard benchmark datasets according to both FID and Recall.

Z. Wang, H. Zheng, and M. Zhou acknowledge the support of NSF-IIS 2212418 and IFML.

References

Appendix A Related work

A root cause of training difficulties in GANs is often attributed to the JS divergence that GANs intend to minimize. This is because when the data and generator distributions have non-overlapping supports, which are often the case for high-dimensional data supported by low-dimensional manifolds, the gradient of the JS divergence may provide no useful guidance to optimize the generator [Arjovsky and Bottou, 2017, Arjovsky et al., 2017, Mescheder et al., 2018, Roth et al., 2017]. For this reason, Arjovsky et al. propose to instead use the Wasserstein-1 distance, which in theory can provide useful gradient for the generator even if the two distributions have disjoint supports. However, Wasserstein GANs often require the use of a critic function under the 1-Lipschitz constraint, which is difficult to satisfy in practice and hence realized with heuristics such as weight clipping [Arjovsky et al., 2017], gradient penalty [Gulrajani et al., 2017], and spectral normalization [Miyato et al., 2018a].

While the divergence minimization perspective has played an important role in motivating the construction of Wasserstein GANs and gradient penalty-based regularizations, cautions should be made on purely relying on it to understand GAN training, due to not only the discrepancy between the divergence in theory and the actual min-max objective function used in practice, but also the potential confounding between different divergences and different training and regularization strategies [Fedus et al., 2018, Mescheder et al., 2018]. E.g., Mescheder et al. have provided a simple example where in theory the Wasserstein GAN is predicted to succeed while the vanilla GAN is predicted to fail, but in practice the Wasserstein GAN with a finite number of discriminator updates per generator update fails to converge while the vanilla GAN with the non-saturating loss can slowly converge. Fedus et al. provide a rich set of empirical evidence to discourage viewing GANs purely from the perspective of minimizing a specific divergence at each training step and emphasize the important role played by gradient penalties on stabilizing GAN training.

Diffusion models. Due to the use of a forward diffusion chain, the proposed Diffusion-GAN can be related to diffusion-based (or score-based) deep generative models [Ho et al., 2020b, Sohl-Dickstein et al., 2015, Song and Ermon, 2019] that employ both a forward (inference) and a reverse (generative) diffusion chain. These diffusion-based generative models are stable to train and can generate high-fidelity photo-realistic images [Dhariwal and Nichol, 2021, Ho et al., 2020b, Nichol et al., 2021, Ramesh et al., 2022, Song and Ermon, 2019, Song et al., 2021b]. However, they are notoriously slow in generation due to the need to traverse the reverse diffusion chain, which involves going through the same U-Net-based generator network hundreds or even thousands of times [Song et al., 2021a]. For this reason, a variety of methods have been proposed to reduce the generation cost of diffusion-based generative models [Kong and Ping, 2021, Luhman and Luhman, 2021, Pandey et al., 2022, San-Roman et al., 2021, Song et al., 2021a, Xiao et al., 2021, Zheng et al., 2022].

A key distinction is that Diffusion-GAN needs a reverse diffusion chain during neither training nor generation. More specifically, its generator maps the noise to a generated sample in a single step. Diffusion-GAN can train and generate as quickly as a vanilla GAN does with the same generator size. For example, it takes around 20 hours to sample 50k images of size 32 × 32 from a DDPM [Ho et al., 2020b] on an Nvidia 2080 Ti GPU, but would take less than a minute to do so from Diffusion-GAN.

Differentiable augmentation. As Diffusion-GAN transforms both the data and generated samples before sending them to the discriminator, we can also relate it to differentiable augmentation [Karras et al., 2020a, Zhao et al., 2020] proposed for data-efficient GAN training. Karras et al. [2020a] introduce a stochastic augmentation pipeline with 18 transformations and develop an adaptive mechanism for controlling the augmentation probability. Zhao et al. propose to use Color + Translation + Cutout as differentiable augmentations for both generated and real images.

While providing good empirical results on some datasets, these augmentation methods are developed with domain-specific knowledge and have the risk of leaking augmentation into generation [Karras et al., 2020a]. As observed in our experiments, they sometime worsen the results when applied to a new dataset, likely because the risk of augmentation leakage overpowers the benefits of enlarging the training set, which could happen especially if the training set size is already sufficiently large.

By contrast, Diffusion-GAN uses a differentiable forward diffusion process to stochastically transform the data and can be considered as both a domain-agnostic and a model-agnostic augmentation method. In other words, Diffusion-GAN can be applied to non-image data or even latent features, for which appropriate data augmentation is difficult to be defined, and easily plugged into an existing GAN to improve its generation performance. Moreover, we prove in theory and show in experiments that augmentation leakage is not a concern for Diffusion-GAN. Tran et al. provide a theoretical analysis for deterministic non-leaking transformation with differentiable and invertible mapping functions. Bora et al. show similar theorems to us for specific stochastic transformations, such as Gaussian Projection, Convolve+Noise, and stochastic Block-Pixels, while our Theorem 2 includes more satisfying possibilities as discussed in Appendix B.

Appendix B Proof

Since N(y;atx,btI){\mathcal{N}}({\bm{y}};a_{t}{\bm{x}},b_{t}{\bm{I}}) is assumed to be an isotropic Gaussian distribution, for simplicity, in what follows we show the proof in uni-variate Gaussian, which could be easily extended to multi-variate Gaussian by the production rule. We first show that under mild conditions, the pr′,t(y)p_{r^{\prime},t}(y) and pg′,t(y)p_{g^{\prime},t}(y) are continuous functions over yy.

where C1C_{1} and C2C_{2} are constants. Hence, pr′,t(y)p_{r^{\prime},t}(y) is a continuous function defined on yy. The proof of continuity for pg′,t(y)p_{g^{\prime},t}(y) is exactly the same proof. Then, given gθg_{\theta} is also a continuous function, it is clear to see that Df(pr′,t(y)∣∣pg′,t(y)){\mathcal{D}}_{f}(p_{r^{\prime},t}({\bm{y}})||p_{g^{\prime},t}({\bm{y}})) is a continuous function over θ\theta.

Next, we show that Df(pr′,t(y)∣∣pg′,t(y)){\mathcal{D}}_{f}(p_{r^{\prime},t}({\bm{y}})||p_{g^{\prime},t}({\bm{y}})) is differentiable. By the chain rule, showing Df(pr′,t(y)∣∣pg′,t(y)){\mathcal{D}}_{f}(p_{r^{\prime},t}({\bm{y}})||p_{g^{\prime},t}({\bm{y}})) to be differentiable is equivalent to show pr′,t(y)p_{r^{\prime},t}(y), pr′,t(y)p_{r^{\prime},t}(y) and ff are differentiable. Usually, ff is defined with differentiability [Nowozin et al., 2016].

where C1C_{1} and C2C_{2} are constants. Hence, pr′,t(y)p_{r^{\prime},t}(y) and pr′,t(y)p_{r^{\prime},t}(y) are differentiable, which concludes the proof. ∎

We have p(y)=∫p(x)q(y ∣ x)dxp({\bm{y}})=\int p({\bm{x}})q({\bm{y}}\,|\,{\bm{x}})d{\bm{x}} and pg(y)=∫pg(x)q(y ∣ x)dxp_{g}({\bm{y}})=\int p_{g}({\bm{x}})q({\bm{y}}\,|\,{\bm{x}})d{\bm{x}}. ⇐\Leftarrow If p(x)=pg(x)p({\bm{x}})=p_{g}({\bm{x}}), then p(y)=pg(y)p({\bm{y}})=p_{g}({\bm{y}}) ⇒\Rightarrow Let y∼p(y){\bm{y}}\sim p({\bm{y}}) and yg∼pg(y){\bm{y}}_{g}\sim p_{g}({\bm{y}}). Given the assumption on q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}), we have

Since ff and gg are one-to-one mapping functions, f(x)f({\bm{x}}) and g(ϵ)g({\bm{\epsilon}}) are identifiable, which indicates f({\bm{x}})\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}f({\bm{x}}_{g})\Rightarrow{\bm{x}}\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}{\bm{x}}_{g}. By the property of moment-generating functions (MGF), given f(x)f({\bm{x}}) is independent with g(ϵ)g({\bm{\epsilon}}), we have for ∀s\forall{\bm{s}}

where My(s)=Ey∼p(y)[esTy]M_{{\bm{y}}}({\bm{s}})=E_{{\bm{y}}\sim p({\bm{y}})}[e^{{\bm{s}}^{T}{\bm{y}}}] denotes the MGF of random variable y{\bm{y}} and the others follow the same form. By the moment-generating function uniqueness theorem, given {\bm{y}}\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}{\bm{y}}_{g} and g({\bm{\epsilon}})\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}g({\bm{\epsilon}}_{g}), we have My(s)=Myg(s)\mboxandMg(ϵ)(s)=Mg(ϵg)(s)\mboxfor∀sM_{{\bm{y}}}({\bm{s}})=M_{{\bm{y}}_{g}}({\bm{s}})\mbox{ and }M_{g({\bm{\epsilon}})}({\bm{s}})=M_{g({\bm{\epsilon}}_{g})}({\bm{s}})\mbox{ for }\forall{\bm{s}}. Then, we could obtain Mf(x)=Mf(xg)\mboxfor∀sM_{f({\bm{x}})}=M_{f({\bm{x}}_{g})}\mbox{ for }\forall{\bm{s}}. Thus, M_{f({\bm{x}})}=M_{f({\bm{x}}_{g})}\Rightarrow f({\bm{x}})\mathrel{\ooalign{\raisebox{1.5pt}{\scriptscriptstyle D}\cr=\cr}}f({\bm{x}}_{g})\Rightarrow p({\bm{x}})=p({\bm{x}}_{g}), which concludes the proof.

Next, we discuss which q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}) fits the assumption we made on it. We follow the discussion of reparameterization of distributions as used in Kingma and Welling . Three basic approaches are:

Tractable inverse CDF. In this case, let ϵ∼U(0,I){\bm{\epsilon}}\sim{\mathcal{U}}(\mathbf{0},{\bm{I}}), and ψ(ϵ,y,x)\psi({\bm{\epsilon}},{\bm{y}},{\bm{x}}) be the inverse CDF of q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}). From ψ(ϵ,y,x)\psi({\bm{\epsilon}},{\bm{y}},{\bm{x}}), if y=f(x)+g(ϵ){\bm{y}}=f({\bm{x}})+g({\bm{\epsilon}}), for example, y∼\mboxCauchy(x,γ)y\sim\mbox{Cauchy}(x,\gamma) and y∼\mboxLogistic(x,s)y\sim\mbox{Logistic}(x,s), then Theorem 2 holds.

Analogous to the Gaussian example, y∼N(x,σ2I)⇒y=x+σ⋅ϵ,ϵ∼N(0,I){\bm{y}}\sim{\mathcal{N}}({\bm{x}},\sigma^{2}{\bm{I}})\Rightarrow{\bm{y}}={\bm{x}}+\sigma\cdot{\bm{\epsilon}},{\bm{\epsilon}}\sim{\mathcal{N}}(\mathbf{0},{\bm{I}}). For any “location-scale” family of distributions we can choose the standard distribution (with location = 0, scale = 1) as the auxiliary variable ϵ{\bm{\epsilon}}, and let g(.)=\mboxlocation+\mboxscale⋅ϵg(.)=\mbox{location}+\mbox{scale}\cdot{\bm{\epsilon}}. Examples: Laplace, Elliptical, Student’s t, Logistic, Uniform, Triangular, and Gaussian distributions.

Implicit distributions. q(y ∣ x)q({\bm{y}}\,|\,{\bm{x}}) could be modeled by neural networks, which implies y=f(x)+g(ϵ),ϵ∼p(ϵ){\bm{y}}=f({\bm{x}})+g({\bm{\epsilon}}),{\bm{\epsilon}}\sim p({\bm{\epsilon}}), where ff and gg are one-to-one nonlinear transformations.

Appendix C Derivations

Appendix D Details of toy example

Here, we provide the detailed analysis of the JS divergence toy example.

We recall the typical example introduced in Arjovsky and Bottou and follow the notations.

Example.

which can not provide a usable gradient for training. The derivation is as follows:

Although this simple example features distributions with disjoint supports, the same conclusion holds when the supports have a non empty intersection contained in a set of measure zero [Arjovsky and Bottou, 2017]. This happens to be the case when two low dimensional manifolds intersect in general position [Arjovsky and Bottou, 2017]. To avoid the potential issue caused by having non-overlapping distribution supports, a common remedy is to use Wasserstein-1 distance which in theory can still provide usable gradient [Arjovsky and Bottou, 2017, Arjovsky et al., 2017]. In this case, the Wasserstein-1 distance is ∣θ∣|\theta|.

Diffusion-based noise injection

In general, with our diffusion noise injected, we could have,

For the previous example, we have Yt′Y_{t}^{\prime} and Yg,t′Y_{g,t}^{\prime} such that,

which is clearly continuous and differentiable.

We show this DJS(pr′,t∣∣pg′,t){\mathcal{D}}_{JS}(p_{r^{\prime},t}||p_{g^{\prime},t}) with respect to increasing tt values and a θ\theta grid in the second row of Figure 2. As shown in the left panel, the black line with t=0t=0 shows the origianl JSD, which is not even continuous, while as the diffusion level tt increments, the lines become smoother and flatter. It is clear to see that these smooth curves provide good learning signals for θ\theta. Recall that the Wasserstein-1 distance is ∣θ∣|\theta| in this case. Meanwhile, we could observe with an intense diffusion, e.g.e.g., t=800t=800, the curve becomes flatter, which indicates smaller gradients and a much slower learning process. This motivates us that an adaptive diffusion could provide different level of gradient smoothness and is possibly better for training. The right panel shows the optimal discriminator outputs over the space X{\mathcal{X}}. With diffusion, the optimal discriminator is well defined over the space and the gradient is smooth, while without diffusion the optimal discriminator is only valid on two star points. Interestingly, we find that smaller tt drives the optimal discriminator to become more assertive while larger tt makes discriminator become more neutral. The diffusion here works like a scale to balance the power of the discriminator.

Appendix E Dataset descriptions

The CIFAR-10 dataset consists of 50k 32×3232\times 32 training images in 10 categories. The STL-10 dataset originated from ImageNet [Deng et al., 2009] consists of 100k unlabeled images in 10 categories, and we resize them to 64×6464\times 64 resolution. For LSUN datasets, we sample 200k images from LSUN-Bedroom, use the whole 125k images from LSUN-Church, and resize them to 256×256256\times 256 resolution for training. The AFHQ datasets includes around 5k 512×512512\times 512 images per category for dogs, cats, and wild life; we train a separate network for each of them. The FFHQ contains 70k images crawled from Flickr at 1024×10241024\times 1024 resolution and we use all of them for training.

Appendix F Algorithm

We provide the Diffusion-GAN algorithm in Algorithm 1.

Appendix G Hyperparameters

Diffusion-GAN is built on GAN backbones, so we keep the learning hyperparameters of the original GAN backbones untouched. Diffusion-GAN introduces four new hyperparameters: noise standard deviation σ\sigma, TmaxT_{max}, TT increasing threshold dtargetd_{target}, and tt sampling distribution pπp_{\pi}.

The σ\sigma is fixed as 0.05 for images (pixel values rescaled to ) in all our experiments and it shows good performance. TmaxT_{max} could be fixed as 500 or 1000, which depends on the diversity of the dataset. We recommend a large TmaxT_{max} for diverse datasets. dtargetd_{target} is usually fixed as 0.6, which does not influence much about the performance. pπp_{\pi} has two choices, ‘uniform’ and ‘priority’. Generally, (σ=0.05,Tmax=500,dtarget=0.6,pπ=‘uniform’)(\sigma=0.05,T_{max}=500,d_{target}=0.6,p_{\pi}=\text{`uniform'}) is a good starting point for a new dataset.

In our experiment, we find StyleGAN2-based models are not sensitive to the values of dtargetd_{target}, so we set dtarget=0.6d_{target}=0.6 for them across all dataset, only except that we set dtarget=0.8d_{target}=0.8 for FFHQ (dtarget=0.8d_{target}=0.8 for FFHQ is slightly better than 0.60.6 in FID). We report dtargetd_{target} of Diffusion ProjectedGAN for our experiments in Table 4. We also evaluated two tt sampling distribution pπp_{\pi}, [‘priority’, ‘uniform’], defined in Equation 6. In most cases, ‘priority’ works slightly better, while in some cases, such as FFHQ, ‘uniform’ is better. Overall, we didn’t modify anything in the model architectures and training hyperparameters, such as learning rate and batch size. The forward diffusion configuration and model training configurations are as follows.

For our diffusion-based noise injection, we set up a linearly increasing schedule for βt\beta_{t}, where t∈{1,2,…,T}t\in\{1,2,\dots,T\}. For pixel level injection in StyleGAN2, we follow Ho et al. [2020b] and set β0=0.0001\beta_{0}=0.0001 and βT=0.02\beta_{T}=0.02. We adaptively modify TT ranging from Tmin=5T_{\text{min}}=5 to Tmax=1000T_{\text{max}}=1000. The image pixels are usually rescaled to $sowesettheGuassiannoisestandarddeviationso we set the Guassian noise standard deviation\sigma=0.05.ForfeaturelevelinjectioninDiffusionProjectedGAN,weset. For feature level injection in Diffusion ProjectedGAN, we set\beta_{0}=0.0001,,\beta_{T}=0.01,,T_{\text{min}}=5,,T_{\text{max}}=500,and, and\sigma=0.5$. We list all these values in Table 5

Model config.

For StyleGAN2-based models, we borrow the config settings provided by Karras et al. [2020a], which include [‘auto’, ‘stylegan2’, ‘cifar’, ‘paper256’, ‘paper512’, ‘stylegan2’]. We create the ‘stl’ config based on ‘cifar’ with a small modification that we change the gamma term to be 0.01. For ProjectedGAN models, we use the recommended default config [Sauer et al., 2021], which is based on FastGAN [Liu et al., 2020]. We report the config settings used for our experiments in Table 6.

Appendix H Implementation details

We implement an additional diffusion sampling pipeline, where the diffusion configurations are set in Appendix G. The TT in the forward diffusion process is adaptively adjusted and clipped to [Tmin,Tmax][T_{\text{min}},T_{\text{max}}]. As illustrated in Algorithm 1, at each update step, we sample tt from teplt_{epl} for each data point x{\bm{x}}, and then use the analytic Gaussian distribution at diffusion step tt to sample y{\bm{y}}. Next, we use y{\bm{y}} and tt instead of x{\bm{x}} for optimization.

We inherit all the network architectures from StyleGAN2 implemented by Karras et al. [2020a]. We modify the original mapping network, which is there for label conditioning and unused for unconditional image generation tasks, inside the discriminator to inject tt. Specifically, we change the original input of mapping network, the class label cc, to our discrete value timestep tt. Then, we train the generator and discriminator with diffused samples y{\bm{y}} and tt.

Diffuson ProjectedGAN.

To simplify the implementation and minimize the modifications to ProjectedGAN, we construct the discriminator as Dϕ(y)D_{\phi}({\bm{y}}), where tt is ignored. Our method is plugged in as a data augmentation method. The only change in the optimization stage is that the discriminator is fed with diffused images y{\bm{y}} instead of original images x{\bm{x}}.

Diffuson InsGen.

To simplify the implementation and minimize the modifications to InsGen, we keep their contrastive learning part untouched. We modify the original discriminator network to inject tt similarly to Diffusion StyleGAN2. Then, we train the generator and discriminator with diffused samples y{\bm{y}} and tt.

Appendix I Ablation on the mixing procedure and T𝑇T adaptiveness

Note the mixing procedure described in Equation 6, referred to as “priority mixing” in what follows, is designed based on our intuition. Here we conduct an ablation study on the mixing procedure by comparing the priority mixing with uniform mixing on three representative datasets. We report in Table 7 the FID results, which suggest that uniform mixing could work better than priority mixing in some dataset, and hence Diffusion-GAN may be further improved by optimizing its mixing procedure according to the training data. While optimizing the mixing procedure is beyond the focus of this paper, it is worth further investigation in future studies.

We further conduct ablation study on whether the TT needs to be adaptively adjusted. As shown in Figure 7, we observe with adaptive diffusion strategy, the training curves of FIDs converge faster and reach lower final FIDs.

Appendix J More GAN variants

To further validate our noise injection via diffusion-based mixtures, we add our diffusion-based training into two more representative GAN variants: DCGAN [Radford et al., 2015] and SNGAN [Miyato et al., 2018b], which have quite different GAN architectures compared to StyleGAN2. We provide the FIDs for CIFAR-10 in Table 8. We observe that both Diffusion-DCGAN and Diffusion-SNGAN clearly outperform their corresponding baseline GANs.

Appendix K Inception Score for CIFAR-10

We report the Inception Score (IS) [Salimans et al., 2016] of Diffusion StyleGAN2 for CIFAR-10 dataset in Table 9 and also include other state-of-the-art GANs and diffusion models as baselines. Note CIFAR-10 is a well-known dataset and tested by almost all baselines, so we pick CIFAR-10 here and we reference the reported IS values from their original papers for a fair comparison.

Appendix L More generated images

We provide more randomly generated images for LSUN-Bedroom, LSUN-Church, AFHQ, and FFHQ datasets in Figure 8, Figure 9, and Figure 10.