On Fast Sampling of Diffusion Probabilistic Models

Zhifeng Kong, Wei Ping

Introduction

Diffusion probabilistic models are a class of deep generative models that use Markov chains to gradually transform between a simple distribution (e.g., isotropic Gaussian) and the complex data distribution (Sohl-Dickstein et al., 2015; Ho et al., 2020). Most recently, these models have obtained the state-of-the-art results in several important domains, including image synthesis (Ho et al., 2020; Song et al., 2020b; Dhariwal and Nichol, 2021), audio synthesis (Kong et al., 2020b; Chen et al., 2020), and 3-D point cloud generation (Luo and Hu, 2021; Zhou et al., 2021). We will use “diffusion models” as shorthand to refer to this family of models.

Diffusion models usually comprise: i) a parameter-free TT-step Markov chain named the diffusion process, which gradually adds random noise into the data, and ii) a parameterized TT-step Markov chain called the reverse or denoising process, which removes the added noise as a denoising function. The likelihood in diffusion models is intractable, but they can be efficiently trained by optimizing a variant of the variational lower bound. In particular, Ho et al. (2020) propose a certain parameterization called the denoising diffusion probabilistic model (DDPM) and show its connection with denoising score matching (Song and Ermon, 2019), so the reverse process can be viewed as sampling from a score-based model using Langevin dynamics. DDPM can produce high-fidelity samples reliably with large model capacity and outperforms the state-of-the-art models in image and audio domains (Dhariwal and Nichol, 2021; Kong et al., 2020b). However, a noticeable limitation of diffusion models is their expensive denoising or sampling process. For example, DDPM requires a Markov chain with T=1000T=1000 steps to generate high quality image samples (Ho et al., 2020), and DiffWave requires T=200T=200 to obtain high-fidelity audio synthesis (Kong et al., 2020b). In other words, one has to run the forward-pass of the neural network TT times to generate a sample, which is much slower than the state-of-the-art GANs or flow-based models for image and audio synthesis (e.g., Karras et al., 2020; Kingma and Dhariwal, 2018; Kong et al., 2020a; Ping et al., 2020).

To deal with this limitation, several methods have been proposed to reduce the length of the reverse process to S≪TS\ll T steps. One class of methods compute continuous noise levels based on discrete diffusion steps and retrain a new model conditioned on these continuous noise levels (Song and Ermon, 2019; Chen et al., 2020; Okamoto et al., 2021; San-Roman et al., 2021). Then, a shorter reverse process can be obtained by carefully choosing a small set (size SS) of noise levels. However, these methods cannot reuse the pretrained diffusion models, because the state-of-the-art DDPM models are conditioned on discrete diffusion steps (Ho et al., 2020; Dhariwal and Nichol, 2021). It is also unclear the diffusion models conditioned on continuous noise levels can achieve comparable sample quality as the state-of-the-art DDPMs on challenging unconditional image and audio synthesis tasks (Dhariwal and Nichol, 2021; Kong et al., 2020b). Another class of methods directly approximate the original reverse process of DDPM models with shorter ones (of length SS), which are conditioned on discrete diffusion steps (Song et al., 2020a; Kong et al., 2020b). Although both classes of methods have shown the trade-off between sampling speed and sample quality (i.e., larger SS lead to higher sample quality), the fast sampling methods without retraining are more advantageous for fast iteration and deployment, while still keeping high-fidelity synthesis with small number of steps in the reverse process (e.g., S=6S=6 in Kong et al. (2020b)).

In this work, we propose FastDPM, a unified framework of fast sampling methods for diffusion models without retraining. The core idea of FastDPM is to i) generalize discrete diffusion steps to continuous diffusion steps, and ii) design a bijective mapping between continuous diffusion steps and continuous noise levels. Then, we use this bijection to construct an approximate diffusion process and an approximate reverse process, both of which have length S≪TS\ll T.

FastDPM includes and generalizes the fast sampling algorithms from denoising diffusion implicit models (DDIM) (Song et al., 2020a) and DiffWave (Kong et al., 2020b). In detail, FastDPM offers two ways to construct the approximate diffusion process: selecting SS steps in the original diffusion process, or more flexibly, choosing SS variances. FastDPM also offers ways to construct the approximate reverse process: using the stochastic DDPM reverse process (DDPM-rev), or using the implicit (deterministic) DDIM reverse process (DDIM-rev). We can control the amount of stochasticity in the reverse process of FastDPM as in Song et al. (2020a).

FastDPM gives rise to new algorithms with improved sample quality than previous methods when the length of the approximate reverse process SS is small. We then extensively evaluate the family of FastDPM methods across image and audio domains. We find the deterministic DDIM-rev significantly outperforms the stochastic DDPM-rev in image generation tasks, but DDPM-rev significantly outperforms DDIM-rev in audio synthesis tasks. Finally, we investigate the performance of different methods by varying the amount of conditional information. We find with different amount of conditional information, we need different amount of stochasticity in the reverse process of FastDPM.

In summary, we make the following contributions:

FastDPM introduces the concept of continuous diffusion steps, and generalizes prior fast sampling algorithms without retraining (Song et al., 2020a; Kong et al., 2020b).

FastDPM gives rise to new algorithms with improved sample quality when the length of the approximate reverse process SS is small.

We extensively evaluate FastDPM across image and audio domains, and provide insights and recipes on the choice of methods for practitioners.

We organize the rest of the paper as follows. Section 2 discusses related work. We introduce the preliminaries of diffusion models in Section 3, and propose FastDPM in Section 4. We report experimental results in Section 5 and conclude the paper in Section 6.

Related Work

Diffusion models are a class of powerful deep generative models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Goyal et al., 2017), which have received a lot of attention recently. These models have been applied to various domains, including image generation (Ho et al., 2020; Dhariwal and Nichol, 2021), audio synthesis (Kong et al., 2020b; Chen et al., 2020; Okamoto et al., 2021), image or audio super-resolution (Li et al., 2021; Lee and Han, 2021), text-to-speech (Jeong et al., 2021; Popov et al., 2021), music synthesis (Liu et al., 2021; Mittal et al., 2021), 3-D point cloud generation (Luo and Hu, 2021; Zhou et al., 2021), and language models (Hoogeboom et al., 2021). Diffusion models are connected with scored-based models (Song and Ermon, 2019, 2020; Song et al., 2020b), and there have been a series of research extending and improving diffusion models (Song et al., 2020b; Gao et al., 2020; Dhariwal and Nichol, 2021; San-Roman et al., 2021; Meng et al., 2021).

There are two families of methods aiming for accelerating diffusion models at synthesis, which reduce the length of the reverse process from TT to a much smaller SS. One family of methods tackle this problem at training. They retrain the network conditioned on continuous noise levels instead of discrete diffusion steps (Song and Ermon, 2019; Chen et al., 2020; Okamoto et al., 2021; San-Roman et al., 2021). Assuming that the corresponding network is able to predict added noise at any noise level, we can carefully choose only S≪TS\ll T noise levels and construct a short reverse process just based on them. San-Roman et al. (2021) present a learning scheme that can step-by-step adjust those noise level parameters, for any given number of steps SS. Another family of methods aim to directly approximate the original reverse process within the pretrained DDPM conditioned on discrete steps. In other words, no retraining is needed. Song et al. (2020a) introduce denoising diffusion implicit models (DDIM), which contain non-Markovian processes that lead to an equivalent training objective as DDPM. These non-Markovian processes naturally permit "jumping steps", or formally, using a subset of steps to form a short reverse process. However, compared to using continuous noise levels, selecting discrete steps offers less flexibility. Kong et al. (2020b) introduce a fast sampling algorithm by interpolating steps according to corresponding noise levels. This can be seen as an attempt to map continuous noise levels to discrete diffusion steps. However, it lacks both theoretical justification for the interpolation and extensive empirical studies.

In this paper, we propose FastDPM, a method that approximates the original DDPM model. FastDPM constructs a bijective mapping between (continuous) diffusion steps and continuous noise levels. This allows us to take advantage of the flexibility of using these continuous noise levels. FastDPM generalizes Kong et al. (2020b) by using Gamma functions to compute noise levels, which naturally extends from discrete domain to continuous domain. FastDPM generalizes Song et al. (2020a) by providing a special set of noise levels that exactly correspond to integer steps.

Diffusion Models

where each of q(xt∣xt−1)=N(xt;1−βtxt−1,βtI)q(x_{t}|x_{t-1})={\mathcal{N}}(x_{t};\sqrt{1-\beta_{t}}x_{t-1},\beta_{t}I) for some small constant βt>0\beta_{t}>0. The hyperparameters β1,⋯ ,βT\beta_{1},\cdots,\beta_{T} are called the variance schedule.

where each of pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) is defined as N(xt−1;μθ(xt,t),σt2I){\mathcal{N}}(x_{t-1};\mu_{\theta}(x_{t},t),\sigma_{t}^{2}I); the mean μθ(xt,t)\mu_{\theta}(x_{t},t) is parameterized through a neural network and the variance σt\sigma_{t} is time-step dependent constant. Based on the reverse process, the sampling process is to first draw xT∼N(0,I)x_{T}\sim{\mathcal{N}}(0,I), then draw xt−1∼pθ(xt−1∣xt)x_{t-1}\sim p_{\theta}(x_{t-1}|x_{t}) for t=T,T−1,⋯ ,1t=T,T-1,\cdots,1, and finally outputs x0x_{0}.

thus one can directly sample xtx_{t} given x0x_{0} (see derivation in Appendix A).

where ϵ∼N(0,I)\epsilon\sim{\mathcal{N}}(0,I), x0∼qdatax_{0}\sim q_{\rm{data}}, tt is uniformly taken from 1,⋯ ,T1,\cdots,T, and xt=αˉt⋅x0+1−αˉt⋅ϵx_{t}=\sqrt{\bar{\alpha}_{t}}\cdot x_{0}+\sqrt{1-\bar{\alpha}_{t}}\cdot\epsilon from Eq. (3). One may simply interpret this objective as a mean-squared error loss between the true noise ϵ\epsilon and the predicted noise ϵθ(xt, t)\epsilon_{\theta}(x_{t},~{}t) at each time-step.

FastDPM: A Unified Framework for Fast Sampling in Diffusion Models

In this section, we generalize discrete (integer) diffusion steps to continuous (real-valued) diffusion steps. Then, we introduce a bijective mapping R{\mathcal{R}} and T=R−1{\mathcal{T}}=R^{-1} between continuous diffusion steps tt and noise levels rr: r=R(t)r={\mathcal{R}}(t) and t=T(r)t={\mathcal{T}}(r).

We start with an integer diffusion step tt. From Eq. (3), one can observe xt=αˉt⋅x0+1−αˉt⋅ϵx_{t}=\sqrt{\bar{\alpha}_{t}}\cdot x_{0}+\sqrt{1-\bar{\alpha}_{t}}\cdot\epsilon where ϵ∼N(0,I)\epsilon\sim{\mathcal{N}}(0,I), thus sampling xtx_{t} given x0x_{0} is equivalent to adding a Gaussian noise to x0x_{0}. Based on this observation, we define the noise level at step tt as R(t)=αˉt{\mathcal{R}}(t)=\sqrt{\bar{\alpha}_{t}}, which means xtx_{t} is composed of R(t){\mathcal{R}}(t) fraction of the data x0x_{0} and (1−R(t))(1-{\mathcal{R}}(t)) fraction of white noise. For example, R(t)=0{\mathcal{R}}(t)=0 means no noise and R(t)=1{\mathcal{R}}(t)=1 means pure white noise. Next, we extend the domain of R{\mathcal{R}} to real values. Assume that the variance schedule {βt}t=1T\{\beta_{t}\}_{t=1}^{T} is linear: βi=β1+(i−1)Δβ\beta_{i}=\beta_{1}+(i-1)\Delta\beta, where Δβ=βT−β1T−1\Delta\beta=\frac{\beta_{T}-\beta_{1}}{T-1} (Ho et al., 2020). We further define an auxiliary constant β^=1−β1Δβ\hat{\beta}=\frac{1-\beta_{1}}{\Delta\beta}, which is ≫T\gg T assuming that βT≪1.0\beta_{T}\ll 1.0. E.g., β1=1×10−4\beta_{1}=1\times 10^{-4}, βT=0.02\beta_{T}=0.02 in Ho et al. (2020); Kong et al. (2020b). Then, we have

Because the Gamma function Γ\Gamma is well-defined on (0,∞)(0,\infty), Eq. (7) gives rise to a natural extension of αˉt\bar{\alpha}_{t} for continuous diffusion steps tt. As a result, for t∈[0,β^)t\in[0,\hat{\beta}), we define the noise level at tt as:

Define 𝒯𝒯{\mathcal{T}}.

For any noise level r∈(0,1)r\in(0,1), its corresponding (continuous) diffusion step, T(r){\mathcal{T}}(r), is defined by inverting R{\mathcal{R}}:

By Stirling’s approximation to Gamma functions, we have

Given a noise level r=R(t)r={\mathcal{R}}(t), we numerically solve t=T(r)t={\mathcal{T}}(r) by applying a binary search based on Eq. (4.1). We have T(r)∈[t,t+1]{\mathcal{T}}(r)\in[t,t+1] for r∈[αˉt+1,αˉt]r\in[\sqrt{\bar{\alpha}_{t+1}},\sqrt{\bar{\alpha}_{t}}], and this provides a good initialization to the binary search algorithm. Experimentally, we find the binary search algorithm converges with high precision in no more than 20 iterations.

2 Approximate the Diffusion Process

One can see this by rewriting Eq. (3): ηs\eta_{s} corresponds to βt=1−αt\beta_{t}=1-\alpha_{t}, γs\gamma_{s} corresponds to αt\alpha_{t}, and rsr_{s} corresponds to αˉt\sqrt{\bar{\alpha}_{t}}. We then propose the following two ways to schedule the noise levels {rs}s=1S\{r_{s}\}_{s=1}^{S}.

We start from the variance schedule {ηs}s=1S\{\eta_{s}\}_{s=1}^{S}. Next, we compute γs=1−ηs\gamma_{s}=1-\eta_{s} and γˉs=∏i=1sγi\bar{\gamma}_{s}=\prod_{i=1}^{s}\gamma_{i}. The noise level at step ss is then rs=γˉsr_{s}=\sqrt{\bar{\gamma}_{s}}.

Noise levels from steps (STEP).

We start from a subset of diffusion steps {τs}s−1S\{\tau_{s}\}_{s-1}^{S} in {1,⋯ ,T}\{1,\cdots,T\}. Then, the noise level at step ss is rs=R(τs)=αˉτsr_{s}={\mathcal{R}}(\tau_{s})=\sqrt{\bar{\alpha}_{\tau_{s}}}.

When ηs=1−αˉτs/αˉτs−1\eta_{s}=1-\bar{\alpha}_{\tau_{s}}/\bar{\alpha}_{\tau_{s-1}}, we have γˉs=αˉτs\bar{\gamma}_{s}=\bar{\alpha}_{\tau_{s}}. Therefore, noise levels from steps can be regarded as a special case of noise levels from variances.

3 Approximate the Reverse Process

Given the same sequence of noise levels in Section 4.2, we aim to approximate the reverse process Eq. (2) in the original DDPM. To achieve this goal, we regard the model ϵθ\epsilon_{\theta} as being trained on variances {ηs}s=1S\{\eta_{s}\}_{s=1}^{S} instead of the original {βt}t=1T\{\beta_{t}\}_{t=1}^{T}. Then, the transition probability in the approximate reverse process is

DDIM reverse process (DDIM-rev).

When κ=1\kappa=1, the coefficient of the ϵθ\epsilon_{\theta} term in the DDIM reverse process is

Therefore, the DDPM reverse process is a special case of the DDIM reverse process (κ=1\kappa=1).

4 Connections with Previous Methods

The DDIM (Song et al., 2020a) method is equivalent to selecting noise levels from steps and using DDIM-rev in FastDPM. The fast sampling algorithm by DiffWave (Kong et al., 2020b) is related to selecting noise levels from variances and using DDPM-rev in FastDPM. Compared with DiffWave, FastDPM offers an automatic way to select variances in different settings and a more natural way to compute noise levels.

Experiments

In this section, we aim to answer the following two questions for FastDPM:

Which approximate diffusion process, VAR or STEP, is better?

Which approximate reverse process, DDPM-rev or DDIM-rev, is better?

We investigate these questions by conducting extensive experiments in both image and audio domains.

Image datasets. We conduct unconditional image generation experiments on three datasets: CIFAR-10 (50k object images of resolution 32×3232\times 32 (Krizhevsky et al., 2009)), CelebA (∼\sim163k face images of resolution 64×6464\times 64 (Liu et al., 2015)), and LSUN-bedroom (∼\sim3M bedroom images of resolution 256×256256\times 256 (Yu et al., 2015)).

Audio datasets. We conduct unconditional and class-conditional audio synthesis experiments on the Speech Commands 0-9 (SC09) dataset, the spoken digit subset of the full Speech Commands dataset (Warden, 2018). SC09 contains ∼\sim31k one-second long utterances of ten classes (0 through 9) with a sampling rate of 16kHz. We conduct neural vocoding experiments (audio synthesis conditioned on mel spectrogram) on the LJSpeech dataset (Ito, 2017). It contains ∼\sim24 hours of audio (∼\sim13k utterances from a female speaker) recorded in home environment with a sampling rate of 22.05kHz.

Models. In all experiments, we use pretrained checkpoints in prior works. In detail, the pretrained models for CIFAR-10 and LSUN-bedroom are taken from DDPM (Ho et al., 2020; Esser, 2020), the pretrained model for CelebA is taken from DDIM (Song et al., 2020a). In these models, TT is 10001000. The pretrained models for SC09 and LJSpeech are taken from DiffWave (Kong et al., 2020b). In these models, TT is 200200. In all models, β1=10−4\beta_{1}=10^{-4}, βT=2×10−2\beta_{T}=2\times 10^{-2}, and all βt\beta_{t}’s are linearly interpolated between β1\beta_{1} and βT\beta_{T}.

Noise level schedules. For each of the approximate diffusion process in Section 4.2, we examine two schedules: linear and quadratic. For noise levels {ηs}s=1S\{\eta_{s}\}_{s=1}^{S} from variances, the two schedules are:

Linear (VAR): ηs=(1+cs) η0\eta_{s}=(1+cs)~{}\eta_{0}.

Quadratic (VAR): ηs=(1+cs)2 η0\eta_{s}=(1+cs)^{2}~{}\eta_{0}.

We let η0=β0\eta_{0}=\beta_{0} and the constant cc satisfy ∏s=1S(1−ηs)=αˉT\prod_{s=1}^{S}(1-\eta_{s})=\bar{\alpha}_{T}. The noise level at step ss is rs=γˉsr_{s}=\sqrt{\bar{\gamma}_{s}}.

For noise levels {ηs}s=1S\{\eta_{s}\}_{s=1}^{S} from steps, they are computed from selected steps {τs}s=1S\{\tau_{s}\}_{s=1}^{S} among {1,⋯ ,T}\{1,\cdots,T\} (Song et al., 2020a). The two schedules are:

Linear (STEP): τs=⌊cs⌋\tau_{s}=\lfloor cs\rfloor, where c=TSc=\frac{T}{S}.

Quadratic (STEP): τs=⌊cs2⌋\tau_{s}=\lfloor cs^{2}\rfloor, where c=45⋅TS2c=\frac{4}{5}\cdot\frac{T}{S^{2}}.

Then, the noise level at step ss is rs=R(τs)=αˉτsr_{s}={\mathcal{R}}(\tau_{s})=\sqrt{\bar{\alpha}_{\tau_{s}}}.

In image generation experiments, we follow the same noise level schedules as in Song et al. (2020a): quadratic schedules for CIFAR-10 and linear schedules for CelebA and LSUN-bedroom. We use linear schedules in SC09 experiments and quadratic schedules in LJSpeech experiments; we find these schedules have better quality.

Evaluations. In all unconditional generation experiments, we use the Fréchet Inception Distance (FID) (Heusel et al., 2017; Lang, 2020) to evaluate generated samples. For the training set XtX_{t} and the set of generated samples XgX_{g}, the FID between these two sets is defined as

where μt,μg\mu_{t},\mu_{g} and Σt,Σg\Sigma_{t},\Sigma_{g} are the means and covariances of Xt,XgX_{t},X_{g} after a feature transformation. In each image generation experiment, XgX_{g} is 50K generated images. The transformed feature is the 2048-dimensional vector output of the last layer of Inception-V3 (Szegedy et al., 2015). In each audio synthesis experiment, XgX_{g} is 5K generated utterances. The transformed feature is the 1024-dimensional vector output of the last layer of a ResNeXT classifier (Xu and Tuguldur, 2017), which achieves 99.06%99.06\% accuracy on the training set and 98.76%98.76\% accuracy on the test set. The FID is the smaller the better.

In the class-conditional generation experiment on SC09, we evaluate with accuracy and the Inception Score (IS) (Salimans et al., 2016). Note that FID is not an appropriate metric for conditional generation. The accuracy is computed by matching the predictions of the ResNeXT classifier and the pre-specified labels in the dataset. The IS of generated samples XgX_{g} is defined as

where p(x)p(x) is the logit vector of the ResNeXT classifier. The IS and accuracy are the larger the better.

In the neural vocoding experiment on LJSpeech, we evaluate the speech quality with the crowdMOS tookit (Ribeiro et al., 2011), where the test utterances from all models were presented to Mechanical Turk workers. We report the 5-scale Mean Opinion Scores (MOS), and it is the larger the better.

2 Results

We report image generation results under different approximate diffusion processes, approximate reverse processes and SS, the length of FastDPM. Evaluation results on CIFAR-10, CelebA, and LSUN-bedroom measured in FID are shown in Table 1, Table 2, and Table 3, respectively.

We report audio synthesis results under different approximate diffusion processes, approximate reverse processes and SS, the length of FastDPM. Evaluation results of unconditional generation on SC09 measured in FID and IS are shown in Table 4. Evaluation results of class-conditional generation on SC09 measured in accuracy and IS are shown in Table 5. Evaluation results of neural vocoding on LJSpeech measured in MOS are shown in Table 6.

We display some generated samples of FastDPM, including image samples and mel-spectrogram of audio samples, in Appendix B. More audio samples can be found on the demo website. Demo: https://fastdpm.github.io. Code: https://github.com/FengNiMa/FastDPM_pytorch

3 Observations and Insights

We have the following observations and insights according to the above experimental results.

VAR marginally outperforms STEP for small SS. In the above experiments, the two approximate diffusion processes (STEP and VAR) generally match performances of each other. On CIFAR-10, VAR outperforms STEP when S=10S=10, and STEP slightly outperforms VAR when S≥20S\geq 20. On CelebA, VAR slightly outperforms STEP when S≤20S\leq 20, and they have similar results when S≥50S\geq 50. On LSUN-bedroom, VAR slightly outperforms STEP when S≤50S\leq 50, and STEP slightly outperforms VAR when S=100S=100. On SC09, VAR slightly outperforms STEP in most cases. On LJSpeech, VAR slightly outperforms STEP when S=5S=5. Based on these results, we conclude that VAR marginally outperforms STEP for small SS.

Different reverse processes dominate in different domains. In the above experiments, the difference between DDPM and DDIM reverse processes is very clear. In image generation tasks, DDIM-rev significantly outperforms DDPM-rev except for the S=100S=100 case in the LSUN-bedroom experiment. When we reduce κ\kappa from 1.01.0 to 0.00.0 (see Table 1), the quality of generated samples consistently improves. In contrast, in audio synthesis tasks, DDPM-rev significantly outperforms DDIM-rev. When we increase κ\kappa from 0.00.0 to 1.01.0 (see Table 4), the quality of generated samples consistently improves. This can also be observed from Figure 8: DDIM produces very noisy utterances while DDPM produces very clean utterances.

The results indicate that in the image domain, DDIM-rev produces better quality whereas in the audio domain, DDPM-rev produces better quality. We speculate the reason behind the difference is that in the audio domain, waveforms naturally exhibit significant amount of stochasticity. The DDPM reverse process offers much stochasticity because at each reverse step ss, x^s−1\hat{x}_{s-1} is sampled from a Gaussian distribution. However, the DDIM reverse process (κ=0.0\kappa=0.0) is a deterministic mapping from latents to data, so it leads to degrade quality in the audio domain. This hypothesis is also aligned with previous result that the flow-based model with deterministic mapping was unable to generate intelligible speech unconditionally on SC09 (Ping, 2021).

The amount of conditional information affects the choice of reverse processes. In audio synthesis experiments, we find the amount of conditional information affects the generation quality of FastDPM with different reverse processes. In the unconditional generation experiment on SC09, DDPM-rev (which corresponds to κ=1.0\kappa=1.0) has the best results. When there is slightly more conditional information in the class-conditional generation experiment on SC09, DDIM-rev with κ=0.5\kappa=0.5 has the best results and slightly outperforms DDPM-rev. In both experiments DDIM-rev with κ=0.0\kappa=0.0 has much worse results. When there is much more conditional information (mel spectrogram) in the neural vocoding experiments on LJSpeech, DDPM-rev is still better than DDIM-rev, but the difference between these two methods is reduced. We speculate that adding conditional information reduces the amount of stochasticity required. When there is no conditional information, we need a large amount of stochasticity (κ=1.0\kappa=1.0); when there is weak class information, we need moderate stochasticity (κ=0.5\kappa=0.5); and when there is strong mel-spectrogram information, even having no stochasticity (κ=0.0\kappa=0.0) is able to generate reasonable samples.

Conclusion

Diffusion models are a class of powerful deep generative models that produce superior quality samples on various generation tasks. In this paper, we introduce FastDPM, a unified framework for fast sampling in diffusion models without retraining. FastDPM generalizes prior methods and provides more flexibility. We extensively evaluate and analyze FastDPM in image and audio generation tasks. One limitation of FastDPM is that when SS is small, there is still quality degradation compared to the original DDPM. We plan to study algorithms offering higher quality for extremely small SS in future.

References

Appendix A Derivations for Diffusion Model

According to the definition of diffusion process, we have

where each ϵt\epsilon_{t} is an i.i.d. standard Gaussian. Then, by recursion, we have

As a result, q(xt∣x0)q(x_{t}|x_{0}) is still Gaussian. Its mean vector is αˉtx0\sqrt{\bar{\alpha}_{t}}x_{0}, and its covariance matrix is (αtαt−1⋯α2β1+⋯+αtβt−1+βt)I=(1−αˉt)I(\alpha_{t}\alpha_{t-1}\cdots\alpha_{2}\beta_{1}+\cdots+\alpha_{t}\beta_{t-1}+\beta_{t})I=(1-\bar{\alpha}_{t})I. Formally, we have

Appendix B Generated Samples in Experiments

B.2 Unconditional Generation on CelebA

B.3 Unconditional Generation on LSUN-bedroom

B.4 Unconditional Generation on SC09

B.5 Conditional Generation on SC09

B.6 Neural Vocoding on LJSpeech